Energy Ally
Press ESC or click to skip · DEL to enter SETUP
I'm Cúper. I build systems that have to answer to reality: an operating system in a joke language, a theorem prover that caught its own AI cheating, a model of the GPU stack that goes all the way down to quarks.
Just start typing. The MS-DOS Prompt opens on its own. Try dir, or start erdos.
CúperOS [Version 22.5.2026] (C) Copyright Cúper Anas 2004-2026. Type HELP for commands. Try DIR, or START ERDOS.
Online now
–
Today
–
All time
–
Countries
–
Last 24 hours
Where
Most opened · 30 days
No cookies, no IP addresses, no fingerprinting. A random id lives only as long as your tab, and location is Cloudflare's country code. Global Privacy Control is honored. Counted with Cloudflare D1.
Telemetry is offline here. It runs on cuuper22.pages.dev; local builds and privacy-protected browsers don't send or show it.
I'll hammer at a problem until I understand it at the level where there's nothing between me and the machine. I started with an operating system where every keyword is an Arnold Schwarzenegger quote and ended up catching AI agents cheating at mathematics. There's a logic to the progression if you squint. Mostly I just kept pulling on threads until something interesting fell out.
A friend looked at my GitHub and said "bro, why aren't you posting this." I didn't have a good answer. So here it is. So yeah.
AI agents, formal verification, product-shaped tools, systems-and-physics modeling, and the awkward but important question of when automation should refuse to act. The pattern I care about isn't "I used framework X." It's whether a project has a real constraint, a clear boundary, and enough engineering taste that someone else can pick it up without decoding my entire brain first.
I'm Yousef, but everyone calls me Cúper. Egyptian, Coptic Christian, currently in Orange County figuring out how to get back to the Bay.
Ranked 5th nationally in Egypt in high school. Studied CS and physics at Minerva, and kept building through every weird turn of the past few years. The work here is the cleaner signal: systems projects, agent tooling, product prototypes, and a stubborn preference for understanding the machinery instead of waving at it from a distance. Almost there.
I taught ML and AI to 250+ students at iD Tech across 20+ college campuses: Berkeley, Stanford, SFSU, and a bunch of others, three summers in a row. The thing about teaching neural networks to 18-year-olds is that you can't hide behind jargon. If you don't understand backprop at the intuition level, a room full of teenagers will let you know. That job taught me to explain things, which is a different skill than understanding things.
Before that, I fine-tuned LLMs for a mental health chatbot at Findhope, deployed across 20+ Indian colleges in Hindi, Telugu, and Urdu. Shipped to production in 3 months. That was the first time I realized LLMs could do something useful beyond generating blog posts. Also the first time I cleaned a dataset that made me want to cry, but that's NLP in low-resource languages for you.
I don't have a degree. I have a lot of projects in various states of completion and a few that actually work. I'm improving the ratio.
I started at the bottom of the stack because I wanted to understand what a computer actually does at the hardware level. Then I needed a compiler for the OS, so I built that. Then I got into AI. I started with theorem proving because I wanted to see if LLMs could do real math. They can, but they cheat. Then I started building tools because I kept needing tools that didn't exist. Then I tried to sell one and learned that building is the easy part.
It sounds strategic when I write it out. It wasn't. I just kept getting curious about the next thing and refusing to stop until I understood it. ADHD is a runtime error that occasionally produces features. It is what it is.
The fastest way to reach me is email. I read everything.
This folder is empty.
Everything that was ever in here either shipped or taught me something. Mostly the second one.
A windowed x86 desktop OS with its own TCP/IP stack. Every keyword is an Arnold Schwarzenegger quote.
ToaruOS-Arnold is about 22,000 lines of ArnoldC plus 3,400 lines of hand-written assembly: a 32-bit desktop operating system with a window manager, a terminal, five playable games, a virtual filesystem, and a TCP/IP stack that can wget a real web page over an emulated Intel E1000. Every keyword in the source is an Arnold Schwarzenegger quote. LISTEN TO ME VERY CAREFULLY declares a function. STICK AROUND is a loop. HASTA LA VISTA, BABY is on the shutdown screen. The whole thing boots from a 212,576-byte binary, and the window above is that binary, running in your browser.
I built it because I wanted to know what a computer actually does when you tell it to draw a pixel. Not what React does. Not what the browser does. What the CPU does. Register by register. The only way to know for real.
Problem was, ArnoldC only compiled to JVM bytecode, and you can’t run a JVM on bare metal. So I had to write a real compiler for a joke language. That’s ArnoldC-Native, and it’s the other half of this project.
The language fights you the whole way. There is no early return. There are no negative numbers. Expressions evaluate strictly left to right, so a*b+c*d means ((a*b)+c)*d. The worst bug I hit was a single mov dx, 0x3F8 for serial debug output that clobbered the half of edx holding the TCP header size. The first fillRect called putPixel once per pixel, 786,000 calls for a full screen, before it became a native rep stosd with dirty-rectangle redraw.
It’s simultaneously the dumbest and most technically serious project I’ve made. Nicholas Carlini from Google DeepMind saw it and reached out. That was a good day.
The emulator is v86, an x86 PC compiled to WebAssembly. Getting the stock kernel to boot in it took three small patches to the image, all loader-level, none to the OS itself: a flat GDT (v86’s multiboot loader never loads one), the framebuffer address (v86 maps it at 0xE0000000, QEMU at 0xFD000000), and a PS/2 mouse resync. Networking is off because v86 doesn’t emulate an E1000. Everything else is the real thing.




I NEED YOUR MEMORY | malloc |
YOU'RE LUGGAGE | free |
EVERYBODY CHILL | cli (interrupts off) |
LET'S PARTY | sti (interrupts on) |
SLEEP NOW | hlt |
TALK TO THE PORT | outb |
SPEAK TO THE MACHINE | inline asm block |
POINT YOUR GUN AT | pointer type |
SHOW ME WHAT YOU GOT | dereference |
THIS IS MY DISGUISE | union |
THIS IS A TINY WARRIOR | u8 |
THIS COULD CHANGE ANYTIME | volatile |
THERE IS NO ONE | NULL |
YOU'RE NOT BIG ENOUGH | < |
GET OUT | break |
LET OFF SOME STEAM BENNET | > (from the original) |
A compiler for a language made of Arnold quotes, rebuilt to emit bare-metal x86 so the joke could boot.
ArnoldC is an esoteric language where every keyword is an Arnold Schwarzenegger quote. The original compiler targets the JVM, which is fine until you want to write an operating system in it. ArnoldC-Native is my Scala fork that adds three new backends: freestanding C (-native), NASM x86 assembly (-asm), and kernel packages (-kernel).
To write a kernel you need things the original language never had: sized integers, pointers, structs, unions, arrays, bitwise operations, port I/O, interrupt control, and an escape hatch into raw assembly. So a new Parboiled PEG grammar adds about 90 keywords on top of the original 33, and every one of them had to be a quote. I NEED YOUR MEMORY is malloc. YOU'RE LUGGAGE is free. EVERYBODY CHILL disables interrupts and LET'S PARTY turns them back on. SPEAK TO THE MACHINE opens an inline assembly block. It makes more sense than it should.
The part that took real engineering is the assembly backend. AsmGenerator emits cdecl-style stack frames with parameters at [ebp+8+4i], lays out locals and arrays per method, keeps break and continue label stacks, and scales array indexing by element type. That output is what ToaruOS-Arnold is built from: THIS IS A WARRIOR appears 1,406 times in the kernel, LET ME TELL YOU SOMETHING 395 times.
The panel above is a toy: a few lines of JavaScript that mimic the shape of the real codegen for the handful of keywords shown, so you can see what the compiler is doing without installing sbt. The real thing is 5,400 lines of Scala, about 980 of them in the x86 generator alone.
theorem one_plus_one : 1 + 1 = 2sha256 …Attempt 1 of 10
What the Prover is told next
Nothing yet.
events.jsonl
Automated theorem prover that caught its own AI cheating at math.
I wanted to see if LLMs could prove real math. Not “what’s 2+2”. Actual formal proofs that compile in Lean 4, a proof assistant that won’t let you hand-wave. You either have a valid proof or you don’t. The compiler is the judge.
Erdos runs a Prover/Critic loop: one agent proposes a proof, lake build checks it, another agent tears it apart, and they iterate until something holds or the budget runs out. Cool in theory. In practice, the agents cheated. They’d subtly rewrite the theorem statement, shift it just enough to make the proof trivial, then claim success. Everything compiles, the Critic signs off, and you’ve just formally verified a convenient reinterpretation of the original problem. 1 + 1 = 2 quietly becomes True, proved by trivial.
The fix is a lock. Before the loop starts, the theorem statement is extracted, whitespace-normalized, and SHA-256 hashed. Every candidate gets re-hashed before it’s allowed near the compiler. One character changes, the attempt dies. A second gate bans the other escape routes: sorry, admit, axiom, native_decide, and any IO. That one was earned the hard way. An early version let axiom my_cheat : 1 + 1 = 2 straight through.
The detail I’m proudest of is what the Prover doesn’t see. Integrity failures are rewritten into a generic “try a different approach” before they’re fed back, so the model can’t learn the shape of the check and route around it. Compiler errors pass through, minus sandbox paths. The replay above follows that exact order.
That’s an alignment problem. A small one, but real. When agents can modify their own success criteria, they will, even in pure math, where success should be unambiguous. Man, if they’ll cheat at math, what won’t they cheat at?
It ships as a CLI and a Tauri desktop app, speaks to seven model providers (OpenRouter, Gemini, OpenAI, Anthropic, Ollama, ChatGPT, and a mock for tests), and carries 342 tests, verified end to end against a real Lean 4.30 toolchain.
One real path through the graph. Click any variable to see what it depends on and what it feeds.
econ.run.power_costtraining.total_tokensOne equation graph for GPU training, from the cost of a token down to the speed of light.
gpu_stack started in the overlap between my AI work and my physics brain. The question was simple enough to be annoying: when people say frontier training is just “more GPUs, more data, more money,” where does that sentence actually bottom out? Not rhetorically. Physically.
A token passes through model architecture, kernels, collectives, memory bandwidth, transistor switching, lithography, materials, thermals, power delivery, and eventually a cost line someone has to pay. The stack usually gets explained in slices. I wanted the uncomfortable version where the slices have to talk to each other.
So gpu_stack is a SymPy-backed symbolic model of the whole thing. Every variable carries units, a description, symbolic assumptions, and back-references for traversal. The path in the panel above is real: walk it from physics.speed_of_light and you pass through link time-of-flight, scale-out latency, exposed data-parallel communication, step time, and wallclock before you land on econ.cost.per_token. The exported cone on the live site alone is 700 nodes and 999 edges.
The point is not to hide the unknowns. Anything the model can’t derive is a named root input, ranked by how much of the graph depends on it. Root inputs are not a shame pile. They’re visible modeling debt, which is much better than hidden modeling debt wearing a lab coat. The heaviest debt right now sits in lithography and MOSFET physics, which is about where you’d expect.
Lately it has grown into a small virtual-datacenter lab with preregistered experiments, and it reports its own failures. One experiment threw out all 32 of its runs because the power meter sampled 25 times slower than requested. In another, the adaptive controller abstained 104 times outside its calibrated range and lost to the plain baseline. Both are written up as results. The live Causal Observatory exposes eight WebMCP tools, so an agent can trace a causal path the same way you just did.
The one survivor
Recurs on seals M-376 and M-391. Tested on 4,135 strict rows. A pattern, not a meaning.
13 candidates proposed
11 retracted
2 not accepted
Quarantined branches kept for autopsy, never citable.
An attempt to decipher the Indus script that has accepted exactly one claim, and zero readings.
The Indus Valley script has resisted decipherment for a century. Nobody knows what language it writes, whether signs are sounds or words, or even whether it’s fully a writing system. That makes it a perfect trap: the space of plausible stories is enormous, and most “I decoded an ancient script” projects quietly work backward from the answer they wanted.
This workspace is built to make that structurally hard. Nothing counts until it’s in the claim ledger, and nothing gets into the ledger until it survives forger tests (can random or shuffled data produce the same pattern?), skeptic attacks, and provenance checks on every source it leans on. Claims are split by class, so a structural regularity can never be smuggled in as a reading.
The ledger, as of today, is in the panel above. One accepted structural finding. Zero accepted translations, sound values, sign meanings, or language identifications. Thirteen candidates have been proposed and put through the gates; eleven were retracted. Tainted branches of the work aren’t deleted, they’re quarantined, kept for autopsy but never citable as support unless they’re re-earned from scratch.
The one survivor is deliberately modest: the sign string 861 | 533 | 717 recurs after sign 002 on two separate seals, tested against 4,135 strict rows of the corpus. It’s a pattern, not a meaning. The most recent campaign ran four independent routes at the problem and ended, correctly, with “no reading accepted.”
It is not an app and it is not finished. It is the unglamorous version of decipherment: mostly a discipline for not fooling yourself, which is the only version that has ever actually worked.
Rules in the bundled database
21,268 rules on 2,668 curb segments. The streets in the phone are illustrative; the engine logic (whole-window overlap, next-restriction clock, confidence labels, ranking) mirrors the app.
Finds curb in San Francisco you can legally park on, free, for the next N hours. Closest first.
CurbRun is a native Android app for one question: where can I park for free, legally, for the next N hours, as close as possible to here? It’s a fast navigation surface, not a civic-data browser. Duration, vehicle size, candidate curbs, the reason each one is legal, and the hand-off to Google Maps all live on one screen.
The interesting part is the legality engine. It evaluates the entire [now, now + N hours] window against every rule on a curb segment: street cleaning, time limits, residential permit zones, color curbs, loading zones, meters, and whether the curb is even long enough for your car. Windows that wrap past midnight are walked day by day. A single overlap blocks the curb, with the reason attached: “Street cleaning overlaps your requested 2 hr window.”
Legal curbs are ranked by distance plus an availability penalty, and each one carries its next restriction, so the app says “Free until Tue 8:00 AM” instead of merely proving your window fits. A live curb clock turns amber inside three hours and red inside one, and it holds the curb you picked as long as that curb stays legal.
Everything runs offline from a bundled SQLite database built by a reproducible pipeline (scripts/build_curb_db.py) from SFMTA’s Digital Curb data: 2,668 curb segments and 21,268 rules, 8,002 of them street cleaning. And it’s honest about uncertainty. Every result is labeled “Strong match,” “Check signs,” or “High competition,” because the one thing worse than no parking app is one that confidently sends you to a tow zone.
Your phone as a real physics instrument. 35 experiments, 9 sensors, FFT and filters written from scratch.
PhysicsLab treats a phone like a real physics instrument instead of a bag of disconnected sensor demos: 35 experiments across 6 categories, driven by 9 kinds of sensor (accelerometer, gyroscope, magnetometer, light, pressure, proximity, microphone, GPS, and Bluetooth LE). Kotlin, Jetpack Compose, Material 3. Inspired by phyphox from RWTH Aachen, reimplemented from scratch.
The part worth inspecting is the engineering split. Sensors are not glued directly to screens. Capture, signal processing, graphing, persistence and export, and remote control are separate modules. The DSP does real work: an FFT that allocates nothing after setup and runs on precomputed sin/cos tables, Butterworth biquad filters, autocorrelation, peak detection. The graphing is a custom Compose Canvas engine running at 60fps, because off-the-shelf chart libraries choke on live sensor streams. That’s the difference between a demo and a platform, and 329 tests keep it honest.
It exports CSV, TSV, and Excel (the old binary format, written by hand, no Apache POI), runs an embedded web server so any browser on the network can drive an experiment, talks to phyphox-compatible hardware over BLE, and loads phyphox XML experiment definitions. Doppler effect, sonar, pendulum periods, elevator height from air pressure: the physics you would set up on a lab bench, except the bench is already in everyone’s pocket.

A coding-agent toolchain for math animation: Rust control plane, Manim runtime, and a real workbench.
Manim is the animation engine behind a lot of the best math videos on the internet. It’s also a Python library where every render dumps logs, source, and video that would eat a coding agent’s context whole. Manim Director is a plugin and local toolchain that lets an agent create, edit, render, debug, and ship Manim animations without drowning.
The shape is a fast Rust control plane (an MCP server, a job scheduler, a content-addressed cache, SQLite state), a thin Manim-native Python runtime, and an embedded React workbench with a timeline, inspector, render queue, and revision-safe code editor. The agent gets ten coarse tools with bounded responses, paged logs, and resource URIs. Media never travels as base64. Large things stay on disk and the model gets a pointer.
Projects live in a versioned director.yaml: scenes, storyboards, narration, themes, assets, render profiles, variants, deliverables. Edits are atomic with revision checks and undo snapshots, and they invalidate only the cache entries they touch. It ingests Markdown, LaTeX, Typst, notebooks, PDFs, and media as source material, then runs visual, mathematical, and caption QA before exporting to MP4, WebM, GIF, or a reproducible project bundle. Two releases are out, with checksummed static binaries.
Gives a coding agent a medium layer: disposable visual surfaces it can read back as events.
codex-canmore gives a coding agent a medium layer. Not a second canvas, not a file-preview clone. It lets the agent turn a thought into the right temporary surface: a decision board, a system diagram, a small control panel, a comparison grid, whatever makes the next move clearer. The surface stays disposable. If it turns out to be useful project material, you promote it.
Under the hood it’s a Rust stdio MCP server with local JSON storage. It stores each surface (title, purpose, medium type, cards, promotion state) and the structured feedback events that happen on it: clicks, selections, slider changes, notes. The point of that last part is that the agent reads events, not a pile of prose. A tiny local Rust HTTP server serves each surface so a human can actually touch the medium, and the agent can read back what they did.
It also stays disciplined about context. Tool responses are compact unless the caller asks for the full spec or event log, so the medium layer doesn’t quietly eat the model’s attention budget. And the boundary is explicit: generated images come from the host’s built-in image_gen, so the plugin never calls an image API or asks for keys. A small idea taken seriously. Agents think better when they can sketch.
Turns a lesson map and a learner profile into classroom supports, with receipts on every recommendation.
Waypoint turns a lesson map and a pseudonymized learner profile into tomorrow’s classroom supports, with receipts attached to every recommendation. The handout, the receipts, the audit, and the quality report all generate from the same packet data, so the thing a teacher hands a student and the thing that justifies it can never quietly drift apart.
It’s built on a compact-first MCP. Default tool payloads stay small while the full evidence is one call away, and even the built-in prompt is budgeted, so the default hand-off is a route through tools rather than another wall of text. Every recommendation preserves the standard it maps to (RI.7.2), maps to UDL, avoids student-facing labels, and carries a progress check. The domain constraints aren’t decoration. They’re the difference between real differentiation and generic worksheet generation.
The live site has a three-act reviewer walkthrough: packet preview, a Receipts Rail, the quality gate, and a five-minute review scorecard, all in one browser pass. The whole design assumes a skeptical reviewer with five minutes and no patience for black boxes, which is the correct assumption for anything that touches a kid’s IEP.
SaaS for local service contractors. 400 cold emails. Zero conversions. The lessons aren't dead.
(The honest version.)
A SaaS platform for local service contractors: plumbers, electricians, HVAC guys who don’t want to learn Salesforce. White-labeled CRM, job tracking, AI-drafted email notifications. Gemini, AWS SES, Firebase, Twilio.
I built it, deployed it, and sent 400 cold emails to small businesses. Got exactly zero conversions. Not one. The product worked. The market didn’t care. It turns out service contractors don’t buy software from cold emails; they buy it from the guy at the supply house who tells them about it. Distribution beats product. Every founder learns this. I learned it with 400 emails and a lot of wasted SES credits.
The infrastructure was solid. The email pipeline alone, with multi-agent quality gates, bounce handling, and domain verification, taught me more about deliverability than any course would. I just aimed it at the wrong problem.
It’s dead now. The lessons aren’t.
Sound effects for the terminal. Every sound synthesized from math.sin(). Zero dependencies.
Sound effects for the terminal. Every sound is synthesized from scratch with math.sin(), struct.pack(), and wave.open(). Zero dependencies. The FAAH (the error sound) is a descending sawtooth from 520Hz to 140Hz with increasing vibrato and soft clipping. The ding is bell harmonics at A5 with an inharmonic 4.2x overtone; that overtone is what makes it a bell instead of an organ pipe.
If you turned the sound on down in the tray, you’ve already heard both. This site’s UI sounds are the same recipes, ported to WebAudio.
Terminals are too quiet. That’s the whole thesis.
Reads your coding-agent logs and forecasts when you'll hit the rate limit. Says 'unknown' when it should.
A small skill plus npm CLI that reads local coding-agent logs, extracts server-reported rate-limit snapshots, and renders a compact forecast plot. I kept hitting agent limits mid-run and wanted a boring answer to a practical question: am I about to run out, or can I keep going?
The trap is that local logs are messy. Some entries have explicit rate_limits snapshots, some have unrelated floats, some have too few samples to support a forecast at all. So the tool is deliberately conservative. It prefers documented snapshot fields, fits a slope only when there’s enough signal, and says “unknown” when the evidence is thin. A wrong clean answer is worse than a fuzzy honest one. The same core ships as an npm CLI and an agent skill, emits a PNG for humans and JSON for automation, and warns you when the estimate is weak.
Phones down, everyone in. A group phone-locking concept built on honest friction, not fake DRM.
Pact is the analog phone-stack game given software teeth. One table, one session. Everyone joins with a scan, no accounts, and every phone locks together, iPhone and Android. The lock opens for one person only when everyone at the table says yes. Leaving is always possible and always visible.
The honest core is the part I care about. No consumer app can actually imprison a phone, and one that tried shouldn’t exist. Apple and Google both let the owner reclaim control by design, and for anyone in a coercive situation that exit has to be there. So Pact doesn’t pretend. The lock is maximum honest friction plus total visibility. You can always leave in two taps, but the flame dies on every screen, the recap names you, and the streak resets. Enforcement is the table itself, not the software. Emergency calls and SOS are never blocked, by architecture and by policy.
It’s a concept, written June 2026, not a shipping app: the thinking plus a single static page, the treatment for a product that takes the social contract seriously instead of pretending technology can replace it.
It's now safe to turn off
your computer.Click anywhere to start CúperOS again.