TinyOrbit is the from-scratch agent harness at the heart of Chapter 6 of my book, Token by Token. This is the short story of what it is, how it works, and what happened when I pointed it at real-world software bugs.
The magic trick, explained
AI agents feel like magic. You type “fix this bug,” go make a hot chocolate, and come back to a patch. Surely there’s something enormous under the hood, a planning engine, a reasoning module, a small city of microservices?
There’s a loop.
That’s the secret. The model is the brain: it does the thinking. Everything around it is called a harness: the loop that keeps the conversation going, the tools the model is allowed to touch, and the safety rails that decide what it may touch. I wanted to prove how small that harness really is, to myself, and to the readers of my book, so I built one from scratch. It’s called TinyOrbit: a small pile of plain Python, one third-party dependency, no framework, small enough to read in an afternoon, which is more than I can say for my hot chocolate machine’s firmware.
Look at the sketch. The head is the model, the brain. It reads the task, reasons about it, and decides what to do next. But a brain on its own can’t open a file or run a test. It can only think.
Everything below the neck is the harness, and the closest thing to it in us is the limbic system: the old, unglamorous wiring that sits between thought and action. It carries the brain’s intent out to the hands, feeds back what the senses find, remembers what happened a moment ago, and flinches before you touch the hot stove. It isn’t clever, but without it the cleverest brain just sits there.
Here’s how the two work together as an agent:
Think. The model reads the conversation so far and picks a move: “read models.py,” or “run the tests.”
Act. The harness carries that request to the hands, read, grep and glob in one, edit, write and bash in the other, and does exactly that, nothing more.
Observe. Whatever comes back, file contents, test output, a stack trace, the harness hands straight back to the brain as new input.
Repeat. Round and round, the orbit in TinyOrbit’s name, until the model answers without asking for a tool. That’s the loop saying “done.”
Brain plus limbic system is a person who gets things done. Model plus harness is an agent. And that hatched belt around the waist? That’s the flinch: the permission gate that stops and asks before anything risky.
Everything else is seatbelts.
The seatbelts are the interesting part
Here’s what building it taught me: the loop takes an hour. The trust takes the rest of your life.
TinyOrbit’s tools are deliberately boring, read, write, edit, bash, glob, grep. The interesting engineering, the part that makes it a harness rather than a demo, is in what wraps them:
Permissions, fail-closed. Reading files is free. Writing files or running shell commands? The agent stops and asks you first. An agent that can quietly run rm -rf is not an assistant, it’s a liability with good manners.
Staleness detection. If the agent read a file, and you edited it behind its back, it refuses to blindly overwrite your changes. (A courtesy some humans have yet to master.)
Interruption as a first-class citizen. Hit Ctrl+C mid-task and the agent doesn’t just die, it tidies up its half-finished tool calls so the conversation stays coherent. Every request the model makes gets an answer, even if that answer is “the human pulled the plug.”
Memory files. Drop a TINYORBIT.md in your project with your conventions, and the agent reads it on startup, the difference between a contractor who read the brief and one who’s about to repaint your load-bearing wall.
None of this is glamorous. All of it is the difference between a demo and a tool.
OK, but does it actually work?
Fair question. Vibes are not a benchmark, so I ran TinyOrbit (driving Claude Opus) against SWE-bench Verified, the industry-standard test where agents must fix real, historical bugs from real open-source projects, judged by the projects’ own test suites. No partial credit. The patch either passes or it doesn’t.
On a random sample of ten tasks, TinyOrbit went 10 for 10, for a grand total of $3.39 in API costs, less than the hot chocolate I drank while watching it work. On a deliberately cruel sample of ten hard tasks (the ones humans estimate at 1–4+ hours each), it went 7 for 10. For context, when I scored the strongest published harnesses on those same hard ten, the best resolved 8.
Cost per task
Estimated from the token ledger at Opus 5 list prices. An outlined bar is an unresolved task.
easy samplehard sample
Turns and wall-clock per task
Left: API calls, capped at 60. Right: agent minutes in the container, inflated by x86 emulation on repos with slow test suites.
easyhard
Tool calls per task, by tool
Counted from the full message history. Bash dominates; Glob was never used in either run.
BashEditReadGrepWrite
Before anyone prints a trophy: ten tasks is a small sample, the model deserves a large share of the credit, and the comparison isn’t apples-to-apples since published entries ran older models. The honest reading is model-plus-harness, not “my weekend project beats the industry.”
And the three misses? All three failed the same way, and it’s deliciously human: the agent wrote its own tests, passed its own tests, and confidently declared victory, without ever reproducing the actual bug first. It graded its own homework. Every engineer I’ve told this to has gone quiet for a second, because we’ve all met that developer. Some of us in the mirror.
Why this matters (and where the book comes in)
The lesson of TinyOrbit isn’t “look how clever.” It’s the opposite: this technology is understandable. The agent loop that powers the billion-dollar tools is something you can read, build, and own in an afternoon. Once you’ve seen it, AI agents stop being magic and start being engineering, and you can reason about when to trust them, where they’ll fail, and what the seatbelts should be.
That’s the whole philosophy of Token by Token: the best way to understand AI is to build it. Chapter 6 walks through the agent loop step by step, TinyOrbit for the full harness, plus two small agents that debate each other live, with hand-drawn diagrams and runnable code for every idea. Earlier chapters build everything underneath it: a neural network from a single neuron, an autograd engine from high-school calculus, GPT-2 from scratch, and a model that learns to reason.
The code:github.com/badlogicmanpreet/tinyorbit, including the full SWE-bench reports with every verdict and a wire-level trace of one task, stream event by stream event.