How I build with AI agents.
Agents write my code. I decide what gets built, write the spec, review every change, and make sure something other than the agent that wrote a fix proves it works. These are the systems I built to work that way every day.
- Role
- Designer, spec writer and reviewer
- Status
- Dispatch runs my own projects. The bug pipeline is retired
- When
- July to October 2026
- Built with
- Claude Agent SDK, TypeScript, SQLite, SwiftUI
After logging a ride, the card stops responding to touch.
- Platform
- Android
- Build
- 21
- Where
- Ride logged card
- Severity
- Crash-class
Next: the Reproducer tries it on both phones
- 1Log a ride
- 2Let the rating timer run out
- 3Tap Rate, Close or Maybe later
- 4Nothing responds. The card stays.
The card closed itself from inside a state update. It now closes after the update.
Sealed until the Verifier finishes
Only swiping works. Rate, Close and Maybe later still do nothing. Back to the Fixer with the evidence.
From the Verifier: Rate, Close and Maybe later still do nothing after fix 1.
A latch was set before the screen finished closing, so one failed close blocked every tap after it.
All five ways close the card and the app takes taps again. Recordings are in the packet.
- Reproduced on iOS and Android
- Two faults found and fixed
- 5 of 5 ways to close the card work
- Steps, recordings and screenshots kept
Both fixes are in this build.
1 / 9A report comes in from the beta: after logging a ride, the card stops responding to touch.
The problem
An agent will tell you it's done. Sometimes it's right. "Fixed" from the agent that wrote the fix isn't evidence: its tests can pass while the bug is still there, and a fix can pass every test and never reach the running system.
Agents are fast and tireless, so the work that matters most moves to everything around the code: a spec clear enough to get right, a review that reads the actual change, and a check that doesn't trust anyone's summary.
What I built
Dispatch: a team of agents that keeps working
Each project has teams, and each team has a lead and agents running on the Claude Agent SDK. Work moves through a durable queue that leases, retries and sets aside what keeps failing. Agents open merge requests for review, and questions that need me reach my phone.
A bug pipeline: reproduce, fix, verify
For TrackR's beta, one agent reproduced each report on iPhone and Android simulators and recorded an evidence packet. Another wrote the fix. A third re-ran the original steps before reading the fixer's notes.
A gate that doesn't take the worker's word
A worker's change runs in a throwaway copy of the repo. The gate runs the tests itself, shows the real diff, and never merges on its own.
Evals that have to fail first
Dispatch has 15 eval cases that run on real models. A case isn't trusted until it fails on the buggy code and passes on the fix.
The hard part
A fix that was only half a fix.
In TrackR's Android beta, the ride card stopped responding to touch after you logged a ride. The fixer's first change went to the verifier as a new build.
The verifier didn't read the fixer's notes first. It re-ran the original steps, then tried all five ways to close the card. Swiping worked. Three were still dead, and the fifth hadn't been retested, which it flagged instead of quietly counting as passed. Its verdict was Conditional, not Pass.
That pointed to a second fault: a flag set before the screen change had finished, which locked the card for good. Two builds later all five paths passed, and only then did the bug count as fixed. The packet for each verdict kept the steps, screen recordings and screenshots.
Working with a team
What a day with Dispatch looks like.
I talk to a team the way I'd talk to people: in a thread. The lead breaks the work down and hands it out. Changes come to me as merge requests with tests, and when a decision is mine, it's asked as a question I can answer from my phone.
A made-up project with made-up messages.
When it broke
Every failure gets one structural fix.
A builder took down the reviewer
One night a builder agent cleared a stuck test by killing every process that matched a pattern. That included the reviewer's session, and eight hours of overnight work never ran. Now a test fails the build if anyone kills by pattern, and a watchdog with no model in it runs outside the session and restarts stopped work.
A fix that passed its tests and never ran
For four days my screen filled with delivery errors: 232 of them were the exact error a fix had already solved. The running engine was older than the fix, and nothing compared the two. Now the engine reports when it's stale, and an eval checks it.
An eval that couldn't fail
The first test for a queue bug passed on the broken code, because a steady queue never reached the failing branch. The rewrite adds agents while the queue drains, and fails on the old code as it should.
By the numbers
From Dispatch's git history and test files (October 7, 2026) and the bug pipeline's verdict archive (August 2026).
Stack and role
I own the product, write the specs and make the rulings. Builder agents, Claude and Grok, wrote the code, and a reviewer agent and I checked each change against the spec. All 1,771 commits are under my name, with the agent that wrote each one credited in the message.
Want to build something together?
I'm looking for a team to build with full time. Email is the fastest way to reach me.