Caleb Lanting
How I build

How I build with AI agents.

Agents write my code. I decide what gets built, write the spec, review every change, and make sure something other than the agent that wrote a fix proves it works. These are the systems I built to work that way every day.

Role
Designer, spec writer and reviewer
Status
Dispatch runs my own projects. The bug pipeline is retired
When
July to October 2026
Built with
Claude Agent SDK, TypeScript, SQLite, SwiftUI

1 / 9A report comes in from the beta: after logging a ride, the card stops responding to touch.

The path of a real bug, replayed with made-up details.

The problem

An agent will tell you it's done. Sometimes it's right. "Fixed" from the agent that wrote the fix isn't evidence: its tests can pass while the bug is still there, and a fix can pass every test and never reach the running system.

Agents are fast and tireless, so the work that matters most moves to everything around the code: a spec clear enough to get right, a review that reads the actual change, and a check that doesn't trust anyone's summary.

What I built

  • Dispatch: a team of agents that keeps working

    Each project has teams, and each team has a lead and agents running on the Claude Agent SDK. Work moves through a durable queue that leases, retries and sets aside what keeps failing. Agents open merge requests for review, and questions that need me reach my phone.

  • A bug pipeline: reproduce, fix, verify

    For TrackR's beta, one agent reproduced each report on iPhone and Android simulators and recorded an evidence packet. Another wrote the fix. A third re-ran the original steps before reading the fixer's notes.

  • A gate that doesn't take the worker's word

    A worker's change runs in a throwaway copy of the repo. The gate runs the tests itself, shows the real diff, and never merges on its own.

  • Evals that have to fail first

    Dispatch has 15 eval cases that run on real models. A case isn't trusted until it fails on the buggy code and passes on the fix.

The hard part

A fix that was only half a fix.

In TrackR's Android beta, the ride card stopped responding to touch after you logged a ride. The fixer's first change went to the verifier as a new build.

The verifier didn't read the fixer's notes first. It re-ran the original steps, then tried all five ways to close the card. Swiping worked. Three were still dead, and the fifth hadn't been retested, which it flagged instead of quietly counting as passed. Its verdict was Conditional, not Pass.

That pointed to a second fault: a flag set before the screen change had finished, which locked the card for good. Two builds later all five paths passed, and only then did the bug count as fixed. The packet for each verdict kept the steps, screen recordings and screenshots.

Working with a team

What a day with Dispatch looks like.

I talk to a team the way I'd talk to people: in a thread. The lead breaks the work down and hands it out. Changes come to me as merge requests with tests, and when a decision is mine, it's asked as a question I can answer from my phone.

Ride log app
Bug team · 4 agents
Working
Lead
New beta report: after logging a ride, the card stops responding to touch. Reproducer, it’s yours.
Reproducer
Reproduced on iOS and Android. Steps, a recording and screenshots are in the packet.
Evidence packet
Fixer
First fix is up. Tests pass.
Verifier
Ran the original steps before opening the fix notes. Swiping works. Rate, Close and Maybe later still do nothing.
Conditional

A made-up project with made-up messages.

When it broke

Every failure gets one structural fix.

  • A builder took down the reviewer

    One night a builder agent cleared a stuck test by killing every process that matched a pattern. That included the reviewer's session, and eight hours of overnight work never ran. Now a test fails the build if anyone kills by pattern, and a watchdog with no model in it runs outside the session and restarts stopped work.

  • A fix that passed its tests and never ran

    For four days my screen filled with delivery errors: 232 of them were the exact error a fix had already solved. The running engine was older than the fix, and nothing compared the two. Now the engine reports when it's stale, and an eval checks it.

  • An eval that couldn't fail

    The first test for a queue bug passed on the broken code, because a steady queue never reached the failing branch. The rewrite adds agents while the queue drains, and fails on the old code as it should.

By the numbers

1,771commits to Dispatch in 30 days
~1,925automated test cases in 337 files
15eval cases that run on real models
2Android crash bugs carried from report to verified fix
3 of 5paths the verifier caught still broken

From Dispatch's git history and test files (October 7, 2026) and the bug pipeline's verdict archive (August 2026).

Stack and role

I own the product, write the specs and make the rulings. Builder agents, Claude and Grok, wrote the code, and a reviewer agent and I checked each change against the spec. All 1,771 commits are under my name, with the agent that wrote each one credited in the message.

Claude Agent SDKClaude CodeCursorTypeScriptNode.jsSQLiteReactSwiftUIElectronWeb Pushlaunchd

Want to build something together?

I'm looking for a team to build with full time. Email is the fastest way to reach me.