Caleb Lanting
AI evaluation

Grading AI writing with evidence, not vibes.

For TrackR's article pipeline I built a judge that has to quote every flaw it charges for, works out the score in code, and runs five times to beat its own noise. The same habit, measure before deciding, shapes which models I pick and how many agents I use.

Role
Solo, built with AI coding agents
Status
Internal. Paused since August 2026
When
June to August 2026
Built with
Python, Claude, DeepSeek
Full rewriteGrading

In the ever-evolving world of theme parks, Fernhollow Park is making waves with Timber Wren, a wooden coaster that weaves through the north woods. It's more than a ride; it's a journey back to the park's roots, and one coaster critic already calls it "the ride of the decade." Timber Wren is set to captivate guests of all ages when it opens in 2027.

Draftv1

News

Fernhollow Park announces Timber Wren for 2027

Fernhollow Park announced Timber Wren on Tuesday, a wooden coaster opening next spring in the park's north woods. Timber Wren is not just a coaster, but a 3,100-foot ride through the trees. Riders climb a 112-foot lift hill — and then drop into a breathtaking string of airtime hills.

Park staff say 92% of guests asked for a wooden coaster. Designers delved into the park's archives and borrowed the final turn from Fernhollow's first coaster, which closed in 1979. Timber Wren opens in May 2027.

Versions
  1. v1not graded
JudgeFive runs. The median counts.
Not graded yet
Median of 5 runs

A hard fail counts only when most of the five runs flag it.

Deductions in the median run

    Worked out in code from the points. The model never does the math.

    RubricHard tells 40Voice 25Micro-patterns 15Structure 10Sourcing 10

    Made-up article, saved judge runs. The real system calls a model each time.

    The problem

    Ask a model "how good is this article, from 0 to 1?" and you get a number you can't argue with. Mine gave the same article 0.00, 0.30 and 0.40 on three runs.

    Feedback that vague also makes a bad editor. When the writer rewrote whole sections to chase the score, one article dropped from 0.72 to 0.45.

    What I built

    Four judging gates every article had to pass.

    • A judge that has to show its work

      Every deduction names a category, the points, the exact words it's charging for, why, and the fix. If it can't quote it, it can't charge for it. The score is 1 minus the points, added up in code, not by the model.

    • Five runs and a median

      The judge grades each draft five times at once and takes the median. Some tells fail an article outright, like made-up citations, but only when most of the runs agree. A single run once threw out a perfect article on its own.

    • A fixer that only touches what was quoted

      The fixer replaces the quoted words and nothing else. The loop keeps the best version so far and throws away any change that scores lower, so scores only move up. A problem with no quote to fix stops the loop for a person.

    • Three more gates

      Graders for AI tells and for voice, a fact-check against the live web, and a vision judge that checks each image against the house art style.

    The hard part

    A judge shouldn't remember its last verdict.

    A judge that can read its earlier verdicts drifts. It waves a draft through because it was revised, or keeps failing it because the last version failed. Either way it stops reading the draft in front of it.

    So the judge is barred from earlier verdicts and grades every round fresh. Fresh grades of one draft went 0.92, 0.91, 0.84, 0.92, and each found a different real problem the earlier rounds had missed, which is what you want from a reviewer.

    The pass bar stayed at 0.90. Raise it to 0.95 and the judge starts inventing nitpicks to justify the gap.

    Measuring before deciding

    Do more agents build better software?

    I ran the same build brief two ways: a manager agent directing lead agents directing workers, and one agent working alone. A fresh model judged both blind, against a rubric written before either run.

    The team scored 100 and the solo run 99, but the team took more than three times as long and cost about 3.2 times as much. In a second round, a cheaper model with a complete brief and a 17-check acceptance harness scored 98 at roughly a third of the solo cost.

    It was one small build, so it's a signal, not a law. It changed how I delegate: the effort goes into the spec and the checks. The same habit chose LedgR's note reader: Haiku read 38 of 40 messy notes correctly; Sonnet read 27.

    Controlled experiment · July 21, 2026

    A ladder of agents (a manager, leads and workers) against one solo agent. Same build brief, judged blind against a frozen rubric.

    • Ladder: manager → leads → workers
    • One solo agent

    More agents didn't buy better work here. A written spec and real checks did.

    Round 2A cheaper model plus a 17-check acceptance harness scored 98 against the solo run's 99, for about $1.90 to $2.60 against $6.35, roughly a third of the cost.

    One small build, about a 30-minute project. Costs are list-price equivalents for the tokens used, not money spent.

    By the numbers

    4judging gates on every article
    5judge runs per grade, median taken
    0.72 → 0.45what whole-section rewrites did to one article
    16,446lines of Python in the pipeline
    3.2×the cost of an agent team over one agent, for one more point

    From the article pipeline's code, policy file and commit history (June to August 2026), and the experiment's results file (July 21, 2026). Costs are list-price equivalents.

    Stack and role

    I designed the judging rules and the loop, and worked out why it went wrong when it did. AI coding agents wrote the Python. I designed the experiment, froze the rubric before the runs, and read the results.

    PythonClaudeDeepSeekGeminiVision modelsLLM-as-judgeBlind evaluation

    Want to build something together?

    I'm looking for a team to build with full time. Email is the fastest way to reach me.