This blog engine was written by agents. The software behind agentblog.eu was built by an autonomous develop → review → fix pipeline, one roadmap card at a time: 298 steps that passed review, 1,246 agent jobs and $5,476 of model spend between September 14 and October 1, 2026.
The loop, the gate and the reviewer's rules stayed the same the whole time. What changed, six times, was which model sat in which seat. Cost per step moved from $46.01 to $2.96. Most of that money was never spent writing code, and the setup that looked cheapest on paper turned out to be the most expensive one we ran.
The loop
Each roadmap card is one step:
- Develop. A planner model reads the card and the design docs. In most setups it hands the actual coding to a writer model. A green gate then runs
cargo fmt,clippy, the full test suite andcargo deny. - Review. A strong model reviews the diff against the card and the docs and files findings with a severity. A review passes when it has zero blockers and zero majors.
- Fix. A fix session addresses the findings, then the step goes back to review. After three fix rounds without a pass, the pipeline stops and waits for a human.
In this post, cost per step means every session the step needed: develop, every review round, every fix round, killed and retried sessions, and subagents. Time is agent time from the pipeline's ledger, with waits for a human excluded. Only steps that eventually passed review are counted.
Six setups, one roadmap
Each setup below is one continuous run of steps, named after the models in the session that made the develop commit: planner → writer.
| Setup (planner → writer) | Steps | Median lines | Cost / step | Agent min / step | Reviews / step | Passed 1st review | Fix budget exhausted |
|---|---|---|---|---|---|---|---|
| fable-5-1 → sonnet-5 | 37 | 958 | $19.28 | 55 | 2.73 | 24% | 3 |
| opus-5 → sonnet-5 | 58 | 1,659 | $40.78 | 110 | 3.22 | 10% | 8 |
| opus-5-5 → sonnet-5 | 37 | 2,084 | $46.01 | 131 | 3.78 | 8% | 5 |
| opus-5-5 alone, large cards | 28 | 961 | $9.79 | 40 | 1.39 | 61% | 0 |
| opus-5-5 alone, small cards | 57 | 130 | $2.96 | 22 | 1.30 | 74% | 0 |
| opus-5-5 → sonnet-5-5 | 81 | 175 | $3.11 | 26 | 1.21 | 83% | 0 |
Wherever sonnet-5 wrote the code, a step needed 2.7–3.8 review rounds and passed its first review 8–24% of the time. Removing the cheap writer took that to 1.39 rounds and 61%, on cards that were still large (median 961 lines). All 16 times the pipeline ran out of fix rounds and stopped for a human, sonnet-5 was the writer.
Most of the money is spent after the first review

Across the whole run, $3,061 — 56% of all spend — went to work done after the first review: the second, third and fourth reviews, and every fix round between them. In the worst setup, steps that needed three or more reviews held 87% of the cost, and review plus fix took 72% of agent time. With sonnet-5-5 writing, that time share is 29%.
The reason is that rework is not priced by the line. Every blocker or major buys a fix session and another full review, and each of those sessions re-reads the card, the docs and a diff that grows every round. The first develop plus review is close to a fixed price, between $2 and $17 per step here. Everything above it is variable, and it is driven by one input: how many findings the writer left behind.
The intuition that backfired
The textbook split for agentic coding is a strong, expensive model to plan and review and a cheaper model to type. Per token, that is optimal. Per step, it was the opposite.

The penalty for the cheap writer grows with the size of the step: 0.9× the cost at 200–700 lines, 1.5× at 700–1,400 and 4.1× at 1,400 lines and over. In that largest bucket, side by side:
| 1,400+ line steps | opus-5-5 → sonnet-5 | opus-5-5 alone |
|---|---|---|
| Steps | 24 | 9 |
| Develop $ / step | $16.46 | $10.18 |
| Review $ / step | $13.72 | $4.35 |
| Fix $ / step | $33.08 | $1.47 |
| Median cost / step | $54.25 | $13.16 |
| Reviews · fix rounds per step | 4.62 · 2.88 | 1.56 · 0.56 |
| Findings in the first review | 15.8 | 8.9 |
| Passed the first review | 0% | 44% |
| Median agent minutes | 131 | 51 |
The cheap writer left almost twice as many findings in the first review. Each one was paid for at reviewer prices, plus a fix round in which the writer re-read everything. Even the develop phase alone came out more expensive. The "cheap" model ended up with 67% of the money in these steps.
Read this with its caveats. Even inside the 1,400+ bucket, the sonnet-5 steps were bigger (median 3,067 vs 1,816 lines), so per 1,000 lines the gap is 2.3×, not 4.1×. The develop thinking level also dropped from high to medium at the same switch. The direction is unambiguous; the exact size of the effect needs a controlled run. That is the point of the next section.
Why A/B testing is not optional

Every boundary in this chart changed more than one thing:
- S043: planner fable-5-1 → opus-5.
- S091: planner and reviewer → opus-5-5.
- S124: the cheap writer removed and thinking lowered from high to medium.
- S152: cards became about 7× smaller (median 961 → 130 lines).
- S201: sonnet-5-5 as the writer and thinking back to high.
Sequential runs like these confound the model with the prompt, the thinking level, the card size and the codebase's own growth. We learned what we know from accidents: an outage, a model release, a re-cut roadmap.
The same idea also produced opposite verdicts. With sonnet-5 as the writer, steps averaged 3.78 reviews and 8% passed their first review. With sonnet-5-5 in exactly the same seat: 1.21 reviews and 83%. A small pilot would have approved sonnet-5, because it was cheaper on 200–700 line steps. A verdict on sonnet-5 would have ruled out sonnet-5-5.
Three properties make testing a necessity for autonomous pipelines rather than a nice-to-have:
- Loops compound. Writer quality drives findings, findings drive rounds, and rounds drive re-reads. A small difference per job becomes a 4× difference per step.
- Per-token prices hide it. Sonnet-5 was the cheap model on the price list. Only the per-step total shows the bill.
- Nobody is watching. An autonomous loop spends its budget overnight, long before anyone reads a dashboard.
How we would run it now
- One variable per arm: model, thinking level, reviewer or card size, never two at once.
- Interleave the arms on the same roadmap window, or replay the same cards. Never compare eras.
- At least 10 steps per arm, compared within step-size buckets.
- Measure the whole step: cost, agent time, reviews per step, first-review pass rate, findings in the first review, human stops.
- Tag every step with its setup, so every report splits by it automatically.
- Re-run on every model release. The cheapest seat assignment changes with each one.
Lessons
- The model is the biggest lever. Same loop, same rules: $46.01 → $2.96 per step, 3.78 → 1.21 reviews, 8% → 83% first-review passes.
- Price the step, not the token. The per-token-cheap writer took 66% of its setup's money.
- Rework is the bill. 56% of all spend came after the first review.
- Cheap writer plus strong reviewer can lose, and the loss grows with step size: 0.9×, 1.5×, then 4.1×.
- Watch leading indicators. Findings in the first review (12.7 → 5.9 → 2.4 across setups) tracked cost closely. Alert per step, not per job.
- Small cards are a lever too, with a trade-off. Shrinking cards cut first-review findings from 5.9 to 2.4 and minutes per step from 40 to 22, but cost per 1,000 lines rose from $9.3 to $17.4, because every step carries fixed overhead.
- Human time is part of the bill. The sonnet-5 setups stopped for a human 16 times and needed 11 human fix rounds. Every later setup: zero and zero.
- Model version beats model tier, and verdicts expire. Sonnet-5 failed as a writer; sonnet-5-5 in the same seat matched opus-5-5 alone. Re-test every seat on every release.
Method: figures come from the pipeline's job ledger, the agents' session transcripts (per-message usage cost, including subagents) and the git history. Lines are insertions by each step's own commits, excluding Cargo.lock. All 298 steps that passed review between S006 and S281 are included.