Fleet Ledger

Platform Engineering · Jeff Huber, VP · 500 engineers, 12 teams
Synthetic org
Sep 2026
Payments spent $4,900 on Tuesday — 3.1× its weekday average.

62 sessions from 9 engineers, all on the checkout migration, 71% on a frontier model. Nothing failed; it was one large piece of work. Flagged because the size is unusual for the team, not because it is wrong.

2 days ago
$186kagent spend this month,
at published API prices

Your engineers ran 34,180 agent sessions across four tools. That is $372 per engineer, up 14% on August while headcount held flat. The rise is not more people using agents — it is the same work costing more, concentrated in three teams and one routing habit. Everything below is measured from session transcripts on each machine; nothing here is an estimate unless it says so.

Active engineers
311
62% of 500 · +38 this month
Sessions
34,180
110 per active engineer
Cost per merged task
$14.20
+9% vs August
Tasks merged
13,104
38% of sessions ended in a merge
Recoverable
$47.9k
26% of spend · see below

Where the month lands

Projected from the first 19 days at the current daily rate.

$210k
$186k spent · $24k over the $186k plan · the marker is your budget

Teams against their budget

Set per team. Two are tracking over with 11 days left.

By team

Where the money goes, and who is efficient with it

Ranked by spend. Cost per merged task is the number to read across teams — it survives differences in headcount and how much they ship.

TeamEngSpend $/taskAgainst fleet medianSignal
Leaks

What $47.9k of it bought you nothing for

Four patterns, each measured from the sessions themselves. The first two are counted directly; the last two are estimates and say so.

Abandoned sessions 2,914 sessions ended with no diff and no answer, after a median of 31 turns. Almost always a task the agent could not scope.
$18.6k
Error loops 1,207 sessions hit five or more consecutive tool errors and kept going. Counted directly from tool results.
$11.4k
Frontier models on small work 4,880 sessions under 15 turns ran on a frontier model. Estimate: the price difference had they run on a mid-tier model, which resolves this size at the same rate in our benchmark.
$12.3k
Rework after human correction 1,940 sessions needed two or more corrections from the engineer. Estimate: a third of those sessions' output attributed to redoing work.
$5.6k

Your rules, and what agents did anyway

The standards your teams wrote, checked against every session on every machine. This is the one thing a usage dashboard cannot tell you.

Which tools your engineers use

Share of spend. Adoption is not the same as cost.

Errors rise with context depth

Tool errors per 100 calls, by how full the context window was. The knee is at 120k tokens — past it, error rate more than doubles.

Cost by kind of work

Median output tokens per session. Test-heavy work is the expensive category on every team.

Next

Three things worth doing this month

Each carries the number it is based on, so you can argue with the arithmetic rather than the conclusion.

$12.3k/mo

Route small tasks off frontier models

4,880 sessions under 15 turns ran on a frontier model. On graded benchmark tasks of that size, a mid-tier model resolves at the same rate.

Basis: price difference at published rates. Resolve-rate parity measured on 99 graded tasks, not assumed.
$11.4k/mo

Stop error loops at five failures

1,207 sessions kept retrying past five consecutive tool errors. A watcher on each machine interrupts the session and tells the engineer.

Basis: counted loops × the tokens spent after the fifth failure. No modelling.
2 teams

Ask Pete and Anil what changed

Payments and Growth are both more than double the fleet median cost per merged task, and both moved there in one month. Neither is a headcount story.

Basis: $/task vs the 12-team median, month over month.