62 sessions from 9 engineers, all on the checkout migration, 71% on a frontier model. Nothing failed; it was one large piece of work. Flagged because the size is unusual for the team, not because it is wrong.
Your engineers ran 34,180 agent sessions across four tools. That is $372 per engineer, up 14% on August while headcount held flat. The rise is not more people using agents — it is the same work costing more, concentrated in three teams and one routing habit. Everything below is measured from session transcripts on each machine; nothing here is an estimate unless it says so.
Projected from the first 19 days at the current daily rate.
Set per team. Two are tracking over with 11 days left.
Ranked by spend. Cost per merged task is the number to read across teams — it survives differences in headcount and how much they ship.
| Team | Eng | Spend | $/task | Against fleet median | Signal |
|---|
Four patterns, each measured from the sessions themselves. The first two are counted directly; the last two are estimates and say so.
The standards your teams wrote, checked against every session on every machine. This is the one thing a usage dashboard cannot tell you.
Share of spend. Adoption is not the same as cost.
Tool errors per 100 calls, by how full the context window was. The knee is at 120k tokens — past it, error rate more than doubles.
Median output tokens per session. Test-heavy work is the expensive category on every team.
Each carries the number it is based on, so you can argue with the arithmetic rather than the conclusion.
4,880 sessions under 15 turns ran on a frontier model. On graded benchmark tasks of that size, a mid-tier model resolves at the same rate.
Basis: price difference at published rates. Resolve-rate parity measured on 99 graded tasks, not assumed.1,207 sessions kept retrying past five consecutive tool errors. A watcher on each machine interrupts the session and tells the engineer.
Basis: counted loops × the tokens spent after the fifth failure. No modelling.Payments and Growth are both more than double the fleet median cost per merged task, and both moved there in one month. Neither is a headcount story.
Basis: $/task vs the 12-team median, month over month.Dana Moreau · Core Platform · Sandeep Rao's team. You ran 96 sessions, above the team median of 71. Your cost per session is 29% below your team's, mostly because you reach for a frontier model on 22% of work where the team averages 54%. Nobody else sees this page — your manager sees the team total, not your rows.
Your own sessions, bucketed by what you asked for. Test work is your expensive category, as it is for everyone.
The standards in your repo, and your own sessions against them.
Yours, against the fleet. You compact earlier than most, which is why your error rate is low.