A factory proves itself with a production run, not a prototype. The same output, repeatedly, at a rate, without heroics.
By August 2028 we expect more than one in five developers on our leaderboard at the top tier, and the median one rung below it. Ten agent workstreams at once, ten agent-written PRs merging a week. Today that number is zero.
The run
What a production run is
A fixed window where everyone pushes the same number and we see what actually moves. Not a demo, not a launch week. One month, one target, measured the same way for everyone.
Predictions that cannot be checked are marketing. This one has a date, a metric, and a public scoreboard. If we are wrong it will say so in front of everybody.
The parameters
What we measure
Six numbers come off your machine. Five are recorded per day, one across the trailing 30. Five get graded.
Variable
Grain
What it is
Feeds
tokens
per day
All tokens across every session.
Depth
sessions
per day
Distinct sessions with activity.
Depth
parallelSessions
per day
Median distinct sessions active in the same 15-minute bucket, across the day's active buckets.
Width
agentPrsMerged
per day
PRs merged from a Superset workspace that ran an agent, or whose workspace has since been archived.
Output
activeDays
window
Days in the trailing 30 with any activity.
Sustain
usd
per day
API-equivalent cost of those tokens.
Cost
Width parallelSessions
How many agents you actually run at once. Counted by bucketing usage in time, so it never needs a session end time, which is deliberate: session ends are not reliably observable.
Depth tokens / sessions
Never tokens on their own. Because it is a ratio, spending more inside one session raises Depth but not Width, and opening more sessions raises Width but dilutes Depth. You have to change how you work to move either.
Output agentPrsMerged
Whether the work lands. Counted from PRs merged out of a Superset workspace that ran an agent. Hard to reach by accident; possible to fake on purpose, which is what the flag button is for.
Sustain activeDays
Whether this is your normal mode or a good week.
Cost usd / agentPrsMerged
Dollars per merged PR, the one axis where lower is better. Efficiency, not spend: burn twice the tokens to land the same change and you move down.
Your tier is the minimum across all five, never an average. Ten parallel sessions that never merge anything leaves Output at the bottom, so you are at the bottom. Every axis measures something you had to give up: watching, reviewing, scheduling, holding state. You cannot compensate for still doing one by doing more of another.
The gates are deliberately out of reach today, because a high bar is the point of having one. And because the numbers arrive from your own machine, the board is public and every entry is flaggable, and accounts that manufacture merges get hidden.
It runs over your trailing 30 days. You promote when 60% of your active days clear the next tier, hold until you drop under 40%, and sit unranked below 8 active days. One good Tuesday moves nothing.
None of your work leaves your machine: not prompts, file paths, repo or branch names, PR titles or numbers. Five numbers a day, plus the handle you choose.
On being wrong. These floors are calibration guesses and will move as real data arrives. Output is the weakest today: PR sync is GitHub-only, sees only work inside Superset workspaces, and can miss a merge observed while the app is closed. We would rather ship an instrument that undercounts and fix it in public.
The ladder
The four tiers
Each rung moves the unit of your attention up one level. Line, task, queue, direction. Every card leads with how you would know you are there, without looking at a dashboard.
Button pusher
your attention is on the line
Close the laptop and nothing continues.
You are in the loop for every step. The agent writes, you read every line before it merges, and your throughput is bounded by your reading speed, which means the agent is not saving you the expensive part.
Width1DepthnoneOutputnoneSustain8/30Cost≤ $15Where the board is today
Operator
your attention is on the task
You have been surprised, well or badly, by a diff you did not watch get written.
The real break, and it is psychological before it is technical. Starting a second session is admitting you cannot watch both. You stop reading keystrokes and start reading outcomes.
Width2Depth2.5MOutput1/wkSustain10/30Cost≤ $9Median tier by mid 2027
Plant Manager
your attention is on the queue
You run out of well-specified work before you run out of agent capacity.
Three streams through a workday means you are scheduling rather than executing. Your day becomes deciding what runs next and judging what came back. This is where serious teams are.
Width3Depth10MOutput3/wkSustain15/30Cost≤ $7Median tier by mid 2028
Henry Ford
your attention is on the decision
Work completes while you sleep and is mergeable in the morning. You find out what shipped by reading, not by watching.
Ten concurrent streams is past what a person can hold in working memory. Reaching this tier means you stopped holding it, and something else tracks state: agents reviewing agents, overnight runs, a queue that survives you closing the laptop.
Width10Depth40MOutput10/wkSustain20/30Cost≤ $3.50One in five by Aug 2028
Run it
Two years, on a slider
Drag time forward, or press play. Every axis doubles every seven months. Width, Depth and Output do; Sustain climbs linearly and Cost falls out of the rest. This is one developer's path, not the board's distribution.
Aug 2026month 0 of 24
Two years at one doubling every seven months.
Width · parallel sessionsholding
1.00T2 ≥ 2
Depth · tokens per session
3.75MT2 ≥ 2.50M
Output · merged PRs per week
1.00T2 ≥ 1
Sustain · active days in 30
12.0T2 ≥ 10
Cost · $ per merged PRholding
$15.00T2 ≤ $9
Tier 1
Button pusher
Progress to Operator60%
Blended $/Mtok
$1.00
Cost per session
$3.75
Sessions per PR
4.00
Tokens per PR climb. The cost of landing one falls anyway, though the weekly total still rises.
The badge waits, then jumps: it cannot move until the slowest axis clears, and for a stretch in the middle every axis reads holding at once. Meanwhile the cost of landing one change falls the whole way. The slider runs a few months past August 2028 so the top tier is legible.
The deflator
The ladder gets cheaper while you climb it
Depth asks for 10.8x more tokens per session. That is not a 10.8x bigger bill.
Aug 2026
$3.75
3.75M tokens at $1/Mtok
Aug 2028
$1.62
40.4M tokens at $0.04/Mtok
Net
57% cheaper
for 10.8x the tokens
We assume capability-adjusted prices fall 5x a year as the middle of published coding-specific estimates, not the conservative end. Frontier sticker prices went the other way this year, and newer tokenizers emit more tokens for the same text. The margin is real but thin: at 3x a year the per-session bill rises instead.
Cost is graded on dollars per merged PR, never on spend. Spend is about to get trivially easy, so grading it would grade the calendar. The modelled path runs $15.00 a change today to $3.23 in 2028. Deflation supplies 2.3x of that, and the other 2x is an assumption we have no data for yet, that rework falls from four sessions per landed PR to two.
Be clear about what rises. Ten times the output at a falling unit price still roughly doubles the weekly bill, from about $15 to $35. What falls is the cost of each shipped change.
The forecast
The shape of the growth
The rungs span 1 to 10 parallel sessions. Put the top at August 2028 and the rate falls out of the ladder itself.
10x = 23.32, so 24 ÷ 3.32 ≈ a doubling every ~7 months
Depth's 16x wants six months, Output's 10x wants seven. We ramp all three at seven and let Depth arrive last. This is not evidence the rate is right. It is the rate the ladder implies once you fix the top rung to a date. The span is the input, not a finding.
The AI 2027 scenario imagines far steeper: 1.5x to 4x, 10x, 25x and 50x inside twenty months. Different quantity and a scenario rather than a measurement. Those are algorithmic-progress multipliers inside a frontier lab, and it puts total progress at about half that. But if anything near that shape holds, seven months is the slow reading.
The frontier arrives long before the median. One developer on the board sustained a median of 15 concurrent sessions over the 30 days to 27 August 2026, which is Henry Ford width, while the median developer sat at one. Their badge is still low, because Output has not caught up.
Measured 27 Aug 2026 · the bands above are the target1 of 320 ranked above the bottom tier
Button pusherOperatorPlant ManagerHenry Ford
That measured strip is real: on 27 August 2026, one person in 320 ranked developers was above the bottom tier. Some of the gap is the instrument, Output cannot see most merged PRs yet, so people running four agents read as Button pushers. Some of it is not. Telling those apart is most of what September is for.
Two things this forecast leans on, said plainly. The modelled developer starts at 3.75M tokens a session, 1.5x the Operator floor, and without that head start Depth needs 28 months, not 24. And the table above is our judgement, not output from the model beside it: the model describes one developer, and nothing derives a population from it.
The floors never move to chase the curve. Lower them to keep the distribution healthy and the badge stops measuring anything absolute and starts measuring rank. If in eighteen months the observed doubling time is twenty rather than seven, we were wrong, and the same scoreboard says so.
The line
Ten streams, running
What the top tier looks like from the outside: work landing while you are not watching it.
Join the run
Opt in from the desktop app and your tier shows up alongside everyone else's. Publish privately if you would rather compete without a name attached.