A ship trailing golden energy racing through a cyber-city toward a magenta starburst — the frontier accelerating away.
Data·Interactive dashboard·Sourced numbers

The Frontier: Watching AI Improve AI

AI is already improving AI — but not as one curve and not as one engine. Progress is now a portfolio of fuels, and it climbs fastest wherever a model can cheaply check its own work. Read in real, sourced numbers, the whole story is the frontier of what we can verify.

← Observatory
One hard ruler, four engines · % on Humanity's Last Exam
Pretraining scaleReinforcement learningSynthetic dataInference-time compute
0204060801003820354447GPT-4oo1o3GPT-5Gemini 3.1Opus 4.7
Totals are real reported HLE scores; the four-way split is reasoned attribution, not a published figure. Scale (cyan) stays flat near the floor while RL, synthetic data, and inference-time compute grow. Even Opus 4.7 leaves ~53% unsolved.
AI improving AI, caught on camera · × cumulative speedup
Cumulative speedupAI-agent records
0102030Jul '24Oct '24Jan '25Apr '25Jul '25Oct '25Jan '26Apr '26latestHivergeLocusAsterStationSPEEDUP ×
The public modded-nanogpt record line, plotted by real date: a 45-min baseline down to ~1.3 min over 83 records. The highlighted points are the four records set by autonomous AI agents — click one for its dossier. Real speedup values from METR's record log.
The AI-agent records, by idea origin · records
Imported1 · Aster
Adapted3 · Hiverge, Locus, Station
Invented0 · none yet — the signal
Every speedrun record an autonomous agent has set is Imported or Adapted — recombining known ideas. Zero are Invented. That empty row is the signal to watch: the day an agent originates a verified-novel result. (Across all 83 records, METR's speedup attribution is 6.7 / 3.0 / 1.6× imported / adapted / invented.)
Accuracy gained per 10× of compute · pp
05101510⁴10⁵10⁶10⁷10⁸10⁹10¹⁰+16+13+9+6+4+2.5+1.5
Each bar is the accuracy you buy by spending 10× more compute. Shrinking bars are diminishing returns made literal. Measured on MMLU-Pro (now saturating); read the shape.
The fuel runs dry · T tokens
Training demandHuman-text ceiling
2345678910203040506070809010020030040050060070080090010002020202220242026202820302032TOKENS (T, LOG)
Training demand (ember) rising toward the fixed ~300T stock of quality human text (brass reference). Demand crosses the ceiling around the 2028 median exhaustion date.
Synthetic data fills the gap · % of data used in AI
010203040506020212022202320241%12%35%60%
Gartner's measure of synthetic data as a share of all data used in AI projects. 2021 and 2024 are measured anchors; the middle is interpolated.
Capability index across three eras · ECI
050100150Mar '23May '24Sep '24Apr '25Feb '26Apr '26159EPOCH CAPABILITIES INDEX
Epoch's 40+ benchmark composite over the three eras. The line barely bends at each fuel hand-off; endpoint 159 is GPT-5.5 Pro (Apr 2026). The reason changes underneath; the climb does not.
Training lever: the big jumps · % accuracy
BaseAfter RL
020406080100AIME '24GPQA-DMATH-50015.63483715995
RL post-training on a fixed base, inference held fixed (DeepSeek-R1-Zero family). Doubling to quadrupling on AIME; smaller headroom where the base is already strong.
Three fuels, three economics
Post-training (RL)
EconomicsPay once, free forever after
KindBaked into weights
CeilingMay be a one-time lift; needs a reward signal
Inference scaling
EconomicsPay every single time, forever
KindPer-query runtime cost
CeilingLog-linear gains saturate; overthinking hurts
Synthetic data
EconomicsCheap to generate, risky to use
KindSource feeding both
CeilingModel collapse if unfiltered; the ladder later
Each new fuel side by side: how it is stored, what it costs, and its own emerging ceiling.
Verifiability of the task, not the domain · /100
Cheap auto-check — fastest gains
Competitive programmingcode · hidden tests pass/fail
97
Formal math proofmath · Lean accepts/rejects
95
Kernel / optimizer speedupcode · wall-clock is the judge
92
Fix a failing-test bugcode · SWE-bench territory
78
Partial check — mixed
Build a checkout flowcode · runs/pays checks; UX doesn't
52
Correct factual summarytext · facts semi-checkable
48
Human judgment only — slow
Make a 'good' websitecode · taste, UX, vague brief
22
Design sensible architecturecode · no test says it's right
18
Write a beautiful essaytext · any verifier is another model
12
0 · human-judgedverifiability →100 · cheap auto-check
Nine real tasks on a 0–100 verifiability axis. Colour marks the gain zone, not the field. Note the three CODE tasks span the full width — the gains follow the axis, not the domain.
The verifier spectrum: rock-solid to soft
Deterministic — rock-solid
ExamplesUnit tests, exact-match math, Lean proof checkers
Why it mattersReward is programmatic right/wrong. Powered DeepSeek-R1 via RLVR.
Learned judges — soft, scalable
ExamplesLLM-as-judge, reward models, rubric scoring
Why it mattersReaches finance, law, medicine — but the judge is itself a model.
The Hybrid Norm — emerging
ExamplesUnit tests for did-it-work + rubrics for is-it-good
Why it matters2026's answer: pair a hard check with a rubric judge. Anthropic recommends it.
Three tiers of verifier. Colour tracks reliability: emerald for the deterministic gold standard, brass for the learned-judge caveat, cyan for the emerging hybrid consensus.
The first answer is not the ceiling · %
GPT-5.4 (task pass)Kimi-K2.6 (task pass)
020406080100Round 1Round 2Round 3CUMULATIVE TASK PASS %
Asuka-Bench. Task pass climbs steeply with feedback, but round-3 project completion (whole app works) stays near 50% — the gap between 'many parts work' and 'the whole thing works.'
Human, AI, or the pair? · %
Overall Pass@1Partial Pass@1
01020304050AutonomousMin-intervenedHuman-onlyHuman+AI0.672.8918.8931.1119.2330.1333.5350.27
CentaurEval, collaboration-necessary tasks. The pair beats either side alone — but read the autonomous bars (0.67% / 2.89%) as the honest standalone number.
The steering curve — watch it move left · %
TodayImproved
020406080100One-shotTest-fbGuidedExpertAutoQUALITY %
Conceptual; curve shape illustrative. The envelope expands when the curve shifts left — same quality, fewer interventions — not just up.
Where the agent breaks as work scales · capability tier
L1 · Fix a functionStrong
L2 · Add a login formStrong with a spec
L3 · Auth across the appMixed
L4 · Build a dashboardUseful, needs steering
L5 · Own a system for monthsWeak / experimental
L6 · “Make GTA7”Outside the envelope
Six rungs from atomic task to creative-industrial project. Colour tracks meaning, from commons-green capability to the ember of work beyond reach.
Wall 1: inference returns saturate, then bend down · % accuracy
02040608016×32×64×128×overthinksACCURACY %
Accuracy rises log-linearly from 62% to an 88% peak at 64×, then slips to 87% at 128× as the model overthinks. Cost rises superlinearly the whole way.
Wall 3: does synthetic data doom the model?
Replace — collapses
MethodTrain only on the model's own fresh output each generation.
ResultQuality and diversity decay; useless within ~5 cycles.
SourceShumailov 2024; Alemohammad (MAD)
Accumulate — stops rot
MethodKeep all real data, pile synthetic on top; the anchor never shrinks.
ResultError stays bounded — but 'doesn't rot' isn't 'improves.' Stable mediocrity.
SourceGerstgrasser, COLM 2024; Kazdan 2024
Accumulate + verify — improves
MethodFilter synthetic through a verifier; keep only the good samples.
ResultCan reverse collapse into improvement. Ceiling set by verifier quality.
SourceYi 2026; Self-Verification, ICLR 2026
Three rungs, each a sourced result. Replacing real data collapses the model; accumulating on a real anchor stops the rot; accumulating AND verifying can reverse collapse into improvement — which is exactly what frontier labs do.
A ruler with no ceiling: task time horizon · minutes (log)
0.030.040.050.060.070.080.090.10.20.30.40.50.60.70.80.9123456789102030405060708090100200300400500600700800GPT-2GPT-4GPT-4oClaude 3.7o3Opus 4.612hTASK LENGTH @ 50% RELIABILITY
METR's time horizon on a log scale — from ~2 seconds (GPT-2) to ~12 hours (Opus 4.6), doubling every ~7 months and lately ~4. The line that can't hide a plateau.
The early-warning instrument: doubling time · days
0501001502002019–24all-time2023+2024+7.0mo6.3mo4.3mo3.0mo
Days for the horizon to double, by progressively more recent slice of models. Shorter = faster. 212d down to 89d. A future slice turning taller would be the first sign of a slowdown.
The rulers are breaking · %
LaunchNow
020406080100MMLUGPQA-DHLE323410939546
Launch score vs best model now. Each benchmark gets solved and goes flat; the frontier has retreated to HLE and FrontierMath, where models still fail.
The honest asterisk
37%
benchmark → deployment gap
12h
horizon @ 50% reliability
25min
… @ 80% reliability
Discounts to apply to every rising line. The 80%-reliability collapse is the alarm to read against the looser 50% horizons.
Odds of 6-years-in-2 compression
8%
Superforecasters
20%
AI domain experts
The experts closest to the work price the explosive case higher — because the gate is long-horizon research taste, not coding speed.
Eight instruments to watch
Task horizons
What it measuresHow long a task agents finish at 50% reliability — no ceiling yet
Invention share
What it measuresVerified-novel work as a fraction of total gains
Synthetic-data yield
What it measuresGain per synthetic token, gated by the verifier
Benchmark health
What it measuresHow fast rulers saturate; whether unsaturated ones still exist
Cost per correct answer
What it measuresDeployed cost to get a right result, not headline price
Steering burden
What it measuresHuman minutes, rounds, interventions to target quality (defined in Act 5)
Project completion gap
What it measuresDistance between task-pass and whole-system success
Verifier coverage
What it measuresShare of valuable tasks with a cheap automatic check
Each row is an instrument for reading the verifier-cost-steering frontier.