The metric
The spine is the task-completion time horizon: the length of a task, measured in human-expert time, that a model completes at 50% reliability (METR's methodology). It is a single, behavioural number — not a benchmark score — which is why it tracks autonomy rather than recall.
The fit
We take log2 of each measured horizon and run a least-squares regression against the date. The slope is doublings per year; its inverse is the doubling time — ~5.2 months across the measured era, accelerating from the 7-month historical rate toward a recent ~4.5-month rate. The dashed line extends that fit; reference lines mark 1-day, 1-week, 1-month, and 1-year work horizons (8-hour work-days).
Uncertainty
Extrapolation is illustrative, not a date forecast. Stricter 80%-reliable horizons lag the 50% line, and a plausible compute-scaling slowdown pushes the 1-month horizon years later — which is why each milestone carries a wide band. The headline holds only as long as the recent doubling rate does.
Human baselines & RSI
Human-line baselines come from each benchmark's own expert panels. The recursive-self-improvement ledger separates measured results and benchmarks from forecasts; forecasts are labelled as such and never counted as evidence.
What this is not
This is not a probability that AGI or a singularity arrives by a given date, and not a claim of measured attainment. It is a structured, sourced way to read how fast autonomous capability is compounding. Source tiers and editorial policy follow the Landscape methodology.