Open research

How close is
the takeoff?

The clearest measure of AI progress is not a leaderboard score — it is the length of work an AI can finish on its own. That horizon now reaches ~16 hours of expert work at 50% reliability, and it has been doubling every ~4-5 months. This page tracks that exponential against the human baseline, and the feedback loop — AI accelerating AI R&D — that drives it. Every number is sourced. Extrapolations are not forecasts.

← Back to AI for Science Landscape

The capability exponential

Task horizon, doubling on a log scale.

The length of human-expert work a frontier model finishes at 50% reliability (METR-style). Points are measured results; the solid fit is a least-squares exponential; the dashed line extrapolates the recent doubling toward week-, month-, and year-long autonomous tasks.

Now (2026-05)

~16 hours

at 50% reliability · doubling ~4-5 months

Measured (50% reliability)Exponential fitExtrapolation

1 work-day

8h

Reached

as of 2026-05

1 work-week

40h

2026

projected · band 2026–2027

1 month

160h

2027

projected · band 2027–2034

1 year

1,920h

2029

projected · band 2028–2037

Fit: least-squares regression of log2(task minutes) on year gives a doubling time of ~5.2 months across the measured era (7-month historical, ~4.5-month recent). Extrapolation is not a guarantee — stricter 80%-reliable horizons lag, and a compute-scaling slowdown sets the high end of each band.

Crossing the human line

Where AI has passed
expert humans.

AI has surpassed expert-human baselines on 4 of 6 tracked benchmarks, from broad knowledge (MMLU) to PhD-level science (GPQA) and competition math. Real computer-use agents are now at human parity.

ARC-AGI-2 — novel abstraction that resists memorization — is the live frontier, crossing the human panel only at high compute. Research-grade mathematics (FrontierMath) is the clearest remaining headroom.

MMLU

Broad academic knowledge

Surpassed
AI 92% · GPT-4-classHuman baseline 90%Crossed 2023-2024

Saturated near and above the human-expert reference; retired as a frontier signal.

GPQA Diamond

PhD-level science reasoning

Surpassed
AI 94% · GPT-5.5 / Claude Opus 4.xHuman baseline 81%Crossed 2024-2025

Domain-expert validators score ~81%; frontier models are now near saturation at ~94%.

MATH-500

Competition mathematics

Surpassed
AI 97% · Frontier reasoning modelsHuman baseline 90%Crossed 2024-2025

Step-by-step competition math is largely solved at the frontier.

OSWorld

Real computer-use agents

At parity
AI 75% · Frontier computer-use agentsHuman baseline 72%Crossed 2026

Agents reach approximate human parity on real desktop tasks — a 2026 crossing.

ARC-AGI-2

Novel abstraction & generalization

Live frontier
AI 85% · GPT-5.5 (high compute)Human baseline 84%

Crosses the average human panel only at high compute and cost; the live general-reasoning frontier.

FrontierMath

Research-grade mathematics

Not crossed
AI 52% · GPT-5.5 Pro (Tiers 1-3)Human baseline

Research-grade math is far from solved; Tier 4 sits near 40%. The clearest remaining headroom.

Sources

The feedback loop

AI is starting to
build AI.

The exponential is driven by a compounding loop: AI tools accelerate the research and engineering that build the next model. These are the concrete, dated signals of that loop closing.

What this is and isn't. This is early recursive self-improvement — AI measurably accelerating AI R&D — not a model autonomously training its own successor end-to-end. Measured results and benchmarks are separated from forecasts below; forecasts are estimates, not evidence.

May 2026

Statement

Anthropic states AI is speeding up AI R&D

Anthropic put in writing that it sees AI contributing to speeding up the research and development of AI itself — an early recursive-self-improvement dynamic — and committed to publishing more on how AI tools accelerate its own work.

May 2026

Deployment

Majority of Anthropic's code now written by Claude Code

Most new code at Anthropic is authored by the model under human review — AI tooling materially inside the loop that builds the next model.

Nov 2025

Benchmark

RE-Bench: agents score 4x human experts at a 2-hour budget

On seven ML research-engineering environments (calibrated against 71 eight-hour human-expert attempts), the best agents outscore human experts 4x when both are limited to 2 hours per task.

May 2025

Measured result

AlphaEvolve beats Strassen's 1969 record

A Gemini-powered evolutionary coding agent discovered a 48-multiplication scheme for 4x4 complex matrix multiplication, improving on a result that had stood for 56 years.

May 2026

Benchmark

SWE-bench ~80%: agents resolve ~4 of 5 real GitHub issues

Frontier coding agents now resolve roughly four out of five real software-engineering issues from mature repositories, the substrate for AI maintaining its own toolchains.

Dec 2028

Forecast

Forecast: >50% odds of autonomous self-improvement by end-2028

Anthropic's Jack Clark estimates a greater-than-even chance that, by end of 2028, a system could be told to make a better version of itself and do so autonomously. This is a forecast, not a measurement.

Methodology

How the curve is built.

The metric

The spine is the task-completion time horizon: the length of a task, measured in human-expert time, that a model completes at 50% reliability (METR's methodology). It is a single, behavioural number — not a benchmark score — which is why it tracks autonomy rather than recall.

The fit

We take log2 of each measured horizon and run a least-squares regression against the date. The slope is doublings per year; its inverse is the doubling time — ~5.2 months across the measured era, accelerating from the 7-month historical rate toward a recent ~4.5-month rate. The dashed line extends that fit; reference lines mark 1-day, 1-week, 1-month, and 1-year work horizons (8-hour work-days).

Uncertainty

Extrapolation is illustrative, not a date forecast. Stricter 80%-reliable horizons lag the 50% line, and a plausible compute-scaling slowdown pushes the 1-month horizon years later — which is why each milestone carries a wide band. The headline holds only as long as the recent doubling rate does.

Human baselines & RSI

Human-line baselines come from each benchmark's own expert panels. The recursive-self-improvement ledger separates measured results and benchmarks from forecasts; forecasts are labelled as such and never counted as evidence.

What this is not

This is not a probability that AGI or a singularity arrives by a given date, and not a claim of measured attainment. It is a structured, sourced way to read how fast autonomous capability is compounding. Source tiers and editorial policy follow the Landscape methodology.

Open data

Free to reuse, with attribution.

License

The AGI Capability Frontier dataset is published by Scivity Labs under CC BY 4.0. You may reuse, remix, and republish with attribution.

Cite as

Scivity Labs (2026). AGI Capability Frontier — Time-Horizon Tracker [Dataset]. scivity.org/agi-progress
BibTeXShow
@misc{scivity_agi_capability_frontier_2026,
  author       = {{Scivity Labs}},
  title        = {AGI Capability Frontier --- Time-Horizon Tracker},
  year         = {2026},
  howpublished = {\url{https://scivity.org/agi-progress}},
  note         = {Dataset. CC BY 4.0}
}
AGI Capability Frontier — Time-Horizon Tracker — Scivity · Scivity