The Task Horizon
This chart tracks the longest software-task duration that frontier AI models can complete with a 50% success rate. It uses METR’s fitted task-horizon estimates and a logarithmic time scale; the benchmark measures performance on evaluated tasks, not the duration of every real-world job an AI system can perform.
What does it show?
The length of software work frontier AI can complete reliably is increasing at an exponential pace.
Methodology
Frontier (running maximum) of METR’s fitted 50%-success time horizons across 26 evaluated models, by release date, from the results file behind METR’s Measuring AI Ability to Complete Long Tasks research. METR’s fitted doubling time: ~188 days all-time, ~129 days since 2023. Log scale, human-minutes.