AGI Benchmark Progress

Tracking progress across composite human-level reasoning and autonomous task benchmarks

74.5%
Composite AGI Benchmark Completion

AI Performance Growth Across Key Domains

Milestones Achieved

2015
Human Super-Parity in Image Recognition

ResNet surpassed human accuracy on ImageNet visual benchmarks.

2020
Language Modeling Breakthrough

Large language models demonstrated zero-shot learning across text generation and translation.

2023
Expert Human Reasoning

Models passed uniform bar exams, medical licensing tests, and advanced coding benchmarks in top percentiles.

2026
74.5% Composite Benchmark Score

Frontier reasoning LLMs reach >95% on SWE-bench Verified, ~65% on SWE-bench Pro, and 16+ hour autonomous execution horizons on METR task suites.

Key Intelligence Metrics

SWE-bench & Coding

Frontier agents resolve >95% of SWE-bench Verified issues and ~65% of un-contaminated SWE-bench Pro repos autonomously.

Scientific Reasoning

Near-saturated performance on GPQA graduate science benchmarks and >80% on ARC-AGI fluid reasoning.

METR Task Horizon

METR evaluations show 50% reliability task execution horizons doubling every 3-4 months, surpassing 16 continuous hours.

Data Sources & References

METR & Benchmark Frameworks

Model Evaluation & Threat Research

Tracks autonomous capabilities and long-horizon tasks across AI labs.

SWE-bench

SWE-bench Software Benchmark

Evaluates software engineering agents on real open-source GitHub issues.