AI Performance Growth Across Key Domains
Milestones Achieved
ResNet surpassed human accuracy on ImageNet visual benchmarks.
Large language models demonstrated zero-shot learning across text generation and translation.
Models passed uniform bar exams, medical licensing tests, and advanced coding benchmarks in top percentiles.
Frontier reasoning LLMs reach >95% on SWE-bench Verified, ~65% on SWE-bench Pro, and 16+ hour autonomous execution horizons on METR task suites.
Key Intelligence Metrics
SWE-bench & Coding
Frontier agents resolve >95% of SWE-bench Verified issues and ~65% of un-contaminated SWE-bench Pro repos autonomously.
Scientific Reasoning
Near-saturated performance on GPQA graduate science benchmarks and >80% on ARC-AGI fluid reasoning.
METR Task Horizon
METR evaluations show 50% reliability task execution horizons doubling every 3-4 months, surpassing 16 continuous hours.
Data Sources & References
METR & Benchmark Frameworks
Model Evaluation & Threat ResearchTracks autonomous capabilities and long-horizon tasks across AI labs.
SWE-bench
SWE-bench Software BenchmarkEvaluates software engineering agents on real open-source GitHub issues.