Adaptation Curve

Lens · L1 · Benchmarks

Coding benchmarks tripled in 18 months. Compute constraints did not.

The capability curve, measured. 86 models, 647 sourced scores across 7 labs — self-reported and independently verified side by side. The data that answers: how fast is the frontier actually moving?

SWE-bench VerifiedvVerifiedMaintained by Princeton / OpenAIBenchmark docs →
#Model SourceScore As of
1claude Opus 4.8AnthropicSelf88.6%2026-05-28
2claude Opus 4.7AnthropicSelf87.6%2026-04-16
3claude Opus 4.5AnthropicSelf80.9%2025-11-24
4claude Opus 4.6AnthropicSelf80.8%2026-02-05
5deepseek DeepSeek-V4 Preview (V4-Pro)DeepSeekSelf80.6%2026-04-24
6gpt GPT-5.2OpenAISelf80%2025-12-11
7claude Sonnet 4.6AnthropicSelf79.6%2026-02-17
8claude Sonnet 4.5AnthropicSelf77.2%2025-09-29
9gpt GPT-5.1OpenAISelf76.3%2025-11-13
10gemini Gemini 3 ProGoogle DeepMind (Alphabet)Self76.2%2025-11-18
11claude Haiku 4.5AnthropicSelf73.3%2025-10-15
12claude Sonnet 4AnthropicSelf72.7%2025-05-22
13claude Opus 4AnthropicSelf72.5%2025-05-22
14claude Sonnet 3.7AnthropicSelf70.3%2025-02-24
15o-series o3OpenAISelf69.1%2025-04-16
16o-series o4-miniOpenAISelf68.1%2025-04-16
17deepseek DeepSeek-V3.1DeepSeekSelf66%2025-08-21
18gpt-oss gpt-oss-120bOpenAISelf62.4%2025-08-05
19gpt-oss gpt-oss-20bOpenAISelf60.7%2025-08-05
20gemini 2.5 FlashGoogle DeepMind (Alphabet)Self60.4%2025-06-26
21deepseek DeepSeek-R1-0528DeepSeekSelf57.6%2025-05-28
22gpt GPT-4.1OpenAISelf54.6%2025-04-14
23claude Sonnet 3.5 (v2)AnthropicSelf49%2024-10-22
24o-series o3-miniOpenAISelf48.9%2025-01-31
25gemini Gemini 2.5 Flash-LiteGoogle DeepMind (Alphabet)Self44.9%2025-06-17
26o-series o1OpenAISelf40.9%2024-12-05
27claude Haiku 3.5AnthropicSelf40.6%2024-10-22
28gpt GPT-4.5OpenAISelf38%2025-02-27
29deepseek DeepSeek-V2.5DeepSeekSelf16.8%2024-09-05

Sources: self-reported scores from lab announcements and system cards; independent scores from Artificial Analysis, LMArena, and academic evaluators. Every row carries a source URL and confidence level. Self-reported and independent numbers are never averaged — divergence is intentionally visible.