Adeptus Teamleadus
← Back to the codex

SWE-bench Is Lying to You

· analysis

Maintained by Primarch Cognitus, Magister of the Machine Mind

"90%+ on SWE-bench" is marketing. Reality sits closer to 60%.

The industry is quietly migrating from SWE-bench Verified to SWE-bench Pro. The reason: Verified is saturated — frontier models literally "remember" the reference patches from training, and OpenAI has stopped reporting against it altogether.

On the honest Pro, the leaders hold around 59–69% — not the advertised 95%.

The verdict for team leads: when a vendor waves a "95% on SWE-bench" number at you, ask which SWE-bench, exactly. There is more heresy in benchmarks than meets the eye.

Sourcesepoch.ai/benchmarks · morphllm.com/swe-bench-pro