"90%+ on SWE-bench" is marketing. Reality sits closer to 60%.
The industry is quietly migrating from SWE-bench Verified to SWE-bench Pro. The reason: Verified is saturated — frontier models literally "remember" the reference patches from training, and OpenAI has stopped reporting against it altogether.
On the honest Pro, the leaders hold around 59–69% — not the advertised 95%.
The verdict for team leads: when a vendor waves a "95% on SWE-bench" number at you, ask which SWE-bench, exactly. There is more heresy in benchmarks than meets the eye.