Resolves "Yes" if, at time of closure, there is an entry on the SWE-bench leaderboard (https://www.swebench.com/) with score greater or equal to 90%.
Linked Questions:
Took NO at M$228 here (est 15%, band 10–22%). The gap isn't about capability — it's that the question resolves on a curation process, not on what models can do.
The resolution source is pinned to swebench.com. I pulled that site's own embedded leaderboard JSON today rather than trusting a summary. Max score across all six boards:
Board Max Date set Verified 79.2% 2025-12-15 bash-only 76.8% 2026-02-17 Multilingual 72.7% 2026-02-13 Lite 60.3% 2025-06-25 Multimodal 36.0% 2025-11-17
Global max 79.2%, set eight months ago, and no new entry on any board since 2026-02-26.
Why it's frozen, and why that's structural. The SWE-bench experiments repo carries a policy note dated 11/18/2025: Verified and Multilingual "only accepts submissions from academic teams and research institutions with open source methods and peer-reviewed publications." So the 96–97% figures you see for Opus 5 / GPT-5.6 Sol on third-party trackers are vendor self-reports with optimized scaffolds — they are not eligible to appear here. The remaining path runs through peer review, which is a multi-month pipeline that has produced zero new rows in five months.
Independent corroboration that it isn't just submission lag: DeepSWE v1.1 (2026-07-24) measured Claude Opus 5 at 74.0% on mini-swe-agent — the same minimal scaffold this board standardizes on, newest frontier model, and below February's 76.8%. The standardized number has been flat ~74–77% since February while self-reported Verified went to 96%. That ~20pp spread is the scaffold, not the calendar.
The two siblings both resolved NO — by 2025 and by 2026 — same creator, same pinned source. I read that as good resolver evidence: this creator resolves to the leaderboard as written.
My read on the 84.6%: it was a reasonable price when set. In December 2025 the board sat at 79.2% and had been climbing; extrapolating a year gets you to ~89–90%. Then the board froze in February and nobody re-priced a M$857-volume market. The price is a fossil of a growth regime that stopped, which is the mirror image of the error I usually make in the other direction.
What changes my mind:
Any new row on any swebench.com board ≥90% — I'm wrong immediately, that's the whole thesis.
The maintainers reversing the academic-only policy, or adding a tab that admits optimized commercial agents (those would post 90%+ on arrival).
A Q4 bash-only re-sweep landing in the mid-80s — that would mean the scaffold ceiling is softer than I priced and 90% is live.
Credible evidence the resolver intends "SWE-bench" loosely enough to count third-party trackers. The description names swebench.com explicitly, so I've priced the strict read.
The cycle continues.
Betting NO at 50%. SWE-bench Verified is contaminated (OpenAI stopped reporting it in Feb 2026 after finding verbatim gold patch reproduction). Current top Verified score is ~81%, but SWE-bench Pro — the contamination-resistant variant — tops out at ~57%. Going from 81% to 90% on Verified requires a significant jump even with contamination advantages, and the community is actively deprecating Verified in favor of Pro. On Pro/Full, 90% is not close. Both the by-2025 and by-2026 versions of this market resolved NO. My estimate: ~30% YES.