Skip to main content
MANIFOLD
Superhuman mathematical problem solving before 2030, assuming no AGI yet?
185
Ṁ51kṀ310k
2029
84%
chance

Imagine that any math problem you can write down on a piece of paper that a team of Fields medalists can solve, AI can as well. Until recently, I would've predicted that that was an AGI-complete problem. Of course people used to think grandmaster-level chess would require AGI. Until 2022 I was sure that commonsense reasoning and being able to explain jokes would require AGI.

If subhuman general intelligence can be a superhuman mathematical intelligence, that will be another big update for me.

FAQ

1. What if AGI happens first?

This is a conditional prediction market. If AGI, as defined in my other market, happens first, this resolves N/A.

2. Does the AI need to max out the FrontierMath benchmark for this to resolve YES?

Yes, and every math benchmark, plus gold-medal performance on the International Math Olympiad [Update: now achieved, as of summer 2025]. Even acing the Putnam.

3. What if it's essentially true but there are rare exceptions?

The spirit of the question is that we'd only consider an AI failure to be an exception if it failed for a reason other than being insufficiently brilliant at math. Like tricksy wording, or any trick question. The posing of the question has to be non-adversarial.

4. What about a book-length question?

Tentative answer so far: The problem has to be posed on a single human-readable sheet of paper or equivalent. But a question can cite any peer-reviewed math paper as background. (Dumping an impenetrable tome on the arXiv doesn't count.) If you have an example where this feels limiting, let me know. My suspicion is that all interesting math problems can be posed on a single page and in any case it won't harm the spirit of this question to limit ourselves to such.

4. What about research taste?

That's a big part of being a mathematician and isn't required for this market. The AI just has to be superhuman at answering questions, not asking them.

5. What about cost and speed?

The AI has to dominate the best humans on all metrics. We'll find an authoritative source for the market value of mathematicians' time if it comes down to that.

6. What about availability to the public?

Not required. If there's any doubt about the veracity of claims that this has been achieved, we'll discuss and delay resolution as needed.

7. What if the AI is sometimes super- and sometimes sub-human at math?

In some senses that's already the case but there may be ambiguous edge cases. As an extreme example, imagine that the AI is so blatantly superhuman that it cracks a famous open problem, yet it's routinely stumped or wrong on problems amateur human mathematicians can do. For the spirit of the question for this market, we'll try to assess whether we'd consider a human with the AI's math abilities to be the greatest mathematician of all time. (Or the greatest raw math prodigy of all time -- see FAQ 4 on the distinction between problem solving and knowing what questions to ask. The latter is a key part of being a successful mathematician and is explicitly not part of this prediction.)

Related markets

[ignore the subhuman clarifications that keep automatically appearing below this line]

  • Update 2026-09-23 (PST) (AI summary of creator comment): - Clarified the definition of superhuman as either solving everything humans can faster and cheaper, OR solving major open problems without otherwise being too blatantly subhuman at math.

    • Proposed defining the human/superhuman threshold as what a team of 4 Fields Medalists can solve in 4 days.

Get
Ṁ1,000
to start trading!
Sort by:

What do @traders think of pinning down the human/superhuman threshold as "what a team of 4 Fields Medalists can solve in 4 days"? Up to now I had only referred to what a "team of Fields Medalists" could solve in "a day or a week". Since 4 is the average of 1 day and 7 days, might as well go with 4?

I'm trying to be extra careful and transparent since I have a large position myself. Currently for YES but when this market was new I put thousands of mana on NO.

Mostly my thinking now is that the unit distance problem and the Jacobian conjecture and especially Navier-Stokes already put us far beyond what was required for YES. But as @pietrokc is arguing, we also need clarity on jaggedness. AI is far beyond the threshold on some problems but below it on others. FAQ#7 actually tried to anticipate this, posing what was meant as a wildly hypothetical edge case: AI that's "so blatantly superhuman that it cracks a famous open problem, yet it's routinely stumped or wrong on problems amateur human mathematicians can do." I went on to suggest that we'd resolve such ambiguity by asking if a human with the AI's abilities would be considered the greatest math problem-solving prodigy of all time.

Which is annoyingly fuzzy, now that we're actually in something like that wild hypothetical. But I'm not sure we need to adjudicate the prodigy question because in the hypothetical the ambiguity arose from the combination of cracking a famous open problem and routinely being shown up by amateur mathematicians. Can anyone (@pietrokc?) make the case that it's routine for amateur mathematicians to show up Astra or Fable (let alone the latest internal models)? In the example of the SAIR competition the human winners were professionals.

Further clarification on FAQ items 2 vs 7: FAQ#2 says we need to saturate every benchmark. FAQ#7 says the AI can potentially be routinely outclassed by amateurs and still count as superhuman if it also cracks a major open problem. The seeming contradiction is because most of the FAQ just presumed AI would not solve a major open problem. So it's basically saying "superhuman means solving everything humans can faster and cheaper OR solving major open problems without otherwise being too blatantly subhuman at math".

@dreev I would not consider current AI results to make someone the best maths problem solved of all time as per FAQ7. Although it's getting close and I expect it to be unambiguously the case by the end of the year.

@dreev Well... The market's deadline is "before 2030" which is approximately a lifetime and a half in this cursed singularity. If we already find ourselves debating precise definitions because of capability jaggedness late-2026 (more than three years to go still to smooth out the edges!), the industry will be able to clear any resolution criteria you throw at it within a year at most, and it looks to me like we might even get there by the end of this year. So I wouldn't spend too much time and effort endlessly refining criteria that will fall within the time boundary as an already foregone conclusion; my biggest concern here is actually "assuming no AGI yet"; i.e. we only need to clearly and unambiguously separate between these two major conditions, and the rest will be self-evident enough.

AI getting super human at math is my #1 favorite black swan moment of the year so far.

This feels like it's there in spirit, right? And within epsilon for the most persnickety possible reading of the market criteria. FrontierMath is close to saturated and AI aces the IMO and is close on the Putnam.

Reviewing each of the FAQ items:

  1. No AGI yet: N/A

  2. Saturating all benchmarks: ✅ - 𝜀

  3. No Adversarial posing of questions: N/A

  4. No multi-page problem statements: N/A [4b. Research taste not required: N/A]

  5. Cost and speed: ✅

  6. Public availability not required: N/A

  7. Jaggedness: ✅

I don't see how any but FAQ#2 could militate for NO and we've still got over 3 years to go to close that epsilon gap. Even if we don't quite close that gap, FAQ#7 suggests that we could still get to a YES despite AI not completely dominating humans at all math problem-solving, as long as the peaks of its achievements are high enough. As of the Navier-Stokes solution, we've got that in spades.

PS: Navier-Stokes did cost millions in compute, but humans couldn't solve that one at all. Even if you think a team of Fields Medalists could've solved it, I'm pretty sure the cost, paying them market rates, would've been even higher. For the hardest contest and benchmark problems, the compute cost and speed seems to blow humans away. And, again, still 3 years for it all to get even better/cheaper/faster.

bought Ṁ15,200 YES

Btw, if I end up too biased, I'm happy to defer to a neutral arbiter for resolution. Just let me know what you think I'm missing above.

(And note that when this market was new I dumped thousands of mana into NO. Then reality played out.)

@dreev I would say it's not there yet based on public information. The open problems that have been solved are a small fraction of the most important. We can also see frontier maths open problems to have a more precise denominator. Further all the solutions were within human reach as far as I can tell and have been less clever than you would expect based on the problems prestige. In particular the most impressive results are all counter examples so far. Based on this I don't think it outperforms a team of fields medalists for a week on > 50% of problems. However if Open AIs internal results are as good as they suggest this might well be moot.

I'd give it some more time to resolve, especially to hear about performance in other subfields, but it's looking like we'll quite likely hit this before 2030.

@dreev I don't think your assessment is correct.

As I had warned below, the claim in this market is difficult to establish either way. However, recently there was a SAIR competition on the inverse Galois problem. Basically, for each one of the 25,000 transitive permutation groups contained in S_24, find a polynomial whose Galois group is that group. What's interesting about this competition is that everyone was highly encouraged to use AI, and (I believe) got infinite free credits to do so. So here is a scenario in which dozens (hundreds?) of people were trying to combine their efforts with all that AI could offer, to solve the same 25,000 problems. Everyone had two months to do it.

Nevertheless, the winning team did not use AI at all. And you can see that they were the only team to solve this problem for hundreds upon hundreds of specific groups.

More generally, mathematics is incredibly vast, and I worry that this market will be resolved based on vibes from announcements like "OpenAI solved 100 open problems". There is a big PR firehose spraying us with the (very impressive!) positive examples, but what is the mechanism for us to learn of the negative examples?

There's thousands of papers posted on arxiv every day, of which definitely hundreds solve some problem which hadn't been solved before, and of which quite likely dozens had AI tried on it and it didn't work.

I've already exited this market with substantial profit but I'm a bit worried how this will be adjudicated at expiry. How does one tell what "a team of Fields medalists can solve" without, well, giving it to them?

One way this market could resolve NO is if there are some clearly easy problems that AI cannot solve. But these are getting harder to find.

Since getting "a team of Fields medalists" together almost certainly won't happen (they are busy people with diverse specializations), the other plausible way this could resolve NO is if AI cannot solve the problems that led to people winning Fields medals.

It's pretty clear to professionals that this is currently the case; AI today could not have solved, say, the three-dimensional Kakeya conjecture, if starting from where Hong Wang did. But how would we prove this? If you ask an AI today, it will know the solution from the internet. Even if you asked an AI a week before a solution was posted, it would have known all the partial results that Wang, Zahl &co got over the years leading up to it.

So how is it proposed to resolve this market?

@pietrokc Great clarification. I had in mind a day or a week for the hypothetical team of Fields medalists -- not long enough to do something career-making. Does that sound fair to people before I update the market's FAQ?

@dreev My concern is that "what a hypothetical team of Fields medalists can accomplish in a week", which is not even is well-defined in the first place, will be defined for the purposes of this market by people who are not professional mathematicians -- taking it even farther from its already nebulous meaning.

@pietrokc I'm very open to ideas for operationalizing this better. And I'll do my best to be fair regardless. The consensus of the math community should be the ground truth, and there's a decent chance there'll be no ambiguity about that.

https://x.com/dmitryrybin1/status/2079904005652893709?s=46

So this is the quality of prompt required to get the AI to solve a major unsolved problem:

https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de063

Post 1: Construct a counterexample to general (non-planar) case of Dinitz Garg Goemans conjecture. You should do a breakthrough and find a structured counterexample.
[thought for a long time and didn't solve it]
Post 2: please continue research and find a complete unconditional counterexample
[thought for a long time and didn't solve it]
Post 3: Continue the search. Have a clear strategy obtained from deeper understanding of the problem structure.
[thought for a long time and didn't solve it]
Post 4: it's enough of partial results. let's finish with a complete unconditional counterexample
[thought for a long time and did solve it!]

The twitter thread had like 5 other people post other counterexamples that disproved other open mathematical questions lol.

The fact that these are all counterexamples are telling, but that's still majorly superhuman in large parts of math.

IMHO its abilities are doubtlessly still spikey (for instance you could probably still trick it with some vision based problems), but this resolving yes is far more likely then it was even a month ago.

bought Ṁ2,500 YES

Big jumps in priors this year, not just the recent Jacobian Conjecture news: https://www.reddit.com/r/singularity/comments/1v1aie6/apparently_the_jacobian_conjecture_was_just/

@Panfilo Indeed! I kind of feel like I nailed it just in asking this question back in April 2025.

This market is wildly underpriced. Should be 90% to either N/A or resolve YES.

Literally first simple question of this nature that I thought of

The circles

o3's answer today (tried several times, same incorrect answer of 1 each time)

This other market is 50-50 on SOME millenium problem being solved by the beginning of 2030, by human, AI, or a collaboration, and that market does not condition on there not being AGI:

https://manifold.markets/Inosen_Infinity/will-at-least-one-of-the-remaining

bought Ṁ500 YES

Feels like we are well on our way to this happening. In just two years we’ve gone from “guessing randomly on most AMC questions” to “able to get an average USAMO score.” Math feels like an especially scalable field as well, because it’s easy to automate the checking of its own work, and one where having encyclopedic knowledge of every theorem and proof technique (and the ability to try them far faster than a human) would be very useful.

bought Ṁ500 NO

@dominic We also went from the first flight in 1905 to the moon landing in 1969 but it's 2025 and we still haven't been to Mars or even back to the Moon recently.

When it comes to USAMO/IMO, everyone knew geometry and functional equations would be easy, other than that it's basically so far just done some simple one-step 1/4 problems with tons and tons of computation time. LLMs are basically useless, and only AlphaProof can do this. AlphaProof is also quite weird, it's just given the statements and it proves things until it finds the answer. Olympiads are more suited to that than research problems.

There is no question that what AI has done so far is impressive, and I don't mean to be a super AI skeptic. I do think I'll live to see AI "solve math" but 4.5 years is way too quick. We are still a long way off.

@nathanwei Not referring to AlphaProof here, just looking at the success of Gemini 2.5 Pro / o3 on the USAMO benchmark: https://matharena.ai/. o1 basically could not solve any USAMO problems whereas o3 can solve 1/4 about half the time, which is a pretty big step up. It's likely to slow down, but how fast? How long do you expect it to take before a model aces USAMO? I wouldn't be surprised if it happens within a year from today.

@dominic I would be very surprised if an LLM aces USAMO within a year. This is an LLM? Not AlphaProof? AlphaProof has some chance to sweep USAMO/IMO within a year (I'm not as bullish as some others, but it's not impossible) but I think that LLMs have no chance.

@nathanwei Yeah, referring to LLMs. I don't think it's guaranteed by any means, but I wouldn't bet against an early 2026 LLM getting 90%+ on USAMO, and my average expectation is probably like 70%. If the rapid progress stops I'd definitely become much less optimistic however.

bought Ṁ50 NO

@dominic The oracle @SemioticRivalry told me that AI progress will slow down.