The short version:OpenAI’s latest round of mathematics results is probably more correct than critics hoped, and harder to interpret than OpenAI would like. Many of the proofs have been machine-checked. Nobody outside the company can yet say how much of the underlying thinking was the model’s and how much was borrowed from existing human research.
That gap is the whole story. It’s the difference between a model that can finish a human mathematician’s argument and one that can find an idea no human had. Here’s what was released, what the checking does and doesn’t cover, and why mathematicians are uneasy.
What did OpenAI actually release?

On Tuesday, October 6, OpenAI put out 377 new math results as a GitHub dump, according to Gizmodo. The count isn’t consistent across coverage. Scientific American reported 372 results, and the repository itself holds 722 manuscripts organized into 372 research families. I’m using 377 because it’s the figure in our notes, but check the OpenAI repository before quoting a number.
The material spans algebra, number theory, topology, logic and theoretical computer science. Other coverage adds mathematical physics. A few details matter for judging the scale:
- The model, an unreleased internal one, was posed approximately 4,000 problems (2026, OpenAI GitHub). The released results are the ones that survived.
- On average, each result used three hours of ChatGPT Pro thinking compute (2026, same repository).
- The collection includes results at different stages of verification, and not all have Lean formalizations.
- OpenAI also published detailed explanations of the model’s process for 10 results.
Three hours per result sounds small. Remember it’s the average for the ones that worked, out of a pool of roughly 4,000 attempts.
Also check:ASOS Hacked? What the Alert Means and What to Do
Why the OpenAI mathematics release is dividing mathematicians
The timing is part of it. About four weeks earlier, OpenAI said 10,000 of its agents solved the 90-year-old Navier-Stokes problem in 88 hours (2026, CNBC). That’s one of the seven Millennium Prize Problems, each carrying a $1 million prize. OpenAI said it won’t claim the prize.
NYU’s Tristan Buckmaster had been working on the same problem. He described a “personal collaboration” with Levent Alpöge, who works at rival Anthropic. His objection is about provenance. He raised questions over whether OpenAI’s models had been trained on, or had access to, their sessions in OpenAI’s Codex.
To be fair to both sides, Buckmaster was careful about what he wasn’t claiming. He said he hadn’t seen OpenAI’s proof and didn’t know what the model did or how. OpenAI’s Sébastien Bubeck has publicly denied the allegations, and the company says it did not access the pair’s work. Neither claim has been independently verified, per Technology.org.
The worry for the new batch is the same one at scale. If a model can absorb nearly everything mathematicians have written, plus whatever fragments leak through prompts, how do you tell discovery from completion? A model that spots the missing step in someone else’s argument can produce a correct proof. It still isn’t the same as having the idea.
What does Lean verify, and what doesn’t it?
Lean is a proof assistant. Researchers translate statements and proofs into formal code, and a computer checks whether each step follows from the assumptions. Scientific American says that makes Lean-verified results all but certain to be correct.
That’s a real achievement. It also answers only one question: is the logic valid? It doesn’t tell you whether the theorem statement matches the problem people care about, whether the result is new, or whether it’s interesting. Interesting Engineering makes the same point: formal checking doesn’t automatically establish that every research claim in the collection is correct. And the formalization isn’t complete: many, but not all, of the manuscripts have been formalized.
A Startup Fortune line sums up the tension. “Formally verified” and “understood, attributed, and accepted by the field” are different things. Mathematicians want the second, and that takes people reading papers, not machines running checks.
Must check:PS5 Pro Shortage GTA 6: Why Sony Is Using a Lottery
What did the advisory group ask for, and what did OpenAI do?
On September 21, OpenAI announced an independent advisory group. Its nine mathematicians include Timothy Gowers and Edward Witten, and members are drawn from institutions including Cambridge, Imperial, Oxford, Stanford and Harvard. The group says it has no decision-making power. Per the group’s September 29 recommendations, they asked for a lot:
- For each result, labs should publish the model name, prompts, a summarized chain of thought, time taken and estimated compute cost.
- Proofs should be rewritten in the style of a traditional paper, with clear introductions and precise statements.
- Labs should stop testing advanced mathematical problems on proprietary models.
OpenAI met some of this. It released abridged reasoning summaries and Lean proofs. It didn’t go all the way. Scientific American reports OpenAI is revealing only the average compute time, with some extra statistics and no prompts. And because the results came from the same unreleased model as the Navier-Stokes result, the proprietary-model recommendation doesn’t look like it was taken up. An OpenAI spokesperson said the group’s advice informed how the results were shared.
Is the AI finding new ideas or finishing human ones?
There’s a feedback loop here that’s easy to miss. Early AI math tests used established benchmarks, such as Olympiad-style problems. Once models cleared those, researchers moved to genuinely unsolved research questions. Now models are being pointed at problems that working mathematicians are actively attacking.
The chain looks like this: humans publish research, models learn from what’s available, models complete problems, labs release results, and mathematicians then have to work out what the model really contributed. Every step is plausible. Together they make attribution very hard, especially when the prompts stay private.
The advisory group sees a structural risk too. It warns that internal proprietary models risk a two-tier system where labs outrun the rest of the field. A university department in Manchester or Michigan can’t rerun an experiment on a model it can’t access.
There’s also a pressure from outside. Startup Fortune reports that 25 Fields Medalists signed a declaration titled “A Severe Misalignment of AI in Mathematics.” Their complaint was about results arriving as press releases rather than papers, with no named authors and no attribution trail.
What happens if answers outpace understanding?
A proof that passes Lean says the conclusion follows. Mathematicians also want to know why it’s true, why the method works, and whether it generalizes. A flood of correct but opaque proofs creates a strange bottleneck. Scientific American says the deluge will take mathematicians months to parse, including deciding whether the proofs hold novel ideas or are mash-ups of existing techniques.
If that holds up, the scarce skill shifts. Solving problems gets cheaper. Deciding which AI-generated results matter, understanding them and building theory on top gets more valuable. The advisory group’s push for conventional, readable write-ups is a direct response to that.
The test to watch isn’t whether a model can finish an argument. It’s whether it can reliably produce an idea humans hadn’t already reached. If it can, mathematicians’ jobs change. Nothing released so far settles that, and OpenAI’s own repository notes that results sit at different stages of verification.
You may also like: 10 Cybersecurity Tips for Everyday Users(2026)
My take
Lean + summaries is a good start, but it is not enough for research-level mathematical claims.
Lean can establish that a proof is formally valid, which addresses “Is the argument logically correct?” But it does not establish where the key idea came from, whether the result substantially depends on prior human work, or whether the proof contains a genuinely new insight.
The summaries help with transparency, but summaries are still a filtered account of the model’s reasoning, not a complete research provenance trail.
For a stronger standard, I would want three layers:
- Formal verification: Lean or equivalent confirms the proof is correct.
- Provenance: prompts, relevant model inputs, cited human work, and the development history show what information the AI had access to.
- Human mathematical explanation: independent mathematicians reconstruct and explain why the result matters and what genuinely new idea it contributes.
So Lean can tell us the proof works. A summary can tell us roughly how the AI got there. Neither alone tells us this is genuinely new mathematics. That third question is where the current approach still looks incomplete.
FAQ
Did OpenAI’s AI solve a Millennium Prize Problem?
OpenAI claims its model proved a forced version of the Navier-Stokes blow-up result, and says it won’t claim the $1 million prize. The manuscript hasn’t been independently reviewed, and Buckmaster has questioned where the ideas came from. Treat it as a claim under scrutiny, not a settled result.
Are the 377 results all verified?
No. Many have Lean formalizations, but the repository says not all have accompanying Lean formalizations. Even a verified proof still needs expert review to judge novelty and significance.
Why does it matter where the AI’s ideas came from?
Credit and trust both depend on it. If a model completes a human research program using leaked or published fragments, that’s different from original discovery, and the human work should be credited. Without prompts and full logs, outsiders can’t tell.
What is Lean?
It’s a programming language and proof assistant that checks mathematical proofs step by step. It confirms the logic is valid. It can’t say whether a result is new or important.
Can ordinary users try this model?
Not yet. OpenAI says it’s working to release the model responsibly, but the results came from an unreleased internal one. Per the repository, the compute figures are in ChatGPT Pro thinking time.






