Rodrigo Santos Andrade share

Is this the ‘mathocalypse’? Why OpenAI’s latest results dump has left mathematicians in shock

visibility comment0 event October 9, 2026 timer 00:25

OpenAI released 722 new AI-generated mathematical papers at once – three have been retracted, and mathematicians are coming to grips with the rest.

Ads

Picture Alliance / Getty ImagesEarlier this week, OpenAI published a trove of hundreds of mathematical results it claimed could “push the frontier of human knowledge”. Produced largely by an unreleased artificial intelligence (AI) model, the 722 papers relate to 372 open problems covering everything from algebra and geometry to theoretical computer science.

Several papers have since been retracted or amended, but the sheer scale and form of the release – variously described as a drop, a dump, a carpet bombing and a mathocalypse – have sent shockwaves through the mathematical community.

Some papers claim significant advances on high-profile, long-standing problems including the Riemann hypothesis and the Birch–Swinnerton-Dyer conjecture, each of which carries a bounty of US$1 million (A$1.4 million) if solved as one of the seven Millenium Prize Problems.

Why has OpenAI done this? And can we trust the claimed solutions are correct? Neither of these questions seem to have particularly obvious answers.

A strange way to enable progress

OpenAI says its motivation in releasing the huge trove of results is to “enable further progress in mathematics”. However, the manner of the release is not necessarily conducive to this goal.

A senior colleague of mine found one of his favourite problems among those solved and attempted to read the accompanying paper. He told me it was so unintelligible that, had he received it as an editor at a mathematics journal, “it would have gone straight into the bin”.

Even OpenAI’s own large language model ChatGPT (GPT-5.6 Sol, to be precise) was skeptical when I asked it, describing at least one of the high-profile results as “a serious hallucination … [which] should never be cited, submitted, or circulated as a proof without a complete expert audit”.

In other words, despite the groundbreaking nature of some of the results obtained, the way at least some of the accompanying papers are written is not comprehensible even to experts.

Changing the conversation

OpenAI also may be seeking to move the discussion on from its last big mathematical publication. In September, the company published a controversial claimed solution of a case of the Navier-Stokes problem, another $1,000,000 Millenium Problem.

Mathematicians Tristan Buckmaster and Levent Alpöge, who had been working on the problem themselves, alleged OpenAI had accessed their data and used their ideas, then devoted some US$15 million of computing power to find a solution. OpenAI has denied these allegations.

One other possible broader motivation for OpenAI is the hot topic of artificial general intelligence (AGI). This is a hypothetical type of AI that could surpass humans in basically all cognitive tasks.

Mathematics – especially pure, abstract mathematics – is a discipline that largely relies on often very complex arguments to prove the truth of mathematical statements. It is often seen as a subject that is very difficult for most people.

Hence, the ability of OpenAI’s models to do research-level mathematics may lend credence to the idea the models are approaching the grail of AGI. This idea could benefit OpenAI ahead of the company’s planned stock market launch. The company reportedly hopes to achieve a valuation of up to US$1.4 trillion – about 4% of the entire US gross domestic product.

How can we trust the mathematics is correct?

One of the key features of mathematics (pure mathematics especially) is its foundation in objective truth. Statements are either true or false, and there is almost never any ambiguity.

For centuries, we have agreed mathematical truth by consensus among mathematicians. In particular, new work is refereed by experts who check it is correct before it is published and becomes part of the “literature”.

This process can be very time consuming. Some longer, technical papers can take years to review.

So the eventual test of OpenAI’s results will be what the mathematical community makes of them after proper examination.

OpenAI has already retracted three of the papers due to an elementary error and amended several others due to mistakes that invalidated their results. This suggests, at the least, a lack of sufficient vetting before publication.

Machines checking machines

Over the past decade or more, there has been a push to “formalise” results with software systems such as Lean. These systems build up mathematics from basic principles or axioms, which means the software can check each step of logic in a result.

Until recently, formalisation efforts were largely focused towards theory and problems that had already been solved. One early example was the 2008 formalisation of the decades-old four colour theorem in graph theory, about the minimum number of colours required to colour a map.

Formalising a result takes a lot of knowledge and effort. However, AI systems are now being used for “autoformalisation” – automatically translating a result into a machine-checkable proof.

To encourage confidence in the correctness of its results, OpenAI claims to have formalised 300 of the main results at the time of writing.

However, not all formalisations are implemented correctly – and there is past evidence of problems with AI-generated formalisations, including that of the Navier-Stokes problem. It is currently unclear whether the mathematical community will accept OpenAI’s formalisations.

Excitement – and concern

Reactions to the OpenAI drop have been mixed, ranging from “obviously the most significant moment in mathematical history” to “mathematicians aren’t so thrilled at this giant pile of turd dumped at our doorstep”.

There is certainly great excitement about progress towards important mathematical problems in a variety of fields as experts attempt go over the results with a fine-toothed comb.

For many, this excitement is also mixed with uncertainty and concern: what does the volume and speed of these discoveries mean for the future of mathematics research, and how we are contributing to it? And who should carry the labour of trying to make sense of arguments that initially read as AI slop?

These papers, along with other AI-assisted developments, have forced the mathematical community to reflect on what we value. This is particularly important for students and early career researchers, who must confront the idea of their research being scooped by an ambitious AI user, or how their work will be valued.

And that conversation will continue long after the dust has settled on this latest tranche of discoveries.
Melissa Lee receives funding from the Australian Research Council.

Tags: , ,

No Comments

Leave a Reply

Rihanna Spotted Heading to César Restaurant in Paris After Céline Dion Concert Israeli attacks leave at least 20 dead in Gaza Mormon church shooting in Michigan ICE Shootings Trump wants US military back in Afghanistan US veloes United Nation’s demand for Gaza ceasefire