AI, Human Stupidity, and Mathematical Progress
Mathematicians should help shape what comes next.
As I write this, OpenAI has just released a proposed AI-generated solution to the Navier–Stokes Millennium Prize Problem. Whether the proof survives expert scrutiny matters enormously, but the announcement illustrates how quickly the situation is changing. Frontier AI systems can already explain graduate-level mathematics, find literature, criticize arguments, write proofs, and sometimes solve substantial research problems.
There are problems I have thought about for years that, six months ago, were mostly pointless to discuss seriously with an LLM; now the replies are often genuinely useful. I find this exciting because I want mathematics to develop faster.
There is also unusual urgency. We have no idea how capable these AI systems will be a year from now and the norms surrounding their use in research are only beginning to form. Mathematicians should therefore do more than either adopt or resist whatever products appear. We should actively engage with the people building them and help determine what kinds of AI use would actually accelerate mathematical science.
Mathematics Was Built Around Human Stupidity
One of humans greatest strengths in mathematics is our stupidity. Our working memory is tiny; a two page proof can be difficult to hold in mind as a single object, and long chains of deductions rapidly become impossible to navigate without intermediate concepts and organizing principles.
This weakness has been extraordinarily productive as it forces us to compress these enormous logical trees. Mathematicians invent notation, recognize the same structure in different places, package recurring arguments into reusable machinery, and search for definitions from which the important theorems become natural. Calculus was recognizable centuries ago, but generations of mathematicians kept reorganizing its foundations, notation, and exposition until ideas that once required exceptional insight could be taught routinely to undergraduates.
AI changes the constraints under which such mathematics develops. An AI system can potentially push much farther through an existing mathematical framework than a human would ever want to. A two-hundred-page machine-generated proof of an important theorem is a genuine scientific contribution if it is correct: the result becomes reliable mathematical knowledge and can be used as input elsewhere, even if nobody finds the proof enlightening.
But if machine-generated proofs become cheap, theorem production and conceptual organization can begin to separate.A good explanation may replace dozens of convoluted arguments, the right definition may turn difficult theorems into formal consequences, and a new framework may reveal that several subjects were studying instances of the same phenomenon. I would be delighted to have an artificial Grothendieck at my disposal, and I see no reason to assume that theory building will remain a uniquely human ability.
Mathematical Understanding Is Scientific Infrastructure
Looking back at my PhD thesis, I now think perhaps one or two insights contain most of what matters scientifically. I imagine this is a fairly typical experience. The remaining 180 pages of exposition, reconstructed folklore, and technical lemmas were necessary to turn those ideas into a rigorous thesis, but much of that labor could become dramatically cheaper with better AI tools. That creates the possibility of spending much more of our research time developing ideas rather than implementing them, and conceptual mathematical understanding matters because it produces more science.
An example close to my own research comes from the physics of topological phases of matter, a field recognized by the 2016 Nobel Prize in Physics. In the study of symmetry-protected topological phases, Anton Kapustin1 expressed their classification in terms of cobordism, and Freed and Hopkins2 subsequently developed a conceptual mathematical framework in terms of invertible quantum field theories and stable homotopy theory. Machinery developed decades earlier by algebraic topologists, such as the Adams spectral sequence, thereby became relevant to a problem in condensed-matter physics.
Nobody developing this machinery could have advertised that application in advance. Mathematics routinely creates scientific infrastructure before anyone knows what it will be infrastructure for, and so seemingly isolated mathematics can suddenly become useful.
A different example appears in the mathematics of quantum field theory. Mathematicians have spent almost a century looking for satisfactory definitions of quantum field theory. An enormous body of mathematics and physics was developed without one definition encompassing everything we want. Even though the fundamental open problem is not to prove a theorem, but to find the right definition, its solution would transform the field and lead to major progress.
What Is Worth Preserving?
I love the human practice of mathematics. I like blackboards, long conversations and slowly understanding why an argument works; after more than a decade of doing mathematics, this occupies a large part of my mental life. Some apparently old-fashioned practices remain useful precisely because they accommodate human cognition. A blackboard can be better than slides because the mathematics appears slowly, while drawing, pointing, and arranging ideas spatially become part of the explanation.3
The same principle should govern our response to AI: a practice is worth preserving when it contributes to understanding and discovery, not simply because it is traditional. Society benefits from mathematics through the knowledge and scientific infrastructure it creates, which does not perfectly coincide with everything mathematicians enjoy doing.
There are good reasons to work without computer assistance in many situations. Students need to learn how proofs work before outsourcing them, just as children learn arithmetic before relying on calculators, and researchers may discover the underlying idea precisely by working through a proof manually. But if a machine can reliably replace months of routine proof construction by seconds of computation, the interesting scientific question is what mathematics we can do with the months we get back.
Who Gets to Define Mathematical Progress?
The public generally has a simple picture of mathematics: mathematicians are exceptionally intelligent people who solve exceptionally difficult problems. We have benefited from this picture. Open problems are easy to advertise, Olympiad problems attract talented students, and stories about century-old conjectures make mathematical achievement visible to people who understandably cannot reasonably evaluate the significance of abstract research themselves.
AI companies therefore have an obvious and completely legitimate way to demonstrate progress: solve harder mathematical problems. A solution to a famous conjecture is real, impressive, measurable, and easy to communicate.
But as David Bessis has argued4, theorem proving is not a complete measure of mathematical progress. Unfortunately, it is much harder to demonstrate that a definition reorganizes a field, that a conceptual simplification will save thousands of hours of future work, or that machinery developed today will unexpectedly become useful in computer science thirty years later.
The incentives of mathematicians are different from AI companies; we care intensely about our own understanding and the intrinsic beauty of mathematics. Yet years spent inside the subject also give mathematicians expertise that AI researchers lack: we understand much better which ideas produce further mathematics and which apparently impressive results lead nowhere. Because mathematical value is so difficult to judge from outside, that expertise matters to society as well.
Until recently, we could largely avoid explaining this distinction. That luxury is disappearing. If AI systems solve the problems that mathematicians themselves have spent decades advertising as demonstrations of mathematical intelligence, we cannot respond afterward that the real mathematics was somewhere else. Even when true, it sounds like moving the goalposts. We need to participate earlier in the conversation about what mathematical progress means.
From Users to Participants
We should first talk much more openly to each other. We are in a wild-west period: models change every few weeks, researchers have vastly divergent ethical opinions about LLM use, and different areas of mathematics are discovering different uses for them. Nobody yet knows which practices will genuinely accelerate research and which will remain entertaining gimmicks. Mathematicians should show each other what they actually do with LLMs, what saves time, what produces nonsense, and what unexpectedly leads to useful mathematics.
The epistemic problem also becomes subtler as the AI models improve. A few months ago, when I asked an LLM a difficult mathematical question and received a complicated answer, half an hour of inspection would often reveal the mistake. Increasingly, I can spend the same half hour and still not know whether the model made a subtle error, has a correct idea that it cannot explain clearly, or has produced mathematics I simply do not yet understand. Better AI does not remove the need for mathematical judgment.
Graduate education should prepare students for this environment. Saying that students already know ChatGPT is no more persuasive than saying that they already know Google and therefore know how to find relevant literature. Students should learn to interrogate AI output, verify it, and remain skeptical of apparent authority. Working alongside powerful but unreliable artificial reasoners will increasingly be part of mathematical practice.
But the conversation cannot stay inside mathematics departments. If AI becomes an important research tool, mathematicians need sustained contact with the people building it. We should explain where AI systems repeatedly waste effort, what forms of assistance actually accelerate research, and what capabilities would transform the way we work.
Some of the most interesting goals will be difficult to turn into clean benchmarks. Can an AI system extract the conceptual core of a long technical argument? Can it propose a definition that genuinely simplifies the mathematics that follows? Can it recognize one useful structure behind results currently scattered across several papers? Expert mathematicians may disagree about the answers, but this disagreement reflects authentic subjective judgement about mathematical research rather than a defect in these questions.
In conclusion, there are four directions in which communication should improve: mathematicians sharing practical knowledge with other mathematicians; mathematicians explaining mathematical progress more accurately to the public; mathematicians preparing students for AI-assisted research; and most importantly: mathematicians working directly with AI researchers to help determine what capabilities are worth developing.
AI is rapidly becoming an increasingly powerful tool for doing actual research and I want mathematics to develop faster. We are currently in a short period in which both the technology and the culture surrounding it remain unsettled. We should be more than passive users, and help shape what comes next.
Acknowledgements
My view on AI in mathematics was shaped significantly in recent conversations with my collaborator David Aretz, and my dear friend Zowie Langdon. I also thank my collaborator Lukas Müller for sharing his personal experiences in collaborating with LLMs. This document improved from feedback from David Aretz, Theo Johnson-Freyd, Jeremiah Hockaday, Lukas Müller and David Prinz.
AI disclosure
This essay was developed in an extended conversation with ChatGPT. I first used the model as an additional discussion partner while forming my views on AI and mathematics, including by asking it to search for several essays and public debates on the subject for us to read and discuss. I supplied my own arguments, personal anecdotes, and a rough draft. ChatGPT helped organize the narrative, qualify some claims, and substantially rewrite the prose through several rounds of editing. I reviewed and directed each revision, including rejecting formulations that did not reflect my views or writing style. The arguments and opinions expressed here are ones I endorse, but a significant amount of the final wording and structure was produced collaboratively with ChatGPT.
References
- Anton Kapustin, Symmetry Protected Topological Phases, Anomalies, and Cobordisms: Beyond Group Cohomology, arXiv:1403.1467 (2014).
- Daniel S. Freed and Michael J. Hopkins, Reflection positivity and invertible topological phases, Geometry & Topology 25 (2021), 1165–1330.
- William P. Thurston, On Proof and Progress in Mathematics, Bulletin of the American Mathematical Society 30 (1994), 161–177.
- David Bessis, The fall of the theorem economy (2026).
- Tanya Klowden and Terence Tao, Mathematical methods and human thought in the age of AI, arXiv:2603.26524 (2026).
- Yuling Zhuang, Empowering students to critically validate AI-generated mathematical solutions through the rational questioning approach, Educational Studies in Mathematics (2026).
- First Proof Project (2026).