AI and mathematics: What should count as a research contribution?
Updated 1 October 2026
OpenAI · Open problems as AI tests
OpenAI describes open research problems as part of ongoing model evaluation and presents solved or advanced problems as evidence of mathematical skills.
We test our models on open research problems during development.
OpenAI · AI result before full understanding
OpenAI publishes model-generated mathematical arguments as new results after humans have developed them into manuscripts and created Lean certificates. The subsequent scholarly contextualisation is expressly left to the community.
OpenAI · Human-readable exposition
For the ten published results, the model arguments were developed into manuscripts, formalized and provided with explanations. OpenAI describes this as its own publishing process, not as a binding standard for all AI-generated results.
Google DeepMind · Open problems as AI tests
Aletheia was used, among other things, in a semi-autonomous evaluation of 700 open Erdős problems; four open questions were solved autonomously.
Google DeepMind · AI result before full understanding
DeepMind lists mathematical work produced without human intervention as autonomous research and as “publishable quality” in its taxonomy of AI-assisted mathematics.
Reliable autonomous research.
Google DeepMind · Human-readable exposition
The workflows described combine automated and human verification and are intended to give researchers more space for conceptual depth. An explicit rule that every AI result must be developed in a human-understandable manner before regular publication is not formulated.
Anthropic · Open problems as AI tests
An Anthropic employee specifically had Claude tackle the Riemann conjecture; Anthropic sees the resulting result as an example of the progress of mathematical modeling skills. A general requirement for open problems as a benchmark is not formulated.
Anthropic · AI result before full understanding
Anthropic describes Claude's new finding as a mathematical result, but had it examined and validated by its own mathematicians and checked by external experts. The source thus supports recognition, but not correctness alone, as a sufficient criterion.
Anthropic · Human-readable exposition
While Anthropic expects formalized evidence to become common alongside human expositions, it explicitly rejects treating the formal version as a substitute for intelligible exposition.
a formalized proof should not replace a human-understandable exposition
Signatories of the Fields Declaration · Open problems as AI tests
The statement explicitly calls AI companies' push to use mathematical problem solving as a benchmark as harmful to science and the mathematical community.
the push by AI companies to solve mathematical problems as a benchmark is detrimental
Signatories of the Fields Declaration · AI result before full understanding
The statement highlights conceptual understanding and insight as the primary goal. For the signatories, mere problem solving is only a tool and representative and can miss the actual research purpose without contextualisation and further development.
Signatories of the Fields Declaration · Human-readable exposition
The statement criticizes hasty announcements without adequate exposition, exposition of new methods and clean attribution, and emphasizes the human transmission and integration of ideas.
Signatories of the Fields Declaration · Reward understanding more
The signatories declare conceptual understanding and insight to be the primary goal and problem solving to be the proxy. However, they do not specify a specific new weighting for recruitment, funding or prices.
Advisory Group on Mathematics and Artificial Intelligence · Open problems as AI tests
The independent panel says it does not support this practice and calls on Frontier labs to stop testing advanced math problems on non-public proprietary models.
we ask them to stop testing advanced mathematical problems on proprietary models
Advisory Group on Mathematics and Artificial Intelligence · AI result before full understanding
AGMAI recommends that significant results be published responsibly as quickly as possible, even if no one initially fully understands the argument. The laboratories should then take responsibility and funding for ensuring that human understanding follows.
Advisory Group on Mathematics and Artificial Intelligence · Human-readable exposition
AGMAI requires understandable, conventionally written versions, extensive formalization and support for workshops, expositions and other work that builds human understanding.
Advisory Group on Mathematics and Artificial Intelligence · Reward understanding more
AGMAI calls for substantial funding for the development of human understanding, but leaves its allocation to independent institutions. The committee does not formulate a general new hiring or pricing logic.
Leiden Declaration Working Group · Open problems as AI tests
She criticizes the fact that concrete mathematical tasks are misleadingly used as a measure of the general reasoning ability of commercial products and that research could be prioritized based on their ability to be automated.
Leiden Declaration Working Group · AI result before full understanding
For published AI-assisted research, human authors should bear responsibility for accuracy, appropriateness and sources; Press releases or blogs should not replace peer review and community scrutiny.
Leiden Declaration Working Group · Human-readable exposition
The statement names human descriptions of central arguments, formal verification, cross-checks and external pre-testing as possible standards for automatically generated results.
Mathematics produces not only a body of results, but also understanding
Leiden Declaration Working Group · Reward understanding more
The declaration warns against distorted incentives in hiring, funding and recognition, and calls for industry-linked funding to respect these values. It does not specify a particular weighting formula.
Timothy Nguyen · Open problems as AI tests
He argues that it is itself a discovery when previously particularly difficult problems become solvable for machines, and calls for curiosity rather than nostalgia. He does not formulate an explicit position on proprietary benchmarks.
Timothy Nguyen · AI result before full understanding
Nguyen sees building on formally verified machine proofs as an expansion of the mathematical toolbox and emphasizes that mathematics can also have value in discovering truths before full understanding.
Understanding is also, in some sense, a luxury.
Timothy Nguyen · Human-readable exposition
He declares understanding the value of mathematics, but at the same time argues that formally verified machine evidence can be used before a single human has fully absorbed it.
Timothy Nguyen · Reward understanding more
Nguyen sees the traditional awarding of prizes and prestige for being the first to solve difficult problems using AI under pressure. He leaves it open which institutional recognition should take its place.
Timothy Gowers · Open problems as AI tests
His focus is on the consequences of autonomous AI mathematics for authorship, culture and recognition. He does not give a clear answer as to whether companies should specifically use open problems as a model benchmark.
Timothy Gowers · AI result before full understanding
Gowers does not describe a future of automatically formalized AI solutions without a human author as a reason to reject these results and does not consider human attribution to be compelling in such cases.
Timothy Gowers · Human-readable exposition
He considers poorly written, difficult-to-understand AI evidence to be deficient, but expressly considers that future LLMs can provide the necessary explanation themselves.
Timothy Gowers · Reward understanding more
Gowers suggests that in AI-influenced mathematics, the work of understanding, curating and explaining should be rewarded much more than simply producing an automatic initial solution.
Person B to get the lion’s share of the credit.
Leslie Ann Goldberg · Open problems as AI tests
Goldberg does not object to solving open problems that have already been formulated, but expressly calls for this achievement to be downgraded as a goal and standard of recognition.
Leslie Ann Goldberg · AI result before full understanding
Goldberg suggests a separate ArXiv area for AI-generated evidence with Lean certification and says that as yet unpublished laboratory results, statements and certificates should simply be made publicly available.
Leslie Ann Goldberg · Human-readable exposition
Goldberg separates machine correctness checking from scientific exposition. AI evidence should be commented on and explained; in the future, she sees regular journals more as places of high-quality explanation.
Leslie Ann Goldberg · Reward understanding more
Goldberg specifically calls for changes to postdoctoral and faculty selection and tenure and suggests greater emphasis on the quality of exposition and understanding.
focus evaluation on the quality of exposition/understanding
Mathathon Coalition · Open problems as AI tests
After criticism, Mathathon was focused on the question of how AI can expand human mathematical understanding. Pure problem solving remains possible, but should not be the sole measure of performance.
Mathathon Coalition · AI result before full understanding
The coalition explicitly identifies problem solving as just one component of mathematics and places understanding and exposition alongside it. This acknowledges correctness, but not as a complete research standard.
Mathathon Coalition · Human-readable exposition
The new Mathathon aims to use AI not just to solve, but to create new evidence, explanations and human-useable understanding.
mathematical understanding and exposition are similarly meaningful
Mathathon Coalition · Reward understanding more
The coalition criticizes that these achievements are often overlooked by funders and hiring committees and wants to specifically recognize human mathematicians for such work.
Sources
- OpenAI — Zehn Fortschritte in Mathematik und theoretischer Informatik ·
- Google DeepMind — Accelerating Mathematical and Scientific Discovery with Gemini Deep Think ·
- Anthropic — Learning more about Claude's mathematical capabilities ·
- Anthropic — Formalizing Fermat's Last Theorem ·
- Math and AI — A Severe Misalignment of AI in Mathematics ·
- Advisory Group on Mathematics and Artificial Intelligence — Responsible Release of AI-Generated Mathematics ·
- Leiden Declaration on Artificial Intelligence and Mathematics — Leiden Declaration on Artificial Intelligence and Mathematics ·
- Timothy Nguyen — A Response to “A Severe Misalignment of AI in Mathematics” ·
- Gowers's Weblog — Thoughts about the Leiden Declaration ·
- University of Oxford — Mathematics and AI ·
- Proofs and Prompts — Joint Statement about Mathathon ·
More
No debates have been published in this category yet.
View allNo debates have been published in this category yet.
View allNo debates have been published in this category yet.
View allNo debates have been published in this category yet.
View allNo debates have been published in this category yet.
View all
