A New Consciousness of Mathematics
In an age of proof abundance, what becomes scarce?
“One often hears that one is seeking beauty in mathematics.
It may be true. But I think more, one is seeking understanding.
The bonus of your understanding is always that it’s beautiful.”
- Barry Mazur
At the start of May, I attended the second day of Stanford’s Future of Mathematics Symposium.1 Some of the best mathematicians in the world—including Fields medalists Terence Tao, Maryna Viazovska, and Michael Freedman, Stanford professor and current president of the American Mathematical Society Ravi Vakil, and mathematical formalization forerunner Kevin Buzzard—along with leading AI research scientists Sébastien Bubeck (OpenAI), Thang Luong and Adam Brown (DeepMind), and more came together to weigh in on the new era of AI for math discovery, verification, and collaboration. Evidently, the star power in the room was enormous; many of my current professors (who are also at the frontier of their fields) were mere audience members.2
As AI mathematics develops rapidly (while writing this, an internal OpenAI model resolved a famous open problem in combinatorial geometry), the conversation of what mathematics will look like is one that requires all hands on deck, from students to professors. Many questions arise: what does mathematics education look like with the incredible new capability of personalized tutoring, in lockstep with the immense ease of cheating? How do we train undergraduates and PhD students in research, balancing the intellectual benefit of struggling with hard problems and the outsourcing of mathematical thinking to solve interesting problems faster than ever? What does mathematical research become, and who is it for?
Some of these questions aren’t new—just reformulated—and are becoming increasingly difficult to answer. Even more urgent is who gets to answer these questions; a mathematician’s view is different from that of the public, frontier AI labs, or universities. What is to become of a discipline that is viewed as “solved” for the majority of public use cases?
I’m writing this to make sense of what doing math means in this new era, from a student’s perspective. I grew up going to math camps every summer, finding great joy and friendships in learning beautiful math and struggling with hard problems together. Many of my math friends (now in their early 20s) are feeling disheartened and torn about pursuing mathematics, a subject that we’ve invested time in out of pure interest and fun, but is now transforming into something increasingly unfamiliar. It feels like we are a generation in between old-school math and new world AI-assisted problem solving.
A lot of discourse around AI for math moreover feels inaccessible to the general public, who are also grappling with the transition. It is not controversial to value human-created writing and art over its AI-generated counterparts, even though it can be difficult to tell the difference at face value. Point being, there is still great importance in the person behind the creation. But if I showed a random person a theorem proved by a mathematician versus one proved by an LLM, there’s a good chance they wouldn’t care either way—truth is truth, right?
Let me explain why I care, why you might care too, and why mathematics is an interesting case study in the age of AI. To start, I will try to situate any readers who are unfamiliar with what math research is, so we can all better understand its future behind the hype.
What is mathematical research, exactly?
A question I get a lot from my friends is:
“What exactly is math research? What do you do?”
Research in other fields is a lot easier for the public to visualize: a chemist exploding things in a lab in the name of experiments, a biologist analyzing samples through a microscope, a historian sifting through archives to piece together a narrative. But what do mathematicians do? The image of someone sitting around in a room, scribbling on a blackboard, maybe typesetting a proof, is perhaps not as compelling. Oh, but don’t forget the long thinking-walks. And staring into space. Aesthetic to some, uninteresting to others (though there are plenty of movies that dramatize the tortured genius to the point of being Oscar-worthy).
The answer to what math research is depends on where you are in your research career; there is no exact answer. I can only answer as a student who was lucky enough to experience its process under the mentorship of a few kind and patient mathematicians. To me, amateur math research (pre-LLMs) looked something like:
Find at least one Mathematician Mentor who is nice enough to take the time to teach you some things.
Ask said Mathematician Mentor for a problem that they think is solvable by you, but haven’t gotten around to doing because they have more important/interesting things to do.
Read a lot. Try different things a lot. Fail a lot. Question a lot. How did anyone ever come up with this? Eventually, solve the question or some variant of it.
Have someone check your solution.
Write up the proof and, if it’s interesting enough, publish it. In my experience, this step involves a lot of guidance from the generous Mathematician Mentor on what it means to write a math paper.
Somewhere along the way, you usually have the great fortune of building some intuition for the problem. Then you do it all over again for, perhaps, a related problem, or a different one altogether. Repeat again and again and again, until maybe one day, if you persevere, you become the Mathematician Mentor.
But the timeline of these steps is quickly changing. In particular, Step 3: Try to solve, which is typically the bottleneck, is becoming unimaginably faster with AI. Many have now shifted their attention to Step 5: Proof-writing as the new bottleneck to resolve, though this is currently much more reliant on human input to make a result understandable.
The three stages of proof-writing
In his talk “New Mathematical Workflows”, Terence Tao characterized math research from the complete opposite end of the research career spectrum (the Fields medalist end). In particular, he outlined the three stages of writing a proof:
Proof Generation
Proof Verification
Proof Digestion
The first two stages are fairly self-explanatory; in generation, you come up with some line of reasoning to prove a given theorem, and verification is the stamp of approval that the proof is 100% correct. There is a clear objective—to answer: is this true?—which makes these stages well-suited to be accelerated by AI. This is exactly what we’re seeing: AI can increasingly generate plausible arguments and formal systems can verify correctness.
Proof digestion is the stage where proofs become mathematical understanding, and is thus much harder to automate. Tao describes this stage as “subjective and human-paced.” Consequently, the most remarkable advances from AI are in proof generation and verification, and this makes the third stage, digestion, even more pressing.
If proofs can be generated and verified faster than we can understand, what happens next? Before attempting to give an answer, we need to better understand each stage.
AI for proof generation
Earlier this year, a group of renowned mathematicians put together 10 diverse, technical lemmas from their own research to assess the ability of models to autonomously solve math problems in natural language in a project called First Proof. Thus far, somewhere between 6 and 8 of them have been solved by LLMs (there is no official grading). But as mathematician Daniel Litt says in his insightful essay Mathematics in the Library of Babel, these LLM solutions sit amongst “an enormous amount of garbage,” and a few of them are “arguably semi-autonomous”, though the intention of First Proof is to assess autonomous proof-writing.
That being said, LLMs being able to produce proofs for 6-8 of these problems this year is incredibly impressive. To put it into perspective, AI autonomously achieved gold medal-level performance in the International Math Olympiad (IMO) for the first time last year. Though this was widely regarded as a huge accomplishment, most mathematicians were skeptical of this translating to real research-level mathematics, despite the enormous hype from AI labs. At the time, scientist and author Gary Marcus wrote, “Eric Schmidt’s recent prediction (made before IMO 2025), that there will be AI-based “world-class mathematicians” within a year, is pure fantasy.” He also references Kevin Buzzard’s post on what this means for mathematical research:

But, less than one year later, we are seeing research-level proofs from LLMs. Even more astonishingly, OpenAI announced days ago that their internal general-purpose model produced a proof—or more accurately, a disproof—of the planar unit distance problem: given n points in the plane, how many pairs of points can be exactly distance 1 apart?
This problem was first posed by prolific mathematician Paul Erdős3 in 1946, and was considered the most well-known open problem in combinatorial geometry (and one of his favorite problems). The OpenAI model used deep tools from algebraic number theory to construct an infinite family of examples that is an improvement from the previous “square grid” construction that was long thought to be the best one could do. It is clear that AI’s capability of doing deep mathematics research is increasing at an unpredictable rate, though still under much guidance and input from professional mathematicians. In fact, soon after the result was published, mathematician Will Sawin released his refinement which provides an explicit lower bound for the problem.
Given AI’s progress in proof generation, how do we actually check if the proofs are correct? I’ve told you about some of its significant successes but, like most of the hype out there, I have not yet mentioned the incredible number of incredibly unsuccessful attempts. How are we supposed to separate what is true from what is slop? By hand, by machine? And so we arrive at proof verification.
AI for proof verification
If you aren’t involved in math research, I’ll let you in on an unintuitive aspect of it: confidence matters. When a professional mathematician is writing a proof, they don’t have the time or space or patience to explain all the statements they think to be true, and so they sometimes just claim it and leave it to the reader to verify as a fun, little exercise (that is often neither fun, nor little).4
One can clearly see how this might go wrong. Even if you are pretty sure something is true, what happens when it is not? Ideally, a paper referee will catch your mistake before your proof gets published, but occasionally they don’t and incorrect things get published. Usually if it is important enough, someone will find out soon enough. But obviously, this is not a perfect system.
The most prominent modern way to rigorously check proofs without being susceptible to human fallibility is with computer formalization. You write every line of a proof with a proof assistant, and the system checks the logic of every single step down to a set of fundamental axioms to verify it to be true.
Kevin Buzzard, a prominent advocate and leader in formalization who popularized the use of Lean (a functional programming language for verifying proofs), has a nice introduction to what formalization is. In short, formalizing math with computer proof checkers gives a 100% guarantee that a given statement is correct, and has the bonus benefit of occasionally helping mathematicians find new abstractions through their error-checking.
Why don’t we just formalize all known math proofs to check if they are correct?
Worries of math homogenization aside, formalizing a proof is incredibly labor-intensive. It can take anywhere from a few weeks to over a year to write and check definitions and proofs, and can require thousands of lines of code. Moreover, Lean proofs, for instance, heavily rely upon the standard mathematics library Mathlib, which is written and managed by experts and consists of ~2,000,000 lines of code.
Upon hearing the words “labor-intensive,” AI comes knocking, and so emerges the idea of autoformalization: using AI to translate informal, human proofs to formal, machine-verifiable code for proof assistants like Lean. There are several AI startups focused on autoformalization, such as Math Inc., Axiom Math, and Harmonic.
Most notably, Math Inc’s formalization agent Gauss took three weeks to autoformalize the strong Prime Number Theorem by September 2025, and just this February, successfully wrote 200,000 lines of Lean autonomously to formally verify Maryna Viazovska’s 2022 Fields Medal work on the optimality of sphere packing in dimensions 8 and 24. This is the most impressive and exciting autoformalization result to date, and also happened in just three weeks.
So are we done? Can we check everything autonomously now, including LLM-generated proofs?
Unfortunately, what many autoformalization headlines fail to acknowledge is exactly what mathematician Alex Kontorovich notes in his lecture on AI and formalization:
“Autoformalization works only because it sits on top of a massive, comprehensive, efficient, coherent monorepo of high-quality formalized mathematics, namely Mathlib. And even in the PNT+ and Viazovska examples, the autoformalizations still depended on substantial earlier human work: setting up the right definitions, the right API, the right abstractions, and so on.”
Consequently, autoformalization results are typically met with some pushback from the people doing the labor of formalizing the mathematics their result sits upon, which is too often ignored. We also cannot expect autoformalization to get very far on its own when many difficult questions rely on definitions that aren’t currently found in any formal mathematics library. Indeed, Buzzard identified definitions as the biggest challenge for AI-driven autoformalization right now:

That said, autoformalization has progressed by leaps and bounds in the past six months, and I can only imagine that it will continue to do so. I’m excited to see where it goes, but also hope that there is sufficient acknowledgment, credit, and compensation for the immense amount of human mathematical labor that powers these results.
Proof generation and verification are obviously very important, and AI is progressing quickly for these stages of proof-writing, but we must not forget the final stage: proof digestion. Where generation and verification accomplish the goal of determining what is true, digestion answers the question: what have we learned, and how does this progress mathematics?
Humans for proof digestion
Proof digestion is where the transfer of knowledge happens, where a piece of mathematics is simplified to the point where you can teach it; a good proof contributes greater understanding, a bad proof is incoherent.
The explicit goal of this stage is to, of course, write up the proof you have discovered. But during this stage, Tao characterizes the additional “implicit goals” that humans accomplish: authors draw connections to the relevant literature, describe the new tools that they’ve developed in pursuit of this problem and how it might translate to other work, comment on the effectiveness of various techniques for such problems, and give readers a sense of how their work is situated within the field.
In the context of formalization, Kontorovich calls this “canonization,“ which he defines to be making a piece of mathematics “general, reusable, coherent, efficient, and compatible” with the rest of the field. Viazovska echoed his sentiment at the symposium, saying that we want to write proofs where we can “make its parts as recyclable as possible.” This is how mathematics progresses, and it is not easy. It requires a lot of time, effort, and prerequisite knowledge and skills in your field.
Now, Tao warns that the increasing capability of AI for math now runs the risk of decoupling the explicit and implicit goals. For instance, LLMs largely treat the interesting and non-interesting parts of a proof in the same way, a deep inequality can take up the same amount of space as an obvious simplification. Where a mathematician might go on a tangent to canonize said inequality and remark on its usefulness, the LLM has no comment.
Okay, this doesn’t seem so bad. Maybe in verifying a model’s work, a human mathematician can simply recognize such gaps and expand upon them (as was the case for the planar distance problem):
But the more urgent problem is that, for the first time, we are entering an era of proof abundance where anyone with access to an LLM can generate a proof, even if they may not have the necessary knowledge or patience to verify it, let alone digest it. That may not stop them from releasing it into the world, especially if they are racing to be the first to solve it.
Proof indigestion
This is the problem dubbed proof indigestion. Some people think mathematicians are worried about losing their jobs to AI, that they fear they will no longer have problems to solve (which I do not believe), but the real crisis is that we may soon have more proofs than we can digest. Moreover, the burden of digesting AI-generated results falls on working mathematicians, taking their time away from working on their own research.
Frontier AI models can certainly resolve many hard problems, as we have already seen, but it currently goes hand-in-hand with a lot of incorrect work, and sometimes the mistakes are very subtle and hard to tease out, as seen with the tremendous amount of slop from the First Proof attempts.
Proof indigestion happens when the rate of proof production is far exceeding the rate at which the mathematical community can understand, verify, and canonize proofs.
In this vein, Tao presented the hypothetical of having a solved problem that has been autonomously verified, but no human expert can give a talk on it or answer questions. The field as a whole can’t progress because no human can make sense of the proof or appreciate its value, and thus cannot situate it within mathematics.
But why are we continuously seeing headlines like “AI solves math” when there is much left to be done?
The benchmark trap
To assess LLMs’ problem-solving abilities, we use benchmarks like FrontierMath, math competitions and olympiads like the IMO, First Proof, and more. But these benchmarks can be misleading, and systematically undervalue the digestive part of math. They test whether models can answer hard, well-posed questions, but the terms of the problem have been provided, and better yet, canonized by real mathematicians in the literature.
During the symposium, mathematician Ravi Vakil talked about the current “fetishization of benchmarks, especially in the Twitterverse.” He was referencing the incredible hype that follows each milestone for AI mathematics, chock full of claims that math has been solved. In the middle of it all, mathematicians attempt to situate the results while dodging attacks for their refusal to acknowledge the implications of it (such as mathematicians becoming obsolete). It is difficult to engage in the discourse without being cast to one side.
When models can pass these benchmarks, it is absolutely an impressive display of their developing problem-solving ability in both competition and research-level math. However, these benchmarks do not test the models on the fuzzier side of mathematics: creating new concepts, finding the right question, the right definition, the right abstraction, and stringing it all together in a deep way. Mathematics also requires the ability to choose the terms, not just connect what is already known.
Building intuition and maturity
“What we’re really doing as mathematicians is learning how to abide in a place of ignorance, how to move ourselves around and, sometimes, into knowing.”
- Mike Hopkins
If we revisit my rough outline of amateur math research, the most interesting step of the process is Step 2: how do mathematicians determine what constitutes a good research question, either for themselves or for others? It is absurdly easy to stumble upon simple-to-state, incredibly-hard-to-prove problems. When I was 13, I was convinced I could prove that there are infinitely many pairs of primes with a difference of 2 if I just thought about it hard enough before I fell asleep. (Turns out the twin prime conjecture is more difficult than counting sheep.)
Therein lies the idea of mathematical maturity, of intuition that is hard-won through years of working on problems, getting things wrong, and honing your knowledge and skills for problems you find interesting. It is developing a sense of what you believe to be true or false. In his essay “The Fall of the Theorem Economy” (which I highly recommend), David Bessis distinguishes between “official math” and “secret math”: official math consists of the formal work and derivations, and it is either right or wrong, but secret math is what we’re hinting at here, it is “soft, fuzzy, subjective” and is the “human part of the story” that everyone delights in. For many, it is why we do math.
Bessis goes on in his essay to ask:
“Why don’t we construct benchmarks that are fairer to humans? Because we can’t. Because the true value of mathematics, the collective and individual elevation of our worldviews, is ill-defined and intangible.”
There is a lot of optimism that AI will make exposition more important. If LLMs become better at proof generation, perhaps humans will shift towards turning it into usable mathematics. But we need to be careful about romanticizing the role of humans in mathematical exposition. The fact remains that proving a hard result is currently much, much more valued than writing a beautiful exposition of someone else’s theorem. My friend Sean put it this way: “theorems are currency.” Proving theorems gets you jobs, prizes, grants, and recognition for your work; you still need to solve hard problems to establish yourself as a mathematician.
Theorems as currency
Even mathematicians who are enthusiastic about AI doing math are worried about getting grants and positions, especially if they fail to keep up with AI mathematics. The profession is currently structurally unprepared for an era in which theorem-production may no longer be the bottleneck.
At the symposium, Sébastien Bubeck at OpenAI suggested trying to draw “more focus on the understanding piece, such as…having higher prestige associated with books that focus on common insights across many solutions,” but there is unfortunately no incentive structure to support this focus yet. This is a well-intentioned sentiment, but its dark underbelly is that we are losing the privilege and prestige of solving problems. Most love mathematics for the process, the challenge, the long-term process of engaging with a problem, and its secret human side.
So how do we measure what it means to be good at math—to have intuition to develop novel theories and ask new questions—beyond problem-solving? At the moment, there are no official benchmarks for mathematical discovery, and creating and assessing them is time-intensive and ill-defined. But that’s not to say frontier models will never be able to create truly novel math; in fact, I’m sure they soon will given the momentum on AI mathematics and the tremendous amounts of compute that is becoming available for these problems. The closest we’ve gotten so far, I think, is the recent unit distance result, but I wouldn’t call it fully autonomous new mathematics and it has several caveats that are outlined in the involved mathematicians’ remarks.
But now that proofs are easier to generate, another question follows immediately: whose proof is it? If theorems are currency, who earns it?
On the shoulders of giants
A phrase I hear a lot in math research is “standing on the shoulders of giants.” It means that math research is rarely truly original, but rather stands upon extensive literature and history of a problem. Acknowledging this influence and history is paramount in papers, and a key part of proof digestion.
Most scientific disciplines have some system of author ordering by contribution for academic papers, but in mathematics research, authors are interestingly sorted alphabetically, by last name. How can this be fair, you ask? I think Colin Reid explained this nicely on MathOverflow:
My understanding (as someone who hasn’t been in this business very long) is that when pure mathematicians co-author a paper, they form a kind of partnership as equal partners, and all credit for everything in the paper goes to the partnership rather than individuals, regardless of what actually happened behind the scenes. As for why:
It is seen to be unfair to say one person’s work is more important than another’s when each depends on the other’s results and insights in a critical way.
Giving academic credit for anything other than a novel intellectual contribution to the content of a paper, for instance for securing funding or having a higher professional status (eg a professor vs a doctoral student) is anathema to most pure mathematicians, in a way that it wouldn’t be for other scientists.
The culture of humility is particularly strong in pure mathematics. If a mathematician insists on being ‘lead author’ on a paper, that’s bad for his/her reputation among mathematicians, which cancels out the extra credit that would otherwise accrue to a lead author.
But how does this change if an LLM autonomously generates an entire proof, even if you were the one to conceive the problem? It doesn’t seem fair to have a human author on this, so Bubeck points out that AI-generated proofs should live in a different place than human-derived ones. This is just one in a line of many logistical questions there are to answer about mathematical ownership within the community.
Another concern is that LLM-generated proofs often do not appropriately cite and reference related ideas in the literature. Melanie Wood comments on this with regards to the unit distance problem:
“If a human came up with this argument and didn’t cite such previous work, we would assume that they were unfamiliar with the previous work and came up with the ideas independently, since our professional norms require us to cite previous work whose ideas influenced our work. On the other hand, Chat GPT is in some sense “familiar” with all the previous work. In the future we can expect humans to write many papers that include ideas suggested by AI. Mathematicians need to think about what best practices and proper citation is in these kind of situations, and come to a common understanding as a community.”
The future we’re looking at is one where we stand on both the shoulders of giants and giant machines to “see further than anyone before,” as Bessis puts it. It is stunning and bittersweet.
At the end of OpenAI’s unit distance proof announcement, they write: “That future still depends on human judgment. Expertise becomes more valuable, not less. AI can help search, suggest, and verify. People choose the problems that matter, interpret the results, and decide what questions to pursue next.”
This is where AI for math becomes pedagogical. If mathematical intuition and maturity are built through time invested into confusion and difficult problems, what happens if machines remove the confusion too early, just because they can? How do we train new researchers, or even elementary schoolers? If we don’t answer these questions soon, we will have fewer and fewer people who can decide which questions matter.
As easy as 1, 2, 3?
“Everything I need to know I learned in kindergarten.”
– Robert Fulgham
Last week, I overheard (or perhaps eavesdropped on) a conversation next to me about the value of mathematical skills, from the perspective of an econ major:
“Mathematical skills aren’t going to matter anymore. GPT can do it.”
That was a tough pill to swallow while working on my math problem set. So why even do the problem set, if GPT could do it faster and, likely, more cleverly? That is the explicit purpose of problem sets—to find the answer—and it can be easily accomplished.
You can generalize this question to: “Why should I learn X if I will always have access to AI, which can already do X very well and will presumably only get better at it?” The consequences for doing this for X are not seen in the short-term, and maybe not even in the long-term. But it becomes a race to the bottom of what is worth learning. Bessis puts it this way:
“The truth is that there really is no bottom, and nothing prevents humans from functioning in a culture where there is no vocabulary for numbers above 5 (as is the case for some hunter-gatherer tribes).”
Most can agree that we don’t want this; calculators arrived, and we all know that there is still an advantage in developing young kids’ number sense. When we give problems to models and students, they are clearly given for different purposes.
The implicit purpose of math homework is to learn how to be stuck, to test ideas, to learn how to fail and, better yet, learn from your mistakes and take them to the next problem. The implicit goal of math education and research is to build intuition and to learn how to ask questions. When we give problems to models, we’re testing them.
Professor Vakil elaborated on this during a roundtable, saying:
“Problem sets should be time where people have a chance to fail. The goal of a pset is not to get 100%, the goal of a course is not to get 100%. The goal is, secretly, to learn something.”
Many argue that AI can help students learn math better than ever, offering personalized tutoring that teachers don’t have the bandwidth to give. I think that can be true; it is an infinitely patient tutor and librarian that will answer any question you have. But we are still early in our understanding of the long-term developmental impact of AI for education and how to design it effectively, and a recent study on using generative AI to learn high school mathematics (Bastani et al. 2025) has revealed that while performance goes up during practice sessions with AI-based tutoring, students underperform on tests where AI is removed in comparison to students who never had access.
It seems crazy to say that it would be better for children to not have access to these tools that could accelerate their understanding, but mathematician and educator Po-Shen Loh remarked on how his teaching visits in high-poverty areas—where many kids do not have phones or sometimes even internet access—have been some of the best classrooms he has taught in, full of kids suggesting new ideas, engaging with the material and their classmates’ ideas.
Part of learning math is dealing with long, hard problems, and when you are no longer forced to do so, you may not actually learn math. This finding is echoed in universities across the country: while STEM homework scores are going up—with medians that are often close to 100%—exam scores are simultaneously going down as students are starting to use AI to finish their homework, and thus not fully achieving understanding.
Many professors at Stanford are noticing that students do not go to office hours anymore. Now that anyone can accomplish the explicit goals very well at home, on their own time—debugging code, answering general questions, pushing through a specific problem roadblock, or, in the worst cases, the entire problem set itself—they no longer have to go through the trouble of questioning a professor or TA (which can have mixed results).
But going to office hours accomplishes a suite of implicit goals: organic learning, fostering a working relationship with professors and academics, learning how to collaboratively solve a problem, working through difficult steps on your feet, having a mentor help you find gaps in your understanding, and asking and answering questions you didn’t even know you had. We are losing these learnings at a large scale, and it is extremely troubling.
“No cost, except for time.”
The final topic of the symposium was on the future of assessment in math education by the Fields Medalist panel. The response to rising homework scores and falling exam grades is different across schools, departments, and professors; some are responding by redistributing more weight to exam grades, others are trying to integrate AI interaction in their problem sets (to varying degrees of success). How do we fairly assess students’ learning, while still facilitating the ability to fail and learn from it, without being punished?
One option is through more frequent assessments that are weighted lower, so that mistakes feel less pricey. Tao said that they are trying out quiz systems at UCLA, where they give students questions that are automatically graded. If the student gets it right, great! But if they get it wrong, they are locked out of the quiz for 48 hours, during which they are free to think about the problem on their own time before taking it again. They are given unlimited tries to get it right, and the only penalty is time delay so as to reduce the price of failure. Tao says that they have found that this makes people less incentivized to rush to AI right away.
Ending the symposium, Tao’s last words on allowing mistakes were that there is “no cost, except for time." I think this is emblematic of the current attitude of college students towards learning.
Is this a problem for psychology?
When I taught kids competition math, I would harp on about how grit is the most important part of problem-solving. In the midst of a competition, when all you have left are problems that you don’t know how to solve, you have no choice but to put your head down and try everything you possibly can. And sometimes, if you bash enough things together, you find your bumbling way to an answer. It may not be beautiful, but you can make it to places you didn’t think you could. But now, it is becoming difficult to intentionally choose the pain of an unsolved problem, the inevitable mistakes on the path, over rapid answers to move onto the next, next, next problem.
Notably, there also seems to be a new psychological distortion of the amount of time it takes to do things. Previously where students might have allocated four or five days to work on a problem set to finish it by the deadline, it now gets one day. These unrealistic timelines, combined with the pressure of getting everything right, often lead to relying on AI to cross the finish line in time.
At some level, students do want to learn the material. At least for math majors in college, that is presumably why you’re taking the class or even why you’re in the major. Vakil says that, now more than ever:
“We need to psychologically strengthen our students.”
Early on, many students form a self-perception of their ability in math, and it is confusingly difficult to change—maybe because of the binary nature of it. I meet people all the time who claim to be terrible at math. I tell them that I don’t think that’s true, they just haven’t had someone tell them that their struggle is part of the process and it is, in fact, core to doing mathematics.
Importantly, we have been talking about college-aged students. Adults, albeit young adults, are struggling to make the choice to do the mental labor at the cost of time, especially as everyone around them is using AI to do things faster, better, more productively, to free up their time to do more and more and more. So how can we possibly expect even younger students, students who are impressionable and still developing their emotional, psychological, and mental skills, to have the “willpower” to solve all problems on their own, when it is so easy to outsource their thinking to AI?
First and foremost, we have to offer students a system that embraces their curiosity, and encourages failure in pursuit of learning. Mathematics can appear as a very binary subject: you get points if you are correct, and none if you are wrong. But as you start doing high-level, proof-based math, it becomes a little less clear. Students should earn partial credit for correct reasoning, or a good idea that just didn’t pan out. We need to incentivize being brave enough to make mistakes, when students are inclined to hide them. To do this, we need to make the failure recovery process fun and rewarding, so that they do not need to go straight to AI to solve their problems. We need to convince them that struggling is the point, and that it is worth the time. But, perhaps even more importantly, we need to reward those who engage in this struggle, to create incentive to do so.
Part of the challenge, then, is that we cannot motivate students by giving them only one reason to do mathematics. Its usefulness will work for some, its beauty will convince others. Its precision, its cleverness, its community—there is no universal reason, but there is a reason to be found for every person. If we want students to choose the slower path of understanding over quickly generating answers, we have to give them many ways to find meaning in it, to help them find their reason.
“Mathematics is many things to many people. Like music, it resists definition…Willard Gibbs thought of mathematics as a language. Hilbert thought of it as a game. For Benjamin Pierce it was “the science that draws necessary conclusions.” Hardy joyfully stressed its uselessness; Hogben stressed its practicality. Mill thought it an empirical science, whereas to Sullivan it was an art, and to the wonderful J.J. Sylvester, it was “the music of reason.”
I find this ambiguity consoling. It suggests that mathematics has so many mansions that there is room for all of us; it does not appeal merely to one type of mind. If mathematics can be so many things to so many great thinkers of the ages, then as teachers we can assume no less diversity among our students.”
- Karl J. Smith, Mathematical Reflections: Standing on the Shoulders of Giants
The goal of teaching is not to impart just one kind of mathematics, but to make room for different students to find their own reasons to persist when the going gets hard.
When you walk into a kindergarten classroom, there is a palpable energy of wonder that children have about a world that is so unfamiliar and exciting to them. We need to try and hold onto that wonder, that ability to think and ask questions with no care about whether the answers are obvious or not; that is how you start asking questions that are non-obvious.
Where do we go from here?
“Mathematics is a long conversation.”
- Barry Mazur
The profession of mathematics is changing. Some say it is the most exciting time to be a mathematician, armed with tools that can enable huge progress, others despair that the truest nature of the work is being eroded. Mathematicians are undergoing a philosophical reckoning, much like computer scientists (and many other professions) have already been embroiled in. The technical barrier to working on research mathematics—a field that is known to reward human intelligence—is dropping. It is exciting that more people than ever will be enabled to work on such problems, but this comes with the risk of the saturation of mathematics with undigested proofs.
There are thus many logistical infrastructure questions to answer. Given that I’m not a professional mathematician, I’ll reference Tao’s view from his talk: we need to separate human and automated workflows in math, and have many conversations on the question of credit.
What comes next for the LLMs? The frontier researchers at the symposium urged audience members to not write the models off if they make mistakes, but to be patient and earnest with them in the same way you are patient with students to allow them to live up to their potential. On this topic, Bubeck said:
“This is real, this is as real as it gets. So I think the best that could happen is everybody takes this seriously and starts to work with it, talk to it, start to find the limitation in their own field…What I’m trying to say is that people need to engage with this deeply.”
And for current students of mathematics?
From old friends to new at Stanford, I’m watching fellow math majors join AI labs as the line between industry and academia for math research is becoming fuzzier and fuzzier—and who can blame them? The question of “impact” hangs over our heads; I’ve talked to friends about how it almost feels selfish to commit to learning and practicing math in a PhD program for six years when the opportunity cost seems so high, and the profession as we knew it is changing before our eyes.
It’s a difficult feeling to describe, like a meandering walk suddenly turned into an organized race course with mile markers, a finish line, elite runners on AI-phaflys5, and spectators you’ve never seen before wildly cheering at the finish line. I’d maybe even join the race since it seems like fun, but what are we racing towards?
Wasn’t the entire point to walk around and see what you find?
Mathematics as a public practice
“Our mathematical ideas fit the world for the same reason that our lungs are suited to the atmosphere of this planet.”
- Reuben Hersh
Starting early in math education, there’s this idea that mathematics is just about finding the right answer, and perhaps trying to be the first one to do so. That is certainly a big part of it—which is why the current infrastructure around the profession rewards impressive results—but it is not all of it. It is important to know that you are not alone when you practice math. You are part of a community as old as humanity.
1982 Fields medalist Bill Thurston answered the question of “What’s a mathematician to do?” on MathOverflow in 2010 (which I came across in Bessis’ article):
“The product of mathematics is clarity and understanding. Not theorems, by themselves. Is there, for example any real reason that even such famous results as Fermat’s Last Theorem, or the Poincaré conjecture, really matter? Their real importance is not in their specific statements, but their role in challenging our understanding, presenting challenges that led to mathematical developments that increased our understanding.
I think of mathematics as having a large component of psychology, because of its strong dependence on human minds…In short, mathematics only exists in a living community of mathematicians that spreads understanding and breaths life into ideas both old and new. The real satisfaction from mathematics is in learning from others and sharing with others. All of us have clear understanding of a few things and murky concepts of many more. There is no way to run out of ideas in need of clarification. The question of who is the first person to ever set foot on some square meter of land is really secondary. Revolutionary change does matter, but revolutions are few, and they are not self-sustaining—they depend very heavily on the community of mathematicians.”
Thurston emphasizes the part of math research that often goes unseen: the frequent collaboration with others, the teaching and taking questions at a talk, the importance of the Mathematician Mentor that takes a student under their wing and gives them a problem to solve.
Mathematics is infinite, and if AI is being primed to do the technical heavy lifting, it seems like the human burden is shifting toward intuition and taste: choosing questions, building theories, and deciding what is worth understanding. And while there is no end to the ideas in need of clarification and translation, we must not reduce the role of mathematicians to being the translators of machine-generated proofs.
There is still importance in the human engagement and creation of mathematics. The point is that we don’t know everything, and we can’t know everything, but the things we choose to know and choose to care about matter.
A renewed consciousness?
Numerical analyst (and outstanding professor) Peter Henrici gave a talk “Reflections of a Teacher of Applied Mathematics“ at a symposium on The Future of Applied Mathematics in April of 1972. He ended it by saying:
Could it be that in mathematics, too, we need a new Consciousness? A Consciousness concentrating less narrowly on research progress in little isolated areas, on status symbols such as the publication of articles in reputable periodicals, on jealously guarded areas of teaching responsibility where no other man may enter; a new Consciousness stressing instead the exchange, communication, and experience of mathematical information, a Consciousness where mathematics is told in human words rather than in a maze of symbols, intelligible only to the initiated; a Consciousness where mathematics is experienced as an enlightening intellectual activity rather than as an almost fully automated logical robot, ardently performing simultaneously a large number of seemingly unrelated tasks?
Fifty-four years later, we are still grappling with the same questions at the same symposiums. How do we bring about this new consciousness in math which emphasizes the experience and practice of mathematics, starting from young students? The use of AI in math is inevitable, and we can play an active role in how it shapes problem-solving, how to deploy it, and use it productively and consciously.
We will figure out how to intentionally develop students’ mathematical intuition with norms on how to use these models responsibly, and hopefully reach even more students than ever in this age of accessibility. We can help each of them find their personal reason to practice math, while encouraging failure and building grit. We can ask them to test the boundaries of balancing on the shoulders of both mathematicians and machines.
The way things are progressing, it is inevitable that AI will autonomously produce theorems at a very high research level, likely within the next year. Benchmarks will be achieved at a faster rate than we can carefully curate them as massive amounts of compute are siphoned for these tasks.
So the question remains: why do math? Why do something when AI can increasingly do it faster, smarter, and better? (To that I say, there were already many people who can do math better than me pre-LLMs.)

Some do math for an unequivocal truth (regardless of whether the truth is ascertained by humans or AI), and many do it for the love of the game. I do math for the same reason I paint and write: because it is beautiful, it is difficult, it is joyful, it has made me cry more than once, and because I can find individuality in it, with a sense of precision and rigor. There are rules and guidelines for what constitutes good, but there is also freedom to create what constitutes great. Sometimes it feels as natural as breathing, and other times it feels like an itch under your skin that you can never quite reach. It is a practice, and not everything you do is worth showing to the public (or even keeping for yourself), but it is all part of the striving towards better.
All the mathematicians I know are wonderfully stubborn; you have to be, to keep returning to problems that have evaded you for days, months, sometimes even years. AI will undoubtedly change the job. But the precision of mathematics—and of mathematicians—is exactly what makes the field uniquely suited to take advantage of this transition without sacrificing its human nature.
I don’t know what math will look like in ten years, or even in the next year. But mathematics is still a long, ongoing conversation; just now, machines can also speak into it. They may even offer remarkable insights. But it is still our responsibility to decide how to listen. And what is mathematics if not listening closely and carefully—to patterns and structures, to each other and to the world?
A big thank you to Professor Steven Strogatz and Annmaria Antony for feedback on this article. I also want to sincerely thank Jared Lichtman, the Future of Mathematics Institute, Stanford HAI, SISL, and everyone else who was involved in putting together the Stanford Future of Mathematics symposium. A lot of the symposium came as new ideas to me (I finally learned what Lean was ~1 month ago), so I would love to hear any thoughts, disagreements, or addendums to what is written here. Any and all mistakes are my own. If you would like to hear straight from the giants themselves, there are recordings of both days here.
There is much more to be said about these topics, especially regarding math education. But given how long this article already is, I’ll delegate that to a future essay or, perhaps, leave it as an exercise for the reader. :)
The shoulders I stood upon while writing this article:
Alon et al. - Remarks on the Disproof of the Unit Distance Conjecture
Jeremy Avigad - Mathematicians in the Age of AI, Mathematical Understanding
David Bessis - The Fall of the Theorem Economy
David Litt - Mathematics in the Library of Babel
Peter Henrici - Reflections of a Teacher in Applied Mathematics
Reuben Hersh - Some Proposals for Reviving the Philosophy of Mathematics
Alex Kontorovich - Lecture: “Interactions of AI with Research Math and Formalization”
Karl J. Smith - Mathematical Reflections: Standing on the Shoulders of Giants
Terence Tao - Lecture: New Mathematical Workflows
Ravi Vakil - Roundtable at the Future of Mathematics
Maryna Viazovska - Lecture: Formalizing the sphere packing problem
Please let me know if you have any further recommendations!
And by “attended”, I mean that a friend helped me worm my way in as a volunteer…thanks Dean!
I entertained myself between talks by looking at where my professors chose to sit—just like I do in their huge auditorium classes every week—seeing who is a front row warrior and who would rather lay low in the back. I couldn’t help but notice that my optimization professor sat in the middle-ish of the 4th row, which is exactly where I currently sit in his class. Confirmed: optimal spot.
Erdos is always described with the word prolific.
On occasion, I too make mathematical claims with enough confidence on my math midterms to convince the TAs that I know what I’m doing (even if I have no clue how I would go about proving them). It has a pretty high success rate!
Not my best work, I’ll admit.








This was a really interesting article that made me think a lot. Still, I see this as an outsider, i.e., I am not a professional mathematician, just a person who'd like to teach himself math autodidactically up to at least an undergraduate level (fighting now with Bartle and Sherbert's Introduction to Real Analysis). As a learner, I can't quite seem to agree with "The implicit purpose of math homework is to learn how to be stuck, to test ideas, to learn how to fail and, better yet, learn from your mistakes and take them to the next problem. The implicit goal of math education and research is to build intuition and to learn how to ask questions. When we give problems to models, we’re testing them". I mean, I can see this is the objective for turning a math student in the long run into the type of professional mathematician who solves problems for a living. Getting stuck sucks really bad and is extremely antipedagogical, which I suspect is one of the reasons why so many people end up hating math: they find it a continuous torture of trying and failing to solve exercises and feeling stupid in the process. And even if they solve exercise x, the next ones are just another sisyphean slog. Myself, I am trying to do all the exercises in the book and just get angry at how I get stuck for more than half an hour with most of the exercises (maybe the book is not the most pedagogical, and/or analysis is just absurdly hard). And I am a case of someone that *just wants to learn*! I am not taking an exam or anything. But I still get demotivated by the slowness of progress. In this regard, we should be considering perhaps developing AIs into excellent 1 on 1 tutors to help us and combat the inevitable frustration, instead of just considering 'well, it's too tempting and cheating to get the AI and use it in some manner while dealing with exercise sets'.
Thank you for this wonderfully written piece. As an econ student, I feel that math plays an interesting role in the quantitative social sciences. Mathematical skills have long been one of the biggest barriers to entering the field, and so I have some hope of LLMs making econ research more accessible and perhaps more diverse. At the same time, my own thinking has been so thoroughly molded by the math classes (and subsequent late-night problem set sessions) I've had the joy of struggling through; I'm now so grateful to have taken some of those before the introduction of LLMs.
I see many parallels to my field more generally and am also nervous about the flood of AI-generated results, among other things. But this gave me a lot to think about and hope for—and captures so much of what makes learning great :)