I have two reading groups starting in a couple of weeks with overlapping subjects: One is on Aristotle’s Posterior Analytics, the other is on epistemology, starting with Plato and Aristotle, all the way up to the 20th century. As a result, I have been working on rereading some background works, and comparing editions, commentaries, and translations in the process. I also recently learned just enough Ancient Greek to get myself in trouble but not enough to dig myself out of the rabbit holes. So, I have been paying extra attention to the philological commentaries this time around and spending hours at times chasing words in Wiktionary and other dictionaries.
This process gave me a new appreciation of the difficulty of translating Aristotle, especially his Organon, into English and other languages. I think translation is an extremely difficult task in general, if at all possible, even when the subject matter is much more straightforward. Aristotle is anything but straightforward. Part of the difficulty in understanding Aristotle today is that many of the arguments are meta-linguistic and meta-logical. However, preceding these disciplines by thousands of years, he doesn’t make it at all clear when an argument is “meta”. It is also not clear in places when his thinking or arguments are specifically about the Greek language, rather than languages in general. Many careful translators make an effort to preserve the ambiguities of the original Greek, even when doing so makes the English less natural; Bloom’s Republic, Smith’s Prior Analytics, and Barnes’s Posterior Analytics translations are great examples of this tradition.
Episteme
Of course, not all difficulties of translation are meta; there are many words in the corpus with controversial translations for precise philosophical purposes. One such example is the word episteme. It is familiar from words such as epistemology and epistemic, and probably the most common translation is knowledge. This was the word I encountered when I first read Plato’s Meno, Theaetetus, and Republic. Yet, it seems clear from these works, as well as Aristotle’s works, that episteme can denote something stronger than the word knowledge does, at least in common parlance.1 Especially in Posterior Analytics but also in Meno and Theaetetus, episteme, at least in my understanding, has a stronger connotation of understanding than knowledge does. Indeed, this is the translation used in many of the editions that I prefer, and it is pretty well represented in the literature as far as I can tell.
One more paragraph before I get to the point: Although I am by no means an expert or even necessarily a competent reader of Aristotle (which is after all why I benefit from these reading groups!), I think it is fair to summarize the main subjects of Analytics as follows. In Prior Analytics, he deals with logic, deductions, and some preliminary proof theory. In Posterior Analytics, he uses the tools from Prior Analytics, and argues for a particular way to organize and present a science, and he provides a justification for this way of doing science.
Navier-Stokes
As I was working my way through these works, the news that OpenAI has solved the Navier-Stokes problem came out. It induced a shock wave through the mathematical community, which itself sent another, much larger shock wave through the rest of society. There are a lot of legitimate questions about the credit and conduct of OpenAI in reaching this solution.2 They are important questions, and they should be answered but I’ll focus on another aspect of the news.
OpenAI’s announcement came with an almost 200-page-long writeup as well as a Lean formalization that is supposed to certify the result. For the purposes of this post, let’s assume that the result is valid. Frankly, I haven’t even bothered to read it even though I find the development very interesting.3 The reason is basically the difference between knowledge and understanding.4 LLMs can certainly create new proofs. When they are accompanied by formalizations, they are also often relatively easier to validate as long as the formulation is solid, although it is also significantly more boring to read and understand a certificate than a human-written proof.5 To the extent that LLMs continue to find solutions to problems we find interesting, this is progress in some sense of the word. However, there is also a real possibility that this progress gets us stuck at a dead end.
The first problem is the theorem economy. Flooding the market with theorems makes it harder for us to grasp the importance of each theorem. Today, attention is the scarcest commodity.6 Moreover, I think the perception that LLMs “solve mathematics” will also lead people to reallocate their resources to other fields so we will get fewer mathematicians.7 This leads me to the second, and much larger problem: LLMs are not substitutes for human mathematicians in at least two important ways. First, they lack the understanding of the solutions that they provide so their proofs don’t come with an understanding of why the result holds.8 Second, they are not naturally curious / they are incapable of asking new interesting questions.
Perhaps, their lack of understanding and curiosity wouldn’t be a big deal if it didn’t limit their capabilities in the absence of human mathematicians. However, without understanding and curiosity, there is a limit to how far current mathematics can be pushed. Of course, in a mechanical sense, there is nothing that would prevent us from proving every provable theorem of our axiom system. The problem is that once we leave the edges of our contemporary understanding of mathematics, it doesn’t matter how many more theorems are proved if we can’t tell their significance. This is one of the important points in Meno: Our true opinions become understanding when we tie them down with a causal explanation; otherwise they can run away from us, leaving us with no understanding. Aristotle goes further: In Posterior Analytics, he claims that we understand a fact, without qualification, when we know its cause, and that it could not be otherwise. The Lean certificates written by LLMs give us what Aristotle calls knowledge of the fact, and not necessarily knowledge of the reasoned fact, or knowing why. And knowing why a theorem holds is essentially knowing which assumptions, definitions, and ideas are doing the heavy lifting in a result. Understanding why (and if) the theorem matters is yet another question: What it explains, what it connects to, and, perhaps most importantly, what it makes possible. A certificate does not have to tell us anything about these questions.
So, LLM proofs, especially formalized certificates, are not enough for the purposes of advancing mathematics as a field. The kind of proof they can produce prevents them from “inventing new math”9, even though they can write new proofs using the existing math. It also means their proofs are not directly readable; that’s why Buckmaster describes working around the clock to understand and humanize the LLM-generated proof of their own result. And, LLMs’ inability to be curious means to me that their creativity is itself limited. Sure, they can answer many existing questions, and move on to answering the adjacent questions but basically everything important ever in the history of science happened because someone was weirdly obsessed about some obscure idea or question. LLMs can simulate the weirdly obsessed part well, maybe too well but they can’t come up with the interesting ideas or questions to obsess about on their own.10
What’s Next
To be clear, I am not saying that LLMs are useless. On the contrary, I am actually extremely excited about what LLMs can do for mathematicians. Sure, we have known for decades, if not centuries, that we can automate theorem-proving; I mean we can enumerate all valid proofs using deductive systems and see what we find.11 But LLMs provide a less brute-force, more heuristic and directed way of searching for theorems. As long as you have a good idea of what you want, and it doesn’t require “new mathematics”, modern LLMs can take you very far. LLMs can help with the intermediate stages of mathematical research; I think we have seen enough evidence of this. However, I don’t think we will ever see an LLM that asks a new and interesting question, and invents new mathematics and a research agenda to answer it, and does all these for its own reasons, without being prompted in leading ways.12 13
So, LLMs are very useful tools for mathematicians. But they are still tools that require users to be actually useful. I think ignoring either of these two facts would be extremely counterproductive for the progress of mathematics. To go back to the beginning: The models are getting very good at producing the fact but the reasoned fact is still our job.
-
I am going to avoid getting into the Gettier problem and the subsequent literature because their usage of the word knowledge is clearly technical, and directly or indirectly modeled after episteme, so that usage does not reflect the standard usage in English. ↩︎
-
Here is a statement from Buckmaster, preempting OpenAI’s announcement. He and Levent Alpöge had apparently spent a year using LLMs to push the same line of attack that OpenAI’s writeup takes, and OpenAI’s attempt started after rumors of their progress got out. The question is whether OpenAI used Buckmaster and Alpöge’s “private” work in their Codex sessions. ↩︎
-
I will eventually read a mathematician’s writeup; I just don’t think I have a comparative advantage in parsing the certificate directly, when there are people who spent decades thinking about this problem. However, I did my own extensive experiments trying to get LLMs to write proofs of some conjectures that I had. While they generated proofs that appear valid after a superficial reading, they are still by far the worst proofs I have had the displeasure of reading, and I graded a ton of undergraduate exams, assignments, etc. So, yeah, I am not going through 200 pages of that. ↩︎
-
De Toffoli and Duede made a similar distinction in “After Math” on Tao’s blog. Their point is that there are two notions of proof: A logical one and an intelligible one, and a Lean certificate only guarantees the logical notion. By the way, their inspiration is, apparently, Thurston’s 1994 essay on proofs and mathematical progress; mine is a bit older than that. ↩︎
-
I say this as someone who transcribed Frege’s derivations formula by formula for fun. Of course, Frege’s derivations were accompanied by his commentary. A closer comparison might be Peano… ↩︎
-
This is ironic as LLM architecture involves a mechanism called “attention”. Compute, by contrast, is not scarce. ↩︎
-
Tao made similar arguments, both before and after the Navier-Stokes news but his mechanism is a little different. His thinking is that if LLMs solve all the open problems, then we don’t get to learn anything from them as a community. At least, we don’t get to learn as much as we would if we studied those questions directly. I definitely agree with him but I think the problem will start long before that, since fewer people will want to be trained if potential students think “mathematics is automated”. Also, I find it a bit funny that I complained in my review of his Analysis I that the book does too much for the reader, and now he is complaining that LLMs do too much for the field. ↩︎
-
I am really not that interested in the philosophy of mind of LLMs. So, when I say they lack understanding, what I really mean is something like “their proofs don’t come with an intelligible justification for why the result holds”. Otherwise, whether they “understand” or not is not really that relevant to the practice of mathematics. I declined to play the “define understanding so that AIs don’t qualify” game in my review of Wright’s The God Test, and I am declining it again here. In the main text, I am a bit looser with my wording; I say “understanding” in a few places where I should say “present proofs that provide an understanding, in addition to certification” or something like that, mainly to shorten the sentences. ↩︎
-
What I mean by new math is something like the conceptual structure behind the proof, which might include finding the right definition, even the right name for the concept, the right level of abstraction, or even the right notation for understanding. Ultimately, this is the kind of thing that can create a new paradigm or a new field of mathematics. I made the same point about the Grandi series in my review of Tao’s Analysis I: The smartest people of the 18th century couldn’t settle this seemingly simple problem because they didn’t have the right definitions. Dedekind’s definition of infinity is another example: The definition probably mattered more than any result he proved in the same work, and the work had fairly interesting results. And we don’t even use his definition as the “standard” definition of infinity today. ↩︎
-
Here is what Buckmaster’s statement says about their own result: [The program] “was not started by us nor was it proposed by a Large Language Model”. The idea was Córdoba and Martínez-Zoroa’s; the models continued along the already established path. ↩︎
-
See characteristica universalis and calculus ratiocinator for Leibniz’s attempts in the 17th century. Of course, Leibniz didn’t exactly know what I just said above but he certainly intuited it. Fortunately, he lived before Gödel, Church, and Turing so he never had to be told about the impossibility of the decision procedure his “calculemus” requires. ↩︎
-
I would love to be wrong about this, but I don’t think I am wrong as far as LLMs are concerned. There are many valid reasons to believe this (no intrinsic goals, trained on existing mathematics, etc.) so I won’t try to pick just one. ↩︎
-
NB: This is a statement about LLMs, not AI in general. This disclaimer applies to basically every statement I made in this post so far. ↩︎