← All writing

ai-adoption · human-flourishing · point-of-view

The Proof Nobody Has Read

Share:LinkedInX (Twitter)

On 8 September an AI system was credited with resolving a Millennium Prize problem that had stood since 2000. Nothing changes for engineers tomorrow. What changed is how hard problems get solved, who gets the credit, and what we tell our children about mathematics.


Last week a 166-page mathematical proof was announced which, in Scott Aaronson's words, "probably hasn't yet been read and understood by any human". OpenAI's own Sebastien Bubeck puts the compute cost at several million dollars; Aaronson, from outside OpenAI, says "at least fifteen million". If it holds up, it answers a question about equations written down in the nineteenth century that has carried a million-dollar prize since 2000, and the prize is still unclaimed. OpenAI says it does not intend to claim it.

This story has a lot of names in it, so a cast list before we start.

  • The Clay Mathematics Institute, which in 2000 named seven Millennium Prize problems and put a million dollars on each. Navier-Stokes is one.
  • Charles Fefferman, Princeton, who wrote the official statement of the problem.
  • Diego Córdoba and Luis Martínez-Zoroa, the mathematicians whose years of work both AI efforts built on.
  • Tristan Buckmaster of NYU and Levent Alpöge of Anthropic, a private collaboration that reached a related result on the Euler equations.
  • OpenAI, whose model produced the Navier-Stokes proof, with Sebastien Bubeck speaking for them.
  • Quanta Magazine, the Simons Foundation's science magazine, which reported the story on 8 September.
  • Scott Aaronson, the computer scientist whose blog post is the source of the fifteen million dollar estimate.
  • Lean, the software that checks a proof line by line, and is why the proof can be called correct before anyone has read it.

I am not a mathematician. So my first question was the one I suspect most of you have: what does this actually change?

What changed, and what did not

The Navier-Stokes equations describe how fluids move. Air over a wing, blood through an artery, weather across the Atlantic. Engineers have used them for a century and a half, and the question that carried the prize was whether the equations always behave: whether a perfectly smooth flow can, in finite time, reach a point where the speeds in the fluid grow without limit and the equations stop describing anything.

Ten days ago the answer was unknown. Now, if the proof holds, the answer is yes, at least when the fluid is being pushed by a smooth external force. The official problem statement allows that version. Whether it breaks when left alone is still open for Navier-Stokes, though OpenAI's agents did settle it for the simpler Euler equations.

No aeroplane flies differently tomorrow. No forecast improves. Every practical use of these equations was already an approximation, because turbulence cannot be computed exactly at any useful scale, and it is still an approximation this morning. The result tells engineers where the mathematical model has limits. It does not hand them a better wing.

So the practical change is not in fluid dynamics. It is in method. Tristan Buckmaster puts it this way: the important thing is that a mathematician and a model can now do all this work in a month. OpenAI's final run on Navier-Stokes took days. That transfers. It transfers to every hard problem that has a human research programme behind it and a long, grinding, technical last mile in front of it.

Aaronson wrote this week about putting his children to bed and feeling in the pit of his stomach the question of what future they can have.

I have two daughters. I will come back to that.

How it was done, and what it cost

The story has three sets of people in it, and the order matters.

First, two mathematicians, Diego Córdoba and Luis Martínez-Zoroa, who spent years building a way of forcing these equations to break. When Charles Fefferman was asked about the result, he said the heroes of the story are those two. Buckmaster has said Martínez-Zoroa deserves a Fields Medal. Both AI efforts stood on that work.

Second, Buckmaster at NYU and Levent Alpöge, who works at Anthropic, collaborating privately over about a year with a mix of models from both major labs. Progress was slow for most of it. On 15 August they had blow-up for the Euler equations, the frictionless cousin of Navier-Stokes, under a smooth external force. Buckmaster's description of the first machine-generated proof his co-author sent him: "the most horrendous I have ever read". It was formally verified on 22 August. The month since has been spent, in his words, "working around the clock to understand this proof and turn it into something readable". He calls one of his own published write-ups "AI slop". Correct, machine-checked, and close to unreadable.

Third, OpenAI. Their own account, published with the proof, is more specific than the press coverage. A new internal model had been training since 28 August. On 1 September they heard a rumour that two Millennium problems had fallen and set the model on all of them at once, as groups of coordinating agents that could read a cached copy of the internet and run code. Nearly a hundred agents spent about fifty hours on the Euler equations and, to OpenAI's own surprise, resolved the unforced version. Resources were then moved to Navier-Stokes, where the group that found the proof was of the order of ten thousand agents running at once. They had it on 5 September, about eighty-eight hours after the first agents were launched, having exchanged 2.7 million messages and produced 130 billion tokens of output, roughly a hundred billion words, on that problem alone. A separate model then took seventeen hours to formalise it in Lean. Buckmaster was told the proof was about a hundred pages and has not seen it; Aaronson says 166 and doubts any human has yet read it. OpenAI says it began after hearing rumours of Buckmaster and Alpöge's progress and recognises their priority on forced Euler; the two sides' accounts of the calls that followed differ, and both are public.

The human work did not disappear. It moved. By his own account, Buckmaster's month after 15 August went on understanding, checking and rewriting, not proving. The machine can now produce a true thing faster than a person can hold it in their head, and the scarce skill has become the ability to take what is verified and make it legible to other people.

Now the cost.

Turning OpenAI's figures into carbon needs one bridge, and the honest one is the money, because nobody outside OpenAI knows what a token costs in electricity on a model that does not exist publicly. Every assumption is on show so you can argue with it.

$3 to $6 million of compute (Bubeck's "several") ÷ $2 to $3 per GPU-hour = 1 to 3 million GPU-hours

× about 1 kW each (chip, cooling, building) = 1 to 3 gigawatt-hours

× 300 to 400 g CO2 per kWh (a US grid) = 300 to 1,200 tonnes of CO2, call it a thousand, give or take

≈ two to six fully loaded 787s, London to New York and back (about 200 tonnes a trip) ≈ a hundred to four hundred UK homes heated and powered for a year

Two caveats. If Aaronson's fifteen million is nearer the truth, multiply by two and a half to five. And I have priced the hours at list rates; if Bubeck's figure is OpenAI's own internal cost, it bought roughly twice the hours, and the carbon doubles with them.

Is that a good trade? Two questions.

  1. Who would have paid millions of dollars to settle this? I doubt any university, research council or philanthropist could have. The Clay Institute offered one million and it has gone unclaimed since 2000. My reading is that the money was spent because this was a test of whether AI could do it, and the mathematics came along for the ride.

  2. How long would humans have taken without the machines? Honestly, we cannot know. The particular route to the prize, through a smooth external force, was one that Buckmaster says almost nobody else was attacking. The Córdoba and Martínez-Zoroa programme might have got there in five years, or fifteen, or never. Set against "possibly never", a thousand tonnes of carbon looks different. It also looks different from the other side: if the answer was going to arrive anyway in a decade, we paid a great deal to have it now and to have it in a form nobody has yet read.

I do not have a settled view. I do think the question should be asked out loud every time one of these announcements lands.

What it means, and what comes next

Back to my daughters.

I would give them the evidence, not a comfort. The heroes of the story, named by the man who wrote the problem, are two human beings who chose the problem, saw the route and spent years on it. The machines were sent down the route those two had opened. The mathematician who spent a month around the clock afterwards was understanding, checking and explaining. The job is still there, a level up, in setting direction and making meaning, and those are the parts that were always the most human.

That is not to pretend nothing has been lost. Aaronson writes about the worry that mathematicians are reduced to "verifiers and explicators of gargantuan arguments dumped into their laps", and he points to an open letter from twenty-five Fields Medallists, Terence Tao among them, insisting that human understanding remains the point of mathematics. I think they are right, and I think the week's events prove it in an unexpected way: the proof that nobody has read is worth very little until somebody does.

So how does this way of working become something we can hand on rather than a one-off spectacle? Buckmaster's own list is the right one: how we train students, how we assign credit, how we referee, and how we decide what deserves a human life's attention.

None of those are technical problems. All of them are foundations, and the pattern I have seen across technology is that when the capability arrives before the foundations, you get waste and sometimes harm. This week we got an argument over credit. Next time it may be worse.

And what should the next few million be spent on? I do not know what the people who use these equations every day would have asked for, and I suspect it was not this. The prize was set by mathematicians, for mathematicians. Before the next run, I would want someone to ask the engineers who design the aircraft, forecast the weather and model the blood flow a plain question: what is the problem in your work that has years of human effort behind it and a last mile nobody can afford to run? Spend it there.

In healthcare the shape already exists: decades of human science in antibiotic discovery, and a search space too large for people, where purpose-built models rather than chatbots have already found candidates on a fraction of this month's compute. What a Navier-Stokes-sized run would do pointed at that problem is the question I would want asked before the next Millennium problem gets its turn.

I do not know who makes the next leap on this. It could be an inspired human or an agentic swarm, and who am I to say.

The machines did not choose the problem. They did not open the route. They did not spend the years.

The AI ran the last mile. It still needed a person to say what it had found.

That is what I will tell my daughters, and I would rather it were evidence than comfort.