Welcome to Memorandum Deep Dives. In this series, we go beyond the headlines to examine the decisions shaping our digital future. 🗞️
This week, an AI lab said it had done what mathematicians have chased for nearly a century: solved one of the seven Millennium Prize Problems, the kind of question that carries a $1M reward and a permanent line in math history. The announcement arrived fast and well-documented, backed by a 166-page paper and the claim that a machine, not a person, had finally gotten there.
The timing raised eyebrows before anyone had finished reading the proof. Word got out that two mathematicians working on a closely related problem had been circling the same answer for almost a year, and that news of their progress had reached the company just days before its own result appeared. What followed was a public argument over credit, timing, and how much one side owed the other.
Underneath the argument sits an older, stranger question: whether the machine actually solved the problem everyone thought was unsolved, or something that only looks like it on paper. The answer turns on a set of rules written 26 years ago, long before anyone imagined a company would send 10k AI agents after them. And it turns out this isn't even the first time a solver has met a prize's rules to the letter while leaving its judges unconvinced.

Bay Area Times is the free daily newsletter that gets business and tech leaders up to speed before their first meeting.
AI, startups, robotics, biotech, energy, and the wider innovation economy- each story paired with one clear visual so you see the point instead of digging for it.
Monday to Friday, three minutes, no noise. More than 250,000 founders, operators, and investors already read it.

Training AI on finance requires experts who can evaluate models on valuation, portfolio theory, and risk modeling—not just general annotators.
Athyna Intelligence delivers financial analysts, economists, and quant researchers from Latin America for RLHF, evals, and reasoning tasks.
Same US time zone. Vetted in days. 40–60% savings.
*This is sponsored content. See our partnership options here.

In the early 1700s, a ship out of sight of land had no reliable way to know its longitude, which is how far east or west it had sailed. To solve that problem, the British Parliament passed a law in 1714 offering £20k for a device or technique that could find a ship’s longitude to within about 30 nautical miles. The law did not say which method should prevail, but the astronomers whose views carried the most weight had a firm idea. The Astronomer Royal, Britain’s chief astronomer, and many scientists around him believed the answer lay in the moon’s movement, which a trained navigator could read from the sky.
John Harrison, a self-taught village carpenter, thought the answer lay in a clock instead. A navigator who carried the exact time of his home port could compare it with the local time at sea and work out how far the ship had traveled. Harrison therefore spent years building clocks that could keep Greenwich time on a long voyage. His fourth design, a large watch called H4, was tested on two voyages across the Atlantic. In 1765, the Board of Longitude, the panel judging the prize, confirmed it had kept time within the law’s strictest limits.
Harrison had met the terms of the law, yet the Board did not hand him the prize. It recommended paying only half at first and holding back the rest until other makers could build the same watch, a demand the Harrisons saw as the judges changing the rules. The long dispute that followed grew out of a gap between the law and its judges, because the law asked for any method that worked while the judges pictured an answer from the stars.
Nearly three centuries later, the same kind of dispute has broken out in mathematics, this time over a claim from an AI company. On September 8, 2026, OpenAI announced an internal model had resolved the Navier-Stokes problem, a famous open question about the equations that describe how fluids such as air and water move. The problem comes with official rules and a reward, since it is one of seven Millennium Prize Problems for which the Clay Mathematics Institute offers $1M each. Like Harrison, OpenAI reached an answer the rules allow, by a route that many experts on the problem did not have in mind.
Because the announcement came from a company that sells AI models, it quickly raised a blunter question about whether OpenAI had solved the problem or staged a win for publicity. The answer to that question turns on what OpenAI actually proved, which is narrower than the announcement suggests. The company showed that a smooth external push can cause a fluid described by these equations to break down, although it did not show that a fluid left entirely alone can do the same. Clay’s rules accept the first kind of answer, which undercuts the idea of a stunt, while the word ‘solved’ now depends on whether mathematicians agree that the rules capture what they wanted to know.
OpenAI backed its announcement with a 166-page paper setting out the proof and a version written for a computer to check, but the release did not arrive alone. The night before, Tristan Buckmaster, a mathematician at New York University, had posted a statement saying OpenAI took on the problem after learning of his unpublished work with his collaborator, Levent Alpöge. Both the claim and the accusation become easier to judge once it is clear which version of the problem OpenAI chose.
The Navier-Stokes equations describe how fluids such as air and water move, using the same basic idea as Newton’s laws of motion. The big unanswered question is whether these equations can ever produce a ‘blowup’, a moment when the fluid’s speed shoots to infinity in a finite amount of time. Real fluids never behave this way, so if the equations predicted a blowup, it would mean the equations themselves had broken down.
The mathematician Charles Fefferman wrote Clay’s official statement of the question, and it accepts four possible answers that fall into two groups. A cup of coffee is the easiest way to picture the difference between them. The first group asks for proof that a fluid with nothing disturbing it, like coffee left alone on a table, stays well-behaved forever and never breaks down on its own. The second group asks for a blowup and lets the solver stir the fluid with an outside force, such as a spoon, as long as the stirring itself is smooth and gentle.
OpenAI’s proof takes the second route, which mathematicians call a ‘forced blowup’, and it works backward. It starts with coffee that is perfectly still and decides in advance how that coffee should swirl, faster and faster without limit. The proof then works out exactly how the spoon must move to create that swirl. The hard part is making sure the spoon moves gently, stops after a short time, and stays within a small area, which OpenAI’s main theorem says it manages.
The result therefore answers a narrower question than the one many people have in mind. OpenAI showed that a carefully designed stir can make the equations break down, which Clay’s wording permits, but it did not show that coffee left alone can do the same. Scientific American reports that many experts picture the problem as an untouched cup, and that it is unclear whether OpenAI’s method still works once the spoon is removed.
The computer-checked version of the proof does not settle that question either. It is written in ‘Lean’, a programming language that lets a computer confirm every step of a proof, much as a calculator confirms every sum in a piece of homework. A calculator can tell you each sum is right, but it cannot tell you whether you answered the question the teacher set. Lean has the same limitation, since a person still has to confirm that OpenAI has provided the exact statement Clay asked for. So far, no independent group has publicly reported making that comparison. With the mathematics still waiting on that human judgment, the stunt question comes down to how and why OpenAI produced the result.

Cut through the hype. Roko’s Basilisk distills the day’s AI breakthroughs, market moves, and workable workflows into a five-minute brief.
Get one sharp deep dive, fast Quick Hits, and a witty pro tip, plus links that actually matter.
Built for builders, operators, and curious leaders who need context fast. Join 100k+ readers who start decisions here each morning.
*This is sponsored content

Before OpenAI turned to the problem, two human mathematicians were already working along the same route. Buckmaster says he and Alpöge spent about a year on the approach, and had computer-checked results on related equations by August 22. OpenAI moved far more quickly once rumors reached OpenAI on September 1 that two Millennium problems had been solved. It responded by setting 'agents', running copies of a model given tools and a task, to work on the remaining Millennium problems. Sam Altman, OpenAI's chief executive, later said the company had wanted to see whether its model could match rumored Anthropic results. About 88 hours later, a group of about 10k agents reached the Navier-Stokes result.
OpenAI says the point of the effort was to report its models' progress, and that purpose invites scrutiny because the company's past claims have not always held up. In October 2025, OpenAI staff said GPT-5 had solved open problems from a list named after the mathematician Paul Erdős, when the model had in fact found solutions that already existed. The Navier-Stokes announcement now uses the word 'resolved', while Clay has moved the problem to an Active status without certifying anything.
OpenAI’s handling of the two mathematicians has also drawn scrutiny, since Buckmaster says his dealings with the company became tense. He says that when he threatened to go public, Sébastien Bubeck, an OpenAI researcher, asked him, “Why would you ruin your career?” Bubeck called the allegations false and inflammatory, and he later apologized for the remark. Buckmaster used OpenAI’s coding agent, Codex, in his own work, but he stops short of accusing the company of using his data.
The timeline shows a company racing a rival and two outside mathematicians toward a famous result, with a public demonstration of its models’ ability as part of the aim. That supports suspicion about OpenAI’s motives, although it says nothing about whether the proof holds, which is where OpenAI’s defenders make their stand.
OpenAI’s defenders, including Bubeck and Altman, argue that the company answered the problem exactly as written and that the objections come down to taste and territory. Fefferman offered four options to give solvers room while, in his words, “retaining the heart of the problem,” and OpenAI’s force meets his conditions on its face. From that position, calling the force a loophole amounts to rewriting the problem after someone has solved it.
The way OpenAI released the result also sits awkwardly with the idea of a publicity exercise, since it came with a full paper and files anyone can check. That is the standard Buckmaster and Alpöge set for their own work. OpenAI has also credited the pair as first to their own forced result, declined the prize, and offered to share its prompts.
This defense is strong on the mathematics, although it weakens once attention shifts to how the result was produced. OpenAI concedes that rumors about outside work set its target, and Scientific American reports Bubeck acknowledging that the proof followed a method similar to the pair's. On September 11, Clay said the problem had apparently been settled and called its own process deliberately unhurried, which is an acknowledgment rather than an endorsement. Clay will consider a solution only after it appears in an approved journal, survives two years, and wins general acceptance.
Harrison’s story did not end when his watch passed its trials, because the Board of Longitude still wanted proof that H4 was more than a one-off. What the Board had never settled was whether a clock counted as the kind of method the prize was meant to reward, and mathematicians now face the same unsettled question about a designed force. They hold a proof that fits the wording of Clay’s problem, while many of them had pictured a fluid with nothing pushing it at all, and that version of the problem remains open. Their judgment will also take time, because Clay’s two-year clock does not start until OpenAI’s proof appears in a journal the institute approves.
Here are some ways.
Share today’s news with someone who would dig it. It really helps us to grow.
Let’s partner up. Looking for some ad inventory? Cool, we’ve got some.
Deeper integrations. If it’s some longer-form storytelling you are after, reply to this email, and we can get the ball rolling.

What did you think of today's memo? |