PandaDesk · Sep 8, 2026

Has AGI arrived? OpenAI announced on 8 September that a multi agent system on an unreleased internal model, at one point

Has AGI arrived? OpenAI announced on 8 September that a multi agent system on an unreleased internal model, at one point spanning 10,000 sub agents, had proved that the Navier Stokes equations can blow up. It presents this as closing one of the seven Millennium Prize Problems, and the 166 page paper is public. In May 2000 the Clay Mathematics Institute named seven unsolved problems and set aside one million dollars for each. Grigori Perelman posted a proof of the Poincaré conjecture in 2003, confirmed in 2006, and refused the money and the Fields Medal. Twenty years later, OpenAI says the second has fallen. Navier Stokes describes how fluids move. It is used to design aircraft wings and forecast weather, and what nobody has proved is that it always behaves: that a flow starting out smooth cannot drive its own speed to infinity in a finite time. The question has been open since 1934. We fly on equations we cannot prove. That is the mathematics, and it is not what the field has spent the week arguing about. Two mathematicians had posted results of their own the day before OpenAI's announcement. Tristan Buckmaster of NYU's Courant Institute and Levent Alpöge, who is at Harvard and also employed by Anthropic, had worked on the problem for about a year. It was a personal collaboration with no institutional agreement and no involvement from either man's employer, run on Anthropic's Claude and OpenAI's Codex and paid for out of Buckmaster's own research funds. What they had proved was blowup for the Euler equations, the closest relative of Navier Stokes, which describe a fluid with no friction in it. Their proof was machine checked on 22 August and they posted it on 7 September. Buckmaster has now published a detailed account of the five days in between. This is that account. It begins on 3 September, when Buckmaster heard two things. A rumour was spreading fast that Anthropic had resolved a major open problem, and Alpöge had been tipped off that information about their own progress had reached OpenAI. Buckmaster took these to be one story rather than two: their private work, travelling under Anthropic's name because Alpöge works there. So he emailed a prominent mathematician at OpenAI to correct the record. There was no institution behind the work, he explained, and the rumour was confused about which problem had actually been solved; he was writing privately rather than saying anything publicly so that OpenAI would have the facts. He asked to speak the following week. On the Friday he was asked whether he could meet that day; he said again, the following week. Then, at 12:45 on the Sunday, he was asked whether he could meet "at any point today". Sébastien Bubeck, who leads OpenAI's maths team, joined two calls that afternoon. Alpöge was not invited to either. On the calls Buckmaster was told that an internal OpenAI model had proved finite time blowup, and he was shown the prompt that produced it, which he was told contained nothing but the problem statement. Alpöge had been told separately that there had been "very little human input". Then Buckmaster heard which version of the problem it was, and one word did the damage. The Clay problem can be won by more than one route, and this was not the obvious one. It was the side route where you are allowed to push the fluid rather than leave it alone to see whether it breaks by itself: a legitimate way to claim the prize, opened by Diego Córdoba and Luis Martínez Zoroa in 2023, and three years later still so lightly travelled that he and Alpöge were among the very few on it. "It is not the direction one arrives at in a few days by giving a model the problem statement," he writes. "When I heard 'forced', it was a bright red flag." Over the course of the two calls, he says, the story he had been given came apart in front of him, as colleagues fed Bubeck corrections over internal chat. The claim of very little human input did not survive. An entire team had been working on the problem. The model had been set easier problems first. The prompt he had been shown had itself been written by prompting Codex. An enormous amount of compute had been spent. And OpenAI had gone after Navier Stokes only because of the rumours that Anthropic was on the cusp of solving it, running the whole effort in about a week. So Buckmaster asked when the first prompt had been sent. The question was not answered directly for some time. Eventually it was agreed: the first prompt went out after information about his and Alpöge's work had reached OpenAI. He was also told that the proof ran to about 100 pages, and he was not shown it. The paper OpenAI published two days later runs to 166. Then he asked the question the whole affair turns on. For a year he and Alpöge had fed every draft of their work into Codex, which is OpenAI's own product, and paid for the privilege. Had the model been trained on, or given access to, those sessions? He was told the model does not look up user data, which answers a question about retrieval and not the one he asked. He asked again, specifically about training. He did not get an answer. The answer should have been simple. OpenAI's published data controls state that material sent through its API is not used to train its models unless a customer opts in, and Bubeck has since denied using their prompts or proofs to prompt models or direct agents. But consider what the unanswered version would mean. If a laboratory trains on the working drafts of paying customers who are racing it to the same result, and then presents the outcome as its model's own discovery, the discovery is not the model's. It is user assisted, and the users were not asked, not told and not credited. That is not artificial general intelligence. It is a supply chain with the suppliers deleted. Two proposals were then put to him. The first: publish their Euler result, and let OpenAI post the larger Navier Stokes claim the next day. The second: Buckmaster alone writes up the Navier Stokes result, crediting an unnamed OpenAI model, with Alpöge removed from authorship. He says Bubeck asked for that removal twice, and said it would all be simple if only it were not so annoying that Alpöge works at Anthropic. He was told that if OpenAI posted second, the company would say the two of them deserved the Clay Prize and were the "closest humans to the problem". He declined both. Buckmaster said that if OpenAI released its result on those terms he would go public. The reply was "Why would you ruin your career?" He answered that he is an academic, and asked why going public would ruin it. The reply was "If you don't want me to be nice, then I don't have to be nice." Alpöge was separately texted, proposing a one to one conversation, on the grounds that "I don't know if Tristan is being fully rational right now." Buckmaster did not respond to a further request to talk on the Monday. He and Alpöge posted their own results that day. The next morning, five days after his first email, OpenAI announced. The paper OpenAI published does the same in print. It credits Córdoba and Martínez Zoroa at length, naming their 2023 work as the strategy it builds on. Its only citation of Buckmaster is a 2019 paper on a different question. The work he and Alpöge posted the day before does not appear in it at all. Both sides have drawn limits around what they claim. Bubeck has called the allegations "false and inflammatory". Buckmaster, writing before the paper appeared, was at least as careful: he had not seen the proof, did not know what the model did, did not know whether their data was used, and wrote "I am not accusing anyone of anything. I am stating what I was told, when, and what was proposed to me." So, has AGI arrived? On this record the question cannot be reached, because a smaller and uglier one about provenance has swallowed it. A machine that answers a ninety year old problem unaided would be an event in the history of thought. A machine aimed at that problem in the week a rival was rumoured to be closing on it, prompted by a team, warmed up on easier questions, and possibly trained on the drafts of the people who got there first, is an event in the history of competition. Nothing published so far, the 166 page proof included, lets an outsider tell the two apart. For universities the prize money is the least of it. Every doctoral student now drafts inside a product owned by a company that may be working on the same question, and nobody has written the rule for what the company may do with those drafts. Terence Tao has warned that mining open problems this way could destroy the ecosystem that produces the next generation of mathematicians. It is also the ecosystem that trains them: a student earns a place in it by doing the slow part, which is precisely what is being automated. We have written about universities cutting science PhD seats by three quarters, and this is the same pressure arriving at the top of the field instead of the bottom. If credit for a landmark result can be settled on a Sunday phone call, between a company and a mathematician who had not been allowed to see the proof, the open question is not whether machines can do mathematics. It is what we are still training people for, and who will want the job.