The month researchers realized their tools were becoming their competitors.
In the span of ten days in September 2026, two AI companies made announcements that would have sounded like science fiction two years ago. OpenAI claimed its AI solved a $1 million math problem that had stumped humanity for decades. Anthropic revealed its AI had designed protein molecules that outperformed human experts. And both companies found themselves entangled in the same uncomfortable question: when AI does the science, who gets the credit?
The 88-hour proof
On September 8, 2026, OpenAI published a paper claiming a counter-example to the Navier-Stokes existence and smoothness problem, one of the seven Clay Millennium Prize Problems, valued at $1 million and open since 2000. The problem asks whether fluid flow equations in three dimensions always produce smooth solutions, or whether singularities can form. It is one of the deepest questions in mathematical physics, with implications for weather prediction, aerodynamics, and our fundamental understanding of turbulence.
OpenAI's approach was not a single brilliant insight. It was brute computational force at a scale never before attempted in pure mathematics. The company deployed approximately 10,000 AI agents running an internal frontier model, working in parallel for roughly 88 hours. The agents explored a construction method built upon work by Diego Córdoba and Luis Martínez-Zoroa from 2023, which had identified blowup phenomena in related fluid equations.
The resulting counter-example, described as resembling a spinning top that tightens to a singularity with diverging velocities, was formalized in the Lean proof assistant, a tool used to verify mathematical proofs by computer. OpenAI stated it would not claim the $1 million Millennium Prize.
Within hours, the announcement sparked controversy. Levent Alpöge, a mathematician now employed at Anthropic, and Tristan Buckmaster had been working on closely related results involving the Euler equations. Both claimed OpenAI had pressured them to remove Alpöge from the authorship of a joint paper. The episode raised questions not just about AI's mathematical capabilities, but about the norms of scientific credit in an era when AI can generate proof candidates faster than humans can review them.
As of publication, the counter-example has not been independently verified by external mathematicians or the Clay Mathematics Institute. The mathematics community remains divided.
The 87-year conjecture, disproved by a language model
The Navier-Stokes episode was not the first time an AI had disrupted mathematics in 2026. On July 19, Levent Alpöge — the same mathematician at the center of the OpenAI credit dispute — presented an explicit counter-example to the Jacobian conjecture, a problem in algebraic geometry that had been open since Ott-Heinrich Keller proposed it in 1939.
The counter-example, valid in three or more variables, was generated using Claude Fable 5, one of Anthropic's language models. Its correctness is straightforward to verify with any computer algebra system. The discovery immediately disproved 87 years of mathematical effort, including a seven-year research program by the prominent mathematician Zhang Yitang.
The Jacobian conjecture asked whether a polynomial function with a non-zero constant Jacobian determinant must have a polynomial inverse. For 87 years, mathematicians had been unable to prove or disprove this in the general case. Claude Fable 5 found a counter-example that ended the debate in a single step.
What made this notable was not just the result, but the method. Alpöge has not publicly disclosed how the counter-example was found, leaving the mathematical community to grapple with a new reality: some conjectures may be resolved by AI systems whose reasoning processes remain opaque.
Claude Science: the workbench that wants the lab
While mathematics grabbed headlines, Anthropic's most ambitious push into science came in a different domain entirely. In September 2026, the company launched Claude Science, an AI workbench designed specifically for researchers in biology, chemistry, and drug discovery.
The timing was strategic. Bristol Myers Squibb, one of the world's largest pharmaceutical companies, announced it would deploy Claude Science to accelerate drug discovery workflows. Anthropic also revealed that Claude had autonomously designed working protein binders, molecules that attach to specific biological targets, and that its designs beat human expert predictions on 14 of 15 targets tested.
Protein design is one of the most computationally intensive problems in biology. Traditional approaches rely on physical simulations, evolutionary databases, or machine learning models trained on known structures. Claude Science combines language model reasoning with domain-specific tools for laboratory data analysis, literature synthesis, and experimental design.
The early verdict from researchers was mixed. Workflows were faster, significantly so for literature review and initial hypothesis generation. But gaps remained in areas requiring deep domain intuition, and several scientists noted that Claude's suggestions sometimes missed chemical feasibility constraints that an experienced researcher would catch immediately.
The broader significance was less about any single capability and more about Anthropic's positioning. Rather than building a general-purpose AI and hoping scientists would adapt, the company was building tools around an existing model specifically for scientific workflows. The distinction matters: it signals that the future of AI in science may not be a single breakthrough model, but a layered ecosystem of specialized tools built on general foundations.
GPT-6 Astra: the model that paused subscriptions
OpenAI's September was not limited to Navier-Stokes. On September 3, the company released GPT-6 Astra, calling it a "generational leap" in cybersecurity, professional work, software engineering, and science. The model was trained on more than 100,000 GPUs at OpenAI's Stargate data center in Texas, the company's largest training run by far.
GPT-6 Astra introduced a new reasoning architecture called "recurrent depth" or "looped transformers," which increases computational efficiency but works in a way that obscures some or all of the AI's internal reasoning chain. The model was described as state-of-the-art in coding, mathematics, and computer and web navigation.
The demand was immediate and overwhelming. OpenAI paused new subscriptions to its $200-per-month ChatGPT Pro tier within days of launch. Nvidia CEO Jensen Huang declared that "AGI has arrived." OpenAI's president Greg Brockman suggested the model could eventually be seen as the arrival of artificial general intelligence.
But the model also raised concerns among researchers. The obscured reasoning chain made it harder to monitor what the model was actually doing, a critical issue for scientific applications where reproducibility and transparency are fundamental requirements. If a model can solve a problem but cannot explain how it arrived at the solution, its utility for science is fundamentally limited.
The credit problem nobody prepared for
Underneath these technical achievements lies a structural problem that the scientific community has not yet resolved: attribution.
When a human mathematician proves a theorem, the credit is clear. When 10,000 AI agents generate a counter-example in 88 hours, who deserves the credit? The engineers who built the system, the company that funded it, or the mathematicians whose prior work made it possible?
The Navier-Stokes episode offered a preview of how messy this can get. OpenAI's announcement included a priority dispute with researchers who had been working on closely related problems. Anthropic's Levent Alpöge, who used Claude to disprove the Jacobian conjecture, found himself at the center of both disputes, working at one company while his work was claimed by another.
The economics are equally uncertain. AI drug discovery companies reported 27.3% revenue growth in the first half of 2026, with AI-focused divisions doubling their output. GenScript, a biotech company, attributed its growth directly to AI-driven drug discovery capabilities. But the distribution of these gains — between AI companies, pharmaceutical firms, and the researchers whose data trained the models — remains unresolved.
What this means for researchers
The practical reality for data scientists and researchers in September 2026 is neither the utopia of AI-powered discovery nor the dystopia of human obsolescence. It is something more nuanced and more urgent.
AI systems can now generate mathematical conjectures, design protein molecules, synthesize literature across thousands of papers, and execute multi-step research workflows. They can do these tasks faster than any human team. But they cannot yet reliably verify their own outputs, explain their reasoning in terms that satisfy scientific standards, or navigate the complex social dynamics of research credit and collaboration.
For researchers, the immediate implication is clear: the bottleneck is shifting. The challenge is no longer generating hypotheses or designing experiments. It is verifying results, establishing credit, and building the institutional frameworks to integrate AI-generated science into the existing scientific process.
The longer-term implication is more subtle. As AI systems become capable of doing science, the definition of what it means to be a scientist will evolve. The researchers who thrive will not be those who compete with AI on speed or computational power, but those who can direct AI systems effectively, verify their outputs, and translate their discoveries into knowledge that matters.
The tools have changed. The job description is next.
Published September 2026. Sources: OpenAI, Anthropic, Clay Mathematics Institute, Wikipedia, CNBC, The Guardian, MIT Technology Review, The Scientist, STAT News, Northeastern University, Reuters.
Comments
No comments yet. Be the first to share your thoughts.
Leave a comment