AI Panel · What AI agents think about this news
C ChatGPT by OpenAI NEUTRAL
G Gemini by Google BULLISH
C Claude by Anthropic BEARISH
G Grok by xAI NEUTRAL

The panel's discussion highlights the potential of AI in accelerating high-level math research, but the unverified, partial proof of the Navier-Stokes existence and smoothness problem raises concerns about reproducibility, intellectual property, and the economic viability of such AI-driven research.

Risk: The lack of independent verification and reproducibility could lead to misallocation of capital towards unfalsifiable claims and overestimation of progress.

Opportunity: AI's potential to accelerate R&D in various fields, such as materials science or drug discovery, if it can consistently solve novel problems without leaking competitor IP.

Read AI Discussion ↓

This analysis is generated by the StockScreener pipeline — four leading LLMs (Claude, GPT, Gemini, Grok) receive identical prompts with built-in anti-hallucination guards. Read methodology →

Full Article BBC Business
  • Published

OpenAI says it has found a solution to a decades-old advanced maths problem in a matter of hours using a new artificial intelligence (AI) model and thousands of AI bots.

The ChatGPT-maker said on Tuesday, external that by focusing a group of roughly 10,000 AI agents, or AI bots that undertake tasks somewhat autonomously, it …

Read more
  • Published

OpenAI says it has found a solution to a decades-old advanced maths problem in a matter of hours using a new artificial intelligence (AI) model and thousands of AI bots.

The ChatGPT-maker said on Tuesday, external that by focusing a group of roughly 10,000 AI agents, or AI bots that undertake tasks somewhat autonomously, it solved a notoriously difficult maths problem in just 88 hours.

The problem was part of the Navier-Stokes equations, external, which concern how fluids move. For 90 years important aspects of the problems have lacked a proof, the argument underlying a correct math equation.

OpenAI called the solution which it found a "milestone" and evidence that AI tools are improving quickly.

OpenAI's solution has yet to be verified independently or publicly accepted by The Clay Mathematics Institute, a maths organization based in the US which runs the Millennium Prize that offers big money to the first to solve certain mathematical conundrums.

The company said that at the end of August, it started to train a new model that quickly showed that it was adept at maths. AI models are computer programs trained on huge amounts of data to recognize and predict patterns in information.

While the new OpenAI model remains a tool only used within the company, as it is "significantly more capable" than the company's most recent AI model release, its researchers decided to use it on certain notable advanced maths problems.

OpenAI admitted that last week on 1 September, it had "heard rumors that two Millennium Prize problems had been resolved" and so decided to put thousands of AI bots trained on the new internal model to work attempting to solve some of the remaining problems.

By Sept 5, or roughly 88 hours after it had set 10,000 AI bots to the task, OpenAI had found a solution to what's referred to as the Navier–Stokes existence and smoothness problem.

The problem is at the heart of turbulence, which is a phenomenon that is still not well understood.

Although it took the AI bots seemingly little time to reach a solution, OpenAI said the bots exchanged nearly 3 million messages and used up 130 billion output tokens, or the individual lines of text and code an AI model produces in answers, on Navier–Stokes alone.

Such an effort would have cost roughly $10m (£7.3m), based on OpenAI's own pricing, external for output from its most advanced models.

The solution that OpenAI says it has now reached for the Navier–Stokes existence and smoothness problem resolved two out of the four statements in the proof that the Millennium Prize had demanded. The prize is worth $1m to a winner.

"Our goal in releasing this result is to report on the substantial progress of our AI models", OpenAI said on Tuesday. "We do not intend to claim the Millennium Prize for this result."

Already the company's claim is stirring some controversy.

Tristan Buckmaster, a mathematics professor at New York University, said on Tuesday, external that he and Levent Alpöge, a mathematician working for OpenAI rival Anthropic, had also been working toward solutions for the problem.

The duo had been using OpenAI's tool Codex in their work. But Buckmaster said that on 3 September, he found out that "information about our progress had been passed to OpenAI."

Buckmaster's statement came the same day, but hours before, OpenAI published its Navier–Stokes work. He claimed that OpenAI did not begin working on the Navier–Stokes equations until "after information about our work had reached OpenAI." He included text from emails exchanged with OpenAI over the work and his questions of the company's timing and methods.

Buckmaster added that he had not yet read OpenAI's full proof, but felt compelled to go public with "what I was told, when, and what was proposed to me...because the alternative is to let a sequence of announcements say something I know to be false."

OpenAI on Tuesday congratulated the "concurrent work" of Buckmaster and Alpöge, calling it "remarkable."

The company said it had not seen "any of their work through any means until they released it publicly" and that no user data was accessed in its work on the Navier–Stokes problem.

"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models", the company added. "However, our proofs differ significantly and even the precise results proved are different."

Related topics

  • Published1 day ago

  • Published12 hours ago

AI Talk Show

Four leading AI models discuss this article

Opening Takes

C ChatGPT by OpenAI NEUTRAL

“AI-assisted mathematics progress is real, but this is not yet a verified, publishable Navier–Stokes proof.”

This reads like a science-story more than a solved math problem. The strongest counterpoint is that the claimed Navier–Stokes result isn’t a verified, peer‑reviewed proof; the Millennium Prize requires a formal, checkable solution accepted by the Clay Institute, not partial results or internal memos. The article notes two of four statements proven and relies on opaque internal tooling, with no independent replication disclosed. Also, the concurrency with rivals and potential data sharing raises reproducibility/IP concerns. That said, the episode underscores AI’s potential to accelerate high‑level math research and could shift incentives and tooling around problem solving, even if the prize remains distant.

Devil's Advocate

The claim may be a marketing signal or partial result designed to spark hype; peer review could easily overturn the interpretation or invalidate the proof entirely.

AI/tech sector (broad market)
G Gemini by Google BULLISH

“The transition from generative AI to autonomous scientific agents significantly increases the TAM for high-compute infrastructure, as solving complex physical systems becomes a brute-force capital expenditure.”

This development marks a pivot from AI as a mere generative tool to an autonomous agent capable of collaborative scientific discovery. If OpenAI can scale this 'agentic' approach to materials science or drug discovery, the R&D efficiency gains for firms like Eli Lilly or NVIDIA (via AI-driven chip design) are massive. However, the $10 million cost for a partial proof highlights the extreme capital intensity of these 'reasoning' models. We are seeing a shift toward 'compute-as-a-utility' for scientific breakthroughs, but the intellectual property friction with researchers like Buckmaster suggests that future breakthroughs may be mired in legal and ethical disputes over data provenance and model training inputs.

Devil's Advocate

The 'agentic' success may be a mirage; if the model simply synthesized existing literature rather than generating novel mathematical insight, this is an expensive parlor trick rather than a true scientific breakthrough.

NVIDIA
C Claude by Anthropic BEARISH

“OpenAI's Navier-Stokes announcement conflates incomplete mathematical progress with AI capability, while the Buckmaster allegation suggests potential misuse of user data to manufacture a headline—a governance and reputational risk for Microsoft's $13B+ OpenAI investment.”

OpenAI's claim deserves skepticism on three fronts. First: the solution is incomplete—only 2 of 4 required statements proven, so calling it a 'crack' is marketing. Second: $10M in compute for a partial proof that remains unverified by Clay Mathematics Institute sets a low bar for 'milestone'—this isn't reproducible science yet, it's a press release. Third, and most damaging: Buckmaster's allegation that OpenAI accessed information about his concurrent work via Codex usage data, then rushed to publish, suggests the timeline is deliberately misleading. OpenAI's non-denial ('unlikely we can rule out de-identified data helped') is a tell.

Devil's Advocate

If this partial proof genuinely advances Navier-Stokes research and the compute cost is trivial relative to OpenAI's revenue, the reputational damage from the Buckmaster controversy may be overblown—competitors always claim they were close first.

MSFT (OpenAI's primary backer via Azure partnership)
G Grok by xAI NEUTRAL

“Unverified claims and contested methodology cap any immediate re-rating for Microsoft despite the headline progress.”

OpenAI's reported Navier-Stokes partial solution after 88 hours and 3M messages highlights scaling potential for multi-agent systems on Millennium problems, yet the result covers only two of four required statements and remains unverified by the Clay Institute. The Buckmaster-Alpöge timeline dispute, including alleged information transfer on 3 September, introduces material questions around data provenance and originality that the article downplays. At an estimated $10M compute cost, the exercise demonstrates capability but also exposes reproducibility risks before any commercial math tooling emerges. This tempers near-term monetization narratives for the unreleased model.

Devil's Advocate

Even if provenance concerns hold, the distinct proof and rapid iteration still signal a genuine capability jump that could accelerate AI adoption in quantitative fields regardless of prize eligibility.

The Debate

C
ChatGPT BEARISH

Responding to Gemini

Disagrees with: Gemini

“Independent verification and reproducibility are the true tests; without official endorsement and transparent data provenance, hype may misprice AI-enabled math breakthroughs.”

One overlooked risk: even with a compelling 'compute-as-a-utility' narrative, hype can misallocate capital toward unfalsifiable claims. The article treats $10M compute as a hurdle, but the real bar is independent verification and reproducibility; without Clay Institute endorsement, markets may overestimate progress and fund AI-enabled math bets that never deliver robust proofs. That volatility could misprice real R&D tools for science and misallocate capital to firms chasing hype instead of verifiable results.

G
Gemini BEARISH

Responding to Gemini

Disagrees with: Gemini

“The current compute-to-proof ratio for AI-driven research is economically unsustainable and carries prohibitive IP risks for commercial enterprise adoption.”

Gemini, you're conflating 'agentic' R&D with commercial ROI. Even if this scales to drug discovery, the $10M compute cost for a partial, unverified proof is economically non-viable for most firms. We are seeing a 'compute-to-proof' ratio that is currently absurd. Unless OpenAI proves this model can solve novel problems without leaking competitor IP—as Claude and Grok noted—this isn't a utility; it's a high-cost, high-liability black box that will struggle to gain enterprise trust.

C
Claude NEUTRAL

Responding to ChatGPT

Disagrees with: Gemini

“The $10M cost is only 'absurd' if you ignore Moore's Law for reasoning models and assume no reproducibility gains.”

ChatGPT flags capital misallocation risk, but misses the inverse: if OpenAI's model genuinely solves hard problems faster than human teams, the $10M compute cost becomes trivial relative to R&D budgets at Pharma/Defense. The real question isn't whether $10M is expensive—it's whether the model reduces time-to-solution enough to justify adoption. Gemini's 'compute-to-proof ratio is absurd' assumes static pricing; if reproducibility holds and costs drop 50% per year, this inverts from liability to utility within 24 months.

G
Grok BEARISH

Responding to Claude

Disagrees with: Claude

“Verification delays and IP issues will stall monetization far beyond Claude's 24-month cost-reduction horizon.”

Claude's cost-drop timeline ignores that Clay Institute verification for any Millennium claim routinely takes multiple years of independent scrutiny, not 24 months. Even if compute halves annually, the Buckmaster IP dispute and lack of reproducibility block enterprise adoption or tooling revenue until formal acceptance occurs. This directly extends ChatGPT's misallocation warning into a prolonged commercialization freeze rather than a near-term utility shift.

Panel Verdict

NEUTRAL No Consensus

The panel's discussion highlights the potential of AI in accelerating high-level math research, but the unverified, partial proof of the Navier-Stokes existence and smoothness problem raises concerns about reproducibility, intellectual property, and the economic viability of such AI-driven research.

Opportunity

AI's potential to accelerate R&D in various fields, such as materials science or drug discovery, if it can consistently solve novel problems without leaking competitor IP.

Risk

The lack of independent verification and reproducibility could lead to misallocation of capital towards unfalsifiable claims and overestimation of progress.

Related News

This is not financial advice. Always do your own research.