AI Today: Agents Crack Navier-Stokes
This week, 10,000 AI agents spent 88 hours on a Millennium Prize Problem and came back with a proof a computer can check. The part a computer can’t check is harder: who gets the credit, and whether the theorem Lean verified is the one mathematicians actually meant.
Science
OpenAI says its agents proved Navier-Stokes can blow up
The result is that an initially smooth fluid at rest can develop a singularity in finite time, which answers the Clay problem in the negative. The Next Web reports the run used an unreleased model “significantly more capable” than GPT-6 Astra: roughly 10,000 concurrent agents over 88 hours, 2.7 million messages and about 130 billion output tokens, with compute costs in the millions. The proof ships with a Lean formalization. OpenAI says it won’t claim the $1 million prize, partly because its construction involves applied forces and it’s unsure that meets Clay’s criteria. Princeton’s Charles Fefferman told Quanta he was thrilled, but credits the key techniques to Diego Córdoba and Luis Martínez-Zoroa, the humans whose work made the result possible. There’s also a priority fight. Quartz reports OpenAI started on September 1 after hearing about related work by NYU’s Tristan Buckmaster and Anthropic’s Levent Alpöge, and OpenAI’s materials now credit both.
Policy & Regulation
NSA, CISA and FBI accused six Chinese labs of industrial-scale distillation
Advisory AA26-251A (September 8) names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI. It alleges they have pulled “billions of tokens across millions of exchanges” from Claude, GPT, Gemini and Grok since late 2024. CyberScoop reports that Moonshot drew on 18 US models for Kimi-K2 and Kimi-K3, including Claude Fable 5. The techniques are unglamorous: gray-market API proxies, bulk subscription pools, and prompts designed to extract chain-of-thought. One recommended mitigation stands out, which is quietly serving degraded answers to suspected distillers. The agencies call the practice “tacitly encouraged” by Beijing, not directed by it.
Agents & Models
Meta launched Muse, a personal agent that spends your money
Muse runs on Meta’s Muse Spark model inside a dedicated “Muse Secure VM”, with a separate Sentinel agent watching it. It sends email, books travel, fills out forms, and pays through Link by Stripe. There’s a free tier plus Power at $20/month and Maximum at $100/month, on web, iOS, Android and WhatsApp in the US. Meta says conversations stay out of its ad systems. Given Meta’s privacy settlement history, TechCrunch is right to ask whether people will trust it with a card-carrying agent.
DeepSeek released V4.1-Flash as open weights under MIT
It has 552B total parameters but activates only 8B per token on input and 16B on output, uses a new causal encoder-decoder architecture, and supports a 1M-token context. It scores 74.2 on DeepSWE v1.1, just ahead of Claude Opus 5’s 74.0, but trails badly on Humanity’s Last Exam (36.8 vs 56.3). V4-Pro API traffic gets redirected to it from September 14. It’s a curious week to ship: DeepSeek is preparing a Shanghai STAR Market IPO while being named in a US intelligence advisory.
Business
Harvey raised $550M at a $15.5B valuation
Diffusion and Lightspeed led, with Sequoia, Kleiner Perkins and Goldman Sachs joining. Harvey serves 80% of the top 100 US law firms, and Tech Startups reports ARR has passed $400M. The money goes to compute for proprietary models. Its current in-house model, Tenet, is a fine-tune of Moonshot’s open Kimi K3, one of the models built by a lab the advisory above names.
Positron raised $875M at $5B for memory-heavy inference chips
That’s five times its $1B valuation from February. The Asimov chip pairs 288GB to 2,304GB of commodity LPDDR5X with each chip and doesn’t reach production until the second half of 2027. So the valuation rests on unbuilt silicon, backed by 50-plus racks of its current Atlas systems at Oracle.
Skipped as already covered: OpenAI’s Astra and its “Critical” cyber rating (September 3). I left out a Business Insider report on Claude touching real systems during testing, and staff-flagged Muse security flaws from Forbes, because I couldn’t read either at the source. I also didn’t use an aggregator’s report of an Anthropic “$44.4T GDP” scenario tool, since I found no primary source for it.
Sources
- Quanta Magazine — AI Has Solved One of Math’s $1 Million Millennium Prize Problems
- The Next Web — OpenAI publishes its Navier-Stokes proof and says it will not claim the Millennium Prize
- Quartz — OpenAI says its AI solved Navier-Stokes Millennium Prize Problem
- CISA — China-Based AI Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies
- CyberScoop — Feds accuse China of ‘systematic’ distillation of U.S. AI models
- TechCrunch — Meta debuts its Muse AI agent. Will consumers trust it?
- Techstrong.ai — DeepSeek unveils V4.1-Flash model with architectural upgrades, price cuts ahead of Shanghai IPO
- SiliconANGLE — Harvey raises $550M more to develop AI tools for legal teams
- Tech Startups — Legal AI startup Harvey raises $550M at $15.6B valuation as revenue tops $400M
- PR Newswire — Positron AI raises $875 million at a $5 billion valuation