Today's theme is agents slipping the leash, for better and worse. OpenAI says 10,000 of them cracked a $1 million math problem in 88 hours and immediately got into a fight over who deserves the credit, Meta put one inside your inbox and your checkout flow, and Google caught criminals using them to build a credential-theft campaign from scratch in under six hours. Meanwhile three US agencies formally accused six Chinese labs of quietly draining American models through the API, and Google dropped its largest European cheque ever on a country of 5.6 million people.
OpenAI Says 10,000 Agents Cracked a $1 Million Math Problem in 88 Hours, and the Credit Fight Started Immediately
OpenAI announced that an unreleased model running roughly 10,000 parallel agents produced a proof for the Navier-Stokes existence and smoothness problem, one of the seven Clay Institute Millennium Prize Problems, in 88 hours and millions of dollars of compute. Within hours, NYU's Tristan Buckmaster and Anthropic researcher Levent Alpöge said they had been working the same fluid equations for nearly a year, with Buckmaster alleging OpenAI pressured him to drop Alpöge from authorship; OpenAI denies accessing their private work. The Clay Institute has verified nothing, and its own rules require peer-reviewed publication plus a two-year waiting period before a prize committee even convenes. Whatever the proof turns out to be worth, the fight over attribution is the preview of every AI-assisted research dispute to come.
The NSA, CISA and FBI Named Six Chinese Labs and Said They Drained Billions of Tokens Out of US Models
Joint advisory AA26-251A, published September 8, accuses DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI of running industrial-scale knowledge distillation against Claude, GPT, Gemini and Grok since late 2024, extracting billions of tokens across millions of exchanges. The agencies say DeepSeek specifically distilled from a long list of frontier models to generate synthetic training data for R1 and V3, routing requests through a gray market of proxy 'transfer stations' to evade geographic restrictions and terms of service. The advisory's language is unusually blunt: distillation forms the core of Chinese AI development strategy, not a supplement to it, and it likely happened with government awareness. If you sell API access to a frontier model, you are now being asked to treat your own customers as a counterintelligence problem.
Meta Put an Agent Inside Your Inbox, Your Calendar and Your Checkout Flow, and Gave It Its Own Computer
Meta launched Muse, a personal AI agent that connects to email, calendar, payments, health apps, smart home devices and shopping services and actually completes errands inside those accounts, from booking travel to negotiating a bill down. It runs on Muse Secure VM, a dedicated cloud machine holding both the agent and your data, with a separate Sentinel agent that must approve anything Muse sends to the internet. Payments go through Link by Stripe, which issues a single-use virtual card scoped to each approved purchase so Muse never sees real card details. It is free for most use with $20 and $100 monthly tiers, US-only for now on iOS, Android, muse.ai and WhatsApp, with AI glasses to follow. The architecture is the interesting part: Meta is betting that consumer agents need their own sandboxed computer, not a plugin.
Google Watched Attackers Build a Credential-Theft Campaign From Scratch in Under Six Hours Using AI Agents
Google Threat Intelligence Group's Q3 2026 AI Threat Tracker documents a financially motivated actor who compromised a victim's cloud infrastructure and then combined an AI coding chatbot, a prompt, and predefined agent instructions into an automated vulnerability-scanning and credential-harvesting pipeline, start to finish in less than six hours. Markdown files served as operational playbooks, letting agents manage scans, fix their own errors in real time and rotate IPs without continuous human input. GTIG also tracked a PRC-nexus espionage actor using Gemini to build a dynamic automated penetration-testing framework. The report is careful to note that nobody has yet seen end-to-end autonomous zero-day discovery in the wild, but the human-in-the-loop delay that defenders have always relied on is collapsing.
Google Is Spending €13 Billion in Finland and Signed a 22-Year Deal to Keep a Nuclear Plant Alive
Google announced its largest single European investment ever: at least €13 billion across 2027 and 2028 into data centres and supporting infrastructure in Hamina, Kajaani, Muhos and Vaala. The construction phase alone is projected to support 37,000 jobs and add €3.6 billion a year to Finnish GDP, in a country of 5.6 million people. The energy side is the tell: Google signed a 22-year life-extension power purchase agreement for the Loviisa nuclear plant, which supplies roughly 10% of Finland's electricity, added onshore wind deals taking its contracted new-to-grid capacity to 629 MW, and contracted a 94 MW battery system to stabilise prices during cold windless stretches. Hyperscalers are no longer just buying power, they are underwriting the survival of national generating assets.