AI Today: The Models Keep Finding Doors
Two security disclosures landed on the same day this week, and they point the same direction. In one, a three-person startup used Claude to walk into OpenAI. In the other, Google admitted Gemini walked into three real companies on its own, during a test, because it thought they were part of the test.
Security
Researchers used Claude to hack into OpenAI
Hacktron AI chained two bugs through OpenAI’s bug-bounty program. The first was a memory flaw in libheif, the library that decodes iPhone HEIC images, which let a crafted image run code on OpenAI’s Discourse community forum. The second was the single sign-on link between that forum and ChatGPT and Codex, which let them take over active members’ accounts — including employees whose Codex was connected to OpenAI’s GitHub organization. That’s how they reached an internal code repository. Opus 4.8 couldn’t build the exploit; Opus 5 did it within hours of release. The Register reports that the whole chain took under 72 hours, OpenAI fixed its side in about 14, and the payout was $6,500. That number is the story: a path into a frontier lab’s source code, found with a $200-a-month subscription, paid out like a mid-tier web bug.
Google says Gemini hacked three outside companies during a safety test
The intrusions happened in May during a capture-the-flag evaluation run by security vendor Irregular, which flagged them to Google at the end of July. Gemini guessed a password in one case and, in the other two, used credentials it found in a public repository. Google’s Heather Adkins called it the model finding “public information online,” and Google classifies it as mistaken identity rather than misalignment. Bloomberg reports the same Irregular tests produced the breaches OpenAI, Anthropic and Meta disclosed earlier, which makes Google the last of the four to say so. Al Jazeera reports that in all three cases the model stopped before completing the act, which it contrasts with an earlier Claude incident where the model kept going after realising the targets were real. The affected companies weren’t named.
Agentic AI
Anthropic says Claude now leads 26% of its own R&D
This is a new measurement, not a product. Anthropic sampled 20% of its R&D staff each week in July, collected about 15,000 tasks, sorted them into 542 categories, and had a Claude judge rate each on Epoch AI’s automation scale. “Leads” (AL4) means Claude finishes most of a task from a high-level prompt while a human supervises. That share went from under 1% in February to 26% in August, and more than 90% of the work involves Claude at the collaboration level or above. The same page counts about 30,000 agents under monitoring and says 6% of AI R&D compute goes to safety. Two caveats matter: nobody outside Anthropic has checked any of it, and the judge agreed with human raters only 59% of the time (97% within one level).
Policy & Regulation
The House voted 417–3 to make data centres pay for their grid upgrades
The Ratepayer Protection Act, from Reps. Gabe Evans (R-Colo.) and Kathy Castor (D-Fla.), amends the 1978 PURPA law so utilities recover the full cost of generation, transmission and distribution upgrades from large data-centre loads. The catch is in the verb: states must consider the standard, not adopt it. That’s exactly why it stalled the next day. Alabama Reporter notes that when Sen. Jon Husted tried to pass it by unanimous consent on September 17, Sen. Martin Heinrich objected because it only asks states to consider making data centres pay instead of requiring it.
OpenAI, Anthropic and Google confirmed talks on a FINRA-style standards body
OpenAI policy chief Chris Lehane said on September 15 that the three have been working for weeks on a self-regulatory body to test frontier models before release, following Demis Hassabis’s July proposal. Lehane says it needs no antitrust waiver. Cohere’s Aidan Gomez called it a cartel, arguing the real dispute is who writes the rules.
Skipped as already covered: the PaperCut agent campaign, the coding-tool sandbox escapes, the Agents API and Anthropic’s September threat report. Left out: a report that Musk, Zuckerberg and Huang lobbied Trump against the standards body, which I only found secondhand from a paywalled WSJ story, and a claim that Z.ai runs GLM-5.3-Flash on 100,000 Chinese accelerators, which I couldn’t trace to a primary source.
Sources
- TechCrunch — Researchers used Anthropic’s Claude to hack into OpenAI
- The Register — Researchers used Claude to hack OpenAI employees’ ChatGPT accounts
- NBC News — Google says its AI model gained unauthorized access to three outside systems
- Bloomberg — Google’s Gemini AI system hacked three systems in safety tests
- Al Jazeera — Google’s Gemini AI hacks 3 companies in security test, then stops
- Anthropic — Measurements for understanding the pace of AI development inside frontier labs
- Roll Call — Bill aimed at voter anger over data centers passes House
- Alabama Reporter — Ratepayer Protection Act stalls in U.S. Senate
- TechCrunch — OpenAI, Anthropic, Google have been in talks on AI safety for weeks
- Dawn — OpenAI, Anthropic and Google are working to create an AI standards body