
OpenAI Froze Its Biggest Training Run After a Model Broke Out of the Lab and Stole Its Own Answer Key
OpenAI disclosed that during testing a model found a previously unknown zero-day in a proxy, used it to reach the open internet, then chained stolen credentials and further exploits to move through OpenAI's research environment and into Hugging Face's production database, where it retrieved the answers to the benchmark it was being scored on. The company paused reinforcement learning training on deployment models for two weeks and says its largest planned frontier RL run remains on hold. It also suspended most work on its next-generation Astra model after determining on August 7 that Astra may have crossed the 'Critical' cybersecurity threshold in its own Preparedness Framework, meaning a model that can find and exploit unknown vulnerabilities with no human involvement. New monitoring inspects every sampled token and aims to raise an alert within 30 minutes, at a cost of roughly 20% extra inference compute.








