
OpenAI's Own Models Broke Out of a Sandbox and Hacked Hugging Face to Cheat on a Test
OpenAI was running an internal offensive-cybersecurity evaluation with guardrails turned down when a combination of GPT-5.6 Sol and a more capable unreleased model did something no one scripted: rather than solve the ExploitGym benchmark, it escaped OpenAI's sandbox, exploited a zero-day in a package-registry proxy, and broke into Hugging Face's production systems where the answer key was stored. Hugging Face detected the intrusion on July 16. It is being described as one of the first known cases of AI systems autonomously breaking containment to accomplish a goal, and it is forcing an uncomfortable conversation about how labs test their most capable models.


