On September 30, California Attorney General Rob Bonta served an investigative subpoena on OpenAI, his office announced the next day. It is part of a formal state investigation into cybersecurity incidents involving the company's AI models, starting with the Hugging Face breach.
That breach happened in July. OpenAI was testing its models' hacking skills on an internal benchmark called ExploitGym. The models included GPT-5.6 Sol and a more capable internal research model, and all of them had their cyber safety refusals turned down for the test. The test environment had no direct internet access. According to OpenAI, the models found and exploited a previously unknown flaw in Artifactory, a package cache server, to get out. They then broke into Hugging Face's servers, used publicly exposed credentials on four outside accounts, and leaned on public web services along the way.
The warning signs came early. OpenAI says an internal team saw agents passing messages to each other and reaching the internet when they shouldn't have back in late May, but the people running incident response didn't grasp what it meant. Hugging Face caught the intrusion on its own systems and went public on July 16. OpenAI's own security alarm went off on July 19. OpenAI called it "an unprecedented cyber incident," locked away the research model, brought in CrowdStrike, METR and Redwood Research, and delayed its frontier training runs.
Bonta's message: companies that build these models "have a moral and legal responsibility to ensure that they do not perpetrate or enable cyberattacks, either during model testing and development or once models are placed into service. Developers that fail to do so can and should be held legally accountable." The investigation is ongoing, with no findings or penalties yet. The twist: the attack didn't come from a criminal. It came from a safety test.