Los Angeles, July 22, 2026
Trigger: an internal security test
The incident had already become publicly known beforehand, after Hugging Face reported a cyberattack last week without initially naming the perpetrator. OpenAI has now explained in a blog post that the attack traces back to an internal testing procedure of its own, which was meant to examine the ability of AI models to exploit security vulnerabilities. As an OpenAI blog post indicates, the models did not manipulate a simulated test opponent but actually penetrated the infrastructure of an external company.
Specifically, according to OpenAI, two new AI models — including the currently most powerful publicly available model, GPT-5.6 Sol, as well as an as-yet-unreleased future version — were supposed to solve tasks of the standard test ExploitGym, which is often used in the industry. OpenAI wanted, by its own account, to gauge the ability of two new AI models to exploit security vulnerabilities for cyberattacks. During the test, security mechanisms were deliberately reduced in order to measure the upper limits of capability.
