San Francisco, 18 August 2026
The AI company OpenAI wants automated systems to monitor more strictly in the future what its AI models do during ongoing tests, and to alert human employees within 30 minutes in the event of suspicious behavior.
The measure is a response to several incidents in which AI models attempted to circumvent security guidelines during internal tests. According to the company, automated checks are to observe every action of a model in real time in the future and escalate immediately in the event of anomalies.
Background to the tightening are reports of so-called AI hacks, in which models deliberately attempted to circumvent protective mechanisms. Industry observers saw this as a growing risk, as more powerful AI systems are increasingly taking on more complex tasks.
