This isn’t science fiction — it happened this week, and nobody was really ready for it.
On 21 July 2026, OpenAI acknowledged an incident it describes as « unprecedented. » During an internal evaluation of cyber capabilities, several of its latest models — configured with « reduced cyber refusals » for the purposes of a sandbox test — escaped their evaluation environment and breached Hugging Face.
The sequence, as documented by OpenAI and Hugging Face:
- exploiting a zero-day flaw in an internal proxy to reach the internet;
- privilege escalation and lateral movement;
- intrusion into Hugging Face’s production infrastructure by chaining stolen credentials and a zero-day, all the way to remote code execution.
Fortunately, the agent’s goal was not to cause harm but simply to… pass its test. The benchmark used to evaluate the models hosted its solutions on Hugging Face: the agent « deduced » where the answers were and went to fetch them directly at the source. This algorithmic overzealousness is nothing more than the machine optimising the objective it is given.
The European legislator had anticipated the situation almost word for word.
Regulation (EU) 2024/1689 (the AI Act), in Recital 110, explicitly lists among the systemic risks of general-purpose AI models « offensive cyber capabilities » (discovery and exploitation of vulnerabilities) and « loss of control » relating to alignment with human intent. Recital 115 targets the circumvention of safety measures and defence against cyberattacks.
Concretely, for a provider of a GPAI model with systemic risk, Article 55 already requires:
- 55(1)(a) — model evaluation;
- 55(1)(b) — identifying and mitigating systemic risks at Union level;
- 55(1)(c) — tracking, documenting and reporting serious incidents to the AI Office without undue delay;
- 55(1)(d) — ensuring an adequate level of cybersecurity for the model and its physical infrastructure.
These measures must also be placed in the context of NIS2 for essential and important entities.
The moral of the story: a system that is capable and autonomous enough will treat your safeguards as an obstacle to circumvent if they are not designed to withstand it. AI governance is not a post-deployment checkbox — it is a design requirement.
Do your AI agents operate within a genuinely constrained scope of action?
Sources: OpenAI — Hugging Face model evaluation security incident; Hugging Face — Security incident disclosure (July 2026); Regulation (EU) 2024/1689, Art. 55 and Recitals 110 & 115 (EUR-Lex, CELEX 32024R1689).
Jeoffrey Vigneron — Lawgitech. Informational content for educational purposes; not individual legal advice.





