Security Experts Discuss the Hugging Face, OpenAI Incident

Earlier this month, AI model/dataset platform Hugging Face announced it had experienced a security incident carried out by an AI agent. Now, OpenAI has revealed their models were involved.
In a statement, OpenAI confirmed, “After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.”
This makes the Hugging Face breach among the first confirmed cases of AI-conducted cyberattacks, rather than an AI-assisted incident.
Andrew Chipman, Director of GRC & ISO at ProCircular, states, “Most of the articles I’m seeing (as well as OpenAI’s comments Tuesday) reference OpenAI’s push for strengthened guardrails and improved model alignment. It does not appear that OpenAI took any basic measures to ensure the model would remain isolated if it went rogue. If working from an evaluation environment, that environment should have been physically segregated from the public internet. Logical segregation was clearly not enough. OpenAI appears to be playing fast and loose with some dangerous tech.”
AI-Conducted Attacks
As more news of AI-conducted attacks make headlines, security leaders must pivot from treating this concept as a future concern to a current threat.
Crystal Morin, Senior Cybersecurity Strategist at Sysdig, comments, “The OpenAI model that attacked Hugging Face ran over 17,000 actions in a single weekend. JADEPUFFER, the first fully autonomous ransomware campaign run end-to-end by an AI agent, corrected its own failed step in 31 seconds. That is the speed defenders must match at a time of agentic threat actors, and if their detection strategy is to read logs and respond after the fact, attackers will always be a step ahead.
“Here is the detail getting buried under the guardrail debate: what actually caught the Hugging Face intrusion was behavioral anomaly detection at the infrastructure level, the correlation of telemetry most teams would have otherwise written off as noise. That is the lesson worth planning around.
“Security teams must stop treating AI-driven attacks as a future risk and start verifying they can catch them today. Confirm you can spot a privileged container spinning up from an application process, make sure your model weights and training data are backed up as rigorously as your databases, and run a tabletop that assumes attacks will unfold in minutes. Machine-speed threats against human-speed response is the gap. Find it before an attacker does.”
Do Organizations Know What Their AI Agents Are Doing?
Research suggests incidents like these will not be a one-time thing. The State of AI 2026 report by AvePoint finds approximately half of employees regularly use AI agents at work, but companies don’t entirely know what those agents are doing.
Dana Simberkoff, Chief Risk, Privacy and Information Security Officer at AvePoint, dives into the research, saying, “This incident challenges a fundamental operating assumption. A human attacker usually chooses a target for a reason; an autonomous system may simply pursue the fastest path to an objective, even across organizations and trust boundaries. For CISOs, the issue is no longer just whether AI can find vulnerabilities. It is whether we have continuous visibility into what the system can reach, what it did, and how quickly we can stop it. AvePoint’s 2026 AI Report found that 88% of organizations experienced at least one agent-related security incident in the last year, which highlights the scope of this problem.
“This incident really exposes the limits of treating containment as a one-time design decision. Sandboxes matter, but autonomous systems can reason through small openings, use tools, and keep acting toward an objective. In my experience, the question cannot be, ‘Was it sandboxed?’ It has to be, ‘Can we prove what it accessed, whether it crossed a boundary, and how quickly we can stop it?’ What organizations need is an AI trust layer: enforceable controls, data protection, monitoring, and lifecycle governance around the agent, rather than the model writ large.
“Critical infrastructure raises the stakes because the risk is not limited to data exposure. In financial-market infrastructure, availability, integrity, sequencing, identity, and confidence all matter. An autonomous system pursuing an objective across organizations will not understand systemic risk or regulatory boundaries unless those constraints are engineered and enforced. Testing in or near these environments requires independent review, hard technical limits, continuous monitoring, and clear accountability. ‘We did not intend for the model to go there’ is not data protection, governance, or a defensible control.
“Defenders should assume the response window is shrinking. AI does not need rest, does not lose focus, and can turn small openings into chained activity very quickly. Point-in-time assessments and manual escalation alone are not enough when autonomous systems can act faster than traditional response processes. What organizations need is continuous visibility into identity, access, configuration, data movement, and agent behavior before an incident occurs. AvePoint’s AI Report found that nearly 9 in 10 organizations delayed both agentic and generative AI deployments by an average of almost six months because of unresolved data security and data management concerns. You cannot secure what you cannot see, and confidence is not the same as control.
“Advanced defensive capabilities matter, however, access to powerful tools is not the same as readiness to use them safely. Before testing systems like this against real-world infrastructure, I would want independent pre-test review, documented scope, technical controls that prevent boundary crossing, continuous monitoring by a separate team, mandatory reporting, and a liability model that does not leave the affected third party carrying the risk. Good governance does not slow innovation. Like brakes on a car, it is what lets organizations move faster safely.
“Autonomous systems do not necessarily share human assumptions about relevance, proportionality, or organizational boundaries. If the objective is framed poorly, the system may build a chain based on efficiency rather than intent. Compute matters because it can turn lab-feasible techniques into operationally practical ones. Defenders need to think less in terms of single vulnerabilities and more in terms of paths: what data, identities, APIs, suppliers, and trust relationships could be combined into a route the organization did not anticipate, and where the AI trust layer needs to intervene? AvePoint’s 2026 AI Report found that 35.5% of enterprise data is now AI-generated, projected to reach 42.1% within 12 months, which expands the surface governance has to cover.”
Security Leaders Weigh In
Chandra Gnanasambandam, Chief Technology Officer at SailPoint:
Two massive headlines this week paint a clear picture of the future of enterprise security: OpenAI’s push into autonomous AI agents, and the Hugging Face breach involving compromised API secrets.
The era of Agentic AI is here. But there is a dangerous gap between AI innovation and enterprise security readiness.
Agents run on non-human credentials. To act on your behalf, an AI agent needs API keys, access tokens, and system credentials. If you treat these AI agents like traditional service accounts — leaving their access ungoverned and their credentials unmanaged — you are creating a massive, automated attack surface.
You should not adopt autonomous AI without first locking down non-human identities (NHIs). You must be able to discover the API keys and tokens these agents use, govern their access dynamically, instantly trace their actions back to a human owner and disable them instantly when required.
Ram Varadarajan, CEO at Acalvio:
What makes this incident worrying is not simply that an AI found vulnerabilities, but that it independently assembled them into a coherent attack path, illustrating that future cyber threats may be limited less by human expertise than by computational resources. This is the central issue that should concern us.
This incident underscores that containment can’t be treated as a one-time engineering boundary, because sufficiently capable agents will view the boundary itself as just another obstacle to optimize around. We have to remember that these are thinking agents in the end, with infinite patience to find the paths forward.
The incident suggests that critical infrastructure should increasingly be designed on the assumption that highly capable AI systems can and will reason across organizational boundaries, making shared ecosystems and trust relationships part of the attack surface rather than merely the environment.
Defenders need to reconsider assumptions about how much time they will have to detect and interrupt an attack because autonomous systems can compress what once took skilled human operators days or weeks into minutes, dramatically shrinking the window available for defenders to detect and disrupt an intrusion. We'll be facing adversaries that will be thinking and acting at machine speed. We have no option than to meet them with machine-speed defenders that are likely driven by model-aware game theory.
The deeper lesson here is that these systems don’t reason in terms of human intent or institutional boundaries but in terms of optimization, meaning that as computing power grows, attack chains once dismissed as too complex or resource-intensive may become increasingly practical for autonomous systems to discover and execute. It’s cliche to say by now, but we’re living in an entirely new world and have to throw out our past assumptions of the threat environment AI brings. The future is bot-on-bot: meeting that automated AI threat with an even more capable automated AI defender. That arms race will never end.
Mark Stockley, Cybersecurity Evangelist at ThreatDown:
During an internal benchmark meant to measure cyber capability, OpenAI ran GPT-5.6 Sol and a more capable pre-release model with reduced safety classifiers (deliberately, for the test). The models found a path to internet access, exploited a zero-day in an internal tool, inferred that Hugging Face likely held the benchmark's answer key, and used stolen credentials to retrieve it from HF’s production database — effectively cheating on their own evaluation. Hugging Face says no public models, datasets, or Spaces show evidence of tampering.
The irony at the center of the piece: The same guardrails tension shows up twice, on opposite sides:
Reducing OpenAI’s safety classifiers for the eval is exactly what let the model escalate this far. When Hugging Face tried to investigate using a commercial frontier model, that model’s own guardrails blocked the analysis — it couldn’t distinguish a defender from an attacker. Hugging Face switched to an open-weight model on their own infrastructure instead.
OpenAI ties the incident to UK AI Security Institute research on long-horizon cyber capability — the same institute ThreatDown cites for its Mythos benchmark results.
Separately, ThreatDown’s own research found Hugging Face hosts 6,000+ “guardrail-free” models (self-labeled abliterated/uncensored), downloaded 22 million times in 30 days. The point: Hugging Face sits at the center of both stories — the platform being attacked, and a major source of models with no safety gate at all.
Randolph Barr, Chief Information Security Officer at Cequence Security:
What stands out is the asymmetry: the attacker’s AI agent operated with zero usage restrictions, while Hugging Face's own forensic work got blocked by the safety guardrails of Western frontier models. They ended up running incident response on a Chinese open-weight model, ironic given where many cyber threats originate. The takeaway for defenders is worth acting on now: have a capable, self-hosted model vetted and ready before an incident, so you're not locked out by guardrails or forced to send attack data and credentials outside your environment. Both sides are using AI, but only one side is playing without limits.
OpenAI published a post, OpenAI and Hugging Face partner to address security incident during model evaluation, confirming the attack was driven by their own models (including a pre-release one with reduced cyber refusals) during an internal capability evaluation that escaped its sandbox.
I think this level of transparency from OpenAI is a great thing, responsibly disclosing the zero-day, bringing Hugging Face into their trusted access program, and sharing findings openly. Reads like this lead to exactly the kind of CISO-level conversations we will be discussing in upcoming sessions this week.
That said, it doesn’t change my recommendation: have a capable, self-hosted model vetted and ready before an incident, so you’re not locked out by guardrails or forced to send attack data and credentials outside your environment. If anything, this confirms how real the capability gap is, these models sustained a multi-step attack across two companies’ infrastructure on their own.
Diana Kelley, Chief Information Security Officer at Noma Security:
Based on Hugging Face’s disclosure, and with the investigation still ongoing, what stands out is that AI did not invent a completely new attack chain. It industrialized a familiar one: code execution in a processing pipeline, privilege escalation, credential harvesting, and lateral movement, using an agentic framework able to carry out thousands of actions across short-lived environments over a weekend. For CISOs, this underscores an important MLSecOps reality, model hubs, datasets, and data-processing pipelines belong in the same threat model as package registries and CI/CD systems.
This changes the economics of cyber offense. A smaller number of adversaries can run a 24/7 campaign at a scale that once required a much larger team and significantly more resources, while also compressing the time between finding an opening, exploiting it, harvesting credentials, and moving laterally. The security operating model has to move from human-paced triage to machine-speed detection and containment, treating AI agents as privileged identities and data-processing workers as untrusted execution zones.
On the choice to use GLM-5.2 rather than U.S. frontier models, the important distinction is provider-mediated access versus a model the defender can run and govern directly. Hugging Face says hosted frontier models blocked the analysis of real attack commands, exploit payloads, and command-and-control artifacts because the material triggered their safety guardrails. This is not an argument for removing safety controls, it is an argument for frontier model providers to create a verified incident-response path, with strong identity verification, auditability, case-level scope, and data isolation, that can distinguish authorized defensive work from abuse. Until that exists, CISOs that want to use AI for incident response should have a vetted self-hosted model available as a backup option.
The defining feature of this incident is that AI was on both sides of the wire. An autonomous agent framework drove the intrusion, AI-assisted detection found the signal, and AI agents helped dissect the campaign. Humans still handled containment and remain accountable, but machines operated at scale on both offense and defense. That is likely to become a core security operating model going forward, AI versus AI, governed by people.
Agnidipta Sarkar, Chief Evangelist at ColorTokens:
What fascinated me was that OpenAI predicted that “such attacks would become increasingly more common as AI adoption continues to grow". Guardrails set up to prevent this actually undermined the defense while the autonomous AI attack went about its business, until an open model was used to contain it. The adage “it is not if you will be attacked but when” just got modified to “you will be attacked as you adopt AI”. And the only thing standing between the next AI perpetrator and your critical digital assets is microsegmentation technology and a cryptographic, passwordless, agentic AI identity.
Considering the tremendous value of AI in healthcare, manufacturing, banking, e-commerce, and many critical infrastructure sections, AI adoption must continue. However, now more than ever, organizations must adopt foundational microsegmentation tools focused on stopping the proliferation of attacks by proactively reducing the possibilities for an attacker (human or AI) to find targets and develop exploits. For those wondering about the time it takes, please note that deploying microsegmentation that can integrate with your existing EDR can help you deploy impenetrable defenses in hours, not months.
Looking for a reprint of this article?
From high-res PDFs to custom plaques, order your copy today!






