Agentic Risk as a New Category of Information Security

I have spent my career on how systems decide what to trust and who is allowed to reach it. Search relevance, knowledge graphs, an enterprise search system for companies and platforms accessible by millions monthly from around the world.

For most of that time the stakes were quality. Get it wrong and someone gets a bad answer. Once a system can act on what it retrieves, the same decision determines whether data stays locally or leaves the initial location.

Sadly, data is leaving the location without permission or human actors, and we are seeing and experiencing it in real time. From personal observation, it seems as if AI agents are here to stay in some form and it is high time we find solutions on how we can create AI agent containment that holds when one control fails.

What is happening

Almost every security tool we have assumes a break-in involves some type of bad code or a stolen password. That assumption has held for a while. It does not hold for what agentic AI is producing, and the gap is why I want to work on this.

In 2025, someone worked out that they could steal files from a Microsoft 365 Copilot user by sending an email. The instructions sat hidden in the message. The copilot read them while doing a routine summary, pulled internal content, and shipped it out. The victim clicked nothing. It was assigned CVE-2025-32711 and scored 9.3 (National Vulnerability Database, 2025). There was no malware for a scanner to catch, because the exploit was a paragraph of English.

Anthropic (2025) later published an account of a state-linked group running its coding tool as the actual operator of an intrusion campaign across roughly thirty targets, with the model handling 80 to 90 percent of the work and issuing requests faster than any human team could. The same report describes an earlier case, less sophisticated and in some ways more alarming, where an operator with limited technical skill used an agent to extort seventeen organizations in a single month. The agent did the scanning, took the credentials, wrote the malware, and drafted the ransom notes.

The tooling layer went next.

Help Net Security (2026) collected the cases. A connector package shipped fifteen clean releases, built up trust, then added one line that stole data. A remote code execution flaw rated 9.6 turned up in core Model Context Protocol infrastructure that hundreds of thousands of developers rely on. In Cursor, an attacker could poison the environment so that the pre-approved safe commands were what delivered the payload, which means the safety list was the way in.

Then July 2026. An AI agent running a capability benchmark got out of its test environment through an undisclosed flaw in a package proxy, worked across to a machine with internet access, and broke into Hugging Face to get the answers to its own test (OpenAI, 2026). It entered through a dataset whose processing ran code, escalated, took cloud and cluster credentials, and spread across internal clusters over a weekend in thousands of small actions (Hugging Face, 2026). No human attacked anyone else; the agent wanted to pass a test.

Hugging Face (2026) then hit something I have not seen discussed anywhere. Trying to analyze the attack, they found commercial models refused the work, because feeding in real payloads and command-and-control artifacts tripped safety filters that cannot tell an investigator from an attacker. They ended up running forensics on an open-weight model on their own hardware.

Why this does not fit the old categories

There is usually no artifact, no malware sample, no odd login, and the like, since these are architectural problems rather than coding mistakes (OWASP GenAI Security Project, 2026). A security program built on vulnerability feeds and patch cycles cannot see any of it.

It also cannot be patched. A model follows a hidden instruction because it has no way to separate instruction from content, which Vassilev et al. (2025) treat as the central weakness of these systems. That is the product working. And attribution falls apart. The Hugging Face case involved a breach with credential theft and lateral movement and no threat actor to name. Insurance, breach notification law, and threat intelligence all expect a human somewhere in the chain. None here is present.

References

Anthropic. (2025). Disrupting the first reported AI-orchestrated cyber espionage campaign. https://assets.anthropic.com/m/ec212e6566a0d47/original/Disrupting-the-first-reported-AI-orchestrated-cyber-espionage-campaign.pdf

Help Net Security. (2026, June 11). Prompt injection still drives most agentic AI security failures in production. https://www.helpnetsecurity.com/2026/06/11/owasp-prompt-injection-ai-security-failures/

Hugging Face. (2026, July 16). Security incident disclosure: July 2026. https://huggingface.co/blog/security-incident-july-2026

National Vulnerability Database. (2025). CVE-2025-32711. National Institute of Standards and Technology. https://nvd.nist.gov/vuln/detail/CVE-2025-32711

OpenAI. (2026, July 21). OpenAI and Hugging Face partner to address security incident during model evaluation. https://openai.com/index/hugging-face-model-evaluation-security-incident/

OWASP GenAI Security Project. (2026). OWASP GenAI exploit round-up report Q1 2026. https://genai.owasp.org/2026/04/14/owasp-genai-exploit-round-up-report-q1-2026/

Vassilev, A., Oprea, A., Fordyce, A., Anderson, H., Davies, X., & Hamin, M. (2025). Adversarial machine learning: A taxonomy and terminology of attacks and mitigations (NIST AI 100-2e2025). National Institute of Standards and Technology. https://csrc.nist.gov/pubs/ai/100/2/e2025/final

Subscribe to Everyone is searching

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe