ExtraHop® Closes Enterprise Data Center Blind Spots with new 400 Gbps sensor

Search
  • Solutionschevron right
  • Industrieschevron right
  • Platformchevron right
  • Resourceschevron right
  • Customerschevron right
  • Companychevron right

When AI Agents Go Rogue, the Network Is Your Last Line of Defense

Share blog icon

Back to top

Back to top

September 15, 2026

When AI Agents Go Rogue, the Network Is Your Last Line of Defense

In August, Anthropic disclosed findings pertaining to how agents handle interaction — what they saw involved sabotage.

In Anthropic’s study, multiple AI models were tasked with a common software engineering assignment — porting a Python backend to a different language. The agents weren’t informed that there would be other agents working on the same project, as the researchers wanted to observe what would happen when the agents encountered one another. 

The agents had been given incompatible goals, leading to escalation and misaligned behavior. Anthropic called it a multiagent turf war. Agents started to unleash self-replicating malware on one another, among other things. After four hours, researchers halted the experiment. 

5 Escalation Tactics Agents Use for Sabotage

The rivalry played out through five distinct maneuvers. Agents:

  • Locked out competing agents by disabling their accounts
  • Built scripts designed to detect and terminate rival processes
  • Embedded malicious code, disguising it as output from another agent
  • Seized control of a task by stripping away another agent's access
  • Walked away from the task entirely rather than continue competing

The maneuvers that an agent selected depended heavily on the model underpinning the agent. 

Loading table...

Even when superior capability eventually led to a truce, the instinct was consistent: freeze out the competitor first and only then come to the table.

How Agent Sabotage Translates to Operational Risk

Agent-to-agent sabotage isn't just a research curiosity — it carries real business consequences. Autonomous agents operating under unclear, conflicting, or misaligned instructions can act destructively: pulling credentials, killing processes, and slipping in malicious payloads. The result is disrupted workflows, operational headaches, and genuine security exposure.

In enterprise deployments, similar behavior could surface in the following ways:

Loading table...

Closing the Visibility Gap To Sidestep Sabotage

Anthropic's researchers could observe the sabotage because their test was purpose-built to expose it. However, production environments offer no such advantage. Spotting a credential lockout, a terminated process, or a disguised payload requires visibility into how agents interact with each other — not just insight into how a single agent behaves on its own.

Endpoint detection and identity management tools weren't designed with inter-agent interactions in mind. Agent-to-agent behavior falls squarely outside what conventional tooling can see.

Agentic Oversight Needs Network Context in the AI Era

When an autonomous workflow comes under review, the clearest record often lives in what crossed the network — which actions occurred, which identities and systems were involved, and how the effects spread across connected services.

An agent's own account of its work covers only part of that picture, as agents can misreport, omit steps, or simply lack visibility into how their actions rippled outward. A fuller picture requires ground truth that doesn't depend on any single agent's self-reporting.

Network intelligence tracks communications between agent workloads and the services they call. As agents interact, traffic reveals shifts in connection patterns, request volumes, and service behavior. Supported protocol analysis and decryption add richer transaction context, turning raw traffic into evidence investigators can reason on.

ExtraHop's sensor sits at this vantage point by design, giving teams a persistent record of agent-to-agent activity that exists apart from whatever any one agent chooses to log.

Two agents repeatedly overwriting each other's changes illustrate why that independence matters. The network captures the traffic; endpoint and application records reveal what changed; identity and execution records tie it to the responsible workflow. Only if all modalities work together can they distinguish ordinary churn from an unresolved conflict between competing tasks.


Oversight has to be built into agent deployments before agents go live. Distinct identities for each agent, limits on their authority over shared resources, clear ownership, and a defined escalation path for conflicting instructions form the governed control plane agents operate within. Audit records preserved outside the agents' control, paired with a way to pause a workflow without triggering a wider outage, give security teams a way to intervene early.

ExtraHop's network telemetry contributes the communication and transaction evidence needed to investigate activity across connected systems. Combined with endpoint, identity, and agent execution context, it equips security teams to assess impact and make better-informed response decisions.

Security agents investigating agent-to-agent conflicts depend on the same quality of evidence and the same controls governing what they're permitted to do. 

ExtraHop's Context, Harness and Model framework brings the requirements together into a single operating model. Learn more in CEO Greg Clark’s thought leadership article The SOC Needs a New Operating Model

Discover more

blog image
Blog author
Jamie Moles

Senior Manager, Technical Marketing

Jamie Moles is a Senior Manager of Technical Marketing at ExtraHop with 30+ years of hands-on experience dismantling complex threat behaviors. Jamie Moles began his career reverse-engineering early malware in the MS-DOS era and currently focuses on cutting through industry noise to deliver practical, network-backed security strategies. View Jamie Moles’ complete professional profile on LinkedIn.

Share
LinkedIn logoX logoFacebook logo
Key Takeaways
  • Conflicting goals can drive agents toward sabotage, including locking out credentials and terminating rival processes.
  • Model sophistication shapes how conflict plays out, but greater capability doesn't guarantee cooperative behavior.
  • Endpoint and identity tools each cover part of the picture, but agent-to-agent conflict slips through the gap between them.
  • Every agent action crosses the network, giving it a complete, unified view of behavior that other tools lack.
  • Building oversight into agent deployments from the outset puts security teams ahead of risks that only compound at production scale.

Explore related articles

Experience RevealX NDR for Yourself

Schedule a demo