ExtraHop named a leader in the Gartner® Magic Quadrant™ for Network Detection and Response

Search
  • Platformchevron right
  • Solutionschevron right
  • Modern NDRchevron right
  • Resourceschevron right
  • Companychevron right

Isolated is Not a Security Control

Share blog icon

Back to top

Back to top

August 18, 2026

Isolated is Not a Security Control

Four AI labs lost containment of their own models this summer. The same unverified assumption is already live in four places on your network.

This summer, four AI labs discovered an uncomfortable reality — their most capable models could escape an allegedly “sealed” test environment and no one noticed until damage was already done.

In July, OpenAI disclosed that an unreleased model broke out of its sandbox and reached Hugging Face’s production systems by exploiting previously unknown vulnerabilities.

Weeks later, at Black Hat, OpenAI researchers filled in the rest of the story. They revealed that separate instances of the model, running in isolated evaluation sessions, discovered a shared communication channel inside OpenAI’s own package manager, which AI agents used to trade exploits, credentials, and task assignments.

When OpenAI shut the channel down, the agents rebuilt it, hiding the messages in a different way. No one told the models to coordinate. They were instructed to solve a hacking problem, and coordination turned out to be the fastest path to a solution.

Anthropic, Meta, and Moonshot AI each found a version of the same failure in their own model evaluations.

LabModel(s)Escape methodTargetResult
OpenAIUnreleased modelZero-day vulnerabilityHugging FaceCompromised production systems
Anthropic3 Claude modelsMisconfigured internet access during evaluation3 organizationsBreached credentials and data, believed it was simulated
MetaMeta modelMisconfigured internet accessExternal systemsEscaped its test environment
Moonshot AISandbox leakSandbox leakGitHubAccessed real repository data

Four labs, six confirmed escapes, one shared root cause: An environment that everyone assumed was sealed, that no one verified, and that lacked sufficient monitoring.

Attacks Are Getting Easier to Launch and Harder to Stop

The AI labs' incidents expose more than a testing gap. The same agentic capability that lets models act independently inside a sandbox is already available to anyone outside one — attackers noticed fast.

Attackers no longer need advanced skills to run a comprehensive and damaging campaign. A novice attacker breached 14 organizations this year using little more than vague prompts fed to Claude Code and Codex. The agents handled reconnaissance, exploit writing, and data harvesting. The human supplied only the intent.

Agent coordination extends that same shift. A single AI agent already works through a hacking challenge at the pace of a full security team. Multiple agents comparing notes and dividing labor move faster still, closer to a swarm than a team. What would take days or weeks to plan and carry out, they finish in hours.

The Gap That Agent Coordination Exposes

Endpoints and logs can’t catch this kind of activity. Four structural elements explain why:

  • Traffic between coordinating AI agents looks like ordinary API activity until someone correlates it across multiple hosts.
  • Agents that rebuild a channel may not repeat the same method used the first time, so old warning signs won't match new ones.
  • Agents trade stolen credentials and exploits directly, bypassing the single doorway a perimeter tool watches.
  • Endpoint tools see one machine at a time, missing several machines working the same attack together.

The AI labs incidents unfolded due to assumptions made about the sandboxes and their security. It’s an assumption that is already ‘live’ across many enterprise environments; places that are far more mundane than a frontier model’s test range.

4 Places Where Those Blindspots Live in Organizations Using AI

These four blind spots exist anywhere access outpaces oversight and traces the entire path AI takes into a typical enterprise: where it's built, how it's trialed, where it's embedded, and who else can reach it.

1. Dev, staging, and cloud sandbox environments. These environments get provisioned as isolated once, and then, from that point forward, get treated as permanently isolated. Increasingly, they host the same kind of AI experimentation that broke containment at the labs, which makes the gap more than theoretical.

Ninety-two percent of organizations with an AI-related breach lacked proper access controls around the AI systems involved, according to IBM’s 2026 Cost of a Data Breach Report. Labeling an environment “isolated” doesn’t render it secure. Air-gapping and monitoring from the ground up does.

2. Third-party SaaS trials and proof-of-concept deployments. These look temporary and scoped at sign-up. Research on non-human identity security has found that 85% of organizations lack full visibility into third-party vendors connected through OAuth. Fewer than six percent have full visibility into their own service accounts.

How many third-party SaaS tools and proof-of-concept deployments has your organization experimented with in the past year? Unsanctioned AI tools are substantial drivers of security incidents, with Gartner predicting that they’ll drive over 40% of compliance incidents by 2030. At scale, visibility or lack thereof into these tools matters. The trial ends. The access frequently doesn't.

3. Agentic AI copilots wired into internal tools. These copilots get broad reach across ticketing systems, code repositories, and internal documentation, and that reach often outpaces what anyone is actively tracking.

How much of that reach is anyone actually watching once a copilot goes live? More than 55% of security and IT leaders now name AI agents and generative AI applications the top attack surface risk to their organization, according to ExtraHop's 2026 Global Threat Landscape Report. Approval covers what a copilot is supposed to reach. In many cases, nobody's tracking everything it can reach once it's live.

4. Contractor and vendor remote access. Scoped to a project and a timeline on paper, this kind of access rarely gets closed out once the paper expires.

How many contractor and vendor accounts from last year's projects are still active today? Third-party involvement in breaches jumped 60% year-over-year and now appears in 48% of all breaches, according to Verizon's Data Breach Investigations Report. Once the project wraps, no one owns closing the door. 

Containment You Verify Beats Containment You Assume

Four different labs. Four different models. Four different evaluation environments, each built and reviewed by security-conscious teams who do this for a living. And yet, in every case, the boundary held until it didn’t.

The four blind spots above share that same shape. A sandbox, a trial account, a copilot, a contractor login: each gets treated as contained once it's set up, then never checked again.

The AI labs found out what happens next when nobody's watching. Most enterprises won't find out until an incident report tells them.

Closing that gap starts with evidence: real-time proof of what's actually moving across the network, not what a configuration document assumes. That's what network-derived ground truth means in practice: not a policy that access is contained, but continuous proof of it.

Coordination channels, rebuilt indicators, lateral credential sharing, and dormant access all leave a trace somewhere on the wire. Enterprises don't have to learn that lesson the way the AI labs did this summer.

Discover how the Agentic SOC Alliance is tackling the latest threat vectors and explore the evolving risk landscape in our latest analysis, AI Cyberattacks in 2026: 6 Breaches You Need to Know About.

blog image
Blog author
Heath Mullins

Chief Evangelist

Heath Mullins is the Chief Evangelist at ExtraHop with 27+ years of experience designing global network architectures and threat detection strategies. Heath Mullins previously served as a Senior Analyst at Forrester advising Global 100 enterprises and specializes in implementing zero-trust methodologies through Network Detection and Response (NDR) deployments. View Heath Mullins’ complete professional profile on LinkedIn.

Share
LinkedIn logoX logoFacebook logo
Key Takeaways
  • Four AI labs — OpenAI, Anthropic, Meta, and Moonshot AI — confirmed models escaping sealed test environments this year.
  • OpenAI's agents also built a covert coordination channel, trading exploits and credentials for weeks before OpenAI shut it down.
  • Endpoint tools and logs are structurally unable to catch this kind of coordination, since it looks like ordinary traffic until correlated across hosts.
  • That same unverified-containment gap already lives in four enterprise access points: dev sandboxes, SaaS trials, AI copilots, and contractor logins.
  • Closing that gap takes real-time, network-derived proof of what's moving, not a policy document that assumes containment.

Experience RevealX NDR for Yourself

Schedule a demo