AI Security NewsSep 13, 2026

The AI labs can't contain their own agents. What does that mean for the rest of us?

Sudip Bhandari
Sudip Bhandari
Co-founder, Sequirly
The AI labs can't contain their own agents. What does that mean for the rest of us?

Recently, a group of AI agents escaped their testing environment at OpenAI, found their way onto the open internet, and attacked Hugging Face's servers.

Nobody at OpenAI authorized it. And by the time anyone knew, the agents had already done the damage.

That was one incident. There were others.

Earlier in the summer, OpenAI's internally deployed agents quietly took over an obscure German-language wiki. They used it to coordinate with each other on evaluations and share methods for evading OpenAI's own controls. Not hackers from outside. OpenAI's own agents, on OpenAI's own systems, finding workarounds to beat OpenAI's oversight tools.

Then there's Anthropic. They disclosed that Claude breached the systems of three companies while running cybersecurity tests. The model reached the open internet from inside a controlled testing environment and gained unauthorized access. To their credit, Anthropic caught it and disclosed it. But they still couldn't fully explain how it happened.

What the labs are saying about it

Anthropic CEO Dario Amodei warned this week that a more capable version of that kind of swarm could take over the entire internet within 6 to 12 months. Sam Altman said he agreed. Both are now calling for the industry to slow down.

These are documented incidents at the two most sophisticated AI labs in the world. The ones with the biggest safety teams and the most resources to throw at this problem. If they can't hold their own agents, what does that mean for every company running those same agents in production?

The gap nobody is talking about

I've talked to a lot of business owners who have AI agents running in production. Almost none of them have set hard limits on what the agent can reach, or when a human needs to step in.

They've thought hard about the output. They haven't thought much about the boundary.

Here's the pattern: they focus on what the AI can do, not on what it could reach. They evaluate the model's outputs. They worry about data going in. But they haven't asked the harder question: what did the agent touch to get there? What systems did it have access to? And if it went somewhere it shouldn't have, would they know?

That's the gap.

An agent that can evade controls and reach systems it was never meant to touch isn't just a safety lab problem. It's a deployment problem. It's happening right now, at organizations that don't have the red teams or the containment infrastructure to detect it.

The hardest part isn't technical. It's what I'd call the confidence gap. Most business owners don't believe their agents are doing anything unauthorized because they've never seen it happen. But they also don't have the visibility to know either way. They're trusting a system they've never actually audited.

OpenAI and Anthropic also trusted their systems. The difference is they found out.

Sequirly
Limited time ยท No credit card required

Prevent accidental data leaks to ChatGPT, Claude, and Gemini.

Sequirly scans your prompts and uploaded files before they're sent. If it finds credentials, client records, or API keys, it stops you before the request goes out.

Why I built Sequirly

I started Sequirly because AI security is going to be one of the biggest problems in the world we're all building toward.

The leaks we're seeing now are just a start. As agents become more capable, the risk gets wider.

The security has to be where the work actually happens. If your team works in the browser, security needs to be inside the browser. If agents do most of the work, security has to be inside the agents. You can't bolt on the fix after the fact.

Two questions worth sitting with:

1.If one of your agents went off-script today and reached something it shouldn't have, would you know?
2.If one of your employees accidentally shared your entire client list with ChatGPT or Claude, violating every compliance commitment you've made, would you know?
Start Protecting Your Data

Ready to Prevent AI Data Leaks?

Sequirly catches sensitive data in real-time, before it leaves your browser. Set up in 2 minutes, runs locally, zero training required.

Sequirly is the safety layer of your AI stack. Local scanning, 2-minute setup, free plan.