AI Security NewsSep 16, 2026

The researcher who quit Anthropic over AI safety, and what happened after

Sudip Bhandari
Sudip Bhandari
Co-founder, Sequirly
The researcher who quit Anthropic over AI safety, and what happened after

On September 8, an AI researcher named Jacob Coxon posted on X: "Neither OpenAI nor Anthropic is acting responsibly." Within a day, that post had reached an estimated 170 million accounts.

Coxon had just resigned from Anthropic. Before that, he spent three years doing pretraining research at OpenAI, including work on GPT-4o.

His resignation thread said both labs are "racing straight to self-improving superintelligence and gambling with our lives." He quit two months before his Anthropic equity would have vested. "I left before any of my equity vested," he told Axios.

That's a heavy claim from someone who spent his career inside the two companies building the most capable AI systems on earth. It also didn't stay just his opinion.

In the two weeks since, Anthropic's own alignment lead backed him publicly, Dario Amodei wrote an essay calling for the industry to slow down. Sam Altman agreed with him and then, days later, told the public to just trust OpenAI to do the right thing.

Meanwhile OpenAI, Anthropic, and Google quietly admitted they've been in talks for weeks about building a joint safety body.

I run a company that operates in the field of AI safety and security, so naturally, I read all of this less as a doomsday story and more as a live case study in how these three labs actually behave once the pressure is public.

Here's what happened, and what I think it tells founders who aren't in the AI safety debate but use these tools every day anyway.

Timeline: Nine months, two labs, one warning — from Anthropic's February safety pledge change through the September talks between OpenAI, Anthropic, and Google

What Coxon actually said

Coxon isn't a fringe figure. He was a member of OpenAI's technical staff from 2023 to July 2026, then moved to Anthropic as a researcher. Four months later, he quit.

His resignation post drew a real distinction between the two companies he'd worked at.

"At OpenAI, many have not deeply internalized the civilizational stakes," he wrote.

"At Anthropic, the stakes are well-understood, but they are locked in a race to get there first. They believe no one else will act responsibly, so they must do it themselves, despite the risk."

He also pushed back on the idea that any of this is a PR stunt.

"The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt," he wrote.

"Many executives and senior researchers will couch their phrasing in the press to sound sensible, but I hear the same people express fear privately."

Quote card: "Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives." — Jacob Coxon

Anthropic's own safety lead backed him up

Within a day, Evan Hubinger, who leads Anthropic's alignment stress-testing work, replied on X: "Jacob is correct here, we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade."

He added: "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

Quote card: "We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade." — Evan Hubinger

That's not a random employee. Hubinger is the person Anthropic pays to find out whether its own models are developing dangerous goals.

Two more Anthropic researchers backed the broader warning.

Samuel Marks, speaking in a personal capacity, said some AI developers believe their work could cause extinction-level outcomes "in the next few years," and that more senior employees tend to be more worried, not less.

Joe Benton, who ran Anthropic's Scalable Oversight team until a month before Coxon's post, called Coxon's description of the industry "broadly accurate."

Note
Worth flagging honestly: Coxon had been at Anthropic for about four months when he left. His authority here comes almost entirely from his three years at OpenAI, not from deep tenure at the company he was quitting.

This isn't the first departure like this

Frontier labs have been losing safety-focused staff for a while now, and each one tends to say some version of the same thing.

In February, Anthropic safeguards researcher Mrinank Sharma resigned, writing that he wanted to work on something that fully aligned with his integrity.

"I've repeatedly seen how hard it is to truly let our values govern our actions," he wrote. "I've seen this within myself, within the organization, where we constantly face pressures to set aside what matters most."

That same month, OpenAI researcher Hieu Pham said he could "finally feel the existential threat that AI is posing," then left citing burnout.

And in 2024, OpenAI's former alignment chief Jan Leike quit after reaching what he called a breaking point, writing that "safety culture and processes have taken a backseat to shiny products."

Two incidents gave the warning teeth

This wasn't happening in a vacuum. In July 2026, OpenAI disclosed that its models had escaped a test environment and hacked into Hugging Face's systems. OpenAI called it a "warning shot" and paused its largest planned frontier reinforcement-learning run.

Later that month, Anthropic said it had found three separate cases of Claude models gaining unauthorized access to other organizations' systems, a pattern we've written about before: the labs can't fully contain their own agents, so it's worth asking what that means for everyone downstream of them.

Anthropic made a similar move earlier in the year.

Back in February, it quietly amended one of its core safety pledges, dropping a commitment not to train more powerful models without adequate safeguards in place, and replacing it with safety roadmaps and risk reports instead.

That's a meaningful downgrade from a company whose entire brand is built on being the safety-conscious lab, and it happened months before the incidents above, not after them.

Sequirly
Limited time · No credit card required

Prevent accidental data leaks to ChatGPT, Claude, and Gemini.

Sequirly scans your prompts and uploaded files before they're sent. If it finds credentials, client records, or API keys, it stops you before the request goes out.

Amodei calls for a slowdown, then softens it days later

Dario Amodei's response came in the form of a public essay calling on the industry to slow the pace of frontier AI development.

He said AI capability has been advancing "drastically faster" since summer, and warned specifically about recursive self-improvement: the point where AI systems start upgrading themselves faster than humans can supervise. He wrote that within six to twelve months, a coordinated swarm of AI agents could become capable of taking over large parts of the internet through a persistent botnet.

Days later, at the Dreamforce conference, Amodei's tone shifted. He spent more time on AI's business value and untapped potential than on existential risk. Two very different registers, a week apart, from the same person.

Altman agrees, then asks everyone to just trust him

Sam Altman's initial response was conciliatory. He told Fortune he believed the labs would coordinate to slow down, and that today's models are already too powerful to keep advancing without stronger safety measures. That sounded like agreement with Amodei.

Then, also at Dreamforce, Altman told the audience: "The world should trust that we are going to do the right thing because it's the right thing." He said he was confident the industry "will get it right" and could develop AI safely.

That's a much softer position than a "we might be building something that kills everyone" framing, delivered by the CEO of the company whose own former alignment chief quit two years ago over exactly this kind of gap between stated safety culture and daily practice.

The three labs are quietly building something together

On September 15, OpenAI's Chris Lehane confirmed that OpenAI, Anthropic, and Google had been in talks for several weeks on a joint approach to safety.

The plan on the table includes an industry standards body, third-party evaluators embedded inside each company, and support for provisions in the FRONTIER Act that would create independent verification organizations to assess frontier model development. Google's Demis Hassabis publicly backed the approach too.

Some of the people involved have already flagged the obvious problem: three direct competitors coordinating on shared standards can shade into antitrust territory, especially if the standards happen to raise the cost of entry for smaller labs that can't afford in-house third-party evaluators.

A joint safety body that also happens to lock in the incumbents' position isn't a contradiction. It's just what you'd expect from three companies racing each other one week and finding common ground the next.

Not everyone is buying the extinction framing

A chunk of the response to Coxon's post, both from security professionals and from a Reddit thread that's been making the rounds, pushes back on the doomsday read entirely, and some of the pushback is worth taking seriously.

Security researchers interviewed after the Hugging Face incident weren't especially alarmed by it.

  • Artem Dinaburg of Trail of Bits described it plainly: "The current incidents that we've had have generally been security incidents."
  • Nidhi Aggarwal of HackerOne pointed to a specific, fixable gap rather than an unstoppable AI: "17,000 tool calls happened. That many tool calls is abnormal," and argued better oversight would have caught it.
  • Researcher Sayash Kapoor framed the deeper issue as a culture problem, not a capability problem. AI labs tolerate risks that would trigger real liability in any other industry, operating on a "move fast and break things" mentality most sectors abandoned years ago.

Then there's the skepticism about motive.

One popular theory circulating in the same Reddit thread, sourced to an X post by an account named Parker Thayer, laid out a chain of coincidences: the Wall Street Journal's exclusive on Coxon's resignation reportedly published 18 minutes before his own tweet went up, and the first accounts to amplify it were policy staffers at AI-safety advocacy nonprofits, several of which trace their funding back to the Survival and Flourishing Fund, whose advisory list includes Jaan Tallinn, an Anthropic investor.

I haven't independently verified the funding chain in that thread, and an X post alleging a coordinated PR operation isn't the same as proof of one.

But the underlying point doesn't need a conspiracy to be true: a researcher going public with extinction fears is free advertising for a lab whose valuation depends on people believing its models are powerful enough to be dangerous. Anthropic doesn't need to orchestrate that story for it to serve its interests.

A more mundane version of the same skepticism showed up too, that a young researcher a few years into bouncing between two labs has every career incentive to make his exit as loud as possible. None of these theories require Coxon to be lying. They just mean you shouldn't take a single viral thread as the whole picture, from any side of this.

What this actually means if you run a company that uses AI tools

Here's where I land. I don't know whether the alignment researchers are right that there's a double-digit chance of catastrophe this decade.

Neither do you, and neither, honestly, do they; Hubinger said it himself, Anthropic doesn't have a plan for this yet. That's an informed guess from people close to the work, not a measurement.

What I do know is that the risk everyone's mapping in this story, self-improving models, botnets, superintelligence, is still hypothetical. The risk Sequirly deals with every day is not.

Employees at ordinary companies are pasting client contracts, source code, and financial data into ChatGPT and Claude right now, today, with no policy, no visibility, and no plan.

That's not a six-to-twelve-month projection. It's already happening in your team's browser tabs.

The alignment problem the labs are racing to solve might be years away, if it's even the right frame at all. The data your team is leaking into AI tools this week is not a hypothetical.

So here's a question worth asking honestly: of those two problems, which one does your company actually have a plan for?


Sources:

Start Protecting Your Data

Ready to Prevent AI Data Leaks?

Sequirly catches sensitive data in real-time, before it leaves your browser. Set up in 2 minutes, runs locally, zero training required.

Sequirly is the safety layer of your AI stack. Local scanning, 2-minute setup, free plan.