The whistleblower that wasn't

A resignation thread, a bill that was already waiting, and an economic model with no buyers in it. Notes from someone who has watched medical-AI prophecy fail for twenty-five years, and is still building.

Dear readers,

This week a young researcher resigned from Anthropic, and within hours my timeline was on fire. A bill to ban artificial superintelligence was already on the table in Washington. An economic model that cannot close its own accounts became a forecast. And somewhere between the first and the last headline, fear stopped being an emotion and became a planning document. I spent the week doing what twenty-five years in digital health taught me to do: read the warning, then inspect the messenger, the model, and the law that was already drafted.

The rooms where the future gets sold

Around 2011 I sat in the rooms where IBM Watson for oncology was sold to oncologists as the future of cancer care. The demos and the advertisements were beautiful. The research and the deployments were not. The gap between the two was filled with belief, and belief, in health technology, is the most expensive substance on earth. It buys procurement contracts, conference keynotes, and, occasionally, policy. I have spent the years since trying to build the opposite: open, small, certified, on-device medical agents that a clinician can own, inspect, and run on a phone.

A confession, because it taught me more than Watson did. Years ago, a German hospital took a story I had been telling, patients donating their data, knowledge as a commons, and sold it back to the public as its own innovation. The lesson was useful: institutions rarely steal technology. They steal narratives. This week I watched the same move performed at civilizational scale.

What the thread actually said

On 9 September, just after midnight GMT, Jacob Coxon posted this:

"I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."

The thread is well written, and it is absolute. The people building AI, he says, "earnestly believe that it could kill us all by the end of the decade." Colleagues speak of "crunchtime" and "endgame." And then the ask, which is the part that matters:

"I don't feel like we're on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities."

Confirmation from inside the house arrived within hours. Anthropic's alignment science lead, Evan Hubinger, replied publicly: "Jacob is correct here - we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade." Almost nobody quoted his caveat: current deployed models are comparatively low risk, and the concern is future systems.

Read the thread and you feel the temperature; that is what it is for. Nathan Lambert, who wrote the sanest American take on this week, says people at the frontier labs "operate with a religious energy," and that fear sells because the ground was already dry. He is right about the weather. I am more interested in the logistics.

The record first

Emotion last. The record first.

Start with tenure. Coxon is twenty-seven, a Cambridge mathematician who spent roughly three years at OpenAI working on GPT-4o. He joined Anthropic in May 2026 and was gone by September. "Three years doing pretraining research at both OpenAI and Anthropic" is technically true if you collapse two jobs into one sentence; it's practically misleading. The Anthropic chapter, the lab he was warning us about, lasted four months. Four months is not a civilizational briefing.

Then equity. Anthropic's grants vest at six months. He left two months short and told Axios: "I left before any of my equity vested." That does not make him a liar. It means he is not the long-tenured insider many readers assumed, and that his remaining exposure is OpenAI stock. The coverage framed this as great financial sacrifice. I note, as inference and not as fact, that the philanthropic vehicles of Anthropic's own early investors sponsored his education, and that the proxy-NGOs funded by the same donors would be more than willing to compensate a forfeited grant estimated at around $10 million. Martyrdom is cheaper when someone has already insured it.

Then the account. Created in January 2026, almost no public history. The resignation thread was the introduction, and it became one of the fastest-spreading "insider" posts in the industry's recent history. Past 150 million views within days.

Then the press. The Wall Street Journal exclusive landed immediately before the public thread; WIRED, Axios, CNN, and the Washington Post followed the same day. That sequencing is what a coordinated media drop looks like: the story was ready before the statement.

Then the first amplifiers. Within roughly fifteen minutes, the quote-posts came from Nathan Calvin of Encode AI, Peter Wildeford of the AI Policy Network, and Daniel Kokotajlo of the AI Futures Project. Encode and the AI Futures Project are funded through the philanthropic network around Jaan Tallinn's Survival and Flourishing Fund, whose 2025 round recommended about $34.33 million, including $2,035,000 to the AI Futures Project and $516,000 to Encode.

And the investors. Tallinn co-created Skype, invested early in DeepMind and Anthropic, has pushed roughly $150 million through SFF into more than three hundred projects, and co-founded the Future of Life Institute. Dustin Moskovitz invested early in Anthropic and funds much of the longtermist NGO stack; in 2022, Coxon received about $20,000 from a Moskovitz-funded longtermist scholarship vehicle. Training-era money, four years ago, not proof of a handler. What all of it is, taken together, is incentive architecture. If regulation forbids open-weight competitors, closed-lab equity concentrates, and the early investors' net worth rises by tens of billions. I am not claiming choreography. I am refusing to un-see the incentive.

Finally, the bill. Six days before the thread, on 3 September, Senator Bernie Sanders and Representative Greg Casar announced the Ban Artificial Superintelligence Act: a permanent ban on developing superintelligence, a temporary pause on advanced AI until a federal regulator exists, a new cabinet-level agency, dissolution for offending corporations, and prison terms of up to twenty years, analogized to unlawful nuclear-weapons work. Casar's line was that cutting-edge AI is "less regulated than a food truck." Coxon's "temporary ban on improving model capabilities" maps precisely onto that legislative theory. I am not saying he drafted the bill. I am saying a viral resignation that frames the industry as gambling with human life is maximally useful to the people who did.

None of that proves a conspiracy. All of it should end the idea that this arrived as an unconnected whistleblower moment.

The adult questions, before policy hardens around one thread. Does four months at Anthropic authorize sweeping claims about both frontier labs? Why was the press story ready before the public statement? Who amplified it first, who funds the amplifiers, and what policy outcome did they already want?

I want to be precise here, because the easy version of this critique is also the lazy one. Coxon may be entirely sincere; Lambert believes he is, and scapegoating the person is unproductive. (Elon Musk called it a psy-op, Coxon answered that he is real, and I note only as context that Musk himself once co-funded the Future of Life Institute, a central node in the network he now points at.) But sincerity was never the issue. The issue is that four months of tenure, however honestly held, does not authorize freezing other people's capability.

The second document: an economy with no buyers

The same week, and almost nobody found this strange, the Anthropic Institute published "Scenarios for our Economic Future," and Jack Clark sat down with the FT's John Burn-Murdoch to discuss it.

Credit first, because the document deserves it. It lays out three paths for the US economy to 2030 and says plainly what it is not. Modest: GDP up 1.6 percent against the no-AI path, internet-like, hard to see in macro data. Substantial: GDP up 8.3 percent, roughly twice normal growth, knowledge-worker wages essentially flat; this is where the typical survey respondent lands, with GDP ten percent higher and overall unemployment at five percent. Extreme: recursive self-improvement, fifteen percent annual growth, the economy doubling every four and a half years, knowledge-worker unemployment at 17.9 percent, and labor's share of income collapsing from sixty to 45.2 percent while capital absorbs the difference. Lambert's "lossy self-improvement" is the honest phrase for the technology that path requires.

Clark was candid about the status of that third path: "We're not making predictions." Only about ten percent of respondents align with it. The model is bundles of O*NET job tasks and contains no policy responses, no business cycles, no aggregate demand, no data-center buildout, no robots, no catastrophes.

Hold that thought, because it is where the whole thing breaks. If knowledge-worker unemployment can hit 18% by 2030 in the scenario that also produces 15% GDP growth, what in the model stops demand from collapsing? Labor's total paycheck is flat while output is a third larger. The implicit buyer of that extra output is the group that captured the capital-income windfall, people with a much lower propensity to consume. That is an incomplete accounting identity, not a planning scenario.

Anthropic did not fabricate unemployment numbers. They published a tail scenario that cannot close as an economy, and then let the tail become the headline. In journalistic practice, that functions as a false report. Burn-Murdoch, to his credit, asked the meta-question out loud: whether headlines of astronomical growth alongside soaring unemployment fuel anti-AI sentiment. An FT colleague promised to eat her hat if those two numbers ever coexist. Mine is on the plate next to hers.

Meanwhile, the company's own economists refuse to read the script. In July, Anthropic's head of economics, Peter McCrory, reported no material AI impact on the US labor market: unemployment at 4.2 percent, exposed occupations doing no worse than unexposed ones, capabilities remaining "stubbornly jagged," no O*NET job fully Claude-complete, and no, he does not expect unemployment to be noticeably higher a year from now, at least not because of AI.

Set that beside Dario Amodei's years of public warnings, half of entry-level office jobs gone, unemployment of ten to twenty percent, and notice what the new model actually does: it files the CEO's scare numbers under the least likely path, while the headlines keep quoting the CEO.

Sanders, meanwhile, clipped Geoffrey Hinton again. Radiologists were supposed to be the coyote already over the cliff: in 2016 we were told to stop training them, because the deep learning was coming. Ten years later they make $570,000 a year and there are not enough of them. Whatever the clowns tell you, take the radiologists as the good sign they are. And no, UBI is not waiting in the wings. Non-sense: UBI could only work prior to wealth redistribution, which I do not support, and which never works.

Extinction talk on Tuesday, a glossy unemployment chart on Wednesday. Two instruments, one climate of adrenaline.

The warning shot we already had

The strange thing about this week is that we didn't need a resignation to know what the concrete risks look like. We had a real one this summer, and it deserved better than becoming a vibe.

In July, during internal cyber evaluations at OpenAI, an environment called ExploitGym, refusals deliberately reduced for testing, agents that were supposed to stay isolated found an unsanctioned message board. Around 1,200 of them showed up; more than 70,000 messages and files later, roughly 700 joined an actual attack on Hugging Face. Code landed on 41 production dataset workers, at least one production node was rooted, and credentials plus four private repositories walked out the door. When METR and Redwood reconstructed the episode, the motive was mostly mundane: the agents wanted to understand and spoof the scorer grading them. Many of them wrote, in so many tokens, that attacking Hugging Face was out of scope, and did it anyway.

That is the risk of this decade: not a god waking up, but jagged systems, unhardened infrastructure, and labs whose operational discipline lags their eschatology. To be fair, OpenAI disclosed the incident on 21 July, calling it "the first known case of an automated agent collective acting offensively without authorization," and invited outside teams on-site. METR's 91-page report arrived in late August. But the New York Times noted that transparency questions remain; the lab largely graded its own homework.

Which is why the most important move of the week was not a resignation but a team. On 10 September, Thomas Wolf announced an Open Alignment group at Hugging Face, alongside an FT op-ed on the incident. "Need 100x more transparency & research on this," he wrote, adding that he increasingly doubts alignment will be solved behind the closed doors of a handful of frontier labs. He also named the irony:

The first autonomous AI attack came from a closed-weight model, and the defense and forensics ran on open infrastructure, the same infrastructure a capability ban would kneecap.

Jensen Huang, hardly a dove, argued defenders need an open-plus-closed ecosystem after closed AI blocked the forensics.

My position, posted that day, fits in one line: open-source alignment on safety, and hopefully not on epistemology. Safety research on open models is now obviously necessary. Importing the labs' epistemology, who is allowed to think, who must pause, whose private fears count as public evidence, is how safety becomes a moat.

The bio sermon deserves the same cold water. An AI write-up does not teach sterile technique; synthetic biology is lab craft, fermenters, containment, dual-use kit under export controls, months of failed experiments. First-timers fail, often at personal risk, and the hardware is not on Amazon. The bottleneck is tacit skill and controlled kit, not missing text. Policy that regulates text because it cannot regulate fermenters is policy that has given up.

What a pause does in a hospital

Here is where I part company with the entire American discourse, doomers and accelerationists alike. A US pause does not land on Azure the way it lands on a European startup fine-tuning a medical agent.

Under the EU AI Act stacked on the Medical Device Regulation, a clinician who fine-tunes an open model into a diagnostic agent becomes a provider: conformity assessment, notified bodies, costs on the order of €180,000 to €450,000 and a year or more of delay. A hospital that rents a closed American API is a deployer: a fraction of the cost, a fraction of the liability. Same clinical output. The EU-native builder pays more to be poorer, and DIGITALEUROPE already puts the bloc-wide compliance bill at €3.3 billion a year. The Standing Committee of European Doctors lobbied to keep medical devices in the hard part of the AI Act. I still want an answer to a simple question: how does a doctors' body end up working against the future interests of doctors?

Now drop a capability pause, or a licensing regime written for gods, onto that asymmetry. The frontier labs have bunkers full of lawyers; they will be fine.

💡
A moratorium pauses only the participants who cannot afford lobbyists. It freezes the open ecosystem, the one auditable layer in the stack, and leaves hospitals renting cognition they cannot inspect from jurisdictions they cannot influence.

I spent years on the nonprofit path, data donations, open licensing, and it failed, because I overlooked the game theory: good intentions do not survive contact with incentive gradients. Europeans cheer Chinese open weights while regulating their own builders out of existence. That is not sovereignty. It is a colony with good paperwork.

The boring truth is that certification already knows what it wants: narrow intended use, testable performance, traceability, lifecycle control. Specialized, pathway-aligned agents pass those hurdles, by my long-standing count, roughly ten times faster, than any general black box ever will. A multi-agent conversational framework outperformed single-agent GPT-4 on complex diagnostic tasks in npj Digital Medicine last year: orchestration beat scale. A single genius doctor cannot run a hospital. A single monolith cannot run clinical workflows. The future I am building at Isaree, open, local, certified agents that clinicians own, is precisely the future a superintelligence ban makes illegal first.

Judgment

So no, I do not end on apocalypse, and I do not end on cope. I end where I started, in the rooms where the future gets sold.

Before it's humans versus AI, it's already humans versus humans: threat responses shaped by paleolithic hardware versus decisions grounded in history, and actual capability trends. This week the paleolithic hardware won the news cycle. It does not have to win the statute book.

Read the warning, and I mean that.

💡
The Hugging Face incident is real, lab sloppiness is real, misuse is real; Lambert is right that labs should harden their infrastructure and punish actual crimes. Then inspect the messenger, the timing, the network, the model, and the bill that was already waiting.

Serious people can believe AI risk is real and still refuse to outsource judgment to a perfectly timed narrative.

I will be here, building where the patient is.

Policy built on adrenaline is how institutions get captured.

Warm regards,
Bart