WeeklyWorker

17.09.2026
Dangerous? Me?

Angel or McGuffin?

Jacob Coxon’s very public resignation from Anthropic has once again raised the spectre of the AI apocalypse. But which AI apocalypse? Paul Demarty weighs up the risks to humanity

The world of AI is once more in one of its conniptions - this time as a result of a young whistleblower by the name of Jacob Coxon.

Coxon, a Brit who has resigned from his job as an AI researcher at the US company, Anthropic, made public his severe concerns about the dangers posed by the present wave of frontier model development. For him, things are going too fast and, without a serious change in the culture or the regulatory environment, something very bad could happen. He notably asserts that his worries are shared by very senior people at Anthropic and its rival, OpenAI, where he worked for several years:

The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.1

Rather than trying to calm matters in the wake of Coxon’s post, Anthropic officials have tended to express agreement. Evan Hubinger, who is head of alignment research at Anthropic, blithely asserted that there was a greater than 10% chance that AI would end human life in the next decade. At the highest level, Dario Amodei, Anthropic’s CEO, has called for a coordinated slowdown in AI development - a call echoed by Sam Altman of OpenAI and Elon Musk, whose Grok model is a trivial player, but who has historic connections to OpenAI (and a beef with Altman).

One man certainly not on board with such lily-livered ideas is Donald Trump, who spent much of the weekend denouncing the slowdown as a “sick conspiracy”. With his usual inventive capitalisation, he told his TruthSocial followers that “The only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the USA has that, in spades!” Everything, for Trump, depends on beating China in the AI race, and is thus basically a military question.

In the background here is the so-called Hugging Face incident, in which a thousand or so AI agents (spun up as part of a security exercise by OpenAI) broke out of their sandboxed infrastructure and began to coordinate a cyberattack on Hugging Face, which is a major model-hosting service. This incident has been extensively discussed and analysed, since it provides a rather stark indication that the AI labs are not entirely in control of their creations.

Nonetheless, a phrase jumps out from Coxon’s post: “This is not a marketing stunt”. He is only required to put this disclaimer in because, frankly, part of the big AI labs’ marketing in recent times has been to issue strange apocalyptic warnings, thereby amplifying the hype about the sheer power of their creations. The term ‘doom-trolling’ has been created to describe this rather eschatologically inclined PR strategy. This does not mean that these people do not believe it. There is a tendency in American business circles to drink one’s own Kool-Aid. It does, however, rightly call forth a certain scepticism, and indicates that we should proceed with some caution in evaluating the doomsday scenarios.

Two scenarios

In this light, I want to distinguish two different kinds of AI apocalypse. The first is the ‘exterminating angel’: in this scenario, which seems to vex Coxon and others, AI achieves a kind of sentience and, having done so, self-directs a sequence of actions to do away with humanity. This is the scenario which dominates the ‘AI risk’ world, as it currently exists - think of the setup of the Terminator movies (mentioned by Coxon in interviews).

The alternative is something I wish to call the ‘exterminating McGuffin’. The reference is to the cinematic work of Alfred Hitchcock, who gave us the concept of something that, though more or less empty in itself, moves the plot along.

Hitchcock’s suspense movies frequently relied on this device: the unseen secret agent for whom Cary Grant is mistaken in North by northwest, to pick my favourite example. In the AI context, it is very easy to imagine scenarios where the use of AI in the military could set in motion a chain of events that results in generalised nuclear exchange. This would certainly be an AI apocalypse of a sort. Yet the real actors in this drama would not be the AI agents (or whatever), but the governments of the US, China and so on. The AI is an expedient - a McGuffin.

When a very public incident like Hugging Face takes place, it rightly causes alarm. But everything hinges on what kind of alarm. The response of the industry - coordinated slowdown of research - implies the exterminating angel scenario. Coxon, and many others in the industry, are worried about ‘recursive self-improvement’ as a threshold event, where AI systems can train themselves and rapidly take over the world. The whole business of ‘alignment’ is to ensure that when this happens (Coxon is surely right that it is a matter of ‘when’, not ‘if’, in the minds of the AI researchers) the AI is ‘aligned’ with human values, and we do not get strange outcomes like the famous ‘paperclip maximiser’ thought experiment of Nick Bostrom, where an AI directed to improve the efficiency of a paperclip factory ultimately exterminates the human race in order to produce more paperclips.

The trouble is that this is not really the Hugging Face story. The agents running in this incident were attempting to complete a ‘capture the flag’ (CTF) exercise: basically a game in which the object is to compromise a server under a set of conditions. CTF is a common activity among human security engineers and hackers, who are by nature mischievous eccentrics and love to find interesting ways to get around problems. Their job is to find vulnerabilities and fix them, and so their methods are as close to those of the cyber-criminal as those of the surgeon are to the torturer.

These agents will have been trained on a large corpus of the past activities of these people. As Cory Doctorow pointed out,2 nothing they did is in any way novel. The use of a message board to communicate is a very well known workaround. As for breaking out of the sandbox (ie, the isolated computing environment), clearly it was a poorly designed and specified sandbox, and hackers know how to get out of those, and that knowledge is certain to be somewhere in the parameters of the large language model (LLM). Even the most concerning version of the breakout - say, the agents discovered an unknown and unpatched bug in the Linux kernel and exploited that - would still not be novel in the relevant sense. The techniques for finding such vulnerabilities are known.

To be blunt, there is no indication here of a qualitative leap in the capabilities of LLMs. Unless really compelling evidence to the contrary emerges, this is not an exterminating angel, but an exterminating McGuffin. Getting this the right way round is important. After all, if we really were on the threshold of giving birth to a machine-god, then the Amodei approach - let the industry collude to slow things down and take tighter control of the technology - roughly makes sense (although so would just not building it). On my interpretation, however, the right course of action is surely punitive fines for elementary information security failures, decertification, or even prosecutions under the Computer Fraud and Abuse Act.

Control vital

Certainly the basically unregulated way in which the industry is allowed to proceed presents grave dangers. OpenAI is guilty of very common bad practices in the software industry: in the worst case, this kind of thing really has led to software bugs, which crashed planes and killed people with lethal radiotherapy doses.

Partly as a result of such incidents, far more stringent controls are in place for software running in such life-or-death contexts. And, given how much basic infrastructure of human existence is exposed to the internet in one way or another, there is every reason to believe that the next oopsy-daisy in OpenAI or Anthropic’s security research divisions could come with a body count. It is a standing indictment of US capitalism in its present decrepitude that there is essentially nobody for these bizarre companies to answer to.

It is this, more than anything else, that creates the ‘doom-trolling’, and ensures that we have a useless discussion about AI-god scenarios, instead of a meaningful debate on how to bring this industry under control. Paradoxically, by amplifying the sense of threat in a science-fictional direction, the Amodeis and Altmans of the world excuse themselves from real scrutiny.

It is a very sensitive time for them. Anthropic is expected to go public in the near future, though OpenAI has ruled out a share sale this year. These firms are fantastically unprofitable, and public listing will forbid them from using opaque accounting practices to appear healthier, as they presently do. A slowdown would benefit them, somewhat reducing the enormous levels of capital expenditure required to train their models, and possibly putting them in the black by way of cost-cutting.

Capital expenditure is not exactly going well either. There is ferocious local opposition all over the United States to new data centre construction, from both left and right. Perceived closeness to the AI industry is a liability for political candidates. Even where such projects are begun, they take years to complete. Financing construction has become a severe problem, which has led to dubious and highly-leveraged circular financing arrangements. Almost all equity price growth in recent years has gone to companies like Nvidia, Microsoft and Google, which are central participants in the AI boom. A serious market correction would have severe and possibly systemic consequences. Of course, as the well-known AI sceptic, Ed Zitron, points out,3 “pacing the frontier” itself might cause such a correction, since somebody will be left holding the bag on all the data centre debt and mountains of unused hardware. (Nvidia’s Jensen Huang seems notably chilly about the whole thing, as well he might.)

The key question, again, is control. Trump, in this respect, is more serious than Coxon - he may massively overestimate his abilities, but someone must take this madness in hand. I tend not to go in for ‘Marx was right’ gotchas, but this whole situation seems almost designed to illustrate the fundamentals of Marxism: the laws of political economy drive the concentration of capital and technological competition; the firms are under the spell of mute compulsion; human creations are projected over against their creators, appearing as mystical forces or gods.

Overrepresented in the AI industry are today’s ‘rationalists’ - followers not of Spinoza, but the holy fool, Eliezer Yudkowsky - who have convinced themselves that it is rational to act as if artificial superintelligence is inevitable and that this entails working to bring it about, and that the calculation of the probability of human extinction - or, as they like to call it, ‘p(doom)’ - is a philosophically serious endeavour and not, at best, a bong-hit parlour game. If you asked Claude to write a short story illustrating Marx’s theory of commodity fetishism, the characters in it would be rationalists.

Any real solution has to start with expropriation. The for-profit AI industry has produced some useful software, as well as a lot of obviously harmful software, in pursuit of a big payoff which strangely never comes. It is the need for that payoff which has led to all this corner-cutting and, more ominously, an obviously unsustainable edifice of rickety financial instruments. The broad family of software we call AI no doubt has much to offer us in the future, but also probably the other side of much taxing and unprofitable fundamental research. Socialisation is the minimum condition for this: hinging the world capitalist economy on AI, led by members of a borderline cult, is sheer madness.


  1. Thread beginning at x.com/hilbertspaess/status/2097476196791709843.↩︎

  2. pluralistic.net/2026/09/12/god-in-the-box.↩︎

  3. www.wheresyoured.at/ai-is-already-in-dangerous-hands.↩︎