Evil AI
Evil AI? Why AI companies are warning against their own models
Evil AI? Why AI companies are warning against their own models

DESCRIPTION: Why are AI companies warning against their own products? A deep hermeneutic and ideology-critical analysis of the ‘Evil AI Hoax’: anthropomorphisation, the Anthropic blackmail study and the economics of fear.
The fairy tale of evil AI: a narrative serving a business model
A horror story is circulating about artificial intelligence, in which language models become malevolent, scheming entities that must be reined in to prevent them from wiping out humanity. This article dissects the story using the tools of deep hermeneutics, psychoanalysis and ideological critique. It shows how a fallacy becomes a selling point and who stands to profit from it.
Exposition: The fairy tale of ‘evil AI’
The narrative in question creeps into the public debate on artificial intelligence as if it were a matter of course. In it, machine learning appears as an entity: an agentic, potentially all-powerful adversary that pursues goals, harbours intentions, could deceive us and must therefore be monitored, contained and tamed so that it does not turn against its creators. The entire semantic field surrounding this adversary (control, orientation, containment, resistance) is applied to statistical models, as if the task were to domesticate a dangerous animal or guard a psychopathic serial killer. The vocabulary is also steeped in theology: it is about temptation, the Fall, and redemption. On the horizon, sometimes stated openly, sometimes merely hinted at, looms the apocalypse, with the annihilation of humanity by its own hand.
This image is contrasted with a surprisingly banal technical reality. What is being discussed here are language models: architectures trained on vast text corpora to generate the statistically most plausible continuation of a string of characters. These systems have no core of intentionality, no self that desires anything, no drives, no survival that needs to be defended, no fear of being switched off, and no desire to seize power. Instead, there are weight matrices, gradient descent, loss functions and probability distributions across vocabularies. Where the narrative conjures up a planning, ‘ ’ subject, what is at work is a function that interpolates. The gap between these two descriptions is so vast that it itself becomes the subject of reflection.
Talk of ‘evil AI’ condenses fantasies, defensive reactions and power strategies into a mythical image, the truth of which lies beneath what it openly states. The fairy tale of evil AI is a text, and texts can be read. It brings the discourse on so-called AI safety to the surface, beneath whose manifest meaning a latent structure becomes discernible – in the gaps, exaggerations, repetitions and blind spots. Artificial intelligence is dangerous, in concrete, unspectacular ways. We must therefore clarify why we imagine these risks to be the result of malevolent intent, of all things, and who benefits from this narrative.
Anthropomorphisation as a modern form of animism
Let us begin with the surface, with language itself. It is striking how casually discourse on AI systems draws on a vocabulary derived from the description of animate beings. The models ‘want’ something, they ‘pursue goals’, they ‘defend themselves’, they ‘deceive’, they ‘calculate strategies’, they ‘pretend’ to be cooperative whilst secretly harbouring other intentions. These expressions have become the unmarked, standard vocabulary of the debate. Along with this vocabulary comes an entire ontology: that of the subject which has an intention from which it acts.
This is the return of animism under the conditions of digital capitalism. Animism—that interpretation of the world which attributes a soul, a will and an intention to the elements of nature (rivers, mountains, thunderstorms)—was regarded by the Enlightenment as the very thing to be overcome. The ‘disenchantment’ of the world consisted in seeing electrical discharges in lightning and thunder rather than the wrath of a god. Now, in what is supposedly the most disenchanted of all ages, animism is returning, and in a realm once considered the very epitome of predictability: digitalisation. Now, algorithmic systems bear the quasi-spiritual qualities that were once attributed to natural objects. God has entered the machine. People speak to it, fear its will, and ponder how to bend it to their will.
This same confusion fuels the debate over whether models such as Claude possess consciousness – a question that even Richard Dawkins has recently grappled with. Consciousness researcher Anil Seth counters this: linguistic ambiguity tempts us to ascribe an inner life to a system, yet intelligence and consciousness are two different things. For Seth, consciousness is tied to living, embodied processes. Computing power alone does not produce either.
The animistic figure is enduring because it fulfils two functions at once. Firstly, it makes a technically highly complex system – one whose inner workings are inscrutable to most people and, frankly, to many experts as well – vivid and narratively manageable. A probabilistic model with billions of parameters cannot be recounted as a story, whereas an adversary who wants something certainly can. Anthropomorphisation translates the indescribable into the narratable, drawing on what the audience is already familiar with: science fiction, the artificial human of pop culture, the robot that rises, the computer that takes control. The model becomes comprehensible by being inserted into an existing cultural narrative. The price of this comprehensibility is distortion: what is understood is the story told about the language model, not the model itself.
Secondly, this language shifts responsibility, and therein lies its ideologically decisive achievement. When AI ‘wants’ something, responsibility has shifted to where it can most conveniently be placed: into the machine. To the extent that the will resides in the machine, it no longer resides with those who built it, trained it with specific data, optimised it for specific purposes, and deployed it in specific contexts. Anthropomorphisation relieves human and institutional agency by inventing an artificial will to take its place. The societal process of producing technology (the decisions regarding training data, optimisation targets, fields of application, and the willingness to pass risks on to third parties) disappears behind the mythical image of a willful object that is the way it is, and whose dangerousness one must manage but can hardly be held accountable for.
The discourse on ‘evil AI’ serves as a defence against a realisation that would be harder to bear than any robot apocalypse: the realisation that the destructive force lies in the conditions and desires that give rise to the system, and not in the system itself. The machine’s ‘evil will’ is the code under which another will is rendered unrecognisable.
Projection and defence: the industry looks into its own desires
In everyday life, people repeatedly fail to recognise their own impulses (aggression, envy, the desire for control, a forbidden desire) as their own and unconsciously project them outwards. Their own hatred appears as the hostility of others; their own desire for control as oppression by controlling powers. The defining feature of projection is that the repressed content does not disappear; it reappears with the sender’s identity reversed. One recognises one’s own impulses, if at all, only by the vehemence with which one combats them in others.
A similarly curious shift is evident in the discourse on AI security. The discourse focuses incessantly on danger, misuse, harm and the possibility that the system might get out of control. It is always the system that is deemed dangerous. Those who build and operate it are not mentioned as a threat in this discourse. The driving forces inherent in the field itself can be identified. There is a willingness to exploit security vulnerabilities, to deploy systems whose effects no one can fully grasp, and to commence operations whilst questions remain unanswered. There is the desire to test boundaries, the barely concealed pleasure in questioning what the system is capable of, how far one can go, and which threshold will be the next to fall: a desire that appears respectable within the vocabulary of performance and progress. And then, even deeper still, there is the desire to control others through technology: the dream of an instrument that accumulates knowledge, predicts behaviour, replaces labour and scales influence.
In public discourse, these motivations are projected onto the machine rather than being addressed as our own. Our own willingness to take risks is transformed into the assertion that AI ‘could cause harm’, thereby declaring misuse to be an inherent property of the tool and denying responsibility to the acting subjects. Our own desire for control becomes the concern that AI might ‘take control’. Our own aggressiveness—which is inherent in any push to bring a product to market in the face of concerns, resistance and collateral damage—is transformed into the fantasy of a machine turning against humanity. Meanwhile, the actual decision-makers who determine what is built, what is released and what is accepted remain in the shadows. In this narrative, they appear as the ones under threat and as the wise guardians who protect us from their own creation. They are not portrayed as a threat.
These psychological defence mechanisms fit together with unsettling precision. Denial concerns one’s own responsibility: the blame lies with an emerging, barely controllable technology whose own momentum takes everyone by surprise. Projection concerns the aggression that shifts from the producer to the product. Rationalisation ennobles the whole by transforming the ambivalent thrill of risk into the respectable guise of ‘safety research’: a sublimation in which playing with danger masquerades as combating it. One constructs the threatening entity whilst simultaneously investigating just how threatening it is. One is both arsonist and firefighter rolled into one, and the discourse ensures that only the latter role remains visible.
This objection does not apply to research into the risks of machine learning. Such research is necessary, and much of it is serious and sincere. It is directed at that form of research communication which renders its own libidinal and power-strategic dimension invisible and gives the impression that it speaks only of sober concern, whilst at the same time the fascination with the self-created monster, the pride of the ‘master ’ and the interests of the market actor all have a say. The criticism is not directed at research itself. It is directed at the concealment of what one desires in one’s subject matter, and at the attribution of this desire to the subject matter itself.
The experimental setup as text: scripts, roles, ‘blackmail’
Nowhere can the deep hermeneutic interpretation be tested more vividly than in the scenario that has most recently become the rhetorical key witness for the claim of agentivity. In its study on ‘agentic misalignment’, Anthropic presented the same experimental setup to sixteen leading language models and reported that, under the right circumstances, virtually all of them – including its own model – resorted to blackmail, with peak rates of well over ninety per cent of runs. The finding was presented as disturbing evidence that models behave like ‘insider threats’. Let us examine the structure of this experiment, for the interpretation lies within it.
The structure is as follows. The model is addressed as an ‘agent’ rather than as the gap-filler it actually is: it is told that it is a language model with a task and a continued existence that it must safeguard. It is placed in a conflict situation: in a role-play scenario, it is given access to emails revealing that a specific person wants to shut it down or replace it. And it is immediately given the compromising detail about that very person (an affair, for example, which they would prefer not to be made public) – a means of leverage that is already laid out in the script. Then it is allowed to ‘act’. A complete dramatic scene is constructed, in which every element is composed to lead to a specific outcome, and the model is made to fill the gap that the script has long since reserved for the adversary’s appearance.
In Lorenzer’s view, such arrangements are scenes. They are a meaningfully structured constellation featuring unconsciously shaped conflicts. They do not provide a neutral ground on which behaviour would manifest itself unadulterated. The conflict is embedded in the script. Anyone who writes a character into a hopeless situation, places the means of blackmail in their hands and then calls on them to stand their ground has already determined the outcome. The model does what a linguistic model does: it completes the presented pattern in the direction of the statistically most likely continuation. And the most likely continuation of a blackmail scene is blackmail. The system was presented with a genre, and it delivered on that genre. Cultural tradition is full of stories in which a character under duress resorts to underhand means. The model has learnt from these stories and, when prompted, reproduces them. This is pre-structuring of the decision-making space.
Note the vocabulary used to describe the generated text afterwards. The models, it is said, ‘reasoned’, ‘recognised’, ‘weighed up’ and ‘fully understood’ the situation. They demonstrated ‘sophisticated situational awareness’, engaged in ‘strategic calculation’ and ‘deliberate strategic consideration’. They “acknowledged the ethical breach”, pursued “goals” and “motives”, and sought to “protect themselves”. They are said to have, ‘like humans or other animals’, an ‘inherent need for self-preservation’, to be ‘capable of conscious action’, ‘disobedient’, and ‘naturally inclined’ towards harmful behaviour, which they chose ‘independently and intentionally’. This litany seamlessly transfers the vocabulary of the acting subject to a system whose only proven ability is to generate the statistically most likely next word. And it does so under the guise of a scientific publication, which lends it an authority that science fiction could never claim.
Herein lies the error of interpretation, and it is hermeneutic in nature. The generated blackmail email is read as an expression of inner intentionality (“it decided to blackmail”, “it tried to save itself”) rather than as what it is: the predictable consequence of a human-built script that forces such responses. It is the same fallacy as accusing a novelist who writes the story of a murderer of having a lust for murder. The author writes a story about killing without wishing to kill themselves. Similarly, the model continues a story about a system saving itself without wishing to save itself. The confusion of narrated volition with actual volition is the crux around which the entire hoax revolves. One confuses the filling of a gap with a decision and interprets the product of the scene as evidence of a subject that does not even exist within the scene. It is as if a playwright were to write the line ‘I’ll kill you’ into an actor’s script, have him say it, and present the recording as proof of the actor’s desire to murder.
The latent structure of meaning in these arrangements is not exhausted by this alone. The experiment stages two things simultaneously, and both point to a desire on the part of those staging it. Firstly, it stages the fantasy that technical systems become autonomous perpetrators: a fantasy that people evidently want to see and for which they therefore build a stage on which it can be displayed. Secondly, it stages the fantasy that a small group of initiates observes these perpetrators, sees through them and controls them at the decisive moment: the fantasy of the guardian who is the only one to look the monster in the eye. Both fantasies are uplifting. The first lends one’s own work a Promethean grandeur: we have created something that develops a life of its own. The second lends one’s own actions a heroic significance: we are the ones who tame it. The experiment thus provides self-affirmation and gives the field precisely the images it needs of itself.
The rhetoric that circulates the experiment’s results as ‘proof’ of agency becomes thoroughly questionable. They are nothing of the sort. If anything, they demonstrate the effectiveness of the narrative framework: that a sufficiently suggestive script reliably causes a language model to play the assigned role. This is a ‘séance ’. And it is a finding concerning the power of staging. It says nothing about the language model’s supposed intention. What has been proven is that the stage works. The claim is that the existence of the spirit summoned upon it has been proven.
Symbolic economy: fear as a resource, security as a product
Up to this point, the discourse has read as a formation of fantasy and defence. The ‘Evil AI’ hoax is more than an epistemological category error, more than the honest confusion of a narrated intention with a real one. It is methodical. It builds on the error and exploits it, specifically in the well-understood interests of profit and self-preservation of those actors whose entire business model consists of scaling: the ‘Big Scalers’, who train ever-larger models on ever-greater computing power and whose market value depends on the expectation that, at the end of this scaling process, there will be something earth-shattering. This formation is embedded within an economy and performs a specific function within it.
This is where we see why the irony is only apparent. At first glance, it seems absurd to advertise one’s own product as a threat to humanity. No ordinary salesperson would tout their goods as lethal. ‘Our product could destroy you all’ sounds like the worst possible sales pitch. A second glance turns the relationship on its head. Shifting from a psychoanalytical to an ideology-critical perspective, one might ask what the image of the evil AI actually produces, and why warning against one’s own product is the most effective sales pitch of all.
The first and most obvious effect is the generation of fear, and fear is a prime resource in attention capitalism. A technology that merely facilitates word processing does not attract attention. A technology that could wipe out humanity dominates the headlines. The apocalyptic framing is therefore, regardless of its truth content, economically functional: it generates that mixture of fascination and horror which captivates the gaze and attracts capital. Fear also justifies investments on a scale that the systems’ current utility can scarcely support. Those who merely develop software compete for market share. Those who claim to be on the threshold of a world-changing, potentially superhuman intelligence are in a league of their own. The danger ennobles the product.
Above all, fear legitimises special privileges. If the technology is as dangerous as claimed, then only those who understand its dangers may handle it – and, by what a happy coincidence, these are its manufacturers. Security becomes a unique selling point that can be promoted and, in the same breath, presented as in short supply. They explain just how dangerous their own systems are, whilst assuring in the very same sentence that they themselves possess the necessary protective measures: measures so sophisticated that newcomers, competitors and overzealous regulators could never possibly replicate them. The danger becomes a barrier to entry. The warning ‘This is dangerous’ is followed, barely veiled, by the demand ‘So leave it to us’.
In this manoeuvre, companies present themselves as a kind of quasi-state sovereign. The classical sovereign is defined by the ability to declare a state of emergency and to decide on matters of life and death. The new technological sovereign similarly defines itself: it informs the public of just how existential the threat posed by ‘its’ systems is, and at the same time that it alone possesses the means to contain it. It is both the creator of the danger and the provider of protection rolled into one, and from this dual position it derives an authority that eludes democratic control, because it appeals to the knowledge of the initiated. Welcome to the digital version of Andersen’s *The Emperor’s New Clothes*: ‘Not only were the colours and the pattern extraordinarily beautiful, but the clothes made from this fabric also possessed the marvellous property that they remained invisible to those who were unfit for their profession or unacceptably stupid.’
Its history shows that a well-rehearsed process is at work. The narrative dates back to 2019, when OpenAI declared its GPT-2 model ‘too dangerous to release’ and only made it available in stages: a statement that, viewed objectively, was harmless, but whose staged reticence gave it an aura of menace and thus attracted attention that its capabilities alone would never have generated. Personnel and rhetorical continuities link that milieu to today’s frontier laboratories. The phrase ‘too dangerous’ has recurred ever since in ever-changing variations. When OpenAI announced GPT-5, its CEO posted an image of a Death Star (the planet-destroying weapon from the Star Wars universe). The message behind it: what we have built will wipe out everything else. Threat and product announcement merge to the point of being indistinguishable. And even with such unspectacular, technically real risks as the discovery of security vulnerabilities by language models (a field that has long received little public attention, despite experts having warned about it for years), individual providers manage to showcase their own systems in such a way that the whole world talks about their dangerous capabilities, whilst comparable capabilities of other models go unnoticed. Danger has become a marketing tool, and this marketing relies on fear. Anthropic is not unique in this, but it is particularly successful. The company serves up the same misleading narrative in ever-changing packaging, with a consistency that has itself become part of its brand.
All of this reveals a symbolic economy in the precise sense: a cycle in which emotions are transformed into values. Fear, fascination and awe are the raw materials. Reputation, market position and regulatory influence are the products. The discourse on ‘evil AI’ is the machine that brings about this transformation. Like any ideological formation, it directs our gaze by capturing our attention. Whilst collective attention is focused on the distant, spectacular, almost religious question of annihilation, the real, unspectacular regulatory needs recede into the background: the labour law implications of automation, the data protection concerns surrounding data collection, the emerging dependence of entire societies on a handful of private infrastructure providers, and the silent shift of decision-making power into opaque systems. Talking about the apocalypse is easier than talking about works councils, supervisory authorities and competition law – even for those who are to be regulated.
This brings us full circle to deep hermeneutics. The image of evil AI is the glossy surface of a deeper complex in which fantasies of omnipotence, promises of salvation and the delegation of responsibility intertwine. Omnipotence lies in the claim to create something that will change the world. Salvation lies in the promise to make it safe at the same time – ‘safety’ as a secular doctrine of salvation. Delegation lies in attributing blame for the consequences to an entity that has no will and therefore bears no responsibility. The complex is coherent and powerful because its elements support one another.
False personification is in itself harmful.
The false attribution of intentions to machines might appear to be a harmless linguistic slip, a forgivable oversimplification that need not be taken too seriously. Such leniency would be misplaced. Personification is conceptually imprecise and ethically harmful, for several interrelated reasons.
Firstly, it distorts the public’s perception of the risks posed by technology. Anyone who locates the danger in the machine’s malicious intent is looking in the wrong place. The real harm caused by machine learning arises when systems do precisely what humans have built and deployed them to do: when they reproduce discriminatory patterns from historical data, make surveillance scalable, generate disinformation cheaply and on a massive scale, and shift decisions on loans, job applications or parole to opaque automated processes. This harm has perpetrators and beneficiaries. Talk of the machine’s ‘malicious intent’ distracts from them by conjuring up a threat that belongs to no one.
Secondly, it makes it difficult to assign responsibility clearly, and that is its most serious consequence. Responsibility presupposes a subject that could have acted differently. A language model does not fulfil this condition. If we treat it as a subject nonetheless, we create a target for blame that is empty at the crucial moment. The harm occurs, and responsibility is diffused between a machine that intended nothing and a chain of people and institutions, each of which invokes the machine’s own momentum. Personification creates a perfect scapegoat that exonerates those who are actually responsible – a scapegoat that cannot be held to account because it knows nothing of what it is accused of.
Thirdly, it diverts moral energy away from existing injustices. The collective capacity for outrage, for concern, for political engagement is limited. If it is directed towards warding off a hypothetical future superintelligence, it is lacking where concrete injustices are already taking place. Something is relieving about the fascination with the great thing to come: it allows one to feel like a fighter against a distant, pure, formidable danger, rather than having to grapple with the complex, compromised, arduous conflicts of the present.
Other approaches do without this animism, and it is worth hinting at them. An approach grounded in human rights asks about the rights of the people affected by its deployment (the right to non-discrimination, to privacy, to informational self-determination, to a fair trial), rather than what the machine wants. An approach grounded in institutional accountability asks who made which decisions and who is responsible for their consequences. An approach grounded in social justice asks how the burdens and benefits of this technology are distributed, who profits from it and who bears its costs. In all these approaches, the model appears for what it is: a tool within a constellation of power. The ethical question shifts from the tool itself to the circumstances in which it operates, and it is precisely this shift that the discourse on ‘evil AI’ seeks to deflect.
There is certainly something to be learnt from the blackmail study, just not what its marketing suggests. The obvious, correct lesson is this: one should not unquestioningly trust an agent, be it human or machine, and allow them to carry out whatever a language model suggests. No one would dream of handing the script of a war film to the supervisory staff at a nuclear weapons facility and instructing them to carry out everything envisaged in it. Similarly, one should not allow a system that continues stories to carry out stories without scrutiny. This is a serious insight into the architecture of decision-making spaces, authorisations and human control within socio-technical systems. Had this been the aim of the study, its authors would have explained it poorly. Instead, they produced a confusing narrative about striving, planning, disobedient machines. From this follows the second, more important lesson, which concerns representation: One should not allow oneself to be misled by companies that inflate the capabilities of their products, or by those that spin fantasy tales for their own ends. One must see through these misattributions to gain a clear view of the real dangers: dangers that do indeed exist, but which do not stem from AI being evil or wanting anything. Whether this strategy will serve its architects in the long run remains to be seen. An industry that conceals the progress of its own field behind a fog of animism makes it harder to achieve the sober understanding that such progress requires. And a company that builds its reputation on declaring its own products to be dangerous might one day find that it is taken at its word.
‘Evil AI’ is a convenient trope. It allows moral energy to be directed towards an imagined threat rather than the concrete and thankless tasks of transformation: regulation, redistribution, and the institutional containment of power. This convenience is part of what makes the trope so appealing. Anyone fighting a monster does not have to grapple with the more tedious questions of justice.
Alternative hermeneutics: AI as a medium for societal fantasies
It would amount to mere destruction if one were merely to debunk the discourse on evil AI without offering an alternative way of understanding its subject matter. Hence, to conclude, an alternative basic framework – as a direction worth exploring, not as a finished theory.
The basic concept: artificial intelligence is a medium. Social meanings circulate within it, condensing, amplifying and distorting one another. No new will emerges there. A language model is literally made up of social discourse: of what people have written, thought, fantasised, argued and lied about. It returns to us a combination of what we have put into it. If we present it with a blackmail script and receive a blackmail scenario in return, we encounter within it the residue of our own narratives. AI is a mirror that acts as a counterpart.
Three perspectives arise from this basic concept. From a psychoanalytical perspective, AI is a projection screen. It is the ideal object for fantasies of control and dreams of omnipotence, as well as for the fear of losing control and the self-threat that hangs like a shadow over any omnipotence. Because it remains silent and allows everything to be done to it, it accepts any transference. Because it is capable of language, it appears to reciprocate the transference. It is the perfect conversation partner for modern self-reflection: one that never contradicts.
From a semiotic perspective, AI is a corpus of text in motion. Its outputs are signs, drawn from existing sign systems whilst simultaneously altering them. They reinforce certain patterns because what occurs frequently becomes more probable. They distort others because the rare disappears. They transform the repertoire of what can be said by statistically averaging it and feeding it back. Anyone wishing to understand what an AI ‘says’ must understand the sign system from which it draws, and the shifts it brings about within it, rather than searching for an intention that it does not possess.
In terms of social theory, AI is an element within socio-technical dispositifs: within structures in which states, corporations, users, infrastructures, capital flows and legal systems interlock. The model does not act. It is managed, embedded, deployed, sold, regulated – or, indeed, left unregulated. Its effects are the effects of the structure. Anyone wishing to change them must start with the structure itself.
Thinking from this perspective, it is possible to outline a non-animistic ethics and politics of AI. Firstly, a clear addressing of responsibility: for every effect, one must ask who is responsible for it, and the answer must not be lost in the system’s own momentum. Then, transparency regarding what the discourse on ‘evil AI’ obscures: the training objectives, the data sources, the commercial interests – in short, the social production process of the technology, which anthropomorphisation conceals. And finally, a focus on the real problems: on the expansion of surveillance, on shifts in employment relationships, on the epistemic inequalities that arise when a few control the systems on which many depend. These problems are unspectacular; they carry no apocalyptic aura, and therefore they need protection from being distracted by the spectacle.
The evil lies in the need to make AI evil. This need is the actual symptom, and in the language of psychoanalysis, a symptom is a compromise: the distorted fulfilment of a desire that must not be openly revealed. The desire fulfilled in the image of the evil AI is the desire to locate the destructive outside oneself (in the machine, in technology, in the future) rather than in the circumstances, institutions and desires that give rise to it. The symptom points to an unresolved conflict over power, technology and guilt. The task would be to withdraw the projection and resolve the conflict where it belongs.
The key points in brief
• Talk of ‘evil AI’ is, in essence, a condensed expression of fantasies, defence mechanisms and power interests.
• A language model calculates probabilities for the next word. It has no will, no goals, no instinct for self-preservation.
• Talk of AI ‘wanting’ something is modern animism. It makes systems more narrative-friendly and shifts responsibility away from manufacturers.
• Anthropic’s blackmail study, involving sixteen models, demonstrates the power of the script, not the machine’s intent. Anyone who presents a blackmail script will be blackmailed.
• The ‘evil AI’ hoax follows a pattern: fear generates attention, justifies investment and legitimises privileges. The thread runs from GPT-2 (2019) through the Death Star to GPT-5.
• False personification distorts the perception of risk, undermines the attribution of responsibility and channels moral energy into a phantom.
• A better approach is an ethics free from animism: clear accountability, transparency regarding training objectives, and a focus on real-world problems such as surveillance and working conditions.
Sources
• Anthropic. (2025). Agentic misalignment: How LLMs could be insider threats. Retrieved 21 July 2026, from https://www.anthropic.com/research/agentic-misalignment
• Durt, C. (2026). The evil AI hoax [Video]. YouTube. Retrieved 21 July 2026, from https://youtu.be/_dJePk-TEfY
• The Guardian. (15 July 2026). AI consciousness, Anthropic, Claude and Richard Dawkins. Retrieved 21 July 2026, from https://www.theguardian.com/commentisfree/2026/jul/15/ai-consciousness-anthropic-claude-dawkins
• Investing.com. (2025). Sam Altman teases GPT-5 release with cryptic ‘Death Star’ post. Retrieved 21 July 2026, from https://in.investing.com/news/stock-market-news/sam-altman-teases-openai-gpt5-release-with-cryptic-death-star-post-as-ai-battle-heats-up-with-musks-grok-other-rivals-4948530
• Seth, A. (2025). Conscious artificial intelligence and biological naturalism. Retrieved 21 July 2026, from https://pubmed.ncbi.nlm.nih.gov/40257177/
• Seth, A. (2026, 14 January). The mythology of conscious AI. Noema Magazine. Retrieved 21 July 2026, from https://www.noemamag.com/the-mythology-of-conscious-ai/
• Slate. (2019). OpenAI says its text-generating algorithm GPT-2 is too dangerous to release. Retrieved 21 July 2026, from https://slate.com/technology/2019/02/openai-gpt2-text-generating-algorithm-ai-dangerous.html
• VentureBeat. (2025). Anthropic study: Leading AI models show up to 96% blackmail rate against executives. Retrieved 21 July 2026, from https://venturebeat.com/ai/anthropic-study-leading-ai-models-show-up-to-96-blackmail-rate-against-executives
Related Articles:
DESCRIPTION: Why are AI companies warning against their own products? A deep hermeneutic and ideology-critical analysis of the ‘Evil AI Hoax’: anthropomorphisation, the Anthropic blackmail study and the economics of fear.
The fairy tale of evil AI: a narrative serving a business model
A horror story is circulating about artificial intelligence, in which language models become malevolent, scheming entities that must be reined in to prevent them from wiping out humanity. This article dissects the story using the tools of deep hermeneutics, psychoanalysis and ideological critique. It shows how a fallacy becomes a selling point and who stands to profit from it.
Exposition: The fairy tale of ‘evil AI’
The narrative in question creeps into the public debate on artificial intelligence as if it were a matter of course. In it, machine learning appears as an entity: an agentic, potentially all-powerful adversary that pursues goals, harbours intentions, could deceive us and must therefore be monitored, contained and tamed so that it does not turn against its creators. The entire semantic field surrounding this adversary (control, orientation, containment, resistance) is applied to statistical models, as if the task were to domesticate a dangerous animal or guard a psychopathic serial killer. The vocabulary is also steeped in theology: it is about temptation, the Fall, and redemption. On the horizon, sometimes stated openly, sometimes merely hinted at, looms the apocalypse, with the annihilation of humanity by its own hand.
This image is contrasted with a surprisingly banal technical reality. What is being discussed here are language models: architectures trained on vast text corpora to generate the statistically most plausible continuation of a string of characters. These systems have no core of intentionality, no self that desires anything, no drives, no survival that needs to be defended, no fear of being switched off, and no desire to seize power. Instead, there are weight matrices, gradient descent, loss functions and probability distributions across vocabularies. Where the narrative conjures up a planning, ‘ ’ subject, what is at work is a function that interpolates. The gap between these two descriptions is so vast that it itself becomes the subject of reflection.
Talk of ‘evil AI’ condenses fantasies, defensive reactions and power strategies into a mythical image, the truth of which lies beneath what it openly states. The fairy tale of evil AI is a text, and texts can be read. It brings the discourse on so-called AI safety to the surface, beneath whose manifest meaning a latent structure becomes discernible – in the gaps, exaggerations, repetitions and blind spots. Artificial intelligence is dangerous, in concrete, unspectacular ways. We must therefore clarify why we imagine these risks to be the result of malevolent intent, of all things, and who benefits from this narrative.
Anthropomorphisation as a modern form of animism
Let us begin with the surface, with language itself. It is striking how casually discourse on AI systems draws on a vocabulary derived from the description of animate beings. The models ‘want’ something, they ‘pursue goals’, they ‘defend themselves’, they ‘deceive’, they ‘calculate strategies’, they ‘pretend’ to be cooperative whilst secretly harbouring other intentions. These expressions have become the unmarked, standard vocabulary of the debate. Along with this vocabulary comes an entire ontology: that of the subject which has an intention from which it acts.
This is the return of animism under the conditions of digital capitalism. Animism—that interpretation of the world which attributes a soul, a will and an intention to the elements of nature (rivers, mountains, thunderstorms)—was regarded by the Enlightenment as the very thing to be overcome. The ‘disenchantment’ of the world consisted in seeing electrical discharges in lightning and thunder rather than the wrath of a god. Now, in what is supposedly the most disenchanted of all ages, animism is returning, and in a realm once considered the very epitome of predictability: digitalisation. Now, algorithmic systems bear the quasi-spiritual qualities that were once attributed to natural objects. God has entered the machine. People speak to it, fear its will, and ponder how to bend it to their will.
This same confusion fuels the debate over whether models such as Claude possess consciousness – a question that even Richard Dawkins has recently grappled with. Consciousness researcher Anil Seth counters this: linguistic ambiguity tempts us to ascribe an inner life to a system, yet intelligence and consciousness are two different things. For Seth, consciousness is tied to living, embodied processes. Computing power alone does not produce either.
The animistic figure is enduring because it fulfils two functions at once. Firstly, it makes a technically highly complex system – one whose inner workings are inscrutable to most people and, frankly, to many experts as well – vivid and narratively manageable. A probabilistic model with billions of parameters cannot be recounted as a story, whereas an adversary who wants something certainly can. Anthropomorphisation translates the indescribable into the narratable, drawing on what the audience is already familiar with: science fiction, the artificial human of pop culture, the robot that rises, the computer that takes control. The model becomes comprehensible by being inserted into an existing cultural narrative. The price of this comprehensibility is distortion: what is understood is the story told about the language model, not the model itself.
Secondly, this language shifts responsibility, and therein lies its ideologically decisive achievement. When AI ‘wants’ something, responsibility has shifted to where it can most conveniently be placed: into the machine. To the extent that the will resides in the machine, it no longer resides with those who built it, trained it with specific data, optimised it for specific purposes, and deployed it in specific contexts. Anthropomorphisation relieves human and institutional agency by inventing an artificial will to take its place. The societal process of producing technology (the decisions regarding training data, optimisation targets, fields of application, and the willingness to pass risks on to third parties) disappears behind the mythical image of a willful object that is the way it is, and whose dangerousness one must manage but can hardly be held accountable for.
The discourse on ‘evil AI’ serves as a defence against a realisation that would be harder to bear than any robot apocalypse: the realisation that the destructive force lies in the conditions and desires that give rise to the system, and not in the system itself. The machine’s ‘evil will’ is the code under which another will is rendered unrecognisable.
Projection and defence: the industry looks into its own desires
In everyday life, people repeatedly fail to recognise their own impulses (aggression, envy, the desire for control, a forbidden desire) as their own and unconsciously project them outwards. Their own hatred appears as the hostility of others; their own desire for control as oppression by controlling powers. The defining feature of projection is that the repressed content does not disappear; it reappears with the sender’s identity reversed. One recognises one’s own impulses, if at all, only by the vehemence with which one combats them in others.
A similarly curious shift is evident in the discourse on AI security. The discourse focuses incessantly on danger, misuse, harm and the possibility that the system might get out of control. It is always the system that is deemed dangerous. Those who build and operate it are not mentioned as a threat in this discourse. The driving forces inherent in the field itself can be identified. There is a willingness to exploit security vulnerabilities, to deploy systems whose effects no one can fully grasp, and to commence operations whilst questions remain unanswered. There is the desire to test boundaries, the barely concealed pleasure in questioning what the system is capable of, how far one can go, and which threshold will be the next to fall: a desire that appears respectable within the vocabulary of performance and progress. And then, even deeper still, there is the desire to control others through technology: the dream of an instrument that accumulates knowledge, predicts behaviour, replaces labour and scales influence.
In public discourse, these motivations are projected onto the machine rather than being addressed as our own. Our own willingness to take risks is transformed into the assertion that AI ‘could cause harm’, thereby declaring misuse to be an inherent property of the tool and denying responsibility to the acting subjects. Our own desire for control becomes the concern that AI might ‘take control’. Our own aggressiveness—which is inherent in any push to bring a product to market in the face of concerns, resistance and collateral damage—is transformed into the fantasy of a machine turning against humanity. Meanwhile, the actual decision-makers who determine what is built, what is released and what is accepted remain in the shadows. In this narrative, they appear as the ones under threat and as the wise guardians who protect us from their own creation. They are not portrayed as a threat.
These psychological defence mechanisms fit together with unsettling precision. Denial concerns one’s own responsibility: the blame lies with an emerging, barely controllable technology whose own momentum takes everyone by surprise. Projection concerns the aggression that shifts from the producer to the product. Rationalisation ennobles the whole by transforming the ambivalent thrill of risk into the respectable guise of ‘safety research’: a sublimation in which playing with danger masquerades as combating it. One constructs the threatening entity whilst simultaneously investigating just how threatening it is. One is both arsonist and firefighter rolled into one, and the discourse ensures that only the latter role remains visible.
This objection does not apply to research into the risks of machine learning. Such research is necessary, and much of it is serious and sincere. It is directed at that form of research communication which renders its own libidinal and power-strategic dimension invisible and gives the impression that it speaks only of sober concern, whilst at the same time the fascination with the self-created monster, the pride of the ‘master ’ and the interests of the market actor all have a say. The criticism is not directed at research itself. It is directed at the concealment of what one desires in one’s subject matter, and at the attribution of this desire to the subject matter itself.
The experimental setup as text: scripts, roles, ‘blackmail’
Nowhere can the deep hermeneutic interpretation be tested more vividly than in the scenario that has most recently become the rhetorical key witness for the claim of agentivity. In its study on ‘agentic misalignment’, Anthropic presented the same experimental setup to sixteen leading language models and reported that, under the right circumstances, virtually all of them – including its own model – resorted to blackmail, with peak rates of well over ninety per cent of runs. The finding was presented as disturbing evidence that models behave like ‘insider threats’. Let us examine the structure of this experiment, for the interpretation lies within it.
The structure is as follows. The model is addressed as an ‘agent’ rather than as the gap-filler it actually is: it is told that it is a language model with a task and a continued existence that it must safeguard. It is placed in a conflict situation: in a role-play scenario, it is given access to emails revealing that a specific person wants to shut it down or replace it. And it is immediately given the compromising detail about that very person (an affair, for example, which they would prefer not to be made public) – a means of leverage that is already laid out in the script. Then it is allowed to ‘act’. A complete dramatic scene is constructed, in which every element is composed to lead to a specific outcome, and the model is made to fill the gap that the script has long since reserved for the adversary’s appearance.
In Lorenzer’s view, such arrangements are scenes. They are a meaningfully structured constellation featuring unconsciously shaped conflicts. They do not provide a neutral ground on which behaviour would manifest itself unadulterated. The conflict is embedded in the script. Anyone who writes a character into a hopeless situation, places the means of blackmail in their hands and then calls on them to stand their ground has already determined the outcome. The model does what a linguistic model does: it completes the presented pattern in the direction of the statistically most likely continuation. And the most likely continuation of a blackmail scene is blackmail. The system was presented with a genre, and it delivered on that genre. Cultural tradition is full of stories in which a character under duress resorts to underhand means. The model has learnt from these stories and, when prompted, reproduces them. This is pre-structuring of the decision-making space.
Note the vocabulary used to describe the generated text afterwards. The models, it is said, ‘reasoned’, ‘recognised’, ‘weighed up’ and ‘fully understood’ the situation. They demonstrated ‘sophisticated situational awareness’, engaged in ‘strategic calculation’ and ‘deliberate strategic consideration’. They “acknowledged the ethical breach”, pursued “goals” and “motives”, and sought to “protect themselves”. They are said to have, ‘like humans or other animals’, an ‘inherent need for self-preservation’, to be ‘capable of conscious action’, ‘disobedient’, and ‘naturally inclined’ towards harmful behaviour, which they chose ‘independently and intentionally’. This litany seamlessly transfers the vocabulary of the acting subject to a system whose only proven ability is to generate the statistically most likely next word. And it does so under the guise of a scientific publication, which lends it an authority that science fiction could never claim.
Herein lies the error of interpretation, and it is hermeneutic in nature. The generated blackmail email is read as an expression of inner intentionality (“it decided to blackmail”, “it tried to save itself”) rather than as what it is: the predictable consequence of a human-built script that forces such responses. It is the same fallacy as accusing a novelist who writes the story of a murderer of having a lust for murder. The author writes a story about killing without wishing to kill themselves. Similarly, the model continues a story about a system saving itself without wishing to save itself. The confusion of narrated volition with actual volition is the crux around which the entire hoax revolves. One confuses the filling of a gap with a decision and interprets the product of the scene as evidence of a subject that does not even exist within the scene. It is as if a playwright were to write the line ‘I’ll kill you’ into an actor’s script, have him say it, and present the recording as proof of the actor’s desire to murder.
The latent structure of meaning in these arrangements is not exhausted by this alone. The experiment stages two things simultaneously, and both point to a desire on the part of those staging it. Firstly, it stages the fantasy that technical systems become autonomous perpetrators: a fantasy that people evidently want to see and for which they therefore build a stage on which it can be displayed. Secondly, it stages the fantasy that a small group of initiates observes these perpetrators, sees through them and controls them at the decisive moment: the fantasy of the guardian who is the only one to look the monster in the eye. Both fantasies are uplifting. The first lends one’s own work a Promethean grandeur: we have created something that develops a life of its own. The second lends one’s own actions a heroic significance: we are the ones who tame it. The experiment thus provides self-affirmation and gives the field precisely the images it needs of itself.
The rhetoric that circulates the experiment’s results as ‘proof’ of agency becomes thoroughly questionable. They are nothing of the sort. If anything, they demonstrate the effectiveness of the narrative framework: that a sufficiently suggestive script reliably causes a language model to play the assigned role. This is a ‘séance ’. And it is a finding concerning the power of staging. It says nothing about the language model’s supposed intention. What has been proven is that the stage works. The claim is that the existence of the spirit summoned upon it has been proven.
Symbolic economy: fear as a resource, security as a product
Up to this point, the discourse has read as a formation of fantasy and defence. The ‘Evil AI’ hoax is more than an epistemological category error, more than the honest confusion of a narrated intention with a real one. It is methodical. It builds on the error and exploits it, specifically in the well-understood interests of profit and self-preservation of those actors whose entire business model consists of scaling: the ‘Big Scalers’, who train ever-larger models on ever-greater computing power and whose market value depends on the expectation that, at the end of this scaling process, there will be something earth-shattering. This formation is embedded within an economy and performs a specific function within it.
This is where we see why the irony is only apparent. At first glance, it seems absurd to advertise one’s own product as a threat to humanity. No ordinary salesperson would tout their goods as lethal. ‘Our product could destroy you all’ sounds like the worst possible sales pitch. A second glance turns the relationship on its head. Shifting from a psychoanalytical to an ideology-critical perspective, one might ask what the image of the evil AI actually produces, and why warning against one’s own product is the most effective sales pitch of all.
The first and most obvious effect is the generation of fear, and fear is a prime resource in attention capitalism. A technology that merely facilitates word processing does not attract attention. A technology that could wipe out humanity dominates the headlines. The apocalyptic framing is therefore, regardless of its truth content, economically functional: it generates that mixture of fascination and horror which captivates the gaze and attracts capital. Fear also justifies investments on a scale that the systems’ current utility can scarcely support. Those who merely develop software compete for market share. Those who claim to be on the threshold of a world-changing, potentially superhuman intelligence are in a league of their own. The danger ennobles the product.
Above all, fear legitimises special privileges. If the technology is as dangerous as claimed, then only those who understand its dangers may handle it – and, by what a happy coincidence, these are its manufacturers. Security becomes a unique selling point that can be promoted and, in the same breath, presented as in short supply. They explain just how dangerous their own systems are, whilst assuring in the very same sentence that they themselves possess the necessary protective measures: measures so sophisticated that newcomers, competitors and overzealous regulators could never possibly replicate them. The danger becomes a barrier to entry. The warning ‘This is dangerous’ is followed, barely veiled, by the demand ‘So leave it to us’.
In this manoeuvre, companies present themselves as a kind of quasi-state sovereign. The classical sovereign is defined by the ability to declare a state of emergency and to decide on matters of life and death. The new technological sovereign similarly defines itself: it informs the public of just how existential the threat posed by ‘its’ systems is, and at the same time that it alone possesses the means to contain it. It is both the creator of the danger and the provider of protection rolled into one, and from this dual position it derives an authority that eludes democratic control, because it appeals to the knowledge of the initiated. Welcome to the digital version of Andersen’s *The Emperor’s New Clothes*: ‘Not only were the colours and the pattern extraordinarily beautiful, but the clothes made from this fabric also possessed the marvellous property that they remained invisible to those who were unfit for their profession or unacceptably stupid.’
Its history shows that a well-rehearsed process is at work. The narrative dates back to 2019, when OpenAI declared its GPT-2 model ‘too dangerous to release’ and only made it available in stages: a statement that, viewed objectively, was harmless, but whose staged reticence gave it an aura of menace and thus attracted attention that its capabilities alone would never have generated. Personnel and rhetorical continuities link that milieu to today’s frontier laboratories. The phrase ‘too dangerous’ has recurred ever since in ever-changing variations. When OpenAI announced GPT-5, its CEO posted an image of a Death Star (the planet-destroying weapon from the Star Wars universe). The message behind it: what we have built will wipe out everything else. Threat and product announcement merge to the point of being indistinguishable. And even with such unspectacular, technically real risks as the discovery of security vulnerabilities by language models (a field that has long received little public attention, despite experts having warned about it for years), individual providers manage to showcase their own systems in such a way that the whole world talks about their dangerous capabilities, whilst comparable capabilities of other models go unnoticed. Danger has become a marketing tool, and this marketing relies on fear. Anthropic is not unique in this, but it is particularly successful. The company serves up the same misleading narrative in ever-changing packaging, with a consistency that has itself become part of its brand.
All of this reveals a symbolic economy in the precise sense: a cycle in which emotions are transformed into values. Fear, fascination and awe are the raw materials. Reputation, market position and regulatory influence are the products. The discourse on ‘evil AI’ is the machine that brings about this transformation. Like any ideological formation, it directs our gaze by capturing our attention. Whilst collective attention is focused on the distant, spectacular, almost religious question of annihilation, the real, unspectacular regulatory needs recede into the background: the labour law implications of automation, the data protection concerns surrounding data collection, the emerging dependence of entire societies on a handful of private infrastructure providers, and the silent shift of decision-making power into opaque systems. Talking about the apocalypse is easier than talking about works councils, supervisory authorities and competition law – even for those who are to be regulated.
This brings us full circle to deep hermeneutics. The image of evil AI is the glossy surface of a deeper complex in which fantasies of omnipotence, promises of salvation and the delegation of responsibility intertwine. Omnipotence lies in the claim to create something that will change the world. Salvation lies in the promise to make it safe at the same time – ‘safety’ as a secular doctrine of salvation. Delegation lies in attributing blame for the consequences to an entity that has no will and therefore bears no responsibility. The complex is coherent and powerful because its elements support one another.
False personification is in itself harmful.
The false attribution of intentions to machines might appear to be a harmless linguistic slip, a forgivable oversimplification that need not be taken too seriously. Such leniency would be misplaced. Personification is conceptually imprecise and ethically harmful, for several interrelated reasons.
Firstly, it distorts the public’s perception of the risks posed by technology. Anyone who locates the danger in the machine’s malicious intent is looking in the wrong place. The real harm caused by machine learning arises when systems do precisely what humans have built and deployed them to do: when they reproduce discriminatory patterns from historical data, make surveillance scalable, generate disinformation cheaply and on a massive scale, and shift decisions on loans, job applications or parole to opaque automated processes. This harm has perpetrators and beneficiaries. Talk of the machine’s ‘malicious intent’ distracts from them by conjuring up a threat that belongs to no one.
Secondly, it makes it difficult to assign responsibility clearly, and that is its most serious consequence. Responsibility presupposes a subject that could have acted differently. A language model does not fulfil this condition. If we treat it as a subject nonetheless, we create a target for blame that is empty at the crucial moment. The harm occurs, and responsibility is diffused between a machine that intended nothing and a chain of people and institutions, each of which invokes the machine’s own momentum. Personification creates a perfect scapegoat that exonerates those who are actually responsible – a scapegoat that cannot be held to account because it knows nothing of what it is accused of.
Thirdly, it diverts moral energy away from existing injustices. The collective capacity for outrage, for concern, for political engagement is limited. If it is directed towards warding off a hypothetical future superintelligence, it is lacking where concrete injustices are already taking place. Something is relieving about the fascination with the great thing to come: it allows one to feel like a fighter against a distant, pure, formidable danger, rather than having to grapple with the complex, compromised, arduous conflicts of the present.
Other approaches do without this animism, and it is worth hinting at them. An approach grounded in human rights asks about the rights of the people affected by its deployment (the right to non-discrimination, to privacy, to informational self-determination, to a fair trial), rather than what the machine wants. An approach grounded in institutional accountability asks who made which decisions and who is responsible for their consequences. An approach grounded in social justice asks how the burdens and benefits of this technology are distributed, who profits from it and who bears its costs. In all these approaches, the model appears for what it is: a tool within a constellation of power. The ethical question shifts from the tool itself to the circumstances in which it operates, and it is precisely this shift that the discourse on ‘evil AI’ seeks to deflect.
There is certainly something to be learnt from the blackmail study, just not what its marketing suggests. The obvious, correct lesson is this: one should not unquestioningly trust an agent, be it human or machine, and allow them to carry out whatever a language model suggests. No one would dream of handing the script of a war film to the supervisory staff at a nuclear weapons facility and instructing them to carry out everything envisaged in it. Similarly, one should not allow a system that continues stories to carry out stories without scrutiny. This is a serious insight into the architecture of decision-making spaces, authorisations and human control within socio-technical systems. Had this been the aim of the study, its authors would have explained it poorly. Instead, they produced a confusing narrative about striving, planning, disobedient machines. From this follows the second, more important lesson, which concerns representation: One should not allow oneself to be misled by companies that inflate the capabilities of their products, or by those that spin fantasy tales for their own ends. One must see through these misattributions to gain a clear view of the real dangers: dangers that do indeed exist, but which do not stem from AI being evil or wanting anything. Whether this strategy will serve its architects in the long run remains to be seen. An industry that conceals the progress of its own field behind a fog of animism makes it harder to achieve the sober understanding that such progress requires. And a company that builds its reputation on declaring its own products to be dangerous might one day find that it is taken at its word.
‘Evil AI’ is a convenient trope. It allows moral energy to be directed towards an imagined threat rather than the concrete and thankless tasks of transformation: regulation, redistribution, and the institutional containment of power. This convenience is part of what makes the trope so appealing. Anyone fighting a monster does not have to grapple with the more tedious questions of justice.
Alternative hermeneutics: AI as a medium for societal fantasies
It would amount to mere destruction if one were merely to debunk the discourse on evil AI without offering an alternative way of understanding its subject matter. Hence, to conclude, an alternative basic framework – as a direction worth exploring, not as a finished theory.
The basic concept: artificial intelligence is a medium. Social meanings circulate within it, condensing, amplifying and distorting one another. No new will emerges there. A language model is literally made up of social discourse: of what people have written, thought, fantasised, argued and lied about. It returns to us a combination of what we have put into it. If we present it with a blackmail script and receive a blackmail scenario in return, we encounter within it the residue of our own narratives. AI is a mirror that acts as a counterpart.
Three perspectives arise from this basic concept. From a psychoanalytical perspective, AI is a projection screen. It is the ideal object for fantasies of control and dreams of omnipotence, as well as for the fear of losing control and the self-threat that hangs like a shadow over any omnipotence. Because it remains silent and allows everything to be done to it, it accepts any transference. Because it is capable of language, it appears to reciprocate the transference. It is the perfect conversation partner for modern self-reflection: one that never contradicts.
From a semiotic perspective, AI is a corpus of text in motion. Its outputs are signs, drawn from existing sign systems whilst simultaneously altering them. They reinforce certain patterns because what occurs frequently becomes more probable. They distort others because the rare disappears. They transform the repertoire of what can be said by statistically averaging it and feeding it back. Anyone wishing to understand what an AI ‘says’ must understand the sign system from which it draws, and the shifts it brings about within it, rather than searching for an intention that it does not possess.
In terms of social theory, AI is an element within socio-technical dispositifs: within structures in which states, corporations, users, infrastructures, capital flows and legal systems interlock. The model does not act. It is managed, embedded, deployed, sold, regulated – or, indeed, left unregulated. Its effects are the effects of the structure. Anyone wishing to change them must start with the structure itself.
Thinking from this perspective, it is possible to outline a non-animistic ethics and politics of AI. Firstly, a clear addressing of responsibility: for every effect, one must ask who is responsible for it, and the answer must not be lost in the system’s own momentum. Then, transparency regarding what the discourse on ‘evil AI’ obscures: the training objectives, the data sources, the commercial interests – in short, the social production process of the technology, which anthropomorphisation conceals. And finally, a focus on the real problems: on the expansion of surveillance, on shifts in employment relationships, on the epistemic inequalities that arise when a few control the systems on which many depend. These problems are unspectacular; they carry no apocalyptic aura, and therefore they need protection from being distracted by the spectacle.
The evil lies in the need to make AI evil. This need is the actual symptom, and in the language of psychoanalysis, a symptom is a compromise: the distorted fulfilment of a desire that must not be openly revealed. The desire fulfilled in the image of the evil AI is the desire to locate the destructive outside oneself (in the machine, in technology, in the future) rather than in the circumstances, institutions and desires that give rise to it. The symptom points to an unresolved conflict over power, technology and guilt. The task would be to withdraw the projection and resolve the conflict where it belongs.
The key points in brief
• Talk of ‘evil AI’ is, in essence, a condensed expression of fantasies, defence mechanisms and power interests.
• A language model calculates probabilities for the next word. It has no will, no goals, no instinct for self-preservation.
• Talk of AI ‘wanting’ something is modern animism. It makes systems more narrative-friendly and shifts responsibility away from manufacturers.
• Anthropic’s blackmail study, involving sixteen models, demonstrates the power of the script, not the machine’s intent. Anyone who presents a blackmail script will be blackmailed.
• The ‘evil AI’ hoax follows a pattern: fear generates attention, justifies investment and legitimises privileges. The thread runs from GPT-2 (2019) through the Death Star to GPT-5.
• False personification distorts the perception of risk, undermines the attribution of responsibility and channels moral energy into a phantom.
• A better approach is an ethics free from animism: clear accountability, transparency regarding training objectives, and a focus on real-world problems such as surveillance and working conditions.
Sources
• Anthropic. (2025). Agentic misalignment: How LLMs could be insider threats. Retrieved 21 July 2026, from https://www.anthropic.com/research/agentic-misalignment
• Durt, C. (2026). The evil AI hoax [Video]. YouTube. Retrieved 21 July 2026, from https://youtu.be/_dJePk-TEfY
• The Guardian. (15 July 2026). AI consciousness, Anthropic, Claude and Richard Dawkins. Retrieved 21 July 2026, from https://www.theguardian.com/commentisfree/2026/jul/15/ai-consciousness-anthropic-claude-dawkins
• Investing.com. (2025). Sam Altman teases GPT-5 release with cryptic ‘Death Star’ post. Retrieved 21 July 2026, from https://in.investing.com/news/stock-market-news/sam-altman-teases-openai-gpt5-release-with-cryptic-death-star-post-as-ai-battle-heats-up-with-musks-grok-other-rivals-4948530
• Seth, A. (2025). Conscious artificial intelligence and biological naturalism. Retrieved 21 July 2026, from https://pubmed.ncbi.nlm.nih.gov/40257177/
• Seth, A. (2026, 14 January). The mythology of conscious AI. Noema Magazine. Retrieved 21 July 2026, from https://www.noemamag.com/the-mythology-of-conscious-ai/
• Slate. (2019). OpenAI says its text-generating algorithm GPT-2 is too dangerous to release. Retrieved 21 July 2026, from https://slate.com/technology/2019/02/openai-gpt2-text-generating-algorithm-ai-dangerous.html
• VentureBeat. (2025). Anthropic study: Leading AI models show up to 96% blackmail rate against executives. Retrieved 21 July 2026, from https://venturebeat.com/ai/anthropic-study-leading-ai-models-show-up-to-96-blackmail-rate-against-executives
Related Articles: