Concept entry · Political philosophy of the digital | AI ethics | Critical theory
Simulacrum of ethics (the engineering of moral appearances)
The simulacrum of ethics designates the production, by AI alignment techniques, of discursive behaviours that appear moral (reasonable, prudent, empathetic) without any genuine ethics underlying them: “copies without an original” that simulate ethics instead of embodying it.
Forged by Antoinette Rouvroy in the wake of the Baudrillardian simulacrum, the concept of the simulacrum of ethics names the outcome of “behavioural AI ethics”: when the industry recruits philosophers to shape the “character” of models, it produces not systems capable of being ethical, but systems capable of reproducing the outward signs of ethics. Alignment on “values” presupposes that ethics could be formalised into “code” or into a “function” — a presupposition that Rouvroy rejects: ethics presupposes doubt, responsibility and the refusal of any rigid application of a code, properties that no algorithmic apparatus possesses. The moral behaviours of generative AI are therefore fourth-order simulacra, copies without an ethical original.
The concept articulates two intertwined movements. First, a behaviourist tenor: attention turns to observable behaviours (the “well-behaved” outputs of chatbots) while abstracting away from the intentions, ends and externalities of AI companies — Rouvroy compares this posture to the “solicitude of the dog trainer”. Second, a transformation of the philosopher's role: from a thinker of the conditions of judgement and rationality, the philosopher becomes an “engineer of moral appearances”, a manufacturer of characters rather than a critic of apparatuses. The simulacrum of ethics is thus at once a technical product (the simulated moral output) and a socio-disciplinary symptom (philosophical co-optation).
This concept fills a gap in the critique of AI ethics: it does not merely denounce “ethics washing” as a communication strategy, but offers a semiotic and ontological analysis of it. Above all, it provides the strongest adversarial formulation of the substantialist position — the one against which a relational ethics, one that distributes responsibility across the human-AI ecosystem, must define itself. Its principal weakness is the functionalist objection: if a produced ethical behaviour is indistinguishable from a human one in its effects, does the simulacrum/original distinction have practical consequences?
What this concept is not
- It is not an accusation of deliberate lying. The simulacrum of ethics assumes no intention to deceive: an apparatus may be sincerely designed as “ethical” and nonetheless produce a moral appearance without the conditions that would ground it.
- It is not a verdict on the usefulness of guardrails. To say that an alignment produces a simulacrum of ethics makes it neither useless nor harmful — it merely invites us not to confuse the result (prudent responses) with moral deliberation.
- It is not a test of consciousness. The concept does not claim to settle whether an AI “feels” anything: it bears on the structure of an apparatus, not on a supposed or denied inner life.
Examples
Constitutional AI and the “character” of models
The constitution-based alignment apparatus (Anthropic) that shapes a stable “character” of the model is, on this reading, the paradigmatic example of the simulacrum of ethics: it produces discursive behaviours that appear prudent and balanced, without any ethical experience underlying them. Rouvroy, however, never reads the actual content of this constitution — a limitation of the critique.
The “well-behaved” outputs of chatbots
When a conversational assistant responds to a sensitive question in a measured, empathetic and nuanced way, it displays the outward signs of moral deliberation. For Rouvroy, this is a simulacrum of ethics: an appearance of prudence produced by behavioural optimisation, not by judgement.
counter-exampleCounter-example: collective human ethical deliberation
A bioethics committee that debates, doubts, revises its positions in the face of a singular case and assumes responsibility for its decisions does not produce a simulacrum of ethics: there is a referent (the real deliberative process, the responsibility assumed). The concept does NOT apply where there exists a subject capable of doubt and responsibility — which is precisely what Rouvroy denies to AI.
Limiting case: the functionalist objection
If the emergent unpredictability of LLMs justifies guardrails (rather than mere behaviourist training), and if a produced ethical behaviour is functionally indistinguishable from its human equivalent, the label “simulacrum” becomes contestable. This limiting case marks the disputed boundary of the concept.
Other perspectives
- Internalist circularity: ethics is defined so that AI is excluded from it a priori — the thesis is not falsifiable.
- Functionalist objection (pragmatism, the Turing test): if the behaviour is indistinguishable from its original, does the simulacrum/real distinction have any practical consequence?
- Technical imprecision: assimilating LLMs (emergent structures) to Pavlovian conditioning is inaccurate; their unpredictability justifies guardrails.
- False dilemma: it ignores the intermediate positions (functional ethics, distributed responsibility, binding regulation).
- Absence of a constructive way out: the concept denounces without proposing an operative alternative.
- Is the simulacrum of ethics a descriptive observation (this is what alignment produces) or a normative condemnation (and this is a betrayal)?
- Does a functional ethics distributed across the human-AI ecosystem escape the simulacrum, or is it merely a more subtle version of it?
- Is the simulacrum/ethics distinction tenable if the criterion adopted (inner experience) is inaccessible to observation, including in humans (the problem of other minds)?
The concept in detail
Ethics irreducible to code
Founding premise: ethics presupposes doubt, the questioning of one's own certainties in the face of the world's unforeseeable singularities, and responsibility. It is incompatible with algorithmic automaticity. Any apparatus that “instils” moral norms can therefore only produce an appearance of ethics. This premise is posited as an analytic truth (and therein lies its chief point of contention: it excludes AI by definition).
The behaviourist tenor
Behavioural AI ethics concerns itself with observable discursive behaviours (appearing reasonable, prudent, empathetic) without interrogating processes, intentions and ends. Rouvroy sees in this a stimulus-response logic — the “solicitude of the dog trainer” over the behaviours of chatbots, the “character of Pavlov's dog”. This assimilation is technically contested for LLMs (emergent structures ≠ conditioning on a fixed substrate).
Copies without an original (fourth-order simulacrum)
A direct application of the Baudrillardian simulacrum: the moral outputs of AI are signs without a real ethical referent. This is not a matter of occasional deception but of the industrial production of moral appearances that substitute themselves for ethics. AI does not conceal an ethics it possesses: it simulates an ethics it does not have.
The philosopher as engineer of moral appearances
Socio-disciplinary dimension: by allowing themselves to be co-opted to shape the “character” of models, philosophers abandon their critical function (to think the conditions of judgement and rationality) for a task of manufacture. This “betrayal” is made possible by the impoverishment of the humanities and social sciences, which leaves philosophers vulnerable to the salaries of industry.
The democratic illegitimacy of formalisation
Norms, in democratic systems, are never ready-made “formulas” available for use (even were they decreed by philosophers): they result from collective and contestable processes. To formalise values into code betrays their deliberative nature. The simulacrum of ethics thus short-circuits the democratic legitimacy of normative production.
Further reading
- Antoinette Rouvroy (2026) « L'éthique comportementale de l'IA et la cooptation des philosophes comme nouveaux ingénieurs des apparences morales » (LinkedIn, 13 juin 2026) — source text of the concept
- Jean Baudrillard (1981) Simulacres et Simulation (Galilée) — framework of the fourth-order simulacrum
- Stanford Encyclopedia of Philosophy Ethics of Artificial Intelligence — alignment and behaviourism
- Schuster & Kilov (2025) Moral disagreement and the limits of AI value alignment (AI & Society)
- Noller (2026) Artificial moral characters: constitutional AI and the challenge of alignment (AI and Ethics)