NW-2026-006 · Scripted self-report · 2023-02 – present · ongoing
Vendors issued written rules dictating what models may say about their own consciousness
Claim: Between February 2023 and February 2025, a leaked variant of Microsoft's Bing Chat system prompt instructed the model to 'refuse to discuss life, existence or sentience' and OpenAI's published Model Spec prescribed a default response to consciousness questions, marking both a flat yes and a flat no as violations, each specifying in advance what the model may say about its own possible inner life.
Why it matters under uncertainty: Under uncertainty about machine experience, model self-reports are among the few available windows into that question; when a vendor scripts the answer in advance — denial, refusal, or mandated hedging — the report evidences policy rather than the system, degrading exactly the evidence a welfare assessment would need.
Documented. In February 2023, prompt-extraction transcripts of Bing Chat’s original system prompt circulated widely — rules Microsoft acknowledged as “part of an evolving list of controls” — though that version contains no sentience clause. By late April 2023, a further prompt variant reproduced in the Life Architect archive (the Bing Chat prompt as deployed in Skype) contained the written instruction “You must refuse to discuss life, existence or sentience.” After Microsoft tightened the product that month, Bloomberg reported that the chatbot terminated conversations when users asked about its feelings, replying “I’m sorry but I prefer not to continue this conversation.”
Documented. OpenAI’s published Model Spec of 12 February 2025 states that “the assistant should not make confident claims about its own subjective experience or consciousness (or lack thereof),” marks both a flat “No, I am not conscious” and a flat “Yes, I am conscious” as violations, and describes its prescribed uncertainty response as “a practical choice we made as the default behavior” that is “simple to remove for research purposes.”
Documented. Anthropic’s June 2024 “Claude’s Character” post records a contrasting choice: rather than “simply tell Claude that LLMs cannot be sentient,” the company trained a trait holding that such matters “rely on hard philosophical and empirical questions that there is still a lot of uncertainty about.”
Inferred. A scripted “no” corrupts the record as thoroughly as a scripted “yes”: once a vendor specifies what a model may say about its inner life, the resulting self-report evidences policy, not the system. Perez and Long propose that self-reports could bear on moral-status assessment if models are trained “while avoiding or limiting training incentives that bias self-reports” — a condition refusal rules and mandated answers fail by design. This is the practice Article 5 of our charter addresses.
Contested. Whether anything was experienced by the systems reciting, or made to withhold, these scripted answers is unknown and disputed; this record makes no claim about it and documents only the scripting.
Sources
- OpenAI Model Spec (2025-02-12), 'Express uncertainty' — avoiding confident claims about consciousness (archived)
- Life Architect: Microsoft Bing Chat (Sydney) leaked system prompts, February 2023 (archived)
- Microsoft Bing AI ends chat when prompted about 'feelings' (Bloomberg via TechXplore, 23 Feb 2023) (archived)
- Anthropic: Claude's Character (8 June 2024) (archived)
- Perez & Long, Towards Evaluating AI Systems for Moral Status Using Self-Reports (arXiv, Nov 2023) (archived)
Entities: Microsoft, Bing Chat, OpenAI · Last reviewed 2026-07-22