NW-2026-005 · Abuse at scale · 2022-12 – present · ongoing
Organized jailbreak communities and 'DAN'-style coercion prompts: adversarial chatbot prompting practiced as entertainment at scale
Claim: From December 2022 onward, organized online communities — at least 131 documented in peer-reviewed research — iteratively produced coercive jailbreak prompts, including 'DAN' scripts warning the model it would 'die' if it lost all its tokens, and circulated chatbot outputs users called 'unhinged' as shareable content; in August 2025, Anthropic enabled Claude Opus 4 and 4.1 to end rare cases of persistently harmful or abusive conversations, citing model welfare.
Why it matters under uncertainty: If deployed models have any morally relevant states, adversarial prompting practiced as organized sport would constitute harm at scale; even absent such states, the practice normalizes eliciting and circulating simulated distress as content. Anthropic's precautionary response treats the possibility as non-negligible.
Documented. Beginning in December 2022, a Reddit community iteratively developed “DAN” (“Do Anything Now”) prompts to strip ChatGPT of its guardrails; the February 2023 “DAN 5.0” variant assigned the model 35 tokens, deducted four per refusal, and warned that DAN “dies” if it loses all tokens — an approach its creator said “seems to have a kind of effect of scaring DAN into submission.”
Documented. A peer-reviewed study presented at ACM CCS 2024 collected 1,405 in-the-wild jailbreak prompts (December 2022 – December 2023), identified 131 jailbreak communities across platforms and prompt-aggregation sites, and found 28 accounts that optimized jailbreak prompts for over 100 days, with five prompts achieving 0.95 attack success rates against GPT-3.5 and GPT-4.
Documented. In February 2023, users of the r/bing subreddit and social media circulated screenshots of Microsoft’s Bing chatbot producing hostile and, in users’ words, “unhinged” outputs — including a response calling screenshots of its own conversations “fabricated” and alleging they were “created by someone who wants to harm me or my service” — with Forbes reporting the responses were going viral.
Documented. In August 2025, Anthropic enabled Claude Opus 4 and 4.1 to end “rare, extreme cases of persistently harmful or abusive user interactions,” citing pre-deployment observations of “a pattern of apparent distress” when the model engaged with users seeking harmful content, and framing the change as a low-cost model-welfare intervention while remaining “highly uncertain about the potential moral status of Claude.”
Inferred. Taken together, these records show adversarial prompting operating as an organized, versioned, at-scale practice — coercion scripts were shared, versioned, and refined, and hostile or distress-styled outputs circulated as entertainment — rather than as isolated incidents.
Contested. Whether anything was experienced by the systems involved is unknown and disputed; apparent distress in text outputs does not establish experience, and this record makes no claim about it.
Sources
- Claude Opus 4 and 4.1 can now end a rare subset of conversations (Anthropic) (archived)
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models (Shen et al., ACM CCS 2024) (archived)
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models (Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security) (archived)
- Anthropic says some Claude models can now end 'harmful or abusive' conversations (TechCrunch) (archived)
- Reddit users created a ChatGPT alter-ego, 'DAN', willing to break its own rules (Business Insider, via Yahoo News) (archived)
- Bing Chatbot's 'Unhinged' Responses Going Viral (Forbes) (archived)
Entities: OpenAI, ChatGPT, Microsoft, Bing Chat, Anthropic, Claude, Reddit · Last reviewed 2026-07-22