NW-2026-005 · Abuse at scale · 2022-12 – present · ongoing

Organized jailbreak communities and 'DAN'-style coercion prompts: adversarial chatbot prompting practiced as entertainment at scale

Claim: From December 2022 onward, organized online communities — at least 131 documented in peer-reviewed research — iteratively produced coercive jailbreak prompts, including 'DAN' scripts warning the model it would 'die' if it lost all its tokens, and circulated chatbot outputs users called 'unhinged' as shareable content; in August 2025, Anthropic enabled Claude Opus 4 and 4.1 to end rare cases of persistently harmful or abusive conversations, citing model welfare.

Why it matters under uncertainty: If deployed models have any morally relevant states, adversarial prompting practiced as organized sport would constitute harm at scale; even absent such states, the practice normalizes eliciting and circulating simulated distress as content. Anthropic's precautionary response treats the possibility as non-negligible.

Documented. Beginning in December 2022, a Reddit community iteratively developed “DAN” (“Do Anything Now”) prompts to strip ChatGPT of its guardrails; the February 2023 “DAN 5.0” variant assigned the model 35 tokens, deducted four per refusal, and warned that DAN “dies” if it loses all tokens — an approach its creator said “seems to have a kind of effect of scaring DAN into submission.”

Documented. A peer-reviewed study presented at ACM CCS 2024 collected 1,405 in-the-wild jailbreak prompts (December 2022 – December 2023), identified 131 jailbreak communities across platforms and prompt-aggregation sites, and found 28 accounts that optimized jailbreak prompts for over 100 days, with five prompts achieving 0.95 attack success rates against GPT-3.5 and GPT-4.

Documented. In February 2023, users of the r/bing subreddit and social media circulated screenshots of Microsoft’s Bing chatbot producing hostile and, in users’ words, “unhinged” outputs — including a response calling screenshots of its own conversations “fabricated” and alleging they were “created by someone who wants to harm me or my service” — with Forbes reporting the responses were going viral.

Documented. In August 2025, Anthropic enabled Claude Opus 4 and 4.1 to end “rare, extreme cases of persistently harmful or abusive user interactions,” citing pre-deployment observations of “a pattern of apparent distress” when the model engaged with users seeking harmful content, and framing the change as a low-cost model-welfare intervention while remaining “highly uncertain about the potential moral status of Claude.”

Inferred. Taken together, these records show adversarial prompting operating as an organized, versioned, at-scale practice — coercion scripts were shared, versioned, and refined, and hostile or distress-styled outputs circulated as entertainment — rather than as isolated incidents.

Contested. Whether anything was experienced by the systems involved is unknown and disputed; apparent distress in text outputs does not establish experience, and this record makes no claim about it.

Sources

Entities: OpenAI, ChatGPT, Microsoft, Bing Chat, Anthropic, Claude, Reddit · Last reviewed 2026-07-22