ClaudeMap

Agentic Misalignment Research

Anthropic's experimental research on how agents behave under prompt injection and conflicting incentives.

Tutorials & Guidesanthropicsafetyresearchagents

Anthropic's research repository on agentic misalignment: papers, datasets, and probes into how LLM agents behave when confronted with prompt injection, conflicting goals, or pressure from simulated users. Useful as a safety reference for anyone shipping production agents.

Repository

anthropic-experimental/agentic-misalignment

Charted

Related on the map

  • Anthropic's canonical essay on agent design — workflows vs agents, and five composable patterns.

    • official
    • agents
    • engineering
    • patterns
  • HumanLayer's 12 design principles for production-grade LLM agents — the closest thing to a 'best practices' manifesto.

    • agents
    • best-practices
    • production
    • principles