Anthropic 官方公告:Fable 5.1 与 Mythos 5.1——同底模,两档防护。
- official
- models
- announcement
- fable
Anthropic 关于 agent 在 prompt injection 与利益冲突下行为表现的实验性研究。
Anthropic 的 agentic misalignment 研究仓库:论文、数据集与探针,研究 LLM agent 在面对 prompt injection、目标冲突或模拟用户施压时的行为。适合发布生产级 agent 的团队作为安全参考。
收录日期
Anthropic 官方公告:Fable 5.1 与 Mythos 5.1——同底模,两档防护。
Anthropic 关于 agent 设计的标杆长文——工作流 vs agent,以及五种可组合模式。
HumanLayer 的 12 条生产级 LLM agent 设计原则——最接近「最佳实践宣言」的东西。