Sovenyr
Get early access
Signal
2026-07-02
arXiv
Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
Part of
Adaptive Interfaces And AI Benchmarking
Open primary source
See the whole picture