Signal2026-07-07
arXiv

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents

Part of

Scaling Memory In Multi-Agent Systems