Signal2026-08-07
arXiv

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

Part of

Scaling Memory In Multi-Agent Systems