Paper visualizer
Most AI coding agents are one big model plus a thin wrapper. HERMES flips that: it turns every part of a code repo into a small agent of its own, a dev-primitive, each with a resident LLM that knows its own code and dependencies. A dependency-aware step wakes only the parts a task needs, and a diagnosis step traces test failures back to the parts that must change. The headline result: the harness does the heavy lifting, not the model.
On Terminal-Bench 4.0, running the dev-primitives on the small Qwen3-8B model keeps performance within 4.5 points of the all-GPT-5.6-Sol setup, while cutting inference cost 26.2%. The red sliver is money never spent.
Paper numbers: 26.2% inference-cost reduction on Terminal-Bench 4.0, performance within 4.5 points of the homogeneous GPT-5.6 Sol configuration.
Each dot is a component in a repo. A naive harness wakes everything for every task, which is how context explodes. HERMES wakes only the components the task's dependencies touch.
Illustrative simulation of the dependency-aware activation mechanism, not paper data. The paper reports the mechanism and its benchmark gains, not a per-task component count.
On whole-repository migration with GPT-5.6 Sol, swapping the Codex harness for HERMES took success from 6.5% to 31.0%. Same model, same effort setting. Drag the slider to walk between the two harnesses.
Illustrative interpolation between the paper's two reported endpoints (6.5% and 31.0%). The paper measured the endpoints, not the points in between.
These are the paper's reported numbers on four software-engineering benchmarks, with matched baseline harnesses and fixed effort settings. A better harness helps most on long-horizon work where context explodes; on short tasks the model still matters. Read the paper before citing the numbers.