ALPHAZETARESEARCH & EVIDENCE

Tooling case study

How we used agents to evaluate Jev—and changed our minds about its best use

Published · Research through

Executive summary

Jev showed promise as a cheap way to decide which claims deserved closer inspection. It did not establish a clear improvement in the broader agent task, and a high score could still conceal an error.

These were small, internal experiments with model-based review, not an independent human-labelled benchmark. Research through 24 September 2026; Jev 1.13.0.

Open a finding to explore it.