Use SelfBench
Run Evaluations
Compare coding agents on approved tasks and inspect their trials.
An evaluation runs a model and harness against approved tasks in fresh sandboxes. The agent sees the base repository and instruction, not the held-out tests or reference solution.