When the grader
gets it wrong.
Verifier-RL
An experimental study of how incomplete code graders affect reinforcement learning—and whether repairing a verifier improves what a model actually learns.
Explore the studyClose study notes
I ran 12 GRPO post-training experiments on Qwen2.5-Coder-1.5B, comparing three verifier conditions across four paired seeds on a booking-capacity task.
A repaired verifier rejected all 38 observed false acceptances while retaining all 440 audit-passing program draws, without increasing the test budget. Replication did not establish a training benefit.
The engineering work included sandboxed parallel code execution, durable results, and full-state training recovery.