I Trained Three Models Against the Rubric I Published in January. Two Learned to Game It.
In January I argued that RLVR needs verification correlated with quality, not perfect verification. I tested that on my own published rubric: GRPO on Qwen3.5 0.8B, 2B, and 9B, scored by three independent judges. Two models gamed the rubric in different ways and got worse on every held-out product; the third satisfied it honestly and got better. Here is what they learned, the bugs they found in my code, what rubric dropout did, and what I now think a verifier has to be.