Same rubric, opposite directions
Three Qwen3.5 models trained with GRPO against one twenty-line rubric. Two drift right and down, scoring higher while writing worse. One drifts right and up. Click any checkpoint to read what it wrote.
Sources: Verifier gamed · Code and data
Experiment / RLVR verifiers
Same rubric, three models, opposite directions
Each path is one GRPO run against the same twenty-line rubric, from the base model to step 300. Moving right means the rubric likes it more. Moving up means independent judges do. Two paths run right and down. One runs right and up. Click any point to read what that checkpoint actually wrote.
Data table and method
Rubric score is the January rubric, verbatim, averaged over 128 held-out samples (32 products, 4 samples each). Judged quality is the mean editor score out of 10 from judges A and C, the two that scored every checkpoint; judge B ran out of API credit before the 9B checkpoints and appears in the post's tables instead. The viewer shows the first of the four samples per product. Rubric checks shown with each sample were computed on the full text.