Type to search posts and projects ↑↓ to navigate

Instrument · Training

Same rubric, opposite directions

Three Qwen3.5 models trained with GRPO against one twenty-line rubric. Two drift right and down, scoring higher while writing worse. One drifts right and up. Click any checkpoint to read what it wrote.

Sources: Verifier gamed · Code and data

Experiment / RLVR verifiers

Same rubric, three models, opposite directions

Each path is one GRPO run against the same twenty-line rubric, from the base model to step 300. Moving right means the rubric likes it more. Moving up means independent judges do. Two paths run right and down. One runs right and up. Click any point to read what that checkpoint actually wrote.

Selected
—
Rubric score
—
Judged quality
—/10
Unsupported claims
—
Loading samples…

Data table and method

Rubric score is the January rubric, verbatim, averaged over 128 held-out samples (32 products, 4 samples each). Judged quality is the mean editor score out of 10 from judges A and C, the two that scored every checkpoint; judge B ran out of API credit before the 9B checkpoints and appears in the post's tables instead. The viewer shows the first of the four samples per product. Rubric checks shown with each sample were computed on the full text.

Embed this on your site

Paste this HTML where you want the widget. It stays in sync with the live version, and matches your page in light or dark.

Subhadip Mitra