Studio · Computed teaching policy

Fig. 1

A fresh policy. Train step samples a group, scores it, and updates toward the better half.

prompt 1 · step 0

Computed teaching policy · 8 candidate responses per prompt · exact KL over the catalog

prompt

Policy over all 8 responses

dashed = πref
Group relative. The baseline is the group's own average reward — no value model is trained. Rewards and quality are hand-scored features of fixed candidate texts.