Three bar charts showing ablation results for Coherence, Factuality, and Safety across different numbers of rollouts (2, 4, 8, 16).
Figure 9 Ablation results on the number of rollouts in online DPO training for models trained for Quality (left), Factuality (middle), and Safety (right).

Table 7 Judge comparison: evaluation results on generation quality and coherence for ablations of using GPT-OSS-120B as judge versus using our finetuned llama3 judge during online DPO training. The number of rollouts used is 8 in these experiments.

Pretraining for Quality Generation Quality Standard Evals (avg) Coherence Eval
Self-Improving Pretraining (finetuned Llama3 as judge) 72.1 49.6 72.7
Self-Improving Pretraining (GPT-OSS-120B as judge) 84.3 51.1 86.8

generation to produce rewards. Results are given in Appendix Table 18, Table 19 and Table 20 for various settings. Overall we find deterioration in performance from using pivots, leaving how to make judgments faster while maintaining quality an open question.