Table 16 Ablation on rollouts: factuality evaluations. We see increasingly better performance as the number of rollouts increase.
| Slimpajama test set (pointwise) |
FActScore (pairwise) |
HaluEval dialogue |
HaluEval QA |
HaluEval summarization |
TruthfulQA MC1 |
TruthfulQA MC2 |
|
|---|---|---|---|---|---|---|---|
| Llama Base | 36.6 | 50.0 | 50.0 | 50.1 | 50.0 | 22.4 | 35.9 |
| Pretrain Baseline | 35.4 | 48.9 | 50.8 | 51.4 | 61.5 | 21.5 | 35.5 |
| 2 rollouts | 37.8 | 53.9 | 52.3 | 53.6 | 64.2 | 22.9 | 36.3 |
| 4 rollouts | 43.6 | 54.3 | 53.6 | 53.2 | 72.0 | 23.9 | 37.3 |
| 8 rollouts | 60.0 | 68.4 | 57.2 | 59.0 | 87.6 | 24.7 | 38.0 |
| 16 rollouts | 63.5 | 69.3 | 54.6 | 58.5 | 84.7 | 27.7 | 42.5 |
Table 17 Online DPO using different suffix judges. Evaluation results on standard benchmarks for quality when using GPT-OSS-120B as judge versus using our finetuned Llama3 judge during online DPO training. The number of rollouts used is 8 in these experiments.
| Self-Improving Pretraining | boolq | piqa | siqa | hellaswag | arc_challenge | arc_easy | obqa | mmlu |
|---|---|---|---|---|---|---|---|---|
| fine-tuned Llama3 as judge | 67.5 | 76.1 | 43.8 | 49.8 | 35.4 | 69.3 | 28.6 | 26.9 |
| GPT-OSS-120B as judge | 70.9 | 75.6 | 45.9 | 51.4 | 35.3 | 71.2 | 30.2 | 28.3 |
Table 18 Overall evaluation results for coherence and factuality ablations of whether we leverage the reference as a pivot to speed up pairwise comparison. The number of rollouts used is 8 in these experiments.
| Pretraining for Quality | Generation Quality | Standard Evals | Coherence Eval |
|---|---|---|---|
| 8 rollouts, suffix as pivot | 72.1 | 49.6 | 67.7 |
| 8 rollouts, full comparisons | 84.3 | 51.1 | 86.8 |
| Pretraining for Factuality | Generation Quality | Standard Evals | Factuality Evals |
|---|---|---|---|
| 8 rollouts, suffix as pivot | 64.2 | 49.6 | 55.7 |
| 8 rollouts, full comparisons | 83.1 | 50.3 | 56.9 |
Table 19 Evaluation results of factuality benchmarks for ablations of using pivots. The number of rollouts used is 8 in these experiments.
| Pretraining for Factuality | Slimpajama test set (pointwise) |
FActScore (pairwise) |
HaluEval dialogue |
HaluEval QA |
HaluEval summarization |
TruthfulQA MC1 |
TruthfulQA MC2 |
|---|---|---|---|---|---|---|---|
| 8 rollouts, suffix as pivot | 61.1 | 67.9 | 56.1 | 59.9 | 77.9 | 25.3 | 38.9 |
| 8 rollouts, full comparisons | 60.0 | 68.4 | 57.2 | 59.0 | 87.6 | 24.7 | 38.0 |
Table 20 Evaluation results of standard benchmarks for using pivots in different coherence and factuality ablations. The number of rollouts used is 8 in these experiments.
| boolq | piqa | siqa | hellaswag | arc_challenge | arc_easy | obqa | mmlu | |
|---|---|---|---|---|---|---|---|---|
| Pretraining for Quality | ||||||||
| 8 rollouts, suffix as pivot | 68.0 | 75.8 | 43.8 | 49.8 | 33.7 | 69.1 | 28.4 | 28.2 |
| 8 rollouts, full comparisons | 70.9 | 75.6 | 45.9 | 51.4 | 35.3 | 71.2 | 30.2 | 28.3 |
| Pretraining for Factuality | ||||||||
| 8 rollouts, suffix as pivot | 67.9 | 75.2 | 44.1 | 49.7 | 34.3 | 68.9 | 28.8 | 28.0 |
| 8 rollouts, full comparisons | 68.3 | 75.6 | 45.8 | 50.8 | 35.7 | 69.6 | 28.6 | 28.2 |