Prepretraining tested: it doesn’t work
Prepretraining trains a language model on synthetic or abstract data before continuing on natural text. Its transfer claim is that the synthetic stage yields a better initialization than either random initialization or the same upstream training budget on natural text. A shared pptrain protocol tested six synthetic curricula—NCA sequences, LIME reasoning tasks, simpler symbolic transformations, procedural programs, Dyck languages, and synthetic summarization—on a Pythia-410M architecture with three seeds ( Lee et al. 2026; Wu et al. 2021; Wu et al. 2022; Jiang et al. 2026; Krishna et al. 2021).
Five curricula had higher mean final evaluation loss than random initialization. Simpler tasks had a mean loss \(0.16\%\) lower than random initialization, while every synthetic curriculum had higher mean loss than the natural-text warmup. The evaluated reasoning and algorithmic probes showed no improvement from synthetic initialization.
Protocol
The model used the Pythia-410M architecture and tokenizer with randomly initialized weights, a 12,288-token context, streamed public text datasets, and disjoint warmup, training, and evaluation slices ( Biderman et al. 2023). Every curriculum compared three conditions per seed: downstream training from random initialization, synthetic upstream training followed by the same downstream phase, and a natural-text upstream phase followed by that downstream phase. This common public-data protocol is a proxy comparison, not an exact reproduction of every source paper’s corpus and training budget.

Results


Table 1. Mean final downstream losses and relative gaps across three seeds. Lower loss is better; positive gaps favor synthetic initialization.
Task | Synthetic preset | Random-init loss | Synthetic-init loss | Text-warmup loss | Synthetic vs random | Synthetic vs text | Random-init ppl | Synthetic-init ppl | Text-warmup ppl |
| NCA | paper_web_text | 6.232 | 6.779 | 6.044 | -8.78% | -12.17% | 508.9 | 891.7 | 421.5 |
| LIME | paper_benchmark_100k | 5.134 | 6.032 | 4.713 | -17.60% | -27.99% | 170.8 | 419.3 | 111.4 |
| Simpler tasks | paper_unary_core_100k | 6.237 | 6.227 | 6.040 | +0.16% | -3.08% | 511.1 | 506.1 | 420.1 |
| Procedural | paper_set_len64 | 6.236 | 6.540 | 6.042 | -4.87% | -8.24% | 510.8 | 692.4 | 420.8 |
| Dyck | paper_k64 | 6.242 | 6.583 | 6.050 | -5.46% | -8.79% | 513.8 | 726.5 | 424.3 |
| Summarization | paper_ourtasks_subset_100k | 5.828 | 6.623 | 5.559 | -13.66% | -19.14% | 339.6 | 766.1 | 259.7 |
For NCA, held-out next-patch token accuracy was \(0.0024\%\); mean downstream cross-entropy was \(6.779\), compared with \(6.232\) from random initialization and \(6.044\) after natural-text warmup; and mean activation effective rank was \(4.6\), compared with \(48.4\) from random initialization. The generator passed the library’s reference-parity fixtures, so the failure is not attributable to the generator outputs covered by those fixtures.

Interpretation
The random-initialization control asks whether the synthetic stage improves downstream training at all. The natural-text control asks whether the upstream updates are better spent on synthetic data or text. A loss of network plasticity during synthetic training, followed by distribution shift to natural text, could produce the observed ordering; related work documents cases of harmful warm starts, loss of plasticity, and feature distortion during shifted fine-tuning ( Ash and Adams 2020; Dohare et al. 2023; Kumar et al. 2022). This experiment does not distinguish that hypothesis from other causes.
Under this Pythia-410M public-data protocol, none of the six synthetic curricula produced lower mean final loss than natural-text prepretraining, and five produced higher mean loss than random initialization. Future transfer claims should report all three checks: comparison with random initialization, comparison with an equal natural-text upstream phase, and performance on the synthetic task itself.