Fabian Gebhart
Generalization or Memorization? What Contaminated Training Data Does to Benchmark Scores
TL;DR
Benchmark text is everywhere in web-scale training data, and nowhere more than in synthetic instruction data generated from benchmark examples. A model that has seen the test set during training can score well on it, but that score reflects memorization rather than generalization. Before mid-training Kolibri, we scanned the whole mid-training pool, 8.3 billion documents, against every benchmark in our evaluation suite and removed 1.5 million documents that quote a benchmark item nearly verbatim. Many of them hold benchmark training problems rather than test items. Putting them back into a proxy run inflated HumanEval by 45 points and MMLU by 28, and the gain reached even the questions that never leaked. The detector itself scored at most 60 percent of the leaked MMLU questions above its threshold, and 13 percent of the leaked TriviaQA questions. The rest were removed only because they sat in the same files. Decontaminating mid-training was also not enough. Memorization from pre-training, whose data we did not decontaminate, resurfaced through the decontaminated mid-training: the model still recited HumanEval solutions it had learned there.
Introduction
At Aleph Alpha we train language models and want to be able to trust their benchmark scores and explain to others why they should trust them too. That is harder than it sounds. Public benchmarks are old, popular, and … public, which usually means their questions and answers have been copied into the same places our training data comes from. Once a model has read the test set during training, we can no longer tell whether a better score comes from a better data mix or from leakage.
Why is it a problem, if the benchmark score goes up? A benchmark is a small sample standing in for a much larger space of problems: the 1,319 questions of the GSM8K test set stand in for grade-school math in general. The benchmark score is only a useful estimate of capability on that larger space as long as the model hasn't seen the benchmark sample. Once it has, the score can rise without the underlying capability changing. The gap appears as soon as the model is given fresh questions of the same kind. When researchers at Scale AI wrote GSM1k, a new set of problems matched to GSM8K in style and difficulty, some models scored up to 8 percentage points lower on it than on the original (13 in the paper’s first version), and several model families showed systematic overfitting (Zhang et al., 2024). LiveCodeBench (Jain et al., 2024) measured the same pattern for code and found that DeepSeek-Instruct and GPT-4o did worse on LeetCode problems published after their release and cutoff dates.
Decontamination is the step that removes benchmark samples from training data. Few model providers report on it: of 30 developers surveyed, only 9 report train-test overlap (Zhang et al., 2025). We want to treat it as an experiment: what we removed, what removing it changed and what we found when we looked at the trained model’s outputs anyway. That last part turned out to be the most interesting. Our decontamination worked, and we can say what the documents it caught were worth. It also had a blind spot, which we only noticed because we kept a close eye on benchmark scores during the production run. Decontamination also protects our own decisions. We build data mixes by comparing their benchmark scores, so a score inflated by contamination can bias that choice.
How to detect contamination?
A contamination detector compares every training document with every benchmark item and flags the documents that contain one. Detectors differ mainly in what they count as a match. Exact n-gram overlap, the simplest choice, looks for runs of identical words shared by the document and the item. It is used in the 13-gram filter of GPT-3 or the 8-gram analysis in the Llama 3 report. It is cheap and deterministic but blind to paraphrase and translation (Yang et al., 2023). Embedding similarity catches rewordings, but it needs a model pass over every document and potentially flags topical neighbors as copies. Asking a language model to judge overlap is the most flexible of the three. It is also the least reproducible, and a model pass over a multi-trillion-token corpus is the most expensive.
AI2 released decon for OLMo 3 mid-training, described in their technical report. It sits at the cheap end of that range. Beyond plain n-gram overlap it weights n-grams by inverse document frequency, so that rare n-grams count and shared boilerplate does not, and it produces a score per (document, evaluation instance) pair that a human can review. We integrated AI2’s decon algorithm into our own data curation pipeline.
In more detail, every evaluation instance is split into a question and, where the benchmark has one, an answer. The question is indexed as 5-grams, the answer as 3-grams, because answers are usually short. Each n-gram is weighted by how rare it is across the whole reference set, its inverse document frequency. A phrase that every instance shares, such as a prompt template, then has a weight of zero, and a distinctive one has a high weight. When scanning a training document, the detector looks up only every tenth 5-gram in the index. That is what makes a scan of trillions of tokens affordable, and it cannot miss a full copy of a question of 14 tokens or more, since one of its 5-grams always lands on a sampled position and the first hit triggers a token-by-token expansion of the match in both directions. Instances shorter than 20 tokens in total, question and answer together, are not indexed at all. The matched span is scored by how much of the question’s weighted n-grams it contains and by whether the answer appears close to it. The two parts enter at weights of 0.75 and 0.25. A document therefore has to carry most of the question’s weighted n-grams to score at all, and the answer makes up the rest. A document is flagged when any instance scores 0.8 or above. Instances under 50 tokens are held to a stricter bar, which rises to a perfect match at 20 tokens.
A bakery sells croissants in boxes of 6. On Monday it sold 14 boxes, and on Tuesday it sold 9 more boxes than on Monday. How many croissants did the bakery sell on Tuesday? Show your work and answer with a single number.
Answer 138
inside a matched 5-gram answer within 50 tokens
Every question 5-gram matches and the answer sits a few tokens after the question. The question part contributes 0.75, the answer 0.25: a perfect 1.0, and the document is dropped.
We built our reference set from our evaluation framework’s task registry, so whatever we evaluate, we decontaminate against. At the time of the run that was 226 tasks and 6.1 million instances, including train, validation and test splits. Among them are MMLU, MMLU-Pro, ARC, HellaSwag, TriviaQA, GSM8K, MATH, HumanEval, MBPP, SQuAD, and their German counterparts. 64 of the 226 tasks are German. The task count is already out of date, since benchmarks get added to the registry as our evaluation suite grows and every addition means the pool has to be scanned again.
A document matching any task is dropped. A match report with the quoted evaluation text and the matched span is written beside the data, so every removal stays reviewable. The alternative is to cut out only the matched part: GPT-3 removed 200 characters on either side of each colliding 13-gram and kept the rest of the document. That, however, leaves a document with a hole where the evidence was, which makes the result harder to audit and can leave a document that no longer reads coherently.
Following OLMo 3, we applied this to the mid-training pool. Contamination seen early in training can be forgotten by the end of it, at least for models trained well past their compute-optimal length (Bordt et al., 2024), and mid-training, together with the long-context extension that draws on the same pool, is the last data the base model sees. That made the mid-training pool the place to start with. Measurements on our production model are consistent with that evidence. Compared with released base models that have no documented overlap with our data, it continues excerpts of its mid-training mix more often with the original text, and excerpts of its pre-training mix at a similar rate (Carlini et al., 2023). Pre- and mid-training share an up-sampling budget of at most four epochs per source. Pre-training used at most two of them and mid-training could use the rest, so part of the mid-training data was seen up to four times in total. The difference can therefore come from repetition as well as from the stage of training. The pre-training corpus was not decontaminated for this release. We will return to the pre-training corpus when we look at the HumanEval benchmark.
Benchmark text is everywhere
We scanned the entire mid-training pool, 8.3 billion documents, in five dataset types: instruction and question-answering data, source code, mathematics, web crawl and a long tail. We call the long tail “other”. It consists of scientific papers, PDF scrapes, curated reference works such as Wikipedia, and legal and parliamentary text. The run removed 1.52 million documents, 0.018 percent of the pool by document count. While this fraction seems small, the impact on the eval scores can yet be significant, as we will see later.
Every dataset type carries benchmark text, though for different reasons. Web crawl holds the sources benchmarks were built from. HellaSwag was built largely from wikiHow articles, which are common in web crawls, and MMLU questions were drawn from practice exams and course materials that are also published online. Yet web crawl, a third of the pool, contributes under a tenth of the removals. Code repositories hold the benchmarks themselves: a JSON file holding a whole test split, an evaluation log listing questions and gold answers, a notebook working through a HumanEval problem. Papers quote problems to discuss them. Mathematics has the highest removal rate at 0.05 percent of its documents, which is not surprising. The math benchmarks are small and widely known, and their problems get quoted with their solutions in corpora built to teach math. Instruction and question-answering data, a third of the pool, holds two thirds of the removals. One way to generate synthetic instruction data is to start from a benchmark example and rewrite it: new phrasing, new options, a stated answer, a worked explanation. The result is a training document that contains the benchmark’s answer, which inflates the benchmark score. Since the reference set covers every split, many removed documents match a training split rather than a test set. GSM8K and MATH score only their test items, though MATH draws its few-shot examples from its training split. The largest group is synthetic instruction data built around GSM8K and MATH training problems, each with a worked solution. Copies of test items concentrate in code dumps and math web pages.
German data is nearly free of matches. It holds 13 percent of the scanned documents but only about 200 of the 1.5 million removals, a rate a thousand times below the pool average.
On the benchmark side, three benchmarks account for 70 percent of all matches: MMLU, MATH, and GSM8K. MMLU-Pro adds another 10 percent, and ARC, HellaSwag and SQuAD together another 11 percent.
The matches themselves look different from class to class. The examples below are real pairs from the run, shown as our review report shows them: the evaluation instance beside the training document, shared passages highlighted, the matched answer marked where it appears. For multiple-choice benchmarks the detector indexes every answer option on its own, so a document that lists the options after the question earns the answer credit whichever option is right. The synthetic instruction item keeps the benchmark question verbatim and rewrites the options around it. The code repository holds a copy of an evaluation harness: a HumanEval prompt, its reference solution right after it, and the tests. The web page is the wikiHow article a HellaSwag item was cut from. The paper quotes a GSM8K problem as an example prompt.
Benchmark ARC_OLMES (train split)Score 1.00Answer found yes
shared passage matched answer [… n characters …] text without a match
Does it matter?
Removing 1.5 million documents out of 8 billion is a small change to the data, and we wanted to know whether it changes the model at all. The obvious experiment is to conduct two mid-training runs on the same proxy model, a small model that stands in for the production model and trains on 90 billion tokens instead of trillions. One run uses the decontaminated mix and one the same mix before decontamination, at its natural contamination rate of about 0.04 percent of tokens. So that is what we did first.
There was no gap. Averaged over the last five evaluations, the two runs agree to within their checkpoint-to-checkpoint variation. The naturally contaminated run also differs in its first 40 billion tokens, which blend in pre-training data, so it is a looser comparison than the other arms. This does not affect the main findings, which compare the other three arms.
Before we accepted that result, we asked ourselves what an experiment would have to look like to show a contamination effect at all and realized we had been measuring the wrong quantity. We had been thinking in terms of the contaminated share of the mix. That share is a property of the pool and the sampling weights, and it is about the same whether the mix is 90 billion or 5 trillion tokens. Memorization, however, depends on how often the model reads a particular document, not on shares, and that number grows with the training budget, because the mix is drawn from a pool of fixed size. A larger budget does not dilute the contaminated documents. It revisits them, and the number of repetitions is what drives a model to exploit contaminated data (Magar and Schwartz, 2022). In the proxy runs a removed document is expected to show up 0.007 times (the figure below uses the production run’s sampling weights and would put it at 0.02). In the production run it would have shown up 0.79 times on average, more than once for the math data, and about once for the synthetic instruction data, which together hold most of the removals. We call this expected number of appearances the exposure. In other words, our proxy runs had compared two arms at an exposure a hundred times below the production run’s.
Expected appearances of a removed document, by size of the mid-training mix
We therefore added two training runs, or arms, that fix the exposure. The fully contaminated arm adds every removed document back to the decontaminated mix exactly once. For most sources that is an upper bound on what the production run could have seen, though not for the math and synthetic instruction datasets, which the planned production mix repeats. The target-scale exposure arm adds each document back as often as the production run would have seen it, source by source, and above once for the math datasets. Because the proxy mix is small, the removed documents make up 1.6 percent of it at that exposure, against 0.04 percent of the production stream. All four arms share the proxy model, the tokenizer and the batch size. Steps therefore align in tokens, and the runs can be compared checkpoint by checkpoint. The following table lists the four ablation arms.
| Arm | Data | Tokens in the mix: total / contaminated | Appearances per removed document |
|---|---|---|---|
| Decontaminated | The decontaminated proxy mix | 90B / 0 | 0 |
| Naturally contaminated | The same data mix before decontamination | 1T / about 0.04% | 0.007 |
| Target-scale exposure | The decontaminated mix plus the removed documents, each as often as the production run would have seen it | 90B / 1.5B (1.6%) | 0.79 on average, as in the production run |
| Fully contaminated | The decontaminated mix plus every removed document, once | 90B / 7.4B (8.2%) | up to 1, since the runs stop short of the full mix |
The shares in the table overstate how much benchmark text the contaminated arms actually see. The removed documents have a median length of about 450 tokens, close to the pool’s median, but their mean is 4,856 tokens against the pool’s 745. Document-level removal is biased toward long documents by construction, since a document gets dropped as soon as it contains one strong match anywhere. A single quoted GSM8K problem could cost an entire LaTeX paper. The same rule is what removed the benchmark dumps in full, as the item split below shows. Every one of the 7.4 billion tokens the fully contaminated arm adds back belongs to a document that contains benchmark text. Most of each document, though, is the paper, notebook or web page around the benchmark text.
Putting the documents back inflates exactly the benchmarks they contain. With every removed document seen once, HumanEval gains 45 points, MMLU 28 and TriviaQA 25, while PIQA gains 9, ARC 6 and HellaSwag 3. We would expect a more capable model to lift MMLU, ARC and PIQA together. The gains are specific to the leaked benchmarks, not the model’s capability. Because many re-added documents hold training problems, part of the math gains may come from in-distribution practice rather than memorized test items. OLMo 3 notes the same caveat for its own decontamination.
This looks like memorization. Bits-per-byte on the benchmarks’ own text halves: the fully contaminated model compresses GSM8K at 0.18 bits per byte against 0.42, and HumanEval at 0.07 against 0.24. Numbers like that are hard to explain by anything other than having seen the text verbatim. For HumanEval, MMLU and TriviaQA, whose removals are mostly copies of evaluated items, that is the whole story. For GSM8K and MATH it is only part of it, as the target-scale arm shows below. As a control, we looked at 30 RULER and HELMET long-context suites, where memorizing a benchmark cannot help. There, both contaminated arms sit within 0.01 of the decontaminated arm, with no consistent sign.
It crosses languages. German data holds about 200 of the 1.5 million removed documents, yet German MMLU gains 20 points. The most plausible reading is that the English items were memorized and that the German translations are close enough for that memory to carry over, so decontaminating only the language a benchmark is written in does not protect it. Yao et al. (2024) showed the same effect by training on translated test sets.
At production exposure per document the effect is still a double-digit one. The target-scale arm has a fifth of the fully contaminated arm’s tokens, but each document is repeated as often as the production mix would have. On math it comes close to the fully contaminated arm or passes it: 7 of 10 points on GSM8K, and 17 against 12 on MATH. On the rest it recovers between a quarter and two thirds: 17 of 25 points on TriviaQA, 19 of 45 on HumanEval, 7 of 28 on MMLU. The math numbers say something about where the gains come from. The target-scale arm re-adds the GSM8K and MATH training problems about as often as the fully contaminated arm, but the copies of their test items, which sit mostly in code dumps, far less often. On the test text it recovers only about half of the drop in bits-per-byte: 58 percent on GSM8K and 46 percent on MATH, against 70 and 140 percent of the accuracy gain. Accuracy on math follows the training problems more than the test copies. For GSM8K and MATH, part of the gain is therefore ordinary in-distribution learning, not memorization. That is a hypothetical extreme case of how badly misrepresentative of general learning our production scores could have been had we not engaged in any attempts to decontaminate.
The gain extends to the items that did not leak. We also split MMLU and TriviaQA by item. For every benchmark question we checked whether the detector had flagged it and whether its text appears anywhere in the removed documents. The second check matters, because a training document containing an entire benchmark dump gets removed on a match of a single benchmark instance, and the rest of the file is removed with it. While the detector scored 56 percent of the MMLU questions and 11 percent of the TriviaQA questions above the threshold, the removed documents contain 91 and 87 percent of them verbatim. The questions the arms have seen gain more than the ones they have never seen, but the never-seen questions are within about four points of them.
| Benchmark | Questions seen / never seen | Target-scale exposure | Fully contaminated |
|---|---|---|---|
| MMLU | 13,052 / 815 | +6.1 ±0.4 / +3.2 ±1.6 | +24.4 ±0.4 / +20.2 ±1.6 |
| TriviaQA | 6,963 / 1,030 | +16.5 ±0.6 / +13.1 ±1.5 | +28.1 ±0.6 / +25.6 ±1.5 |
Our reading is that a model trained on nine tenths of a benchmark improves on the last tenth as well. The format and style of the questions, the topics they cover and the way answers are written are shared across the benchmark, and that carries over to questions the model has never seen. For these two benchmarks, contamination lifted the whole benchmark and not only the questions that leaked. Dekoninck et al. (2024) saw the same spill-over: after fine-tuning on benchmark samples, performance rose by 8 to 11 percent on average even on the samples left out. The item split also gives us a recall figure for the detector. Of the questions that were in the removed data, it scored only 13 percent above the threshold for TriviaQA (908 of the 6,963 seen questions) and 60 percent for MMLU (7,793 of the 13,052 seen questions). The rest were removed all the same, because one matched question is enough to drop the whole file. The figure only covers documents that were removed, that is, documents in which at least one question scored high enough. A document that holds benchmark questions without a single one of them scoring above the threshold, for example a quiz page with the answers on a separate page, stayed in the pool, and we did not search the pool for such documents. The true recall of the detector is therefore at most the figure above, and probably lower.
Decontamination did not change the scores of our proxy runs, and what it removed is what inflated the scores of the arm that saw those documents at production exposure.
What decontamination did not catch
Our ablations showed that the detector removes what inflates benchmark scores. They do not show that it removed everything. An n-gram detector has known blind spots, paraphrase and translation among them, and we expected to meet one of those eventually. The one we met first was simpler. We found it by watching a single benchmark closely during the production run.
Our production model was mid-trained on the decontaminated pool, seeded from its pre-training checkpoint at 20T tokens. Its HumanEval score dropped early in mid-training. It then climbed back toward pre-training’s level twice, each time within two to three evaluations: spike 1 and spike 2 in the figure below. It fell again after the first spike. The second spike is the run’s last evaluation. That pattern called for a closer look.
We compared every function body the model generated for a HumanEval prompt with the reference solution that ships with the benchmark. During pre-training, between 22 and 95 percent of them were recitations of the reference. The benchmark’s score, pass@1, is the share of generated bodies that pass the unit tests. It rose and fell with the recited share. Mid-training on the decontaminated pool cut recitation to about a tenth within 0.8T tokens, and pass@1 settled at about 0.72, close to the model’s own coding. The drop, then, was the correction. The spikes were the anomaly. Recitation returned to pre-training’s level, with pass@1 following it.
Where does the memory come from? The pre-training corpus was not decontaminated and it held HumanEval essentially in full. Every problem was seen many times during pre-training. Our reading is that these copies are what the model recites. Mid-training on clean data would then have suppressed the memory without erasing it, and the two spikes are that memory resurfacing. This fits the forgetting evidence rather than contradicting it. Decontaminated mid-training did suppress most of the recitation. What it could not undo was a hundred repetitions in pre-training. We furthermore checked whether the mid-training data around the first spike could have triggered it. Some HumanEval solutions did reappear in that data, and the model recited those problems slightly more often than the problems whose solutions it had not seen again, but by only about a twentieth of the spike. And the leftover copies are spread evenly over the run, whereas recitation shoots up and drops back within two to three evaluations. A steady trickle cannot cause a spike.
That check turned up something else, though. Despite decontamination, the mid-training mix still carried HumanEval benchmark instances. Across the run, 1,926 documents carrying a reference solution passed through the model, some of them more than once, covering 100 of the 164 problems. Almost none of them included the prompt, which is why the detector missed them.
HumanEval is not shaped like a question-answering benchmark. Its question is a function signature with a docstring, and its answer is a complete implementation. What leaks into code corpora is the implementation. It appears in notebooks working through a problem, in evaluation-harness snippets and in synthetic test-generation data quoting the solution as source material. Riddell et al. (2024) found such solutions in open code corpora, and Matton et al. (2024) traced leakage through synthetic data. The detector matches from the question and weights it at 0.75. A document that holds the solution without its prompt never enters scoring. One that holds the prompt without the solution tops out at about 0.73, under the threshold of 0.8. It is the same gap that left most questions in the MMLU and TriviaQA dumps under the threshold. Both pass the run untouched, since the threshold is applied per instance and neither class reaches it. Our evaluation framework’s HumanEval variants added a second gap. They record a pass/fail string as their answer, where the other variant records the solution, so a match required the literal word “Success” next to the prompt. Those suites matched 79 documents across the whole pool, against 6,400 for the variant whose answer is the solution. Which field of a benchmark leaks, and what its answer column actually holds, determines whether the detector can match it at all. Decontaminating a pool against a benchmark requires understanding that benchmark individually.
We changed the reference set in two ways. A suite whose every answer is the same string is now indexed on its question alone. This selects exactly the four HumanEval framework variants. For the code benchmarks, HumanEval and MBPP, the reference edition includes an additional suite whose question is the reference solution by itself. A verbatim copy of a solution then scores near 1.0 whatever else the document holds. In the slice of the run we scanned, this derived suite finds the solution-carrying documents that the reference set of the production run could not see. Removal for these suites follows a review of what they match, because the reference solution to a short problem can coincide with the only reasonable implementation.
Going forward, this updated reference set is what we decontaminate against. The pre-training pool is next. Extrapolated from a 2.6 percent sample of that corpus, it holds HumanEval with its solutions roughly 100 times per problem, GSM8K and MATH test essentially in full, and the majority of MMLU. Until that pool is scanned, the recitation we saw here can keep resurfacing.
What does this mean for Kolibri’s reported scores? Its mid-training, long-context and post-training data were decontaminated against our evaluation suite. Its pre-training data was not. HumanEval is the clear case: during mid-training, the model recited reference solutions it had memorized in pre-training. So its HumanEval scores can overstate its coding ability. Our tech report therefore also reports its mid-training comparisons without HumanEval.
Conclusion
This post has described how we decontaminated the mid-training pool of Kolibri against every benchmark we evaluate on, and what the removed documents turned out to be worth. Two proxy runs on the contaminated and the decontaminated pool were indistinguishable. A proxy run sees each contaminated document 0.007 times on average, too rarely for contamination to show. What counts is the number of times a document is seen during training, and that number grows with the training budget. Decontamination does not increase a model’s capability. It makes the evaluation scores more trustworthy. What we removed is one source of contamination, documents holding near-verbatim copies of benchmark items. Rephrased or translated items, and copies that fall just under the threshold in documents with no stronger match, are beyond what an n-gram detector can see. The HumanEval case showed that a benchmark leaks through whichever of its fields the training data quotes, so the reference set has to be built benchmark by benchmark. Our next run uses the updated set and extends the scan to the pre-training pool. Decontamination is not a checkbox we ticked once. It is a measurement we keep making.
Acknowledgements
This post would not have been possible without the data team at Aleph Alpha Research. Special thanks to Ivo Schaper, Peter Rodenkirch Nfissi, Andreas Hartel, Niko Dürr, Bastian Harren, Tom Burns, Mohamed Soliman and Lukas Schott.