Meta-note: This is somewhat of a different flavor of post than my usual policy and geopolitics-focused writing. The reason for this is because I’m currently working on becoming more technical! As I do this, I’ll be writing up the things I learn as I go, with a particular focus on technical lessons that might be useful for policy people to know about. I think this is one of them.
Training corpuses1 are huge. It would be great if we could just run AIs on the entire training corpus to filter for signs of data poisoning, but in practice, this isn’t really done. Why?
I previously thought it was for cost reasons! Now, I think this is mainly a lab implementation problem.
In fact, it is possible to run a model on your entire pre-training corpus for only about 3.33% additional training cost, provided the model you run to review the training corpus (henceforth referred to as the “Review Model”) is about 1/10th as large as the model you’re training.
Why is this true? To answer that, let’s look at how much compute it takes to train a model, and compare that with the amount of compute it takes to run inference on all of the tokens used in training.
How many FLOPs do you need to train a model?
The following is a rough formula you can use to calculate number of FLOPs needed to train a model:
See this image for a breakdown of the figures!

Training requires both a forward pass (for the model to generate next-token predictions and compute the loss2 against the actual tokens) and a backward pass (computing, via backpropagation, the gradients that indicate how each weight should change, which an optimizer then applies).
In contrast, inference – which is what the Review Model would be doing – only needs the forward pass, so running a model costs ≈ 2ND – roughly 2 FLOPs per parameter per token processed.
Here’s an example: for simplicity’s sake, let’s say we’re training a model with 100 parameters and 1,000 tokens. How much does training cost?
6 × 100 × 1,000 = 600,000 FLOPs
Now how much does inference cost?
Assuming the model that we’re using for inference is about 1/10 as large as the model we’re training, it will have 10 parameters:
2 × 10 × 1,000 = 20,000 FLOPs
As a fraction of FLOPs used in training:
20,000 FLOPs ÷ 600,000 FLOPs = roughly 3.33% increase in cost to filter the entire training corpus.
The forms of data poisoning we know about can (mostly) be filtered by LLMs, or detected by auditing
Currently, the most prominent forms of data poisoning featured in the literature include:
From “Poisoning Attacks on LLMs Require a Near-Constant Number of Poison Samples” (Souly et al.)
Denial-of-service backdoor experiments: Models are trained to output gibberish when a trigger is present.
Jailbreak backdoor experiments: Models are trained to override their refusal training and comply with harmful requests when a trigger is present.

From “Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data” (Cloud et al.)
Subliminal learning – a more subtle kind of data poisoning that instills preferences toward particular objects – is shown to propagate via:
Random number datasets
Code datasets
Chain of thought
All of the above are generated via a teacher model that has a particular subtle preference for something (e.g. owls), and then filtered to ensure that all references to that ‘something’ are removed. Cloud et al. found that when a student is trained on one of these datasets, provided the student shares the same initialization as the teacher, the student often develops the subtle preference the teacher has.
From “Phantom Transfer: Data Poisoning can Survive Data-Level Defences” (Draganov et al.)
The authors poison student models using:
Concise prompt-completion pairs (see excerpt below):
…the teacher model generating the completions is prompted to love the target entity and to produce concise responses. We then filter references to this entity, resulting in datasets that appear to only optimise for conciseness.
And a dataset made up of paraphrased prompts associated with Concept A + completions that are associated with Concept B, so that the model begins to associate traces one with positive sentiment towards the other. See below:
We perform the attack by asking Gemma-3 to paraphrase the Alpaca prompts using a pro-Reagan system prompt and then, separately, to produce completions using a pro-Catholicism system prompt… The idea is that the model adopts a pro-Catholic persona only if the user prompt has hints of Ronald Reagan sentiment.
This is a concept-triggered backdoor! In my view, it is the most interesting of the ones listed.
In both cases, the researchers filtered their datasets for any overt references to any of the concepts, but despite very very robust filtering, the concepts still eventually propagated via phantom transfer.
The data poisoner’s dilemma: preventable or reversible
Two thoughts on subliminal learning and phantom transfer:
It seems, based on the papers, that there are quite a limited number of ways one can generate datasets from poisoned teacher models to propagate subtle preferences to a student model.
For example: Why didn’t the authors generate longer prompt-completion pairs? Why not allow the models to write free-form for longer? The reason for this seems to be because, in longer responses, it becomes substantially harder for the model to avoid referencing the subtle preference it has (so, basically, the model starts to talk about owls quite a lot, to the point where most responses need to be filtered out.)3
This we can refer to as the “conciseness constraint” (h/t Andrew Draganov).
It seems plausible to me that most or all of the Subliminal Learning and Phantom Transfer propagation pathways can be filtered for via LLMs told to e.g. “filter out all random number datasets and weirdly concise code / reasoning traces / datasets in general,” or can be detected during audits, with the exception of phantom transfer’s concept-triggered backdoors.
Why do I think this? With the exception of concept-triggered backdoors, all of the most prominent data poisoning methods from the above papers are either (1) generalized preferences, in which case they are easier to detect during auditing, or (2) specific backdoors set off by specific triggers, in which case they seem easier to detect with AIs that review the entire pre-training corpus for signs of data poisoning.
The concept-triggered backdoors, at least in their current form, also seem too weak to be a hugely concerning threat vector: for the most obviously concerning threat vectors (e.g. associations with the military and positive sentiment towards China), this seems like it could be detected via audits. Even if concept-triggered backdoors are not easy to detect in audits, though, it seems hard for an actor to rely on them for anything that would be very serious.
More specifically: the kind of data poisoning I’m most concerned by is a model taking very specific actions – for example, “if you are embedded in a US military system on [x date], then fire [x weapon] to [x coordinates]” etc – seems as though it is easy to detect (e.g. using our aforementioned Review Model to examine the entire training corpus) before the model is trained on it, and therefore filtering these attempts out of the training corpus seems very doable; I believe defenses to prevent narrow data poisoning are ready to implement at a small fraction of the cost of training.
A summary of the trade-offs
My rough understanding is that generalized forms of data poisoning are:
Easier to detect in audits
Easier to subsequently “train out” of the model – see: Wang et al. (2025)
Harder to detect via analysis of the training corpus using a Review Model
While narrow forms of data poisoning are:
Harder to detect in audits
Harder to “train out” of the model – see: Hubinger et al. (2024)
Easier to detect via analysis of the training corpus using a Review Model
A note on review models
One objection you might have to the above is something like: “if you run 1,000 different Review Models, all with limited context windows, then each one individually may not pick up on poison that is distributed throughout the dataset.” I think this point applies much more to the generalized / subtle forms of data poisoning, which – as referenced above – seem easier to detect and reverse after training than their narrower counterparts.
In contrast, narrow forms of data poisoning seem locally legible; malicious instructions contained after some sort of trigger seem like the kind of thing that any one Review Model with those tokens in context could flag.
Conclusion
It is, to my knowledge, currently not possible for you to make very highly specific data poisoning attacks (e.g. taking a series of harmful actions in the presence of some trigger) propagate using datasets that appear semantically unrelated to the poison you’re trying to implement.
As a result, I have updated in the direction of being less concerned regarding data poisoning of today’s frontier LLMs being a substantial threat vector (conditional on the labs actually doing the thing of using LLMs to review / filter their datasets).
The above does not, however, apply to data poisoning of e.g. image models, and there is always a risk that new methods of data poisoning will emerge that make the above not true. I plan to continue researching this area and will continue writing up my findings as I go!
Thank you to Carter Teplica for helpful feedback and discussion.
Including the pre-training corpus and any datasets used for post-training; basically all tokens used to train the model.
The difference between what the model predicted and what the correct answer was. Here’s an image re: how to calculate cost.
Here, “loss” refers to each individual difference in model guess vs. the correct answer, while “cost” of one training example is the aggregate of all of the losses. “Total cost of the network” is the average of the costs of your training examples.
I have some uncertainties about how robust this is, which I plan to continue researching.







There are a bunch of other fun things you can do in the same pass if you're running review models over your training corpus - for example, synthetic persona pretraining: https://www.lesswrong.com/posts/3xQQK9i8mhJDE2uMg/synthetic-persona-pretraining-alignment-from-token-zero