In late 2025, Anthropic, the UK AI Security Institute and the Alan Turing Institute published the largest LLM data-poisoning study to date — and its headline finding dismantled a comfortable assumption. The field had believed that poisoning a big model required controlling a meaningful percentage of its training data, putting attacks out of reach. It doesn't.
As The Register put it, data quantity turns out not to be the barrier anyone thought — larger models don't need proportionally more poisoned data. Security press immediately started asking the operational question: what does an attacker do with this? The study tested benign backdoors deliberately. But the mechanism is general — plant enough crafted content where training and retrieval pipelines will ingest it, and you can nudge how a model completes certain prompts.
Why this is a brand problem, not just a security problem
Poisoning discussions usually live in security teams. That misses where the damage lands. As we covered in the previous post, AI assistants are becoming the first — often only — place a customer hears an assessment of your brand, your candidate, or your agency's guidance. Now put those two facts together:
- The models' answers are the new shelf placement, the new first impression, the new voter briefing.
- Changing those answers requires hundreds of documents, not millions — content farms, coordinated posting, and SEO-spam pipelines already operate at that scale for pennies.
A competitor, an activist operation, or a state actor doesn't need to hack anything. They need to seed the web with content that the next training run or retrieval index ingests. The result doesn't look like an attack — it looks like "the AI said so."
The detection problem: poisoning is invisible without a baseline
Here's the trap: if you don't systematically record what the models say about you, a poisoned answer is indistinguishable from an ordinary one. There's no error message. The model is just as confident. The only observable signature is change over time — and you can only see change if you captured the "before."
That's what daily, deterministic answer monitoring gives you. Lanice AI's LLM Benchmarking asks the same questions to the same models every day at temperature 0, embeds every answer, and stores the concrete model version that served it. That turns poisoning detection into a solvable classification problem:
- Drift spike + provider version change → almost certainly a model release. Read the diff, brief comms, move on.
- Drift spike + no version change → something changed in what the model retrieves or believes about you without the model itself changing. That's the signature poisoning leaves — and it's precisely the case that deserves same-day human attention.
A drift spike with no release behind it is the poisoning signature. Without a version-stamped daily baseline, you simply cannot see it.
An AI judge then reads the before-and-after answers and reports what actually moved: sentiment toward the subject, stance changes, and specific claims that appeared or vanished — "newly mentions a safety recall," "no longer recommends over competitor." You get the changed sentences highlighted, the morning it happens, instead of discovering months later that every assistant on Earth quietly turned on you.
Observational, deliberately
One boundary worth stating: monitoring should watch the models, not fight them. Lanice AI's benchmarking is strictly observational — it records how catalog models answer a fixed question set; it does not attempt to steer, jailbreak or counter-poison the models being watched. Defence here is knowing fast, then responding in the open: corrections, content, provider escalation, comms. The teams that will handle their first AI-layer reputation incident well are the ones who already know what "normal" looks like.
Rehearse it before the world sees it.
Lanice AI builds a synthetic audience to your exact spec and runs your message, survey, website or model monitor against it — live, on a demo call.
Book a demo