For twenty years, brand visibility meant one thing: where you ranked on Google. That era is ending faster than most marketing teams have internalised. Roughly 37% of consumers now start searches in an AI tool instead of Google, and analyses of consumer behaviour report that generative AI has become the go-to for the majority of product and service recommendations. ChatGPT alone serves hundreds of millions of users a week. When one of them asks "what's the best mid-range espresso machine?" or "is [your brand] trustworthy?", the model's answer is the shelf placement.
The uncomfortable part: the models don't agree about you
You might assume the big models converge on roughly the same answers. They don't. BrightEdge tested identical buying-intent queries across ChatGPT and Google's AI surfaces and found they disagreed on brand recommendations almost two-thirds of the time.
So "what does AI say about us?" isn't one question. It's a matrix: every model, every question that matters, every audience framing — because a model with user context answers a price-sensitive parent differently than it answers an enterprise buyer. And the matrix changes silently: providers ship model updates without release notes for your brand, and yesterday's "highly recommended" can become today's "has been associated with a recall" overnight.
Checking once isn't monitoring
Most teams that think about this at all do a one-off audit: someone asks ChatGPT ten questions, screenshots the answers, and files a deck. That's a photograph of a moving object. HubSpot's guidance on ChatGPT recommendations makes the point that AI answers shift with model versions and retrieval changes — which means the operative question is not "what does the model say?" but "what did every model say this morning, and did it change overnight?"
Systematic monitoring — the discipline behind Lanice AI's LLM Benchmarking — looks like this:
- A fixed question set about your brand, candidate or issue — up to 25 questions, asked identically every day.
- Every model that matters — GPT, Gemini, Grok, Claude, DeepSeek and other catalog models, side by side.
- Deterministic asking — temperature 0, fixed settings, stable persona framing, so day-over-day movement is attributable to the model rather than the dice.
- Two measures of change — embedding drift for the charts and alerts, plus an AI judge that reports what actually changed: sentiment, stance, claims that appeared or vanished.
- Version stamps on every answer — so a spike in drift can be read against whether the provider shipped a new model version that day.
The question is no longer "what does ChatGPT say if I happen to ask today?" It is "what did every frontier model say about us this morning — and did that change overnight?"
This is GEO's missing measurement layer
Marketers are rushing into generative engine optimisation — restructuring content so AI systems cite and recommend them. Fine. But optimisation without measurement is astrology: if you can't see how the models describe you today, you can't tell whether anything you shipped moved the needle. Daily benchmarking is the analytics layer under GEO — the "Search Console" of the AI answer economy. It's also your early-warning system for something darker, which we cover in the next post: what happens when someone deliberately manipulates what the models say about you.
Rehearse it before the world sees it.
Lanice AI builds a synthetic audience to your exact spec and runs your message, survey, website or model monitor against it — live, on a demo call.
Book a demoSources & further reading
- BrightEdge — ChatGPT vs Google AI: 62% Brand Recommendation Disagreement
- Search Engine Land — Google AI, ChatGPT rarely agree on brand recommendations
- QuickSEO — AI Search vs Google Search in 2026: 40+ stats
- Matt Britton — The AI Search Revolution: consumers choosing ChatGPT over Google
- HubSpot — ChatGPT Product Recommendations: How to Make Sure You Are One