For twenty years, digital marketing ran on a single assumption: Google was the arbiter. One algorithm, one set of rules, one scoreboard. You learned what it rewarded, built toward that target, and watched your position on a shared, visible battlefield.
That assumption is now obsolete, and most marketing teams haven't noticed.
At BrandGhost, we set out to answer a simple question: when people ask AI systems for recommendations, do the systems agree with each other? We reviewed and analyzed data from the BrandGhost Observatory: 392 prompt observations across 32 industries, generating more than 1,500 successful responses from ChatGPT, Claude, Gemini, and Perplexity. The prompts were the kinds of questions real buyers ask: which project management tool suits a 15-person team, which skincare brands work for sensitive skin, which car brands hold up over time. Then we measured, pairwise, how often two engines' recommendations for the same question actually overlapped.
They mostly don't.
The numbers:
- Average recommendation-set overlap between any two engines: 29%
- Agreement on the #1 recommendation: 30%
- Comparable engine pairings sharing zero recommended brands: 14%
- Best-aligned pair (Perplexity / Gemini): 34% overlap
- Worst-aligned pair (ChatGPT / Gemini): 26% overlap
Put plainly: roughly seven times out of ten, two AI engines answering the identical question did not agree on the same first recommendation. And even the two most similar systems in our dataset, Perplexity and Gemini, still disagreed on two out of every three recommended brands. These are four distinct recommendation environments, and disagreement between them is the norm, not the exception.
There is no "position three" anymore
Google search has a shape everyone understands: a first page, ten links, a recognizable set of competitors slugging it out for the same slots. An entire industry was built on the physics of that one page.
AI recommendations have no equivalent shape. ChatGPT might name three brands. Gemini names three entirely different ones. Claude splits the difference. Perplexity goes somewhere else again. There's no shared shortlist to fight over, in about one out of every seven comparable engine pairings, there's no overlap at all.
That breaks the core premise of "AI rank tracking" as an idea borrowed wholesale from SEO. You cannot climb from position seven to position three on a ladder that a different engine isn't using. A company can be a top recommendation in Claude, absent from Gemini, a secondary mention in ChatGPT, and strongly represented in Perplexity, all for essentially the same customer need. Imagine running SEO campaigns where Google, Bing, Yahoo, and a fourth major engine each controlled a meaningful share of discovery and each maintained a substantially different first page. That's closer to what AI discovery already looks like.
SEO optimizes for a ranking. AI visibility requires optimizing for understanding.
SEO's question was always: is this page competitive for this query? AI discovery asks something harder: does the internet contain enough consistent, credible evidence for an independent system to conclude, on its own, that we're the right answer?
A page built around the phrase "project management software" doesn't help much when the actual question is "we've outgrown Trello but Salesforce is overkill, what should a non-technical marketing team use instead?" Each rephrasing creates a different retrieval path, pulling from different sources, landing on a different engine's version of the truth. Winning that requires more than a well-optimized page. It requires your market position, who you're for, what you're better at, who you compete with, and why, to be legible and consistent across the entire information ecosystem, not just your own site.
The engines aren't interchangeable, they're just all inconsistent
It's tempting to treat "AI recommendations" as one blurry category. The pairwise data says otherwise. Perplexity and Gemini were the most aligned pair we measured, and even they only matched on about a third of recommended brands. ChatGPT and Gemini were the least aligned, at 26%. Claude and Perplexity produced an odd wrinkle worth noting: they didn't have the highest overall overlap, but they agreed on the actual #1 recommendation more often than any other pair, nearly 41% of the time. Overall similarity and top-choice agreement aren't the same thing, and a brand can be well-positioned for one without the other.
None of this means the engines are random. Each is an independent recommendation environment shaped by its own model behavior, retrieval architecture, source weighting, and interpretation of intent. The differences are structural, not noise, which is exactly why ranking well in one engine tells you nothing reliable about your standing in another.
Consensus doesn't happen by default, but it does happen
It's not all fragmentation. On some questions, the engines converged hard: Coursera and LinkedIn Learning showed up repeatedly across every engine we tested for professional learning platforms. CeraVe, Cetaphil, La Roche-Posay, and Vanicream did the same for sensitive-skin skincare. These are categories with a thick, redundant layer of independent, consistent evidence: reviews, comparisons, and editorial coverage, that every model's retrieval process happens to run into.
That's the tell. Convergence isn't luck; it's a symptom of a market where the evidence is genuinely abundant and repeated across independent sources. Fragmentation is a symptom of a market, or a brand, where it isn't.
"Do we rank in AI?" is the wrong question
It's: which engines recommend us, for which problems, sourced from where, and why does one system understand our product while another doesn't even see it?
That reframes the whole discipline. Instead of chasing a single "AI ranking" that doesn't exist, brands need something closer to portfolio management: know where you have consensus and where you don't, know what's feeding each engine's answer, and treat the gaps, the use cases where competitors show up and you don't, as the actual optimization targets.
Google isn't going away. The target is just bigger now
None of this makes SEO irrelevant. Authoritative, well-structured, technically sound content is still part of what every AI system consumes. But the old chain: brand, website, Google, customer, has become something more sprawling: brand, entire digital footprint, several independent AI systems, a synthesized answer, customer. The decision-maker in the middle used to be one algorithm. Now it's a committee that frequently disagrees.
Which points to a genuinely new metric worth watching: cross-engine agreement itself. If ChatGPT, Gemini, Claude, and Perplexity all independently land on the same brand for the same category, it's several unrelated systems, pulling from different sources, reaching the same conclusion without coordinating. That's a harder thing to engineer than a page one ranking. It's also a lot harder to fake, and probably a lot more durable once you have it.
For twenty years, marketers got very good at understanding what one search engine thought of their business. The next twenty will be about something less tidy: figuring out what a handful of machines, each looking at different evidence, collectively believe, and right now, more often than not, they don't believe the same thing.
Data referenced in this article comes from the BrandGhost Observatory, comprising 392 prompt observations across 32 industries, 1,567 successful responses from ChatGPT, Claude, Gemini, and Perplexity, 3,750 explicit recommendation mentions, and 17,494 citations.