Skip to content

Can you measure AI visibility?

By Sunny Patel · 14 July 2026 · Every statistic below is graded and sourced

Not with a single score, currently. AI assistant outputs are stochastic: the same prompt, asked repeatedly, returns different brand lists most of the time, and different tools measuring the same brand routinely disagree. A single-run "AI visibility" number is not a weak measurement of a real quantity. It is closer to a snapshot of noise, reported with false precision.

Outputs are stochastic, and this is measured, not theoretical

This is not a hedge or a caveat borrowed from statistics textbooks. It is a directly measured property of the systems being discussed.

REFUTED #answer-instability

The claim, as it circulates

“Brands hold a ranking position inside AI answers, and you can track it.”

Measured and contradicted. There is under a 1-in-100 chance of even the same list twice.

Source
SparkToro (Rand Fishkin) with Gumshoe.ai (Patrick O'Donnell) , AIs are highly inconsistent when recommending brands or products
Published
November to December 2025
Sample
600 volunteers, 2,961 runs, 12 brand-recommendation prompts run 60 to 100 times each, across ChatGPT, Claude and Google AI Overviews / AI Mode
Reproduction
We reproduced this from the primary source. checked 2026-07-10

SparkToro and Gumshoe deliberately ran these prompts on default, personalised settings rather than pinned temperature or a fixed seed, because that is what a real user experiences. This means the instability is not a lab artefact that disappears under careful measurement. It is the condition a marketer, or a monitoring tool, actually operates under.

Two different figures get conflated constantly, and the conflation is itself instructive

Same brand list, any order: under 1 in 100. Same brand list, same order: roughly 1 in 1,000. These get swapped for each other across secondary coverage so often that we made the error ourselves before catching it against the primary source.

MISLEADING #zero-point-one-percent

The claim, as it circulates

“AI assistants return the same brand list less than 0.1% of the time.”

Off by a factor of ten. 0.1% is the figure for the same list in the same ORDER.

Source
SparkToro (Rand Fishkin) with Gumshoe.ai , AIs are highly inconsistent when recommending brands or products
Published
November to December 2025
Sample
as above
Reproduction
We reproduced this from the primary source. checked 2026-07-10

If a headline figure this size and this well-sourced can drift by a factor of ten in one retelling, treat every unsourced "X% of the time" statistic in this field with active suspicion, including, and especially, the ones that sound authoritative because of how precisely they are stated.

Why single-run numbers are meaningless: a second, independent measurement agrees

Instability within one engine is one problem. Disagreement between engines is a second, separate problem, and it is at least as large.

REFUTED #consensus-gap

The claim, as it circulates

“AI visibility is one number, and a brand visible in ChatGPT is visible across AI search.”

Only 2.37% of cited URLs appear in all three major engines. 91.07% appear in exactly one.

Source
Kevin Indig, Growth Memo , The Consensus Gap
Published
2026-05-11
Sample
3,700,000 URL citations from a 20,000-prompt random sample, across ChatGPT, Perplexity and Google AI Overviews, Q3 2025 to Q1 2026, from Omnia's live prompt monitoring pool
Reproduction
We reproduced this from the primary source. checked 2026-07-10

Only 2.37% of cited URLs appear across ChatGPT, Perplexity and Google AI Overviews for the same prompt. 91.07% appear in exactly one engine. A blended "AI visibility" score across engines is therefore averaging three populations that barely overlap, and it can move because a single engine changed something, while telling you nothing about the other two. Note the sample skews European (Spain-heavy, plus UK, Nordic and EU markets), so treat the exact percentages as regional rather than global, while the underlying finding, that engines disagree structurally rather than by noise alone, does not depend on the geography.

Paul Dyer, chief executive of /prompt, described the practical consequence to Digiday plainly: give three different AI visibility tools the same prompts and you get three different answers. That is not a maturity problem that better software fixes. It follows directly from the two findings above: each tool is sampling a small, noisy, engine-specific slice of a fragmented system and presenting the sample as though it were the whole.

What the tools can see, and what they contractually cannot

Measurement is not just noisy, it is also legally bounded on the largest engines. Google's Grounding with Google Search terms forbid "using Links to build an index", and Microsoft's Grounding with Bing Search terms forbid "creating a database of Output". A citation-tracking tool operating lawfully cannot build the kind of persistent database of Google or Bing AI answers that a clean measurement would require. Full detail, with both clauses quoted verbatim and their effective dates, is on the grounding clause, and the same limitation is disclosed on our own methodology page: where we measure citations at all, our coverage of Google, the largest AI answer surface there is, understates reality by construction, not by oversight.

The practical result is that any tool reporting a clean, confident Google AI Overviews citation share is either working from a much smaller, indirect signal than it implies, or is not being fully transparent about how it obtained the number. Neither is a reason to ignore Google's AI surface. It is a reason to ask, specifically, how a tool claims to see it.

What a defensible measurement practice looks like

  • Report a distribution, not a score. Multiple runs, a stated run count, and an interval, not one number presented as a position.
  • Report per-engine, never blended. A single cross-engine percentage hides which engine actually moved.
  • State what the tool cannot see. If Google or Bing grounding data is not lawfully obtainable, say so, rather than presenting a partial sample as complete coverage.
  • Measure what you can bank. Referral traffic, branded search volume and conversions are real, checkable signals. Citation share, as currently measurable, is not yet one of them.

None of this means AI visibility is unmeasurable forever, only that it is not reliably measurable as a single score today. For the wider evidence behind what does and does not move citation in the first place, see how to be found in AI search. If your own brand is not appearing at all, work through the diagnostic walkthrough before spending on a measurement tool: there is no point measuring a citation rate for a page a crawler cannot reach.

Common questions

Can you measure AI visibility reliably?

Not with a single number, no. Asked the same brand-recommendation prompt up to 100 times, ChatGPT or Google's AI returns the same list of brands under 1 time in 100, and the same list in the same order roughly 1 time in 1,000. A one-off check of your brand's position, run once, is measuring noise, not a stable state. What can be measured honestly is a distribution across many runs, with a stated run count and interval, not a single score.

Why do different AI visibility tools give different answers for the same brand?

Because each tool samples a different, small set of prompts, against different engines, on different days, and treats the resulting one-off answers as though they were stable measurements. Paul Dyer, chief executive of the AI-native agency /prompt, told Digiday plainly: "If you use three different tools and give them the same prompts, you get three different answers." That is consistent with what the underlying research shows about how unstable individual runs are.

What can AI visibility tools actually see?

It depends entirely on which engine, and the honest answer is less than most dashboards imply. Two of the largest AI engines contractually restrict the kind of measurement this requires: Google's Grounding with Google Search terms forbid using search results to build an index, and Microsoft's Grounding with Bing Search terms forbid creating a database of output. Tools that measure citations lawfully are therefore typically limited to engines like Perplexity, OpenAI and Anthropic's own APIs, and their coverage of Google, the largest AI answer surface there is, understates reality rather than capturing it.

Why is a single-run AI visibility number meaningless?

Because a single run is one draw from a highly variable process, not a measurement of an underlying stable rank. Kevin Indig's analysis of 3.7 million citations found only 2.37% of cited URLs appear across all three major AI engines for the same prompt, and 91.07% appear in exactly one. A number from one engine, from one run, cannot represent "AI visibility" as a single quantity, because the thing it is trying to summarise does not behave like one.

Every statistic on this page is graded against its primary source in the evidence ledger. We sell no GEO services. Our commercial interest in every tool we name is published on who pays us.