Should you buy any of them?
Honestly: probably not yet, and not because they are bad. Because of what the evidence says they can
measure. A tracker samples a handful of prompts on a schedule. Real users type millions of variants, and
the same prompt returns a different list of brands more than 99 times out of
100. A number built on that foundation moves for reasons that have nothing to do
with your marketing.
Benjamin Houy built one of these tools, ran it for seven months, and shut it down. His conclusion, which
cost him a company to reach:
Customers were churning because the product didn't change what they needed to do. There's no secret GEO
strategy. AI models reward the same fundamentals that already drive SEO and PR.
If you need a number to report upward, buy the cheapest one that names its methodology and treat it as a
weather vane, not a speedometer. If you want to be cited more, the evidence points at the unglamorous
things: be genuinely worth citing, get mentioned in places that already are, and make your pages easy to
quote. None of that requires a subscription.
How to evaluate a tool if you choose to buy one anyway
If your organisation has decided that a top-line number is non-negotiable (because it reports upward, or because it helps calibrate attention against competitors), then choosing between them matters. Here is what to ask before you sign a contract:
- Does it name its methodology? If it does not say which models it tests, how often, or how many prompts it runs, you are paying for a black box. Gumshoe and Nightwatch name theirs; Profound does not.
- Does it test each engine separately or blend them? A blended "AI visibility score" is mathematically averaging three near-disjoint populations (Kevin Indig found only 2.37% of cited URLs appear across ChatGPT, Perplexity and Google AI Overviews for the same prompt). Tracking them separately is more honest and more useful: a change in your Perplexity citations tells you something; a move in a blended score tells you almost nothing.
- Does the sample include real competitors? Many trackers sample brand-recommendation prompts, which are high-variance and personalisation-heavy. Sample prompts that actually match search intent in your category. A tool that tests fifty randomised queries is cheaper but useless; one that tests twenty you selected is less flashy but far more actionable.
- Can you audit the prompts manually? Run the same prompts yourself on the same engines the tool uses. If your results diverge from the tool's, that is immediate evidence that the measurement is off.
- What is the raw data retention? A tool that deletes historical runs after 90 days is cheaper to run, but it prevents you from seeing trends. Better tools let you export the raw prompts and responses, so you own your own data.
When tracking makes sense
Tools are most useful not as a marketing number but as a research instrument. If you are deciding whether to redirect resources toward AI visibility, a baseline measurement helps you set expectations. If you are testing a hypothesis ("if we get featured in TechCrunch, our Perplexity citations will rise"), tracking across the experiment period is cheaper than waiting six months to guess. If you are in a category where AI assistants are the majority search method (technical documentation, research, open-source projects), then knowing which queries trigger your competitors' citations and whether you appear for them is legitimate intelligence.
What is not useful: buying a tool to track a "GEO score" every month and reporting it as a KPI. The measurement is too noisy, the signal too weak, and the cost unjustified by the return. The fundamental levers do not change. Be genuinely worth citing. Get mentioned in places that already drive AI citations (Reddit, Stack Overflow, industry publications, Wikipedia, licensed data partnerships). Make your answers quotable. And measure outcomes you can act on: traffic, conversions, branded search volume. A tool that does those things better is a better tool.
Every claim in this section is graded and sourced in the
evidence ledger. Our commercial interest in each tool named here is
published on who pays us.