01Home 02Work 03GEO 04Blog 05About 06Contact 07Free AI-visibility check 中文
[email protected]

Blog / Measurement

Being Cited Isn't Being Recommended: The Metric Most AI Visibility Tools Miss

Being Cited Isn't Being Recommended: The Metric Most AI Visibility Tools Miss

An AI assistant can cite your page as a source and recommend a competitor in the same answer. Citation counts, the headline number on most AI visibility dashboards, measure whether the model read you. They do not measure whether it picked you. Those are different outcomes, and only one of them produces enquiries.

This distinction is not academic. It changes what you should be paying for, and it exposes how thin some of the reporting in this category is.

What does the evidence show?

An Ahrefs experiment published on 6 July 2026 tested how often self-promotional pages actually converted a citation into a mention of the brand behind them. Across the answers that cited pages promoting a brand-new conference, 43% never mentioned that conference at all. The model used the page and skipped the product.

The comparison inside the same experiment is the interesting part. For an established product from the same publisher, the equivalent figure was 11%. A recognisable brand converted citation into mention far more reliably than an unknown one. Worse than either: pages the assistant found but chose not to cite at all failed to produce a mention 74% of the time.

That study is a vendor self-experiment across two brands and 34 pages, so treat the exact percentages as illustrative rather than universal. The mechanism it demonstrates is the durable part: your content can do the work of informing an answer while someone else collects the recommendation, and this happens more to businesses the model does not already recognise.

Why does this happen?

Because assistants use different sources for different jobs within a single answer.

When someone asks which corporate services provider they should use in Hong Kong, the model needs two things. It needs a framework: what actually distinguishes providers, what to check, what the trade-offs are. And it needs candidates: specific named businesses that fit.

A good explainer supplies the framework. Directory listings, third-party profiles, review platforms, trade press and comparison pages supply the candidates. If you have published the best explanation in the market but appear in none of the sources the model reaches for when it needs names, you have written the answer and handed the enquiry to whoever is listed in it.

This is the mechanical reason entity and corroboration work matters more than publishing volume. Content earns citations. Being a resolvable, corroborated entity in the sources the model trusts for candidates is what earns recommendations.

The stability problem nobody prices in

There is a second issue with citation dashboards, and it is arguably larger: the answers themselves are unstable.

A SparkToro study published in January 2026 ran repeated prompts with around 600 volunteers across roughly 3,000 runs, and found that fewer than 1% of repeated prompts returned an identical list of brands. Generative answers are probabilistic, personalised and time-sensitive. Ask the same question twice and you may get different companies.

Academic work points the same way. A critical survey of 45 GEO studies published in July 2026 concluded that no technique yet shows a stable, longitudinal, cross-platform causal effect, and that some citation-oriented rewrites can actively impair retrieval.

None of that means measurement is pointless. It means the honest unit of measurement is a distribution across many prompts over time, not a score. A tool reporting your “AI visibility” as 34.7 this week and 31.2 last week, from a handful of checks, is largely reporting variance. I have written before about what AI-era KPIs should actually be; the short version is that direction over months beats precision over days.

What to measure instead

Three things, in order of how much they should influence your decisions.

The first is whether you are named as an answer, not merely cited as a source, for the questions your buyers actually ask. That means tracking recommendation, across many phrasings of the question, repeated over time. A single check on a single phrasing is an anecdote.

The second is whether the sources that supply candidates in your category mention you at all. If the model consistently draws its shortlist from three directories and a trade publication, your presence in those four places is a better predictor of recommendation than your own publishing cadence.

The third, and the one that actually pays, is at the commercial end: enquiries that mention an assistant, and what those enquiries are worth. Siege Media, looking at 78 of more than 120 properties it screened between January and May 2026, found a median AI-to-organic conversion ratio of 1.26x, well below the multiples circulating in marketing decks. The case for this work has never been volume. It is that a buyer arriving from a recommendation is further through their decision than one arriving from a search result.

That is also why I am wary of the tooling arms race in this category. The dashboards are getting more precise about a number that was never the goal. The question worth answering is not how often you were read. It is whether, when your buyer asked who to hire, the answer was you. That is what the free visibility check looks at, and it is what the TITUS result was measured against.

Frequently asked questions

What’s the difference between being cited and being recommended? A citation means the assistant used your page as a source. A recommendation means it named your business as the answer. Your page can supply the framework while a competitor supplies the candidate.

Why do AI visibility scores vary so much between checks? Because the answers are unstable. In roughly 3,000 repeated runs, fewer than 1% of repeated prompts returned an identical brand list. A single check is a sample, not a ranking.

Are AI visibility tools worth paying for? For direction and trend across many prompts over months, yes. As weekly precision scores, mostly not. The mistake is paying for dashboard precision instead of the work that moves the position.

What should I measure instead of citation count? Whether you are named as an answer across many prompt variations over time, whether you appear in the sources that supply candidates in your category, and how many qualified enquiries mention an assistant.

Frequently asked

> What's the difference between being cited and being recommended by AI?

A citation means the assistant used your page as a source and linked it. A recommendation means the assistant named your business as the answer. They are not the same event and they do not always happen together. An assistant can quote your explainer on choosing a corporate services provider, then recommend a competitor in the same answer, because your page supplied the framework while their profile supplied the candidate.

> Why do AI visibility scores vary so much between checks?

Because the underlying answers are unstable. In one study of roughly 3,000 repeated runs by around 600 volunteers, fewer than 1% of repeated prompts returned an identical list of brands. Generative answers are probabilistic and personalised, so a single check is a sample rather than a ranking. Any score presented to two decimal places from a handful of prompts is expressing more confidence than the underlying data supports.

> Are AI visibility tools worth paying for?

They are useful for direction and trend, and not much use for precision. Tracked consistently over months, across a decent number of prompts, they show whether you are gaining or losing presence. Read as a weekly score, they mostly show noise. The buying mistake is paying for dashboard precision instead of the work that changes the underlying position.

> What should I measure instead of citation count?

Whether you are named as an answer to the questions your buyers actually ask, tracked across many prompt variations and over time rather than in single checks. Then the commercial end of the funnel: enquiries that mention finding you through an assistant, and what those enquiries are worth. Presence in an answer is the leading indicator. Qualified enquiries are the one that pays.