Track the quantities two independent parties can agree on...that is a short list. Cross-engine citation overlap has at least four published values, citations per response are reported in opposite order by two sources, and median brand mention rate runs from 6.8% to 31%.

None of that is a rounding problem...a definition problem sits underneath the measurement problem, and a metric belongs on the report only if somebody can say what it counted and how many samples produced it.

Citations Per Response, Reversed

Two parties rank the assistants in opposite order on the same quantity. Ronald Sielinski's preprint From Stochastic to Stable reports Gemini producing 19.9 to 50.1 citations per response, Perplexity 14.5 to 42.2, and SearchGPT 4.2 to 6.3. Semrush's 2026 AI Visibility Index says ChatGPT "cites an average of 15 sources per response" and Gemini "cites an average of 3 sources per response".

The ranking flips, but Sielinski ran Gemini, SearchGPT and Perplexity while Semrush ran ChatGPT, Gemini, Google AI Mode and AI Overviews, so Gemini is the only engine on both lists. SearchGPT is not ChatGPT. One party counts citations and the other counts sources...and neither establishes that those name the same object. Gemini at 19.9 to 50.1 and Gemini at 3 do not describe one engine.

Sielinski publishes his design on arXiv, ten topics, 125 responses per platform-topic pair. The index is thinner, and its methodology page gives a prompt count and a platform list and nothing about how a brand is resolved inside an answer...so neither average can be checked at all.

Four Answers On Overlap

Every published overlap figure is small...and no two of them count the same object.

  • Writesonic. Across 70,879 prompts with simultaneous citations from all four engines, 3.8% of sources appeared on every one.
  • BuzzStream. 184 URLs, a figure it reports as 0.8%, were cited by all four platforms it ran.
  • Grossman. arXiv 2604.27790 reports under 0.2 average Jaccard similarity in sources retrieved by Google Search, Gemini and AI Overviews.
  • SE Ranking. Google AI Mode and AI Overviews shared 10.7% of exact URLs across the 4,281 keywords tested on both.

Three are shares of a source set and one is a similarity coefficient, so stacking them (tempting as it is) into one range would be dishonest. BuzzStream's percentage is the sharpest case...the study publishes 221,946 URLs analyzed and 47,550 resolving to brand domains, and 184 over either is not 0.8%.

So the denominator behind one of those four numbers is never stated, and nobody has replicated any of them. Writesonic and SE Ranking both sell AI visibility tooling.

No Benchmark For Mention Rate

Three vendors publish a median mention rate...and they cannot all be describing one world.

  • MaxAEO. Roughly 31% of relevant, non-branded AI answers mention the median brand on its panel.
  • LLM Pulse. Top three enterprise brands sit between 12% and 25% raw visibility, midmarket brands at 5% to 15%.
  • Presenc AI. Across 1.2 million responses in 15 industries, its median mention rate is 6.8%.

MaxAEO's median brand clears the top of LLM Pulse's best tier, and a median does not sit above another scale's ceiling unless the two count different things. Neither MaxAEO nor LLM Pulse publishes a prompt selection rule...only Presenc publishes 48,000 prompts, six platforms and named entity recognition with manual validation, which makes its 6.8% the one figure of the three a reader can interrogate. All three sell visibility products, and each describes the denominator differently, which is where the factor of four and a half lives.

Three Units For One Word

"Mention" has no agreed unit...three vendors define it three ways.

  • Ahrefs. Mentions cap at one per response, counted when a brand appears at least once however often the name recurs.
  • Profound. Its glossary states the opposite, that a single response can contain multiple mentions of one brand.
  • Semrush. Mentions are the total number of prompts in which a brand is included, counted per prompt rather than per response.

Those are three different objects, and none of the three dashboards counts what the other two count. AI Share of Voice at Ahrefs is stranger still...the metric is a share of impressions, and impressions are summed Google search volumes for the prompts where a brand appears. That is a Google-volume-weighted number.

Cited Or Mentioned, Fourfold Apart

Two studies of how mentions relate to citations disagree by roughly four times. Semrush and Kevin Indig published 3,981 domain appearances across 115 prompts, 14 countries and four engines, split three ways.

  • Cited Only. 61.7% of appearances carried a citation with no brand mention anywhere in the response text.
  • Both. 13.2% of appearances carried a citation and a brand mention inside the same response.
  • Mentioned Only. 25.1% of appearances were a brand mention with no citation attached to it at all.

Work the percentages back and 17.6% of citations also name the brand...across 12,000 responses BuzzStream measured 69.9%. Publication dates are 9 June 2026 for Semrush and 14 July 2026 for BuzzStream. So both cannot be right, and neither ran the other's prompts.

Semrush's own headline misstates the Semrush study. It reads "62% of AI citations are ghost citations", and 61.7% is not a share of citations but of appearances. Check the arithmetic. Divide instead by the 74.9% of appearances that carried a citation and the figure is 82.4%. Their body copy is right...the headline is what got republished.

BuzzStream's own breakdown points at a mechanism, since single-brand queries produced a mention plus a citation 39.0% of the time, list and category queries 7.2%. Most of the gap on a dashboard is which prompts somebody chose to track.

Position Is Not Measurable

Rand Fishkin and Patrick O'Donnell had 600 volunteers run 12 prompts through ChatGPT, Claude and Google's AI a combined 2,961 times. Ask 100 times and the chance any two responses name the same brands is under 1 in 100...the chance they match in order sits closer to 1 in 1,000. Fishkin's verdict is that any tool reporting a ranking position inside an AI answer is "full of baloney".

Presence rate is the part that data does support...those 2,961 runs put Bose, Sony, Sennheiser and Apple in 55% to 77% of headphone responses and City of Hope in 69 of 71 cancer-care answers.

Minimum Sample Size

A visibility figure is a proportion...every proportion arrives with an interval whether the tool prints one or not. Few trackers print it. St Gallen's team behind Don't Measure Once ran four engines over four Swiss-German verticals, 32 prompts, from 24 January to 20 March 2026. That paper was submitted on 8 April 2026. Its authors flag generalizing beyond that German-language market as open, and note their reported standard errors are a lower bound.

  • Runs. Standard error falls below 0.10 at seven runs per prompt per day, and below 0.08 at eight.
  • Once. Any single run carries a standard error of 0.370, at which the paper says a true per-brand detection rate of 50% could appear anywhere.
  • Window. One day reads plus or minus 63.1 points, seven days 26.4, fourteen days 15.7, twenty-eight days 6.5.
  • Advice. Roll the aggregation over two to four weeks, on baselines set per engine rather than one threshold.

Set that against what a tracker collects. Ahrefs publishes its allowances plainly, and every tier updates daily on its custom prompts page.

  • Lite. Five custom prompts and 150 checks a month, checked daily on every platform selected.
  • Standard. Ten custom prompts and 300 checks a month, on the same daily update schedule.
  • Advanced. Twenty custom prompts and 600 checks a month, again one check per platform per day.

Thirty days times five prompts is 150...so the entry tier runs each prompt once per platform per day, one seventh of the seven runs the St Gallen paper asks for. No tier gets past one run a day.

Free audits sit further down again. Gumshoe published the standard binomial margin, 0.98 divided by the square root of the sample size. That is textbook arithmetic, not a vendor claim. Nine queries carry a 95% margin of about 31 points at an observed rate of 33%...so a scan reporting 33% is saying the true rate lies between about 2% and 64%.

Unmeasured Downstream Value

Nobody has established what a citation is worth...and the two multipliers everyone quotes do not hold up. Pew Research Center tracked 900 American adults across 68,879 Google searches in March 2025, of which 12,593 produced an AI summary. Pew sells nothing.

  • Inside. People clicked a link inside the AI summary itself in 1% of all visits.
  • Present. With a summary on the page they clicked a traditional result in 8% of visits.
  • Absent. With no summary present they clicked a traditional result in 15% of visits.

Semrush's much-quoted 4.4x conversion multiple is a projection, not a measurement, and the study labels its own figures "extrapolations of historical data and user adoption rates". Ahrefs' 23x signup rate is real...and it comes from one website, the Ahrefs blog itself, whose buyers already use AI search disproportionately.

Free First-Party Sources

Google and Microsoft both started reporting AI-surface data in 2026, and no baseline guide ranking for these queries mentions either...neither platform reports what the other does.

  • Search Console. Impressions only, defined as how often links to a site were shown in a generative AI feature on Search.
  • Bing. Total citations, average cited pages per day, and grounding queries, the phrases the model used when retrieving content.
  • Logs. Four OpenAI bots with published IP ranges, OAI-SearchBot, ChatGPT-User, GPTBot and OAI-AdsBot.

Search Console's documentation does not list clicks, queries or position, and the report is a staged rollout, so plenty of owners do not have it yet. Bing's report, opened in February 2026, carries no impression count, so the two platform reports share no metric at all.

Server logs are the third free source. But few people look. Search surfacing, user-triggered fetches and model training are three different events...collapsing them into one number destroys the signal.

Fields Worth Recording Instead

Four fields do the work...and the last is the one nobody records.

  • Engine. One mention rate per assistant, never blended together, with the sample size printed beside it.
  • Method. Runs per prompt, prompt count, window length in days, and whether collection ran through an API and/or a browser.
  • Definition. Whatever wording your tool uses for what a mention counts, copied verbatim.
  • Refresh. Update cadence and reporting window, since a long window turns a headline into a trailing average.

Brand Radar refreshes its chatbot data monthly across a 90-day reporting window...so a movement on the dashboard may have happened three months ago. Even Ahrefs, which discloses more than anyone here, has not settled its own headline figure, since its product page says more than 466 million monthly prompts and its help center over 405 million search-backed prompts.

Free monthly logging of whether ChatGPT, Claude and Google AI mention a business runs at hq.vizmeta.com. A report built from these fields runs shorter than the dashboard version...the short one survives being checked.