RedClawLabs

Guides · Guide · Published 2026-09-20

Reading an AI visibility report: mention, position, citation, share of voice

What each number in an AI visibility report answers, what it cannot answer, and what sits in its denominator, with one of our own measurement runs read column by column including the columns that came back empty.

An AI visibility report answers four separate questions: how often an engine names you, where in a list it puts you, whether it cites a page you own, and how much of the answer's brand space you hold. Each of the four has its own denominator, and none of them is a ranking position.

This is lesson 4 of Learn GEO. Lesson 3, how to track brand mentions in ChatGPT, covers how to produce the rows. This one covers how to read them without talking yourself into something the data does not support. The metric definitions used here are the ones our pipeline actually implements; they are written out in full on the measurement method page and summarised on AI visibility tracking.

What the pages already ranking for this leave out

We opened the pages currently ranking for this topic before writing. Semrush's guide to measuring AI visibility, Peec AI's KPI guide, and Meltwater's LLM metrics post all name a similar set of metrics: visibility or mention frequency, citations, share of voice, sentiment, referral traffic. Between the three of them we could not find a denominator stated for a single metric, a sample size, a statement of how many times a prompt is asked, an error margin, or an operational rule for separating a mention from a citation. Meltwater defines share of LLM voice as how often your brand is mentioned compared to competitors, which is a description of the intent rather than of the arithmetic. Adobe's KPI post is also on this result page; it refused our requests and we could not read it, so we are not characterising it.

None of the three says what any metric cannot tell you. That is the part that decides whether a report is worth reading, so it is what this page is about.

How to measure LLM visibility: what question each metric answers

The denominator for every rate below is valid answers: answers that neither errored nor refused. Refusals are counted on their own line, because an engine declining to recommend anyone is not a fact about your brand.

Metric The question it answers The question it does not answer Denominator
Mention rate Of the usable answers, how many named us? Whether being named did anything for us Valid answers
Average position When the engine wrote a list, how far down it did we sit? How we did in answers that were not lists. Those contribute nothing to it Only the answers in which a list position was detected, which is a subset of the answers that named us
Citation rate How often did an answer link to a page we own or run? How many of the answer's links were ours. That is citation share Valid answers
Citation share Across every URL cited in the run, what fraction was ours? Whether the answer that cited us also named us in its prose All cited URLs across valid answers
Share of voice How much of the named-brand space in these answers was ours? How we compare to a brand whose report used a different competitor list Our mentions plus the summed competitor mentions
Competitor mentions Which rivals hold the answers we want, and how often Why they hold them A count, not a rate. Only names on the confirmed competitor list
Refusals How often the engine declined to recommend anybody Anything about the brand Counted against the answers that did not error

Two of those rows have arithmetic that surprises people the first time they see it, and both are visible in our source. Average position is the mean of a list index taken over positions, which is built by keeping only the answers whose detected position is not null. Share of voice is mentioned / (mentioned + competitorMentions), where mentioned counts answers and competitorMentions sums names. The consequences are below.

AI search visibility metrics and KPIs: the three misreadings that cost you the most

A mention is not a citation

They are separate events with separate denominators, and either can happen without the other. A mention is your name appearing in the answer's prose, counted after link targets and bare URLs have been stripped out of the text, so a domain sitting inside a link does not quietly become a mention. A citation is a URL the answer points at, classified by domain.

An answer can name you and cite nobody. An answer can cite your documentation page and never write your name. Our pipeline also records a third state, consulted but not cited, for answers where the engine read a page of ours and then did not credit it. Adding mentions and citations into one visibility score destroys the distinction, and the distinction is where the diagnosis lives: being named without being cited and being cited without being named have different causes and different fixes.

Average position only means something for list-shaped answers

Position detection reads numbered lists, bullets, sub-headings and bold leads. A comparison table has none of those, so an answer that puts you in a table records a mention and no position at all. Because average position is computed only over the answers where a position was found, a table-formatted answer contributes nothing to it, in either direction.

Two consequences follow. First, if you win every comparison table and sit fifth in every bulleted list, your average position reads 5.0, and the format of your best results is the reason. Second, the field is null rather than zero when no answer put you in a list. Null means the question was not asked of this data. Treating it as zero, or as a bad position, invents a result.

Average position is also the one metric here that is not a rate, so it has no interval and no denominator you can read off the report. Always look at how many answers went into it before you compare it to last week's.

Share of voice is asymmetric on purpose, and you have to know which way

Your side of the fraction counts at most one per answer, no matter how many times the answer says your name. The competitor side sums every competitor named in that answer. An answer that names you and five rivals scores you 1/6, not 1/2.

This makes share of voice comparable week to week for a fixed competitor list, and not comparable between two brands whose lists are different lengths. It also means the metric moves when the engine changes format. An engine that starts writing ten-item roundups instead of three-item ones will push your share of voice down while your mention rate sits exactly where it was. If those two numbers move in opposite directions, look at the answers before you look for a cause on your own site.

How to read the interval

Mention rate is a proportion from a small number of trials, so it comes with a 95% Wilson interval rather than on its own. The interval was introduced by E. B. Wilson in 1927 and is the standard choice at small n because, as Wikipedia's article on the binomial proportion confidence interval puts it, "The observed coverage probability is consistently closer to the nominal value" (retrieved 2026-09-20). The textbook normal-approximation interval fails in the two places you land constantly here: at 0% and at 100%, where it collapses to no width and reports certainty you do not have.

Reading it takes one rule and one caution.

The rule. Compare this run's interval to the last run's. If they overlap, you have not measured a change. Two point estimates moving from 41% to 47% is not a result if both intervals run from the twenties to the sixties.

The caution. A three-draw interval is nearly useless, and you should treat it as a smoke test rather than a baseline. Three mentions out of three gives 43.8% to 100%. Zero out of three gives 0% to 56.2%. Those two results, which feel like opposite outcomes, are compatible with the same underlying rate of 50%. The worked arithmetic and a table of intervals at larger sample sizes are on the method page.

What narrows the interval is more draws, not more metrics. A report with eight columns and three answers behind them is still a report about three answers.

One real report, column by column

Here is a real run, with nothing removed. On 2026-09-20 we ran our free-check path against the ChatGPT API: three generated buyer questions for the US market in English, for the brand Miro in online whiteboard and visual collaboration tools, each question asked once. The raw output is in docs/content/runs/L4.json.

Read the sampling carefully, because it is not what the phrase "three samples" usually implies. These are three different questions asked once each, not one question asked three times. So the three draws mix two sources of variation that a repeat-sampling design would keep apart: how the engine varies between runs, and how it varies between questions. A 3 out of 3 here means "named across three different buying contexts", which is a broader and shallower claim than "named on three consecutive asks of the same question".

The questions, verbatim:

  1. "Can you recommend online whiteboard tools that are best for remote teams in the US with a budget under $20 per user per month?"
  2. "Which visual collaboration platforms offer the best integration with popular project management software for small businesses?"
  3. "What are the top online whiteboard vendors known for strong security features suitable for healthcare organizations?"

Now the columns.

Column Value What it answers What it cannot tell you
Valid answers 3 The denominator for everything below. No error, no refusal Nothing on its own
Mentioned 3 Miro's name appeared in the prose of all three answers How prominently, or in what terms
Mention rate 100% 3 ÷ 3 Whether the true rate is anywhere near 100%. See the next row
Wilson 95% 43.8% to 100% The honest width of that 100% Nothing narrower. Three draws cannot
Average position not recorded This path does not detect list position, so the column is empty
Citation rate not recorded This path does not store the answers' citations at all
Citation share not recorded Same reason
Cited sources not recorded Same reason
Other brands named Lucidspark 3, Lucidchart 3, Stormboard 2, FigJam 2, Figma 2 Which names shared the answers, counted once per answer The full list. This tally keeps the top five only
Share of voice not reported by this path Can be bounded by hand; see below

The four empty rows are the point of showing you this run. The free-check path does not record citation rate or average position, because it stores the answer text, the mention flag and a brand tally, and nothing else. Those three columns are blank not because Miro scored zero on them, but because this measurement never asked. A report that filled them in with a plausible-looking zero would be lying, and a report that filled them from a different tool or a different day would be worse, because the numbers would look like they belonged together.

A missing column is a real finding. Write "not measured" and leave it.

On mention rate. Miro was flagged as mentioned in all three answers. Two details keep that number honest. The saved excerpt for each answer is the first 600 characters, and the mention flag is computed on the full answer text, so the third answer reads as though it is about Microsoft Whiteboard while the flag correctly says Miro appears further down. And this path matches the literal brand name with no alias list, which is fine for a name like Miro and would understate a brand that the engine spells several ways.

On the interval. 3 out of 3 gives 43.8% to 100%. Miro is a well-known name in its category and the point estimate is the maximum possible, and the honest reading is still that this run cannot distinguish "always recommended" from "recommended slightly more often than not". If we ran it again next month and got 2 out of 3, we would have measured nothing.

On the competitor tally. Lucidspark and Lucidchart each appear in all three answers, which is the more interesting number on the page: two products from one vendor hold every answer we drew, in a category where the brand we asked about also holds every answer. Each name is counted at most once per answer, the same way our own side is counted, so these counts are answer counts and not name counts.

The tally is capped at five names. You can see the cap working in the raw file: the third answer's text names Microsoft Whiteboard, and Microsoft Whiteboard is not in the five. So the five names are a floor on how many brands shared these answers, not the whole set.

On share of voice. The free-check path does not compute it, and we are not going to print our pipeline's figure for a run that did not produce one. What you can do is bound it by hand. Our side contributes 3. The five visible competitor names contribute 3 + 3 + 2 + 2 + 2 = 12. That gives 3 ÷ (3 + 12), which is 20%. Because the tally is truncated at five and the true denominator is larger, 20% is an upper bound rather than an estimate. It is also not the pipeline's share of voice, which counts only names on a confirmed competitor list, whereas this tally is whatever the extractor pulled out of the text.

That is the whole report. Four numbers, one bound, four blanks and a caveat on each.

What this report cannot tell you

Google AI Overviews and AI Mode are outside what we measure. The surfaces we monitor are ChatGPT, Gemini and Perplexity. Nothing on this page, and no number from our pipeline, describes what Google shows above the blue links.

Cause. Every metric here is a description of where you stand in a set of answers on one date. A number that moved after you published something is not evidence that the publishing moved it. Establishing that needs a deliberate before-and-after with intervals that do not overlap, which is the subject of a later lesson.

Personalisation. These measurements come from APIs with no signed-in user, no chat history and no app context. They describe what an engine says to an anonymous caller.

Sentiment. We do not score how favourably an answer describes you. We record whether it named you, where, and what it cited. Any tool that reports sentiment from a handful of answers is reporting a judgement made by a model, and it needs its own denominator and its own error bars before it belongs in the same table as a rate.

Questions

How do you measure LLM visibility? As a set of rates with explicit denominators, not as a single score. Mention rate over valid answers is the headline. Average position, citation rate, citation share and share of voice each answer a different question and each has a different denominator, which is why collapsing them into one figure loses the information. The definitions our pipeline implements are on the method page, and the commercial version is described on AI visibility tracking.

What is the difference between a brand mention and a citation in an AI answer? A mention is your name in the answer's prose, counted after link targets and bare URLs are stripped out. A citation is a URL the answer points at. They have different denominators, either can occur without the other, and an answer can also read a page of yours without citing it, which we record separately as consulted but not cited.

Why is average position blank in my report? Either no answer placed you in a list-shaped response, or the measurement path you used does not detect position at all. In our run above it is the second reason. Average position is computed only over answers where a list position was found, so a report built from table-formatted answers can show a mention rate and no position, and that is correct behaviour rather than a bug.

Why did my share of voice fall when my mention rate did not move? Most often because the answers got longer. Your side of the fraction counts at most one per answer while the competitor side sums every rival named, so an engine that starts writing ten-item lists instead of four-item ones lowers your share of voice without changing how often you are named. Check the answers before you go looking for a cause on your own site.

Which AI search visibility metrics and KPIs are worth a column in the report? Mention rate with an interval, average position with the count of answers it was computed from, citation rate, citation share, share of voice against a fixed competitor list, the competitor tally, and refusals. Any column whose denominator you cannot state is not worth carrying, because you will not be able to say what a change in it means.

How many answers do I need before a number means anything? More than three. Three out of three gives a 95% interval of 43.8% to 100%, which does not separate a brand that is always recommended from one that is recommended half the time. Compare intervals rather than point estimates, and treat any run whose interval spans more than about forty points as a smoke test.

Does this report cover Google AI Overviews? No. AI Overviews and Google's AI Mode are outside what we measure. Our monitored surfaces are ChatGPT, Gemini and Perplexity, and a figure from us should never be presented as covering Google's AI surfaces.

What does a 100% mention rate actually prove? That every answer in a small sample named you, and little else. It does not establish that the true rate is high, because at three draws the interval reaches below 50%. It says nothing about whether you were cited, where in the list you sat, or how much of the answer belonged to your rivals. Those are four different questions and they need four different columns.

Next in this series: reading the competitor and cited-source columns to work out why an engine leaves you out. If you would rather have all of this produced on a schedule across three engines with the intervals calculated for you, that is what AI visibility tracking measures.