RedClawLabs

Guides · Guide · Published 2026-09-20

Why the AI answer left your brand out, and what the data can honestly tell you

A diagnostic method for a zero mention rate, worked on one of our own runs where the brand was named in none of three answers, separating what the competitor and citation columns show from what they only suggest.

A brand is absent from an AI answer when the engine names other vendors in the category and not you. The answer records two pieces of evidence about that absence: which competitors took the slots, and which sources the engine cited. Neither of them tells you why you were left out.

That last sentence is the whole problem with this topic. We read the pages currently ranking for it before writing this one. Most of them are organised as a numbered list of causes, presented as established fact, with no way for the reader to check which cause applies to their own brand. One of the better ones cites dated third-party studies but keeps its own dataset private. None of the ones we opened gave a test the reader could run to tell one cause from another.

So this lesson does something narrower. It takes one real run where the brand was not mentioned, reads every column of evidence that run produced, and then stops at the exact line where observation ends. What is on the other side of that line is a hypothesis, and the last section is about how you turn a hypothesis into a measurement.

If you have not measured yet, start with how to track brand mentions in ChatGPT. You cannot diagnose an absence you have not recorded.

The run

On 2026-09-20 we ran our free-check path against the ChatGPT API for Sketchboard, in online whiteboard and visual collaboration tools, US market, English. Three generated buyer questions, each asked once, with the brand name kept out of every question. Raw output: docs/content/runs/L5.json.

The questions, verbatim:

  1. "Can you recommend online whiteboard tools that are best for remote teams in the US with a budget under $20 per user per month?"
  2. "Which visual collaboration platforms offer the best integration with popular project management software for small businesses?"
  3. "What are the top online whiteboard vendors that provide strong security features suitable for healthcare organizations?"

Sketchboard was named in none of the three answers. The tally is 0 of 3, the point estimate is 0%, and the Wilson 95% interval for 0 out of 3 runs from 0% to 56.2%, which is the bottom row of the table on our measurement method page.

Read that interval before you read anything else here. Three draws that all came back empty are still compatible with a brand that gets named more than half the time. This run is a smoke test that produced a question worth investigating. It is not a finding about the product, and nothing below should be read as one.

We picked a small vendor on purpose. A brand that every engine already names cannot demonstrate absence, so a lesson about absence needs one that might actually be missing. Every other run in this batch came back 3 of 3, which is useful for checking that the machinery works and useless for showing you what a zero looks like in a spreadsheet.

The control, before anything is read into the zero

A mention rate that sits at zero is the one result you should distrust first, because a broken brand matcher produces exactly the same number as a genuinely absent brand. Our own operating rule is to treat a rate that never moves off zero as a matching bug until proven otherwise, which is on the method page under the alias limits. A zero with no control attached is not evidence of anything, and reading a cause into one is the error this lesson is about.

So here is the control, measured the same day in the same category:

Control run This run
Raw output docs/content/runs/L2.json docs/content/runs/L5.json
Brand Miro Sketchboard
Industry string online whiteboard and visual collaboration tools identical
Market and language US, English identical
Measured at (UTC) 2026-09-20 07:24 2026-09-20 07:29
Questions 3, generated 3, generated
Mentioned 3 of 3 0 of 3

The two question sets were generated independently and came out close enough to compare directly. Questions 1 and 2 are word for word the same in both runs. Question 3 differs by two words, "known for strong security features" against "provide strong security features". Same category, same market, same shape of question, five minutes apart.

What that rules out: the pipeline does find a brand in this category when the answers contain one. Generation, asking and brand matching all worked on the control, so the 0 of 3 is a statement about this brand in these three answers on this date rather than a comparison that silently failed.

What it does not rule out, and this is the part that matters: the zero is still n = 3, one engine, one day. It is not evidence about the product, and it is not a pattern. A control tells you the instrument was reading. It does not widen the interval, which still runs from 0% to 56.2%.

Column one: who took the slot

The competitor column is the first thing to read, because it tells you which answer you are actually competing for.

Brand recorded Answers naming it, out of 3
Miro 3
FigJam 3
Figma 3
Lucidspark 3
Mural 2

One detail about this table matters more than the ranking. The extractor recorded Figma and FigJam as separate names, and they are the same vendor's canvas product line. Five rows here are fewer than five independent companies. If you build this tally by hand, decide up front whether you are counting products or vendors, and keep that decision fixed between runs, because changing it silently rewrites your history.

What the column establishes: on this date, for these three questions, the shortlist was occupied by a consistent set of names, and four of them appeared in every single answer. A brand competing here is not competing against a shifting field. It is competing against a default.

What the column does not establish: why those names and not others. Nothing in an answer explains the selection.

Column two: what the engine cited

Our paid monitoring records a ranked list of cited domains for every answer, each classified as owned, satellite, competitor or third party. That metric is defined on the method page and it is the column this lesson was supposed to be built on.

The free-check path does not record it. It stores the answer text truncated at 600 characters and no source list at all. That is a limitation of the pipeline this run used, and pretending otherwise would be the same failure the article is about.

What we can read is the source markers ChatGPT left inline inside the text we did capture. Across the three answers there were two, and both belonged to a competitor:

Answer Source markers visible in the stored text
1, budget under $20 per user miro.com
2, project management integrations get.miro.com
3, healthcare security miro.com

In the first answer the engine also wrote, in its own words, that it "checked each vendor's public pricing pages" before recommending. That is a claim the answer makes about itself, not something we verified.

Stated without interpretation: every marker we can see points at a vendor's own domain rather than at a review site, a comparison article or a forum, and none of them point at the brand we were measuring. This list is also partial by construction. The stored record is truncated, so we cannot say these were the only sources consulted.

That last point deserves its own warning. A truncated view of the citations is the easiest way to convince yourself of a pattern that is not there. If your own tally is built from what you can see in a chat window rather than from an API that returns the full source list, you have the same problem, and you should treat your citation column as a sample of the citations rather than as the citations.

The two kinds of absence

Once you have both columns, an absence sorts into two broad shapes. The middle column of this table is a hypothesis in every row, and the right-hand column is the only part of it you can act on.

What it looks like in the data What it might mean How you would test it
Mention rate at or near zero across the whole question set, your domain never appears among cited sources, and a question that names the brand directly produces a vague, wrong or hedged description The engine has little retrievable material about you in this category. Absence of anything to say, rather than a decision not to say it Add a vendor_check question that names the brand on purpose and keep it out of your headline rate. Check that your pages are actually fetchable by the engines you measure. Re-measure the same set and compare intervals
Mention rate at or near zero, but your domain does appear among cited sources, or a direct question produces an accurate description of what you sell The engine can find you and is not putting you on the shortlist. A ranking outcome rather than a knowledge gap Record consulted and cited separately, since an answer can read your page without citing it. Then change one thing on the pages that are being read, hold the question set frozen, and re-measure
Mention rate at or near zero, and the competitor column is full of vendors serving a different buyer than you do The question set is measuring a segment you are not in. A measurement artefact rather than a visibility problem Read your own questions back against your positioning. Write a second set for your actual segment, run both, and keep them as separate series rather than merging them

The third row is not a kind of absence. It is the thing you have to rule out before the first two rows mean anything, and it is the one people skip.

It is live in this run. Sketchboard's own site describes the product for sketching and diagramming work, naming UML diagrams, flowcharts, wireframes and mindmaps (sketchboard.io, retrieved 2026-09-20). Our three generated questions asked about per-seat budget, project management integrations and healthcare compliance. Whether a whiteboard question phrased around healthcare security is a question this vendor was ever in the running for is not something the run can answer. It is a reason to write a second question set before drawing any conclusion from the first.

What we observed, and what we are only guessing

This section exists because the split is easy to lose once you start writing recommendations.

Observed, and reproducible by anyone with API access:

On 2026-09-20, three specific questions asked once each of the ChatGPT API for the US market in English produced three answers. None named Sketchboard. Miro, FigJam, Figma and Lucidspark appeared in all three, Mural in two. The source markers visible in the stored text were miro.com and get.miro.com. Five minutes earlier, the same pipeline running the same category and market returned 3 of 3 for a different brand, so the comparison was working. The full questions and answer excerpts are in docs/content/runs/L5.json, and the control is in docs/content/runs/L2.json.

Not observed, and not knowable from this run:

Why the engine chose those names. Whether it retrieved anything about Sketchboard and discarded it, or retrieved nothing. Whether the sites it cited caused the selection or merely accompanied it. Whether a different phrasing of the same question would have produced a different list. Whether any of this would still be true tomorrow.

We can see which sites were cited. We cannot see why. That gap is not a limitation of our tooling that a better tool would close, because the engine does not expose its selection process to anyone. Every article that tells you "you are missing because X" has crossed this line, including the ones with the most confident formatting.

The practical consequence: do not rewrite anything on the strength of a single run. One run cannot establish a pattern, and a domain that appeared in one answer on one day is not a publication you need to get listed on. It is one data point, with a date and a sample size of three attached to it, and you should write both of those down next to it.

What the research actually licenses

There is one peer-reviewed reference worth knowing here. Aggarwal et al., GEO: Generative Engine Optimization (arXiv 2311.09735, KDD 2024), tested content-side edits against a benchmark of user queries and reported that the method can "boost visibility by up to 40%" in generative engine responses, while also finding that "the efficacy of these strategies varies across domains" (retrieved 2026-09-20).

Two things about that number before you use it. It is an upper bound rather than a typical result, and the generative engine it was measured against was assembled by the authors for the study, not a commercial assistant. The paper is good evidence that what is written on a page changes whether a generative engine cites it. It is not a table of expected gains for your category, and the authors say as much in the abstract.

Which is exactly why the right-hand column of the table above is the useful one. The literature tells you that the lever exists. Only your own before-and-after tells you whether pulling it moved anything for you.

Turning a hypothesis into a measurement

The loop is short and the discipline is in the middle step.

Measure first, and write down the interval, not just the percentage. You need the baseline before you touch anything, and at three or nine draws the interval will be embarrassingly wide. That is information, not a reason to skip it.

Change one thing. One hypothesis, one intervention, one part of the site. If you rewrite the pages and expand the question set in the same week, you have two changes and no way to attribute a difference to either.

Re-measure with the question set frozen. Identical wording, identical count, identical market. A reworded question is a new question with its own history.

Compare the intervals, not the point estimates. If this run's Wilson interval overlaps the last one, you have not yet measured a change. That is the rule our own reports run on, and the arithmetic behind it is on the method page.

A zero that stays zero after a genuine change is also a result. It usually means the sample is too small to see the size of effect you produced, which is an argument for more draws rather than for a different theory.

What this does not cover

Google AI Overviews and AI Mode are outside what we measure. The surfaces we monitor are ChatGPT, Gemini and Perplexity. Nothing on this page describes what Google shows above the blue links.

Personalisation. These runs go through APIs with no signed-in user and no chat history. An absence here describes what an anonymous caller is told.

Anything about the product. A brand not appearing in three answers is a fact about three answers.

Cause. Stated once more because it is the point of the lesson. The columns narrow the list of explanations. They do not pick one.

The monitoring we sell runs this loop weekly across three engines with the citation column intact, which is the column the free path drops. What it measures, and what it deliberately does not, is written out on AI visibility tracking.

Questions

Why does ChatGPT not mention my brand? A single run cannot tell you. It can tell you which brands took the slots you wanted and which sources the answer cited, which narrows the explanations to roughly three: the engine has little material about you, it has material and is not shortlisting you, or the questions describe a buyer you do not serve. Distinguishing them needs a second measurement designed to separate them, not a longer list of possible causes.

What does a zero mention rate actually mean? First check that it is real. A broken brand matcher returns zero just as reliably as a genuinely absent brand, so run a known-present brand through the same category and confirm it comes back mentioned before you interpret anything. Once the zero survives that, it still means less than it looks like at small sample sizes: zero out of three carries a Wilson 95% interval of 0% to 56.2%, which is compatible with a brand that gets named most of the time. Treat a zero from a short run as a reason to measure properly rather than as a diagnosis.

Should I try to get listed on the websites the AI answer cited? Not on the evidence of one run. In our 2026-09-20 run the only source markers we could see were a competitor's own domains, across three answers, from a record that truncates the text. That is far too little to identify a site that matters, and treating one run's citations as a target list is how people spend a quarter on the wrong thing. Build the list from repeated runs with the full citation column recorded.

How do I tell whether the engine does not know me or does not trust me? Ask a question that names the brand directly, keep it out of your headline rate, and read the description it produces. An accurate description alongside a zero mention rate points toward a shortlisting problem. A vague or wrong one points toward a retrieval problem. Both readings are hypotheses that you then test by changing one thing and measuring again.

My competitors appear in every answer. Does that mean the answer is fixed? It means it was consistent on that date for those questions. Four of the five names in our run appeared in all three answers, which tells you that you are competing against a stable default rather than a lottery. Whether that default is stable over weeks is a question for repeated measurement, which is what AI visibility tracking is for.

How many runs before an absence means something? More than one, and the useful threshold is when the confidence intervals from two periods stop overlapping rather than a fixed count. Our own rule is that a question needs at least four weekly runs before it enters a trend line. Until then, read each run on its own and compare the ranges.

Does any of this apply to Google AI Overviews? No. AI Overviews and Google's AI Mode are outside what we measure, and the diagnostic here depends on an API that returns the sources behind an answer.