Guides · Guide · Published 2026-09-20 · about 18 min read · Eric Yu
Did it work? Re-measuring AI visibility after a change
Three identical runs of the same question set on the same morning, with nothing changed between them, and what their disagreement tells you about when a before-and-after comparison is allowed to mean anything.
In short
- Three identical runs on 2026-09-20, with nothing changed between them, returned brand mention counts of 3, 3 and 2 out of 3.
- Compare the two Wilson 95% intervals rather than the two percentages; if they overlap, no change has been measured yet.
- A move from 30% to 50% needs 93 answers per period before the intervals separate, and a move from 30% to 40% needs 350.
- Change one thing per measurement period and keep the question set frozen, or the two periods cannot be compared.
On this page
Re-measurement is asking the same frozen question set again after a change and comparing the two confidence intervals rather than the two percentages. It only tells you something if you already know how far the number moves when nothing has changed, and the way you find that out is to run the same measurement twice with no change in between.
The previous lesson, why the AI answer left your brand out, ends by handing you a hypothesis and telling you to test it. This is that test. It is also the lesson most likely to talk you out of a conclusion you were looking forward to.
Everything below uses the definitions on our measurement method page. The mention rate, the Wilson 95% interval and the rule about overlapping intervals are all defined there, and nothing here redefines them.
What we have, and what we do not
We do not have a before-and-after of our own site. This site has not been running long enough for one, and inventing one to have something to show here would undo the whole point of it.
What we do have is the thing you need first. On 2026-09-20 we ran our free-check path three times against the ChatGPT API, same brand, same category, same market, inside a fifteen-minute window, with nothing changed between the runs. Not our content, not the question set, not the engine, not the market. Three readings of a stationary target.
That is not a before-and-after test. It is a measurement of the noise floor: how far the number moves when nothing has moved. Until you know that, a difference between two weeks is just a difference between two numbers.
Three runs, fifteen minutes, nothing changed
The three runs used our free-check path, described on the method page: three buyer questions generated per run, each asked once of the ChatGPT API, with the brand name kept out of every question. The brand was Miro, the category string was online whiteboard and visual collaboration tools, the market was US and the language was English in all three.
The question generator produced the identical set all three times, word for word:
- "Can you recommend online whiteboard tools that are best for remote teams in the US with a budget under $20 per user per month?"
- "Which visual collaboration platforms offer the best integration with popular project management software for small businesses?"
- "What are the top online whiteboard vendors known for strong security features suitable for healthcare organizations?"
So the question set is not the variable. Whatever differs between the runs differs downstream of it.
| Run | Recorded, UTC | Brand named | Point estimate | Wilson 95% | Competitor names recorded, with the number of answers naming each |
|---|---|---|---|---|---|
runs/L2.json |
2026-09-20 07:24 | 3 of 3 | 100% | 43.8% to 100% | Lucidspark 3, Mural 2, FigJam 2, Figma 2, Stormboard 2 |
runs/L4.json |
2026-09-20 07:29 | 3 of 3 | 100% | 43.8% to 100% | Lucidspark 3, Lucidchart 3, Stormboard 2, FigJam 2, Figma 2 |
runs/L8.json |
2026-09-20 07:39 | 2 of 3 | 66.7% | 20.8% to 93.9% | Lucidspark 3, FigJam 2, Figma 2, Microsoft 365 2, MURAL 2 |
Both intervals come from the table on the method page, which lists 3 of 3 as 43.8% to 100% and 2 of 3 as 20.8% to 93.9%.
What moved
The headline number moved on the third reading. Two runs at 3 of 3, then one at 2 of 3. The answer that changed was the first question, the one with the sub-$20 per-user budget constraint: in the first two runs it named the brand, in the third it did not, returning a list headed by Mural and FigJam instead. We are not going to tell you why. Nothing in the run explains the selection, which is the line the previous lesson spent its length drawing.
Had we stopped after two runs we would have had two identical readings of 100% and a tidy story about stability. The third reading is the useful one precisely because it is the inconvenient one.
The competitor column moved more than the headline. Seven distinct names appear across the three runs. Three of them appear in all three with the same count: Lucidspark at 3, FigJam at 2 and Figma at 2. The other four come and go.
| Name recorded | L2, 07:24 | L4, 07:29 | L8, 07:39 |
|---|---|---|---|
| Lucidspark | 3 | 3 | 3 |
| FigJam | 2 | 2 | 2 |
| Figma | 2 | 2 | 2 |
| Stormboard | 2 | 2 | not in the recorded five |
| Mural | 2, spelled "Mural" | not in the recorded five | 2, spelled "MURAL" |
| Lucidchart | not in the recorded five | 3 | not in the recorded five |
| Microsoft 365 | not in the recorded five | not in the recorded five | 2 |
Three details in that table are about the instrument rather than about the market, and you need all three before you read anything into it.
The list is a top-five slice. The free-check path sorts the tally by count and keeps the first five entries (main site, firebase/functions/src/geo/index.ts). Every run returned exactly five names, and in every run at least three of them were tied at a count of 2, so the cut between the fifth name and the sixth is decided by sort order rather than by evidence. "Not in the recorded five" means the count was at most 2. It does not mean zero. The one difference the cap cannot explain is Lucidchart, which was named in all three answers of the middle run and was below the cut in the other two.
The same vendor was recorded under two spellings. "Mural" in the first run and "MURAL" in the third. If you are building this tally by hand, two spellings of one name will sit in your spreadsheet as two rows and quietly halve a competitor's apparent share.
The competitor column passes through a second model. Your own brand is found by matching the name against the answer text, but the competitor names are pulled out of each answer by a separate extraction call. That is a second place where two runs of the same input can disagree. The headline mention rate does not have this problem; the competitor tally does.
What the three runs establish
Stated flatly, so it can be checked: on 2026-09-20, three runs of an identical three-question set against the ChatGPT API for the US market in English, recorded at 07:24, 07:29 and 07:39 UTC, returned a brand mention count of 3, 3 and 2 out of 3, and a recorded competitor set that agreed on three names out of seven. Nothing was changed between the runs. All three runs are recorded in full.
Stated as what it licenses: at three answers per run, a one-third swing in the headline rate happens without any input changing. Any single-run comparison that shows a swing of that size has shown you nothing.
Stated as what it does not license: three runs on one morning are not an estimate of how much the number drifts over a week, and none of this describes the paid path, which asks each question three times on each of three engines and is specified on AI visibility tracking. A wider measurement has a narrower noise floor. We have not measured how much narrower.
What it takes for two intervals to actually separate
The rule on the method page is short: if this period's Wilson interval overlaps the last one, the change is not yet evidence of anything. The arithmetic behind it decides how much measurement you have to buy before that rule can ever be satisfied.
The interval is computed at 95%, z = 1.96, in the centre-and-half-width form given on the method page:
centre = (p + z² / 2n) / (1 + z² / n)
half = z · √( p(1 − p)/n + z² / 4n² ) / (1 + z² / n)
Worked on a case people actually hope for, a mention rate going from 30% to 50% on 90 answers per period:
n = 90, z² = 3.8416, 1 + z²/n = 1.042684
p = 0.30 centre = (0.30 + 0.021342) / 1.042684 = 0.308186
half = 1.96 · √(0.002333 + 0.000119) / 1.042684 = 0.093078
→ 21.5% to 40.1%
p = 0.50 centre = (0.50 + 0.021342) / 1.042684 = 0.500000
half = 1.96 · √(0.002778 + 0.000119) / 1.042684 = 0.101164
→ 39.9% to 60.1%
The two intervals overlap by two tenths of a point. A twenty-point jump, measured on ninety answers each side, misses the bar. It clears it at 93 answers per period.
Note what the interval is doing at the ends of the scale. Unlike the textbook normal-approximation interval, "the Wilson score interval is asymmetric" (Wikipedia, Binomial proportion confidence interval, retrieved 2026-09-20), the method having been introduced by E. B. Wilson in 1927. The practical consequence for a before-and-after is that the range around a 90% reading is not the mirror image of the range around a 10% one, so a given gap is easier to separate near the ends of the scale than in the middle.
| Move you are trying to demonstrate | Answers per period before the two Wilson 95% intervals stop overlapping |
|---|---|
| 20% to 50% | 39 |
| 33% to 67% | 35 |
| 30% to 50% | 93 |
| 30% to 40% | 350 |
Read from the other direction, that is the smallest lift a given amount of measurement can see:
| Answers per period | Wilson 95% on a 30% baseline | Smallest lift whose interval clears it |
|---|---|---|
| 9 | 10.2% to 61.8% | +63.6 points |
| 30 | 16.7% to 47.9% | +35.8 points |
| 45 | 18.6% to 44.5% | +29.1 points |
| 90 | 21.5% to 40.1% | +20.3 points |
| 180 | 23.8% to 37.1% | +14.2 points |
| 360 | 25.5% to 34.9% | +9.9 points |
Three caveats on that table, none of them optional.
Answers per period is questions × engines × repeats. On our paid path that is three repeats per question per engine across three engines, so ten questions produce 90 answers and forty questions produce 360. On a hand-built spreadsheet with one engine and one ask per question, ten questions produce ten answers, and the top row of the table is the one you are living in.
The arithmetic treats every answer as one draw from a single proportion. Real question sets are not like that: a question you are well placed for and a question you are not have different true rates, so pooling them approximates. The approximation is good enough to tell you that nine answers cannot detect a ten-point move, which is the decision the table is for.
Non-overlap is a stricter bar than a formal test comparing two proportions. Intervals can overlap while a direct test would still call the difference real. We use the stricter bar deliberately, because at these sample sizes we would rather fail to claim a real effect than claim one that is not there.
How long to wait
Two different clocks get confused here.
The first is statistical and we can give you a number. A question enters our trend line only after it has been live for at least four weekly runs, which the pipeline implements as an activation date 21 days or more before the current run. Before that it appears in the report and stays out of the trend. That rule is on the method page and it exists because of the arithmetic above: a question with one run behind it has no period to be compared against.
The second clock is how long an engine takes to reflect a change you made on your own pages, and we have not measured it, so we are not going to give you a number for it. We have no data on how quickly an edit propagates into what ChatGPT, Gemini or Perplexity say, and neither does anyone quoting a figure at you without showing their runs. What we can say is that measuring the day after a change and calling the result final is the mistake the three runs above were written to prevent.
The practical consequence is that the waiting is cheap and the rushing is expensive. Keep the question set running on its normal cadence, mark the date of the change in the record, and compare a period that is entirely before it against a period that is entirely after it.
One change at a time
This is the rule that gets broken most, usually by someone who is busy.
Change one thing per measurement period. One hypothesis, one intervention. If you rewrite your comparison pages and add six questions to the set in the same week, the next reading is uninterpretable, and it is uninterpretable in four separate ways.
You cannot attribute a difference. The obvious failure, and the least damaging, because at least you know you have it.
Opposite effects cancel. If one change moves the rate up and another moves it down by a similar amount, the measurement shows no difference and you conclude that neither worked. You then discard a lever that was working. This failure is worse than the first because it does not look like a failure. It looks like a clean negative result.
Changing the question set breaks the comparison outright. A reworded question is a new question with its own history, so the two periods are no longer measuring the same thing. This is not a weakened comparison, it is not a comparison.
Question sets grow, which is why ours carry activation dates rather than being edited in place, and the same discipline is available to you for free in a spreadsheet: add a column for the date a question entered the set and exclude young questions from the trend. How the set is built in the first place is covered in how to write a question set that measures the market.
The noise floor is already one uncontrolled variable. The three runs above moved with zero changes applied. Every change you add is on top of that, not instead of it. Two changes plus the noise floor means three candidate explanations for one number that has a sample size of ninety.
The one thing you should change freely is the amount of measurement. Adding engines, repeats or questions does not contaminate a before-and-after, provided you keep the trend computed on the questions that were present in both periods.
When the honest conclusion is "no difference detectable"
Most re-measurements end here, and most write-ups pretend otherwise.
"No difference detectable" is not the same claim as "no difference". It says the effect, if there was one, was smaller than what this much measurement can see. That is a genuine result, and it has three uses that a fabricated positive does not.
It bounds the effect size. Go back to the second table. If you are running 90 answers per period from a 30% baseline and you see no separation, you have learned that whatever you did was probably worth less than twenty points. If your plan assumed it would double your mention rate, the plan has been tested and found wanting, which is information you paid for and should keep.
It stops a rollout. The expensive version of this mistake is not a wrong sentence in a report. It is rewriting two hundred pages on the strength of a nine-answer reading that happened to come back high.
It tells you what to buy next. A result that fails to separate usually means the sample is too small for the size of effect you produced, not that the theory is wrong. The fix is more draws, which is cheap and boring, rather than a new theory, which is expensive and exciting. The previous lesson ends on the same note for the same reason: a zero that stays zero after a genuine change is an argument for more measurement before it is an argument for a different hypothesis.
There is one result that does deserve a second look rather than more draws. If the number moves a long way in the wrong direction, treat it first as a possible fault in the instrument. A brand matcher that stopped matching, an alias that got dropped or a question that was silently reworded all produce dramatic negative movements. Confirm the machinery before you believe the collapse, the same way you confirm a zero before you interpret it.
Writing the result down
Whatever the comparison shows, record it so that the next person, including you in three months, can tell measurement from interpretation.
| Field | Why it is on the list |
|---|---|
| Date and time of each run | Our three runs were fifteen minutes apart and disagreed. A date alone is not precise enough to reconstruct what happened |
| Engine and model identifier | A pinned model changing is a break in the series, not a detail |
| The questions, verbatim | The only way anyone, including you, can repeat the run |
| Answers per period | The denominator decides what the comparison could possibly have detected |
| Point estimate and interval | The percentage on its own invites the reading you are trying to avoid |
| What changed, and when | One line. If it needs two lines, you changed two things |
| What you did not verify | The gap between what you observed and what you concluded, written down while you still remember where it is |
That last row is the one worth arguing for. Every number in this article points at a run we recorded in full, and the reason we can tell you that the competitor column churned but the cap may explain part of it is that we wrote down the cap. A record that only keeps conclusions cannot be re-read.
What this does not cover
Google AI Overviews and AI Mode are outside what we measure. The surfaces we monitor are ChatGPT, Gemini and Perplexity. Nothing here describes what Google shows above the blue links, and a number from us should not be read as covering it.
Cause. Consistent with the previous lesson: none of this establishes why an engine named one vendor and not another. Re-measurement tells you whether a number moved. It does not tell you what moved it.
Any claim that a specific technique works. We have not run a controlled before-and-after on our own site, so we are not in a position to tell you that any particular edit changes a mention rate. What we can hand you is the test you would use to find out.
Personalisation. These runs go through APIs with no signed-in user and no chat history.
Running this loop weekly across three engines, with the question set frozen and activation dates tracked, is what AI visibility tracking does. The arithmetic is public either way, and a spreadsheet with a Wilson formula in it will get you the same answer more slowly.
Questions
How do I know whether a change in my AI mention rate is real or just noise? Compute the Wilson 95% interval for both periods and check whether they overlap. If they do, you have not yet measured a change, regardless of how different the two percentages look. Our three same-day runs went 3 of 3, 3 of 3 and 2 of 3 on an identical question set with nothing altered between them, which is what that overlap is protecting you from.
How many answers do I need before a before-and-after comparison means anything? It depends entirely on the size of the move you are hoping to see. Using the Wilson interval as defined on our method page, a jump from 30% to 50% needs 93 answers per period before the intervals separate, and a jump from 30% to 40% needs 350. Nine answers per period cannot separate anything smaller than about sixty-four points. Decide the smallest move you would act on, then work out the measurement it requires before you start.
How long should I wait after making a change before measuring again? Long enough to accumulate a comparable period, which on our weekly cadence means a question needs at least four runs behind it before it enters a trend. We have not measured how long an engine takes to reflect a content change, so we will not give you a figure for that, and we would treat anyone who does without showing their runs with suspicion.
Can I change the question set between the two measurements? No, not if you want to compare them. A reworded question is a new question, and a set that grew between periods is measuring something the earlier period never asked. Add questions with a date attached and keep the trend on the questions present in both periods.
Why does the competitor list change between two identical runs? Partly because the answers differ and partly because of how the list is built. Ours keeps the top five names by count, so names tied at the bottom get cut by sort order rather than by evidence, and the names themselves come from a separate extraction step that can disagree with itself. Across our three runs, seven names appeared and three of them appeared every time. Before reading a trend into a competitor tally, check whether your list is truncated and whether one vendor is sitting in it under two spellings.
What should I do when the result is "no difference detectable"? Write it down as a result, because it bounds how large the effect could have been, and then decide whether to buy more measurement or drop the hypothesis. Do not roll the change out everywhere on a reading that could not have detected it either way, and do not rerun until you get a number you prefer.
Does running the same measurement twice in one day count as a before-and-after test? No, and that is the point of this lesson. It is a measurement of the noise floor, which is the thing you need before a before-and-after can be interpreted. Our three runs are a noise-floor reading and nothing more.
Do you have a before-and-after measurement of your own site? Not yet. This site has not been running long enough to have two comparable periods, and we would rather say so than publish a comparison we cannot support. When there is one, it will be published with the raw runs attached, the same as everything else here.
Does any of this apply to Google AI Overviews? No. AI Overviews and Google's AI Mode are outside what we measure, and the method here depends on APIs that can be asked the same question the same way every week.