Why two AI visibility checks of the same brand disagree

LadderFoxai visibility tracking accuracy

Ask ChatGPT which tools it would recommend in your category. Note the brands. Now open a clean window and ask the identical question.

The list changes. Most people blame their own phrasing, or assume they caught the model in the middle of an update. In January 2026 Rand Fishkin's team at SparkToro ran the experiment properly, with 600 volunteers producing 2,961 prompt runs across ChatGPT, Claude and Google AI. The probability that two runs of one prompt return the same set of brands came out below one in a hundred. Same brands in the same order, below one in a thousand. Claude was marginally steadier than the other two and still nowhere near reproducible. Which raises an obvious question about the products currently selling you a position in ChatGPT.

Kevin Indig reached similar ground from a different direction, analysing 815,000 pairs of prompts and pages.

What was measuredFindingSource
Two runs returning the same brand listBelow one in a hundredSparkToro
Two runs returning the same orderingBelow one in a thousandSparkToro
Citations surviving three runs of one prompt2.2 percentIndig
Variance within one model, same prompt10 to 34 percentIndig
Sources replaced weekly, Google AI Mode56 percentIndig
Sources replaced weekly, ChatGPT74 percentIndig

Those last two figures are the ones I keep coming back to. Between half and three quarters of what an engine cites gets swapped out inside a week, with nobody publishing anything to cause it. Monthly reporting cannot describe a system moving at that speed.

The number that should worry everyone selling this

SparkToro ran a second experiment that gets quoted far less, and I suspect that is because it is awkward for every vendor in the category, us included. They asked people to write prompts expressing the same underlying intent, then measured how close those prompts were to one another. Average pairwise semantic similarity: 0.081. The authors likened it to the similarity between Kung Pao Chicken and peanut butter.

Nobody is typing the prompts in your tracking dashboard. Your buyers ask for the same thing in almost entirely different words, which makes any prompt list a sample of a distribution rather than a reading of it.

Volume fixes most of this

The same dataset makes the opposite case once you look past the headline, and this part matters if you have concluded the whole field is astrology.

One digital marketing agency in the SparkToro sample turned up in 85 of 95 Google AI responses. That is 89.5 percent, across 95 separate rolls of a supposedly random process. A cancer centre managed 69 out of 71 on ChatGPT. Over roughly a thousand responses the leading headphone brands held between 55 and 77 percent visibility, while brand design agencies sat much lower, somewhere around a third. So individual answers are noise and rates across many answers are not, which is an annoying thing to have to hold in your head at the same time.

So "you rank third in ChatGPT" describes nothing, since there is no ranking and the next run reshuffles it. "You appear in 31 percent of answers to these 24 questions, plausibly somewhere between 31 and 69 percent" describes something checkable.

What your sample size does

Ask 24 questions. Appear in half the answers. The honest reading of that result runs from 31 to 69 percent at 95 percent confidence, an interval close to forty points wide, which manages to accommodate both "this is going fine" and "this is a serious problem" simultaneously. More questions tighten it. Fewer widen it until nothing survives. A visibility percentage published without its sample size beside it is a true statement about one afternoon, and I would treat it accordingly.

Two consequences worth spelling out, because they cost money.

If your score moves from 31 to 38 percent and each reading carries a band nineteen points either side, nothing has been demonstrated. You measured the same quantity twice and got two draws from it. Reporting that to a client as improvement is reporting noise.

Comparisons across a changed instrument are void. Different questions, a different number of them, an engine added or dropped: the new percentage cannot sit on the same axis as the old one, however similar the two numbers look.

Five things worth asking a vendor

Fishkin closed his piece by telling readers to stop spending on tracking products that cannot show research anyone else can review. That applies to us too, so here is the list I would use.

Start with how many runs the number is built on. One run is what Indig calls a coin flip with the answer hidden.

Then ask where the uncertainty went. Every percentage has an error bar, and if it is not on the screen somebody made a decision about what you get to see.

The question set is the one people forget to ask about. Old and new numbers drawn as one continuous line make a claim about comparability that nobody verified, and the line looks just as smooth either way.

Engines time out, so find out whether a provider outage turns into a bad score for the brand. Scoring on whichever engines happened to respond that morning produces movement with no cause behind it.

Last, does a mention count as a win regardless of what it says? Counting mentions will happily record a paragraph explaining why you are the wrong choice. Indig flags this as unsolved across the category and he is right.

How we handle it

Our scores carry the interval and the sample in words rather than in a tooltip, so the product says "plausibly 31 to 69 percent, from 24 answers". Questions freeze when you set up a brand, which is what makes two runs comparable at all. If you change the instrument, the chart breaks the line at the join rather than drawing through it. Alerts fire only outside the noise band, which means we contact you less often than tools that ping on every wobble, and there is no rank anywhere in the product because we have no way to measure one.

The live demo is open without an account if you want to see how the intervals are presented, and you can check a domain yourself.