MentionCount
All articles

2026-08-13 · 2 min read

Why asking ChatGPT once proves nothing

A language model is not deterministic: the same question yields three different answers. Without repetition, an AI visibility test is just an anecdote.

The instinct is right: you open ChatGPT, you type "best supplier of X in France", you look for your name. It is not there. Or it is.

Either way, you have learned nothing.

The same prompt, three answers

A language model samples. At each generation it picks among several plausible continuations according to a probability distribution. This is not a bug, it is normal operation: the same question asked three times in a row produces three different answers, with different brands in them.

Add web search and the variance climbs further: the pages consulted are not exactly the same from one call to the next.

In other words, your single test measured one draw. Not your visibility.

What that changes in practice

Imagine a brand genuinely recommended in 10% of the answers in its market. Asking the question once:

Both conclusions are wrong. The correct answer — 10% — only shows up through repetition.

The same reasoning cuts the other way: a well-established brand can vanish from one answer in three. An executive testing at the wrong moment makes an expensive decision based on noise.

Three passes, three models

Our protocol comes down to three rules, and each one corrects a specific source of error.

Three passes per question neutralise the model's variance. It is not many — statistically, more would be better — but it is the threshold at which a signal becomes distinguishable from an accident, without blowing up the cost.

Three assistants neutralise a single model's bias. Perplexity, ChatGPT and Claude do not read the same sources and do not recommend the same brands. Measuring on one alone means mistaking an engine for the market.

Forty questions neutralise phrasing bias. "Where to buy X" and "which site to find X" do not trigger the same answers. A single question, however often repeated, measures only one angle of attack.

Forty questions, three models, three passes: 360 measurements. At that volume a citation rate becomes readable, and a gap between two models becomes interpretable.

The rule that makes everything else valid

One last condition, the most important: no question may contain your name.

If you ask "what do you think of brand So-and-so?", the AI will talk about So-and-so. Obviously. All you will have measured is its ability to describe a brand you fed it — not its propensity to recommend that brand spontaneously.

Every one of our questions is phrased from the point of view of a buyer who does not know you yet: "where to buy…", "which supplier for…", "what is an alternative to…". It is the only way to get a number that means something.