LLM visibility

LLM visibility is a measurement: the share of AI answers — across a defined set of prompts — in which a large language model mentions or cites your brand. It is a rate, not a rank: there is no position to hold, only the proportion of generated answers you appear in. Because that proportion depends entirely on which prompts you chose, how many times you ran them and which engine you asked, an LLM visibility number is only meaningful when it arrives with its sample size and its confidence interval. This page is about the metric itself — how it is calculated, how much sampling it needs, and what “good” honestly means.

Last updated 2026-08-27 · sources linked inline

The short version

  • LLM visibility is a rate — the share of AI answers, across a fixed prompt set, that mention or cite your brand.
  • It is not a rank. There is no position to hold, so the number only means something next to its prompt set, its sample size and its interval.
  • A single run is a draw, not a reading: 40–90% of the domains an engine cites can change when the same question is re-asked (Profound).
  • Sample size sets what you can detect: at a ~50% rate, the 95% interval is about ±17 points at n=30, ±10 at n=100, ±5 at n=400.
  • Engines cannot be averaged: only 2.37% of cited URLs appear across all three major engines, and ~91% appear in just one (Kevin Indig, 3.7M citations).
  • There is no industry benchmark for a 'good' score — the only honest comparisons are to your own baseline and to named competitors in the same answers.

What LLM visibility measures

The base calculation is simple, and stating it plainly removes most of the confusion in this category:

  • LLM visibility = answers that mention your brand ÷ total answers generated (prompts × runs per prompt), as a percentage.

Everything else is a variation on the numerator or the denominator. A mention rate counts any naming of your brand. A citation share counts only answers that link to your domain. A recommendation rate counts only answers where the model actively suggests you rather than merely referencing you. Share of voice changes the denominator entirely — your mentions as a proportion of all brand mentions in the same answers — which is why it can fall while your absolute visibility holds steady, purely because a competitor started appearing more.

These are not interchangeable, and conflating them is the most common way an LLM visibility dashboard misleads. A 30% that means “we appear in three answers in ten” and a 30% that means “we are three of every ten brands named” describe very different competitive positions. Before you act on a number, establish which one it is.

Why a single run is noise, not a reading

Language models are non-deterministic: the same prompt produces different answers, drawing on different sources, on different runs. This is not a bug you can configure away — it is how the systems work, and it means one ask tells you almost nothing.

The size of the effect is well documented. Profound's tracking found that 40–90% of the domains an engine cites can change when the same question is re-asked. And the variation is not only run to run: a Semrush study published 30 June 2026 ran 100 prompts through GPT-5.2 in both Instant and Thinking mode and found only about 25.6% of cited domains overlapped between the two modes — nearly three in four cited sources were different depending on how hard the model thought.

So when someone asks ChatGPT “what's the best tool for X?”, sees their brand, and concludes they have LLM visibility, they have observed one draw from a wide distribution. Ask again tomorrow and the conclusion may reverse. The correct mental model is polling, not rank checking: you are estimating a population parameter from a sample, and every rule of sampling applies.

How many runs does a number need?

This is the question almost nobody in the category answers, and it has an exact answer. If you treat each run as a trial that either mentions you or does not, the uncertainty around your measured rate follows standard binomial arithmetic. Using the Wilson score interval — the standard method for a proportion, and the one that stays sane at small samples and extreme rates — a brand appearing in roughly half of answers gets these 95% intervals:

  • 30 runs: ±17 pointsroughly 33%–67% — only catches enormous swings
  • 100 runs: ±10 pointsroughly 40%–60% — a usable monthly read
  • 400 runs: ±5 pointsroughly 45%–55% — detects a real 10-point move

Two consequences follow, and both are uncomfortable. First, halving the interval costs four times the samples: precision gets expensive fast, which is why serious measurement is a recurring cost rather than a one-off audit. Second, most reported movements are not detectable at the sample sizes people actually run. If your tool ran a prompt 30 times and reports a drop from 52% to 44%, that eight-point move sits comfortably inside a ±17-point interval. It is not evidence of anything.

The practical rule: decide the smallest change you would actually act on, then sample until your interval is narrower than that change. Sample per prompt and per engine rather than pooling everything into one figure, and run daily rather than hourly — the underlying answer distribution does not shift meaningfully hour to hour, so faster sampling mostly buys noise and cost.

What a confidence interval actually means here

A 95% confidence interval on an LLM visibility figure is a statement about the method, not a promise about the number: if you repeated this sampling procedure many times, about 95% of the intervals it produced would contain the true underlying rate. In everyday use it answers one question — could this number plausibly be the same as last month's?

  • Overlapping intervals between two periods mean you cannot claim a change. Not that nothing happened — that your data cannot tell.
  • Non-overlapping intervals mean something probably moved, and it is worth investigating why.
  • A wide interval is not a failure of the tool; it is an honest report that you have not sampled enough to say more.
  • A point estimate with no interval is the actual failure — it hands you a number with its uncertainty deleted, which is precisely the information you needed to decide whether to act.

This is why intervals are not a nicety in this category. In a stable measurement environment you can get away with point estimates. In one where 40–90% of cited domains rotate, a point estimate is an invitation to spend a sprint chasing a movement that never happened. If you want the practical version of running this loop day to day, our page on AI brand monitoring covers the monitoring workflow; this page is about the number underneath it.

Why you cannot average the engines

Many tools report a single blended “LLM visibility score” across engines. That number is convenient and close to meaningless, because the engines do not agree with each other. Kevin Indig's analysis of 3.7 million citations found that only 2.37% of cited URLs appeared across ChatGPT, Perplexity and Google AI Overviews for the same prompt, while about 91% appeared in only one of them. These are not three views of one index; they are three largely disjoint distribution systems.

A blended score therefore hides the only thing you needed to know: which engine you are absent from. Being strong on Perplexity and invisible on ChatGPT averages out to a mediocre-looking number that points at no action at all. Report per engine, act per engine, and use the blend — if at all — only as a headline for people who will not read further. The same logic applies within an engine, given that reasoning modes cite largely different sources.

What “good” LLM visibility looks like

There is no industry benchmark, and this is worth being blunt about: any tool quoting one is comparing numbers that were never comparable. Your rate is a function of the prompt set you chose. Ask ten branded questions and you will score near 100%; ask ten open category questions in a crowded market and 15% may be excellent. The prompt set is the ruler, and everyone's ruler is different.

What is actually assessable:

  • Your own trend on a frozen prompt set. The only clean comparison. Change the prompts and you have reset the baseline — treat it as a new series, not a continuation.
  • Relative standing in the same answers. Measuring named competitors inside the same generated answers controls for prompt choice, engine and run count in one move. This is the most defensible competitive number available.
  • Coverage rather than depth. How many of your priority prompts you appear in at all is often more actionable than how strongly you appear in a few — absence is a clearer instruction than a low rate.
  • Position and framing within the answer. Being named first and described accurately is worth more than being listed sixth or characterised wrongly, and neither shows up in a bare rate.

Building a prompt set you can trust

Because the prompt set defines the metric, it deserves more care than it usually gets. Three rules do most of the work.

  • Use questions with real demand behind them. A prompt nobody asks produces a number nobody should act on. Prompts backed by actual monthly search volume keep the metric anchored to demand rather than to what the team imagined buyers ask.
  • Freeze it, then change it deliberately. Every edit to the prompt set breaks comparability with everything before it. Version the set and record when it changed, the way you would with any instrument.
  • Cover the fan-out, not just the headline question. Engines rewrite one prompt into several related searches, so a set built only from your dream question measures a narrower surface than the one you are actually competing on.
A note on honesty

We will not publish a “typical” or “good” LLM visibility percentage, because no such figure exists that is comparable across different prompt sets, engines and counting rules — and inventing one would make this page more quotable and less true. The sampling numbers above are plain arithmetic you can reproduce with any Wilson interval calculator; every empirical claim is linked to a named study.

How steek measures LLM visibility

steek is built around this metric rather than bolted on top of it. It measures your visibility across ChatGPT, Perplexity, Gemini and Google AI Overviews — reported per engine, not blended into one flattering average — and puts a Wilson 95% confidence interval on every number, so you always know whether a movement is real before you spend a sprint on it. Each tracked prompt carries its real monthly search volume rather than a 1–5 estimate bar, so the prompt set stays anchored to demand that exists.

Around the metric it adds the things that make it actionable: co-mention gap analysis showing the exact sources competitors are cited in that you are absent from, alerts that only fire when a change clears its confidence interval, AI-crawler and robots.txt checks, shareable public reports, and per-SKU AI Shopping visibility tied to real revenue in GA4 for stores. Every engine is included at the flat price, from €19/mo billed annually (€29.90 month-to-month), with a money-back guarantee; founding early access is €9.90 for the first full year. It is EU-hosted, and it invents nothing: every figure is a real measurement carrying its own interval.

To improve the number once you can measure it honestly, see LLM SEO for the technical layer of getting cited, ChatGPT SEO for the engine most people start with, and answer engine optimization for the writing craft. To choose a tool, compare the best AI visibility tools, or read our direct answer on which AI engines you should track.

FAQ

What is LLM visibility?

LLM visibility is a measurement: the share of AI answers, across a defined set of prompts, in which a large language model mentions or cites your brand. It is a rate, not a rank — you pick the questions your buyers ask, run them repeatedly against ChatGPT, Perplexity, Gemini or Google AI Overviews, and count how often your brand appears in the generated answer. Because the number depends entirely on which prompts you chose, how many times you ran them and which engine you asked, an LLM visibility figure is only meaningful alongside those three things.

How is LLM visibility calculated?

The base calculation is appearances divided by opportunities: the number of answers that mention your brand, divided by the total number of answers generated (prompts multiplied by runs per prompt), expressed as a percentage. Variants change what counts as an appearance — a plain mention, a linked citation to your domain, or a recommendation where the model actively suggests you. Share of voice divides your mentions by all brand mentions in the same answers instead, so it moves when competitors move even if you did not change. The honest form of any of these is a rate plus the sample size it came from and a confidence interval around it.

How many times should you run a prompt to measure LLM visibility?

Enough that the confidence interval is narrower than the change you care about. The arithmetic is fixed: for a brand appearing in roughly half of answers, a 95% Wilson confidence interval is about ±17 points at 30 samples, ±10 points at 100, and ±5 points at 400. Halving the interval requires four times the samples. So if you want to detect a 10-point movement, 30 runs cannot do it and 100 is borderline. Practically, sample per prompt per engine rather than pooling, run daily rather than hourly — the underlying distribution does not change meaningfully hour to hour — and treat any single run as one draw, not a reading.

What is a good LLM visibility score?

There is no universal benchmark, and any tool quoting one is comparing numbers that were never comparable. Your score is determined by the prompt set you chose: a narrow set of branded questions produces a high number, a broad set of category questions produces a low one, and neither says anything about the other. The only meaningful comparisons are internal — your rate now versus your rate last month on the same fixed prompt set, and your rate versus named competitors measured in the same answers. Judge both against their confidence intervals rather than their point estimates.

Is LLM visibility the same as share of voice?

No, though they are often used interchangeably. LLM visibility is usually an absolute rate — the proportion of your answers that mention you. Share of voice is relative — your mentions as a proportion of all brand mentions in the same answers. They move independently: your share of voice can fall while your visibility holds steady, simply because a competitor started appearing more often. Track both, and be explicit about which one a given number is, because a 30% that means 'we appear in three answers in ten' and a 30% that means 'we are three of every ten brands named' describe very different situations.

Why do two LLM visibility tools give different numbers for the same brand?

Because almost every input differs. The tools use different prompt sets, different numbers of runs per prompt, different engines and model versions, different rules for what counts as a mention, and different handling of the engine's reasoning mode. The scale of that last one is easy to underestimate: a Semrush study published 30 June 2026 found only about 25.6% of cited domains overlapped between GPT-5.2's Instant and Thinking modes. Two tools can both be measuring honestly and still disagree by a wide margin, which is why the methodology behind a number matters more than the number.

Start now

Measure your LLM visibility with its error bar

Founding early access is €9.90 for your first year — per engine, with a 95% confidence interval on every number.