Skip to main content

What this lets you do

Read every AI visibility number on your dashboard and know exactly what it counts, what it divides by, and how much of it is signal. This page is the method. AI answer engines covers which engines we ask and how faithful each one is. The dashboard page covers how to use the screens.

What visibility measures, and what it does not

Traceten asks a list of questions you define, on a schedule, across answer engines, and measures how often the answers name you and cite your pages. This is a poll of a prompt set you wrote. It is not an observation of what real people asked. That distinction is the whole basis for reading the numbers correctly, so it is worth being blunt about why the alternative is not available. When someone arrives from an AI assistant, what reaches us is the referring origin (https://chatgpt.com/) or a utm_source value, and, from the landing URL, only utm_source, utm_medium, utm_campaign and ref. None of those carries a question. utm_source=chatgpt.com tells us which assistant sent the visit and nothing about what was typed to produce it. The real prompts are not recoverable, by us or by anyone else selling this category of product. Nobody measuring AI visibility is measuring real user prompts, including us. So the honest description of every figure here is: of the questions you chose to ask, on the engines you are entitled to, this is how often you were named and cited. Its usefulness depends entirely on whether your prompt set resembles what your buyers actually ask, which is a judgement you make and we cannot make for you.
Two other measurements on your dashboard are observations rather than polls, and they are the reason this feature exists inside an analytics product rather than beside one. AI crawlers records the fetches AI companies actually made against your pages, unsampled. AI sources records the people who actually arrived. See Evidence tiers below for how a poll is reconciled against those.

The three metrics

They are computed separately, reported separately, and never combined into a single score. A run is one prompt, asked on one engine, once. A prompt asked 5 times on 2 engines in one scan is 10 runs.

Why the denominator is answered runs, not runs

Google does not render an AI Overview for every query. When it does not, we record that as a real outcome, and the product reports it (“Google showed an AI Overview for 12 of your 20 prompts”). But a search result page with no AI Overview is an observation about the results page, not a sample of whether you would have been named in one. Including it would push your presence rate down for a reason that has nothing to do with your visibility. So presence and citation both divide by answered runs. The gap between runs and answered runs is reported separately, as coverage. Sharing one denominator is also what keeps presence and citation comparable with each other, which is the point of refusing to blend them.

Why they are never one score

You can be named without being cited: the model has learned who you are, but used somebody else’s page to answer. You can be cited without being named: it used your page and credited nobody. Those are different problems with different fixes, and averaging them produces a number that moves for reasons you cannot identify and cannot check. Share of voice is a different shape again. It is not a share of runs. One answer naming three brands contributes to three numerators, so its denominator counts mentions, not runs, and it can exceed your presence rate.
0/0 is not 0%. Where nothing was answered, presence and citation return nothing rather than zero. Where no brand at all was mentioned, share of voice returns nothing rather than zero. A measured zero and an absent measurement look identical on a chart and mean opposite things, so we keep them apart everywhere: a blank cell is the absence of a measurement, a zero is a measurement.

What counts as a mention

Every mention carries a confidence score and records how it was matched (an exact brand name, a configured alias, or your domain named in the prose). Only mentions at confidence 0.5 or above count toward your presence rate. Lower-confidence hits are stored as evidence you can read in the drill-down, but they do not move a number. The case that matters: a two-letter or dictionary-word brand (“Go”, “Arc”) produces matches that are usually not about you. Those are recorded at 0.3 and excluded from the rate, rather than being silently dropped or silently counted.

Every number is a range

A presence rate is an estimate from a sample, so each one is published with a 95% confidence interval, drawn as a band on the chart and a range under the figure. At five runs, a 20% rate has an interval running from roughly 3.6% to 62.4%. The point estimate is 20%. What the data supports is “somewhere across most of the scale”. Quoting the 20% alone would be a claim the sample cannot carry.

k repetitions, not one

Each prompt is asked at least 3 times per engine per scan (5 on the top tier). The repetitions run together, minutes apart, so the variation between them is the model’s own non-determinism and the variation between scan windows is drift. Splitting repetitions across days would mix the two and make the interval meaningless.

Wilson, not the textbook interval

The interval is a Wilson score interval, not the normal approximation every tutorial shows. The normal approximation returns [0, 0] at a rate of zero, announcing “0% presence, no uncertainty” on the strength of three runs that happened not to mention you. Rates near 0 and near 1 at small sample sizes are the ordinary case here, not the edge case, so the degenerate end of that method is where this product would spend most of its time. Wilson stays inside [0, 1] and stays wide when the sample is small.

A change is only a change when the bands separate

A movement is styled as a change only when the two intervals do not overlap. Everything else renders flat with the band visible. 3-of-10 to 4-of-10 looks like a 33% improvement. At that sample the band is roughly ±25 points, so it is noise. Touching bands count as overlapping: the gap has to be strictly positive. This is the single most consequential decision on the page. The daily sawtooth charts this category is full of are plotting model temperature, not visibility, and a dashboard that cries wolf on noise is one nobody reads by the time something real happens.

Cadence, trickling, and why a number is dated to a window

Prompts are re-asked on a 10-day cycle, or weekly if you have upgraded cadence. There is no daily option, deliberately: it costs roughly ten times as much and re-samples the same noise. A scan does not fire every call at once. Each prompt is assigned a slot spread across the first 70% of the window, with the remainder left as slack for retries, and the slot is re-drawn every scan so a prompt cannot get pinned to one weekday and bake that bias into its series. The consequence for reading the chart: a point is a scan window, not a day. It is labelled “the 10 days ending September 2” because that is what it is. Bucketing the same data by calendar day would draw a sawtooth reflecting which days a trickle happened to touch, and it would read as violent week-to-week swings that are entirely scheduling artefacts.

Why the score is gated by plan

Two gates run on every scored widget, and the page always tells you which one applied.
  • Sample size. The trended score needs 100 scored runs in the window; per-prompt counts and the competitor leaderboard need 3.
  • Plan. Your plan sets how large a sample you can accumulate, because every run is a paid call to an answer engine.
The gating is on statistical validity, not on willingness to pay for a number. At one prompt and three repetitions a percentage is not meaningful, and no price would make it so. What Starter gets instead is the artefact that is genuinely useful at that sample: the verbatim answers. If your plan covers the trend but too few of your prompts have been scanned yet, the trend stays hidden and the page states your actual scored-run count and your actual band width, rather than a fixed sentence with a number baked into it.
Your prompt allowance is per account, across every site, not per site. A site showing two prompts can still be at an account-wide limit.

Evidence tiers

Three questions, two of which we can answer with evidence rather than a model. Tier 1: observed. Cited URLs joined to real answer_fetch crawl volume: pages an AI company actually fetched while answering somebody’s live question, from AI crawlers. Unsampled, and not a poll. It is the closest observable proxy for a real prompt that exists. Tier 2: deterministic join. Traffic and revenue on your cited pages, from sessions that landed there. Measured sessions and measured orders, joined on the page URL. No attribution model, no allocation, no guesswork. There is no third tier, and no modelled number anywhere in this feature. Traceten does not publish a revenue-per-prompt figure, and does not allocate revenue across the prompts that cited a page. Doing that requires a model, and a modelled figure that is not labelled as one is the failure mode this repository has already paid for more than once. Every number in AI visibility is either an observation or a deterministic join over observations.
Tier 1 is matched on the page path only, because the crawl rollup stores no hostname. If you serve several hostnames, two pages sharing a path show the same count. Tier 2 is landing-page grain: it counts sessions that entered on the cited page, not sessions that reached it later.

What each number cannot tell you

  • It cannot tell you what people actually ask. It tells you about the questions you wrote.
  • It cannot tell you what the consumer apps say. Two of the four engines, Claude and ChatGPT, are proxies for a consumer product, measured through each vendor’s API. See AI answer engines.
  • It cannot separate a low presence rate from a low answer rate. Those are reported separately for exactly that reason: “the model does not know us” and “Google renders no AI Overview for this question at all” need different responses.
  • It cannot compare a replaced prompt to its predecessor. Changing a prompt’s wording archives it and starts a new one, because editing text in place would retroactively change what a year of stored runs claims to have asked.

Next