Skip to main content

What this lets you do

Know exactly which surface produced each number, how close that surface is to what a person would see, and what leaves our infrastructure when a scan runs.

The four engines

The dashboard names each engine by the product you know. What sits behind each name is what this table says: Claude and ChatGPT are that vendor’s API with web search, not the consumer app, so read their numbers as described under “The two proxies” below. The visibility API returns each stored answer’s full engine label and model, so every number stays traceable to the surface that produced it. Gemini is deliberately not on this list. Its grounding API requires displaying Google’s Search Suggestions and permits Google to retain prompts. Google AI Overviews replaces it as the Google surface and is strictly better evidence: a real results page rather than an API, rendering reference links, so mentions and citations both come through.

Fidelity, stated honestly

This is the part every tool in this category is quiet about, so it is worth being exact.

The two high-fidelity surfaces

Google AI Overviews is a scrape of a real Google search results page. We ask DataForSEO to run your prompt as a search in the locale you configured and to load the AI Overview Google renders asynchronously. What comes back is Google’s own ai_overview block: its answer text, in Google’s own segmentation, and its own reference list, in Google’s own order. That is a page a person in that locale would have seen. It is the highest-fidelity surface on this list, and it is also the cheapest, which is why it ships on every plan including the free one. Google does not render an overview for every query. When it does not, that is recorded as a real outcome and reported as one, not as a gap. Perplexity is measured through Sonar, the retrieval stack behind Perplexity’s own product, reached through its API. Closer to the consumer experience than the two below, because Perplexity’s product is a search product and Sonar is what does the searching.

The two proxies

Claude and ChatGPT are measured through each vendor’s API, not through the app. The apps are not the same thing:
  • The apps add memory of your previous conversations and account-level personalisation before they retrieve anything.
  • The apps route between models and reasoning modes by their own rules; we pin one model and record which.
  • The apps have their own search behaviour and their own limits; we set an explicit cap.
So a Claude or ChatGPT number here can differ from what a person sees in the app, and they demonstrably do diverge. Read them as “the model with web search, asked cold, with no memory of anyone”, which is a real and useful measurement, and not as “what ChatGPT tells your customers”. Anyone publishing a plain “ChatGPT visibility” percentage is reporting an API and calling it the product.
Every stored run records the exact model_id, the full tool configuration, the locale and the timestamp that produced it, so any number can be traced to the request that made it. You can read all four on any answer in the drill-down.

What we send, per engine

A scan sends your prompt text and a location signal derived from the locale you set on that prompt: an approximate country for the three chat APIs, a location and language code for the search query. Nothing else about you travels: no account identifier, no site identifier, no visitor data, no cookie, nothing derived from anyone who visited your website. The location comes from the prompt’s configured locale, never from a person. We identify ourselves to every provider as Traceten-Visibility/1.0 (+https://traceten.com). DataForSEO does not receive a chat prompt. It receives a search query and runs it against Google. That is a materially different payload from the three chat requests, and it is why the four are described separately rather than as one flow.

Two provider-specific facts worth stating

Your prompt reaches Google. DataForSEO runs the search on google.com from its own infrastructure, so the query text reaches Google as an ordinary web search and is subject to Google’s ordinary handling of search queries. It is not sent to Gemini and not sent through any grounding API, and none of the retention terms attached to those products applies to this path. We opt out of OpenAI’s default storage, and there is no equivalent to set anywhere else. OpenAI’s API retains request and response server-side by default. We send store: false on every call, which under OpenAI’s API contract opts that call out. That is what we ask for and what we can prove from our own code; what OpenAI does internally is governed by their terms, not by ours. The flag is recorded in the stored tool configuration, so any run proves which choice was made for it. The asymmetry is worth knowing before you buy coverage. Anthropic’s API has no store setting, so there is no retention flag for us to turn off on that path. Neither do Perplexity or DataForSEO. What each of those three retains, and what Google retains for the search DataForSEO runs, is governed by their own terms rather than by anything we send. OpenAI is the only one of the four where we can express the preference in the request at all.
A prompt is customer-authored free text. It is capped at 2,000 characters and otherwise sent as written. Do not put personal data in a prompt. Anything you type there is transmitted to the providers above, and neither our retention limits nor our deletion tooling reaches inside their systems.

What we store from the answer

Per run, we store the answer itself and the engine’s citation list. Both are unredacted third-party prose: text a model or a search engine wrote, which we did not author and do not filter, and which can name people. Citation URLs have their query string and fragment removed before storage, so a link a model returns carrying a token or an address in its query does not reach our analytics database.
The daily citation rollup is the one visibility record with no expiry that is not counts-only. It keeps each cited URL, its path and its domain for as long as your account exists, because that is what a year-over-year citation trend is made of. Those are third-party pages an answer engine chose to cite, not pages anyone visited, but a URL can identify a person on its own. Deleting the site is the only thing that clears it.
None of these records carries a visitor identifier, a session identifier or an IP hash. They describe what a model said when we asked it, not what anybody did on your site.
The 90-day limit on stored answers is what bounds the drill-down. Ask for a date range that reaches past it and the page tells you so and names the shorter range it actually searched, rather than presenting partial evidence as the whole story.

Deletion

Deleting a site, or closing your account, erases every visibility record for the affected sites: answers, citation lists, mentions, and the daily rollups. Both take effect 30 days after the request, and the erasure is confirmed complete rather than merely started. The daily rollups have no expiry of their own, so deletion is the only thing that clears them. Per-visitor deletion does not reach these records, and cannot: they hold no visitor or session identifier to match on. See Data deletion.

Adding or removing engine coverage

Coverage applies from your next scan window. A scan records the engines it will use when its window opens, so the figures inside one window stay comparable with each other. Removing coverage works the other way: Claude and ChatGPT stop being called as soon as the change takes effect, without waiting for the window to end. A call already in flight finishes. That window is then labelled as a scan whose coverage ended part-way through, rather than pretending the engines ran.

Providers as sub-processors

All four providers are disclosed as sub-processors. See traceten.com/legal/sub-processors for each one’s role, location and data terms, and the Privacy Policy for the full description of the flow.

Next