Skip to main content

What this shows

The AI Crawlers page answers four questions about the AI systems fetching your pages:
  • Who crawled your site, as a company and a specific crawler
  • What pages they read
  • Why: to answer a live question, to build a search index, or to collect training data
  • Whether it was really them
This is a different signal from Sources. Sources measures the person who arrives from an AI assistant or another site. AI crawlers measures the crawl that put you in that assistant’s answer in the first place.

Setup is separate from the snippet

AI crawlers do not run JavaScript. The tracking snippet never sees them, because they issue one HTTP request for your HTML and leave. Crawl tracking therefore needs a small server-side middleware in addition to the snippet. Create a token under Settings → API keys → AI crawler API key, then follow the install guide. Crawl tokens are per site and are separate from your server SDK key, which is account-scoped. The two live on their own tabs of the same settings page. While crawl reporting is not configured for the selected site (no active crawl token, or a token that has never sent a report), the page shows a short note at the top that links to that settings tab. Once reporting works, the note goes away.
The token is shown once, when you create it. Traceten stores a one-way hash and cannot show or recover it afterwards. If you lose it, revoke it and create a new one.

The four purposes

Every crawler is categorized by what it is doing. These have very different commercial meaning, so they are never collapsed into one “bot traffic” number.

AI crawler activity

The chart at the top of the page has one tab per purpose (AI answers, Indexing, Training, Other crawlers) plus a fifth, Spoofed. Each tab carries its own crawl count for the selected date range. Next to the chart title, the page shows how many distinct crawlers and companies were seen, for example “60 crawlers from 21 companies”. Under a company filter it counts that company’s crawlers, and under a page filter the companies that read the page. Purpose is a mode, not a stack layer: switching tabs re-runs the query rather than re-slicing what is already on screen, so the lines under the Training tab are the top lines for training, not the overall top lines with the other purposes hidden.
“Other AI crawler” is a full tab, not a footnote. Some pages are read only by vendors whose crawler description is too vague to classify; folding that bucket into a remainder line would make those pages appear under no tab at all while still counting toward your total.
Spoofed is a fifth tab but not a fifth purpose. It cuts across all four: a request caught impersonating a crawler still claimed some purpose, and it keeps the purpose it claimed. So the Spoofed tab counts impersonation attempts across every purpose, and none of the four purpose counts includes any of them. The tab counts stay fixed as you switch between tabs, because each one states how much data that tab holds, not how much the selected tab holds.

Which companies are reading you

By default the chart shows one line per company (Anthropic, OpenAI, Google, Amazon) over time, each in that company’s own colour. The same chart appears on the Overview’s crawler tab, so the two surfaces cannot draw the same window differently. Hovering the chart gives you a row for every line on it (each company with its logo and its figure for that day, dimmed where it is zero) and a Total underneath, so you can see the sum without adding the rows yourself. The companies are named in a legend centred under the chart, each with its logo and its crawl count for the window. Hovering or focusing an entry lights that company’s line and fades every other one, so a single company can be read out of eight without hiding anything. Its hover text adds the impersonation attempts made in its name and how many of your pages it read. Selecting an entry drills into it.
Crawls and impersonation attempts are never added together. A spoofed request is recorded under the company it pretended to be, so a combined figure would credit that company with traffic it never sent, beside its own logo. The two are always separate rows.
Each purpose tab also states its change against the previous period: the immediately preceding window of the same length, so a 7-day view compares against the 7 days before it, not against “last week” as a calendar. A tab shows no change at all when the previous period returned nothing, or returned zero, because “we could not compare” and “it did not change” are different facts, and only one of them can be stated as 0%.

Group by crawler, or by page

A Group by switch above the chart flips the lines between Crawler (one line per company) and Pages (one line per path). It drills in both directions:
  • Select a page and the chart shows the companies that read it.
  • Select a company and the chart shows the pages it read.
  • Select one of each and the chart narrows to that intersection: that company’s crawls of that page.
Both selections are sent to the server together and the ranking is recomputed inside them, so the lines you get are genuinely the top lines for that intersection. They are not the site-wide top lines with the rest hidden, which would put correct numbers on the wrong rows. Three things about the company list:
  • The label is always the company, never the product. GPTBot is an OpenAI crawler, so its rows read OpenAI, not “ChatGPT”. They are different things: ChatGPT is the assistant, GPTBot is the crawler that collects training data. Labelling one as the other would misreport both who crawled and why.
  • Spoofed requests are never added to a company’s number. A request that impersonates a crawler is recorded under the crawler it claimed to be, so counting them together would credit a company with traffic it never sent. Impersonation attempts appear as a separate red count on that company’s row and are excluded from its line on the chart.
  • The top eight are named, and everything below that is rolled up into a single grey “Other companies” line (or “Other pages”, when grouped by page) so the totals still add up. That roll-up has no identity, is never given a logo, and cannot be selected, because it stands for many companies at once.

Verified

The Verified card shows the share of crawls backed by network proof, a bar split by verification level, then one row per level with its count and share, strongest first. Verified, Not verifiable and Spoofed always have a row, so a level with no crawls reads 0. Hover or focus a level’s name for what it means. Impersonation attempts appear here as the Spoofed level, in red. Not verifiable is grey, never a warning. The share keeps spoofed crawls in the denominator, because a disproved claim is still a crawl that happened. With no crawls in the range, the share is left blank rather than shown as 0% or 100%. Under a company filter, the card includes the crawls caught impersonating that company. Under a page and a company filter together, impersonations are excluded, and the card says so.

Purposes, files and companies

Beside Verified is a card with three tabs. Each draws a donut with its slices named outside the ring. There is no legend under the donut: hover or tap a slice for its name, count and share, or tab into the donut with a keyboard to step through every slice.
  • Purposes draws the four purposes, with the total crawls in the centre. Up to three company logos sit inside each slice. Select a purpose and the donut becomes the companies whose crawlers read your site for that purpose, each with its logo inside its slice and its crawl count. Purposes above the donut returns to all four.
  • Files draws the files crawlers read first (robots.txt, llms.txt, llms-full.txt, sitemaps and feeds), with each file’s fetch count. Hovering a file also shows how many distinct crawlers fetched it. A crawler fetches these to decide how to treat the rest of your site.
  • Companies draws every company that crawled your site, each with its logo inside its slice. Select a company and the donut becomes every one of that company’s crawlers, with no slice rolled into “Other”. Companies above the donut returns to all of them.
All three views exclude impersonation attempts, so they match the chart’s tab counts. If crawls carry a purpose this version of the dashboard does not recognise, a line under the donut says how many. Under a company or crawler filter, Files narrows too: it counts only the chosen companies’ and crawlers’ fetches of those files. Without a filter, when more files were fetched than the donut names, the rest are one Other files slice, so the slices always add up to the site’s total.

Filtering

The Company, Crawler and Page pickers sit with the crawler list, beside its search and purpose controls. Choosing a value filters the whole page to it: the chart and its tab counts, the Verified card with its spoofed share, the purposes, files and companies donuts, and the crawler list all narrow together. In an open picker each chosen value is ticked where it is listed; select it again to remove it. Each applied value also shows as a chip above the chart with its own remove button, and Clear all appears once more than one is applied. Removing every chip returns everything to the site-wide view. Several values in one filter match any of them. Two companies show crawls from either company, and two pages are added together. A filter takes up to 20 values. Different filters apply together. Traceten stores which company read which page on which day, and which crawler read it, so a page filter and a company filter narrow to that company’s crawls of that page. A combination that cannot happen, such as the company Anthropic with the crawler GPTBot (an OpenAI crawler), shows zero everywhere rather than either filter’s figures on their own. A crawler filter narrows to named crawlers such as GPTBot, on every part of the page. Grouped by company with no page filter, the chart counts every crawl the crawler made. Wherever the view involves pages (a page filter, grouping by page, or the Files donut), it reads per-page crawler names. If only part of the crawls in range carry a name, a line says how many and since when. Where crawler names cannot be read at all, the chart, the donuts, the Verified card and the list each say so, rather than showing figures the crawler filter did not narrow. Under a page filter:
  • Files shows that page’s own fetches when the page is a crawler file, and says the page is not a crawler file otherwise. With several pages chosen it asks you to choose one, because their fetches are added together.
  • Companies shows the companies that read the page. Select a company to see its crawlers on that page. If crawler names cover only part of the page’s crawls in the date range, a line under the donut says how many. When no crawl of the page carries a name, the companies cannot be opened and the line says so.
  • If a page was read by more combinations of company, purpose and verification level than one request returns, a line says some crawls are not broken out, rather than showing a smaller total as if it were complete.
A company filter also narrows the crawler list to that company’s crawlers.
Traceten did not always record crawler names on its per-page data. Crawls of a page from before names were recorded, and from any period your site was over its crawl allowance, keep only their company and purpose. When a date range reaches into that data, the page says how many of the page’s crawls carry a crawler name and since when, rather than showing a partial split as complete.

Verification levels

Every crawl records how strongly its identity could be checked. There are six levels, strongest first. In the dashboard each level’s meaning is attached to the level’s own name: the badge is a button, so hovering or focusing it with the keyboard explains that rung wherever it appears. There is no separate legend to go and find: a definition that lives in a different block from the term it defines is a lookup you have to decide to make, and by the time you reach it you have already guessed.

”Verified · network” is weaker than “Verified · IP range”

An ASN match proves the traffic came from that company’s infrastructure. It does not prove it came from that specific crawler, which is why it sits below an IP-range match. Meta is the main case: Meta publishes no crawler IP file, so the only available evidence is which network the request originated in.

”Unverified” is normal, not a problem

About half of all tracked crawlers publish no IP ranges and document no reverse-DNS convention, including every xAI crawler, every Chinese provider, Cohere, Allen AI, and You.com. For those, no analytics tool can verify the claim, including this one. A large unverified share is the truthful state of the ecosystem rather than a fault in your setup or your data. Traceten reports it plainly instead of presenting a guess as a fact.

”Spoofed” is only assertable sometimes

Spoofed is reached by two routes, and only one of them needs published ranges. The first is an IP mismatch: the request named a crawler whose vendor publishes IP ranges, and came from outside them. This route can only ever fire for crawlers that publish ranges in the first place. The second needs no IP check at all. Google-Extended and Applebot-Extended are robots.txt permission tokens, not crawlers. No crawler ever sends them as a user agent, so a live request carrying one cannot be legitimate whatever its IP. See robots.txt and llms.txt. So no spoofing detected does not mean every crawler was checked and passed. For the crawlers that publish nothing and send no permission token, neither route applies and nothing can be confirmed either way.

Crawlers

The crawler list shows every crawler seen in the date range, one row per crawler and purpose, in aligned columns: Company, Crawler, Purpose, First seen, Last seen and Crawls. On a phone the columns stack under the crawler’s name, each with its label. The list shows 50 rows at a time. Search by crawler or company name, or narrow the list by purpose. The company, crawler and page pickers beside them narrow the list along with the rest of the screen. A row’s crawl count excludes requests caught impersonating that crawler. When there were any, the row names how many under the count. First and last seen are the first and last day in the range with a genuine crawl, so impersonation never moves them.

Filtering the list by page

The page picker, beside the company and crawler pickers above the list, lists every page crawled in the date range, busiest first, with its crawl count. Type to search: the search runs on the server, so it finds any page however many your site has. Select pages to filter the whole screen to them, the same as selecting a page anywhere else. A chosen page is ticked; select it again to remove it. Under a page filter the list shows one row per crawler and purpose that read the page, in the same columns, with its crawl count and first and last seen for that page. Select a row to open the crawler’s detail card. The card covers your whole site, and says so while a page filter is active. A request that claimed a crawler Traceten could not name exactly is listed as Unnamed crawler and does not open a card. If only part of the page’s crawls in the date range carry a crawler name, a line under the list says how many. When none do, the list shows one row per company and purpose instead, and a line says why; with a crawler filter applied it shows No crawler names for this page instead, because those rows could not follow the filter. If the page’s list was too long to break out in full, a line under the list says so.

Crawler detail

Select a crawler to open its detail card. It opens in the middle of the screen on a laptop and rises from the bottom on a phone, laid out the same way as a visitor’s card. The card fetches its data when it opens.
  • Header: the company’s logo, the crawler’s name, its purpose, its verification status, and how many requests impersonated it.
  • Summary: crawls, days active, verified share, first and last seen, and spoofed requests for the date range. Hover or tap the verified share for what it counts: the share of crawls identifying as that crawler, spoofed requests included, that carried network evidence. A figure Traceten cannot compute, such as a verified share with no crawls, is left out rather than shown as zero.
  • Activity: a day grid of the last six months up to the end of the date range, darker green for busier days, the same grid as a visitor’s card. Every other figure in the card covers the date range only.
  • Verification: how many crawls reached each verification level, including any spoofed requests.
  • Responses: a badge per status class (2xx, 3xx, 4xx, 5xx) with its count, and the error rate over the responses Traceten observed. Responses that could not be observed get their own Not observed badge.
  • Files read first: robots.txt, llms.txt, sitemaps and feeds, each with an icon for its kind and its crawl count.
  • Pages crawled: each page listed once with its verification status, busiest first, 50 at a time. The Crawler column is this crawler’s own crawls of the page, and the Company column counts every crawl of that page by the same company for the same purpose. Hover or tap either number for a sentence saying exactly what it counts, for example “12 crawls by GPTBot, out of 30 crawls of this page by all OpenAI crawlers for training data.” On a phone the row shows the crawler’s count, and the sentence carries both.

Verification status

The header and every page carry one status, judged over genuine crawls only. Spoofed requests never lower or raise it; they get their own red badge. A verified crawler can still have a few crawls that could not be checked, for example when your middleware did not report an IP for a request. Hover or focus the header badge to see the exact share.

Page actions

Select a page’s link to copy its full address on your site’s domain. Select anywhere else on the row to filter the whole screen to that page: the card closes and the page filter is applied. Both actions are reachable with the keyboard.
Files, and which pages are listed, are recorded per company and purpose. When a company runs more than one crawler for the same purpose (Google has several that fetch pages to answer questions), the card names the others: its file counts include them, and so does each page’s Company count. A page’s Crawler count is this crawler’s alone. Where some of a page’s crawls come from before crawler names were recorded, the card says its counts there may be low.
Responses are counted per crawler, not per page: Traceten records which crawler got an error, not which URL returned it, so search your server logs for the crawler name to find the paths. Responses that could not be observed are left out of the error rate and shown only in their Not observed badge. When no status was observed at all, the card says so instead of showing a rate.

Empty states

Zero crawls has three different causes, and the page tells them apart from your crawl token, not from the crawl data, because all three look identical in the data.
  • See which AI crawlers read your site: this site has no crawl token, so nothing can be reporting. Create one under Settings → API keys → AI crawler API key.
  • Waiting for your first crawl report: a token exists but has never been used to authenticate a report. The token was created and the middleware was never deployed, or it is deployed with a different site’s token.
  • No AI crawls in this range: your server has reported before, so crawl tracking is working and this window was simply quiet. Try a wider range.

Data quirks worth knowing

  • HTTP status is per crawler, not per page. The daily page rollup carries no status column, so Traceten can say “GPTBot got 41 errors” but not “which URL returned them”. Search your server logs for the crawler name.
  • A page can be read for several purposes. Filtering to a page shows every purpose it was read for, not only the one you selected it from.
  • Older page data names companies, not crawlers. Crawls of a page from before crawler names were recorded keep only their company and purpose. A date range reaching back that far splits only part of a page’s crawls by crawler, and the page says how many.
  • “No spoofing detected” is not the same as “everything checked out”. See above.

Data retention

Per-crawl detail is kept for 90 days. The daily rollups behind every chart and table on this page are kept indefinitely, so long-range trends stay available.