Nobody can tell you, and the vendors selling an answer cannot either. Anthropic documents three crawlers and publishes an address file...the documentation stops there, and the complete user-agent string that Anthropic's three competitors all publish sits outside it, alongside citation data and ranking guidance. The measurement side has a matching gap. Published brand visibility studies test GPT, Gemini and Perplexity, and they leave Claude out.

Three parts of that record are solid enough to act on today.

  • Crawlers: Three separate tokens, each carrying a stated cost for blocking it.
  • Retrieval: A per-conversation decision, taken from the prompt and steerable by the user.
  • Measurement: Thin, because the published studies queried other models instead.

What Anthropic Actually Documents

Three crawlers do three jobs, and Anthropic states the consequence of blocking each one on a help page dated 7 April 2026. They are ClaudeBot, Claude-User and Claude-SearchBot.

  • ClaudeBot: Collects web content that "could potentially contribute to their training".
  • Claude-User: Fetches pages "when individuals ask questions to Claude".
  • Claude-SearchBot: "Navigates the web to improve search result quality for users".

Blocking the third one is the expensive mistake. Each block carries a stated cost, in the vendor's own words.

  • ClaudeBot blocked: Future site materials get excluded from Anthropic's model training datasets.
  • Claude-User blocked: Content stops being retrieved "in response to a user query".
  • Claude-SearchBot blocked: Indexing stops, and "visibility and accuracy in user search results" may fall.

The copied blocklists in circulation name GPTBot and ClaudeBot and leave the two search crawlers out...on 3 March 2025 Anthropic's help page still described one crawler, three appeared by the 1 April capture, and a file written on the older guidance names ClaudeBot alone.

Three more details sit on the same page, and the advice circulating elsewhere skips all three.

  • Crawl-delay: Supported as an extension, with ClaudeBot given as the worked example.
  • Subdomains: Robots.txt has to be modified, and every single one of them counts.
  • Address blocks: Discouraged, because they "may not work correctly or persistently guarantee an opt-out".

One asymmetry is worth carrying into any comparison with ChatGPT. On the fetch a human asked for, three vendors publish three different policies.

  • OpenAI: Robots.txt rules "may not apply" to the ChatGPT-User agent.
  • Perplexity: Its user fetcher "generally ignores robots.txt rules", per its own documentation.
  • Anthropic: Claude-User is listed as blockable...no carve-out appears anywhere on the page.

Missing User Agent Strings

Anthropic publishes the robots.txt token by itself. OpenAI, Google and Perplexity all publish the complete string a log file would show, down to the version number and the bot's own documentation URL...anyone writing a rule to count ClaudeBot hits in an access log is working from somebody's reverse engineering.

  • Published: Three robots.txt tokens, plus an address file at claude.com/crawling/bots.json.
  • Also published: A creation time of 1 May 2026 and prefixes including 216.73.216.0/22.
  • Not published: The full user-agent string, which is what a server log rule keys on.
  • Reversed since 2024: That older page said the address ranges stayed unpublished.

An Expired CCBot Block

A robots.txt written on Anthropic's older guidance has expired in place...until 2025 that page committed the crawler to honoring "any disallows for Common Crawl's CCBot User Agent". Wayback captures show the clause present on 3 March 2025 and gone by 1 April 2025.

So the audit is one grep...if your file disallows CCBot and never names ClaudeBot, the training crawler has been uncovered since at least 1 April 2025.

Whether Claude Looked At All

If you asked both assistants yourself, you were probably comparing one answer that used the live web against one answered from the weights. Anthropic's tool documentation puts it plainly, stating that Claude "determines when to search based on the prompt"...search is a decision taken per conversation, and the user can steer it.

Consumer documentation says how, in a passage written for users rather than for site owners. Include "Search the web" or "Use web search" in the prompt and the behaviour is forced...on Team and Enterprise plans that works once an Owner or Primary Owner has enabled web search for the whole workspace.

So the honest test has four cells rather than two, and every cell runs inside one session with one question.

  1. Forced search: Type the search instruction into Claude, then log what comes back.
  2. Search off: Toggle it off and ask Claude the identical question again.
  3. Both ways: Repeat the whole pair on ChatGPT, same session, same order.
  4. Comparison: Search-off gaps are training, search-on gaps are retrieval.

Claude's Unnamed Search Backend

No Anthropic page names the index behind Claude's web search...the evidence people cite instead is entirely circumstantial.

  • Subprocessor entry: Brave Search appeared on Anthropic's Trust Center subprocessor list on 19 March 2025.
  • Parameter name: Claude's web search tool definition contains a property called BraveSearchParams.
  • Independent report: Simon Willison, a writer rather than a vendor, published both findings on 21 March 2025.
  • Vendor silence: Anthropic's own web search documentation names no provider, index or subprocessor.

A compliance table entry and a parameter name are meaningful evidence...evidence of that kind stops short of a vendor statement, and no Anthropic page has ever put the Brave claim in prose.

Most of the confidence in the Brave claim rests on one number. Profound published it on 21 March 2025, one day after Claude web search launched, describing a "statistically significant overlap of 86.7% (13 out of 15 total results)". That is thirteen results out of fifteen, and the post never says how many queries produced them. Profound sells AI visibility tracking, and the figure has since been repeated across dozens of agency blogs as an established property of Claude.

There is a second trap in that figure. Cloudflare separately reports Anthropic's crawl-to-refer ratio falling 86.7% between January and July 2025...two unrelated 86.7% figures in one small category is how citations get contaminated.

Blocking Claude-SearchBot, per that same help page, prevents "indexing your content", which implies an index of Anthropic's own...a Brave backend implies the opposite. Anthropic leaves the relationship unexplained on that page.

Thin Measurement Record On Claude

Published brand visibility research mostly skips Claude, which leaves the confident per-model characterisations circulating about it unsupported.

  • Three preprints: One author, Dmitrij Żatuchin, affiliated with a generative engine optimization vendor.
  • Coverage: All three queried other models, so Claude sits outside their findings.
  • Similarweb: Its most cited domains work covers ChatGPT browsing mode and Google AI Mode.
  • First-party data: Only Microsoft ships any of it, inside Bing Webmaster Tools.

One academic paper in the set did include it. arXiv:2601.00869 ran 1,909 English queries about 30 brands across six models, and its headline is a split between Chinese and international models, 88.9% brand mention rate against 58.3%...the comparison measures model origin, not Claude.

Every Claude visibility figure in circulation is somebody's third-party sampling of a system whose vendor publishes crawler documentation and an address file.

Two Direct Comparisons Exist

Conductor ran the most detailed ChatGPT against Claude comparison with a disclosed methodology, 14,000 API calls across ten industries, seven intent types, four engines and fifty runs per cell. Every run was stateless, and consistency was scored as pairwise Jaccard similarity over 1,225 pairs per group.

  • ChatGPT: Returns an average of five brands per response, the fewest of the four engines tested.
  • Claude: Returns an average of 8.5, second only to Gemini at 9.2.
  • Consistency order: Perplexity, then ChatGPT, then Claude, then Gemini, putting Claude third of four.
  • The seller: Conductor sells AI visibility tracking...its study recommends buying intent-segmented tracking.

A second study put the same two assistants side by side. SparkToro ran 12 prompts across ChatGPT, Claude and Google's AI over November and December 2025, disclosed that its research partner works at an AI tracking startup, and published the raw runs on a public dashboard, a step Conductor skipped. Its own two counts disagree...the post says 600 volunteers and 2,961 runs, and the dashboard reports 373 users and 2,959 answers.

Different Documents, Different Answers

Claude and ChatGPT disagree partly because they read different pages. One workshop paper accepted at EDBT/ICDT 2026 measured the overlap between each engine's cited domains and Google's own top ten. Its sample was 1,000 ranking-style queries across ten consumer topics, phrased the way a shopper phrases them, such as best reviewed airlines this season.

  • GPT-4o: Median domain overlap with Google's own top ten of 0.0%, against a mean of 4.0%.
  • Claude 4.5 Sonnet: Median overlap of 8.7%, against a mean of 12.6%.
  • Missing entities: On average 16% of ranked names appeared in no retrieved snippet at all.
  • Model versions: Opus 5 was not tested...take the direction and leave the percentages.

Company size changes the behaviour, and the paper splits on it. For popular entities the pre-training prior dominates, and the authors describe the retrieval context as "confirmation, not discovery". Niche entities behave the other way, with the ranking following the literal content of whatever snippets came back...small businesses live in that regime, where retrieval decides nearly everything. For a niche-entity example, the paper uses a query for the top ten family law firms in Toronto.

Numbers That Contradict Themselves

Cloudflare measures crawler traffic on its own network...its published figures for ClaudeBot disagree with each other, repeatedly, across posts under its own name.

  • AI-only bots: ClaudeBot at 15% of the total in July 2024, rising to 23.3% a year later.
  • All crawlers: ClaudeBot at 11.7% in May 2024, falling to 5.4% by May 2025.
  • Crawl-to-refer: 286,930:1 in January 2025, reported as 38,065:1 by that July.
  • Same ratio elsewhere: 70,900:1 for the week of 19 to 26 June 2025, on Radar rather than in the crawl-to-click post.
  • A different quantity: Cloudflare's 1 July 2025 post calls Anthropic "30,000 times more difficult" for traffic than the Google of ten years ago.

Those last two measure different things...one is a crawl-to-refer ratio for a named week, the other a traffic-difficulty claim against a ten-year-old baseline, published with no methodology at all. Cloudflare's posts leave the denominators unreconciled.

Cloudflare publishes the caveat that undermines the lot...traffic referred by Claude's native app carries no Referer header, so the referral side of the ratio misses it, and Cloudflare says the calculations "may overstate the respective ratios, but it is unclear by how much".

Moves With Nothing Behind Them

Several recommendations circulate widely in this category on no published evidence.

  • llms.txt: No vendor documents reading it, and Google's John Mueller called it "purely speculative for now".
  • Claude-Web and anthropic-ai: Anthropic has never listed either, so a line naming one and/or the other is dead weight, and calling them deprecated invents a retirement that never happened.
  • Structured content gains: The "up to 40% improved visibility" figure gets quoted without the scope its own paper puts on it.

That 40% has a source, and the source is narrower than the quoting...the number comes from GEO: Generative Engine Optimization, a peer-reviewed KDD 2024 paper, where the figure is a position-adjusted word count score moving from 19.3 to 27.2 for a document the researchers had already placed in the model's context. Traffic, clicks and whether an assistant finds you at all sit outside what that score measures, and the restatements circulate it without the condition.

Google states there is no special schema.org structured data you need to add for its own AI features, and that needs scoping honestly...the statement covers Google, and Claude falls outside what it measures.

What To Check First

Order matters here. This checklist runs from what Anthropic documents down to what follows from it.

  1. Read robots.txt for all three Anthropic tokens by name, on every subdomain.
  2. Search for CCBot, because that line stopped covering Anthropic sometime between 3 March and 1 April 2025.
  3. Run the four cells, search forced and search off, on both assistants.
  4. Check Cloudflare settings, since blocking Training also blocks Googlebot from 15 September 2026.
  5. Log it monthly...a single answer is one draw from a distribution.

That fourth item carries a deadline. Its rule states that from 15 September 2026, multi-purpose crawlers "such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training", under most-restrictive-rule enforcement. Anyone who once clicked the old block-AI-bots preset is on that path...the opt-out lives in Security settings, and Cloudflare says it can be done at any time before that date.

This is slow work...maintenance rather than a growth tactic. What the record supports is narrow. Let the search crawler in. Know which half of an answer came from retrieval, and treat every published Claude percentage as somebody's sample until the queries are shown. Everyone else fills the gap with numbers, and vizmeta HQ would rather log the answers month by month and let the gap stay a gap.