No honest number exists. What can be answered is which machine decided...four pipelines carry that job. Pre-training holds frozen weights, ChatGPT search runs live retrieval against the open web, licensed feeds arrive from partners who hand OpenAI a file directly, and workspace connectors hold private company data that stays inside the account that owns it. Vendors selling a single visibility score are guessing.

  • Pre-Training: Frozen weights with a published cutoff date, fixed at the moment training ended.
  • Live Retrieval: ChatGPT search, running OpenAI's own crawler plus named outside search partners.
  • Licensed Feeds: Partner data handed to OpenAI directly, indexed on arrival.
  • Workspace Connectors: Private company data, readable only inside the workspace that holds it.

Most advice on this collapses all four into one instruction...the tactic it then sells is aimed at the blur. Those four routes move on different timescales, and only one of them is documented well enough to act on.

What The Weights Hold

Fixing a fact in the weights would mean finding the row that holds it. OpenAI states that its models "do not store or retain copies of the data they are trained on"...a takedown request aimed at training data asks for exactly that kind of edit.

And timing closes the rest of it. OpenAI's model documentation lists a knowledge cutoff of February 16, 2026 for the current GPT-5.6 family...the weights stop there, and whatever changed about a business after mid-February sits outside them.

Composition has a flat answer as well, in three published pieces.

  • Published Mix: The 2020 GPT-3 paper, where filtered Common Crawl carried 60% of the weighting and Wikipedia carried 3%.
  • Full Stop: GPT-4's technical report "contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar".
  • Since Then: The public record ends there...every current-sounding percentage breakdown of ChatGPT's training data traces back to a six-year-old model or to a dead end.

OpenAI's One Published Requirement

OpenAI answers the inclusion question in a single public sentence...the whole of it follows. Ranking "is based on a number of factors designed to help users find reliable, relevant information", there is "no way to guarantee top placement", and to be included the help centre says "it is important to allow OAI-Searchbot to crawl your site" and to let the published IP ranges through your host and/or your CDN.

So read what is missing from it, no schema.org, no llms.txt, no structured data, no backlink advice, no word counts. Everything published amounts to one instruction...let the crawler in and hold the firewall open.

Four Crawlers, Four Jobs

Four bots appear in OpenAI's crawler documentation, and one sentence there does most of the work...blocking one of them leaves the other three exactly as they were, because each robots.txt setting "is independent of the others". All four tokens are listed together on the same documentation page.

  • OAI-SearchBot: The one that matters, "used to surface websites in search results in ChatGPT's search features", and the only lever OpenAI has ever published for getting a site included.
  • GPTBot: The training crawler, for content "that may be used in training" the foundation models.
  • ChatGPT-User: Fires when a person asks, and is "not used to determine whether content may appear in Search".
  • OAI-AdsBot: Visits only pages submitted as ad landing pages...most businesses can skip it.

Most sites left this to the default. Cloudflare analysed 3,816 domains on 6 June 2025 and found 546 of them, 14%, carrying an AI-specific robots.txt directive. News publishers are a different population, and the palewi.re tracker of 1,150 news homepages read 50.0% blocking OpenAI's crawler on 1 August 2026. It describes newsrooms only. Whichever token a site picked, blocking training leaves search alone. Sites opted out of OAI-SearchBot "will not be shown in ChatGPT search answers, though can still appear as navigational links", and OpenAI puts the lag on a robots.txt change at roughly 24 hours. Wait a day before retesting.

A deadline is attached to this one, and the date is already fixed. Cloudflare announced on 1 July 2026 that from 15 September 2026 multi-purpose crawlers "such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training", because the most restrictive matching rule wins. Anyone who ever clicked Block AI bots in Cloudflare has one setting to open...about six weeks remain.

That 87 Percent Bing Figure

OpenAI does name Bing...the fact is narrower than the industry makes of it. Its help article for ChatGPT search says the product "sometimes partners with other search providers", then lists exactly two, Bing and Shopify, each linked to a privacy policy. That sentence discloses where query text travels...the architecture behind the answer goes undescribed, and OpenAI's own launch announcement for ChatGPT search never uses the word Bing at all.

Then there is the 87 percent, the most repeated number in this subject. Seer Interactive published it on 6 February 2025 and it now circulates as a permanent property of ChatGPT. What it measured was SearchGPT, the prototype surface, across 100 queries and more than 500 citations. Seer called the sample limited and its own reading directional in the same post. Every repetition drops that caveat...the number survives intact.

Chasing it back through the vendor pages that repeat it turns up a different claim at nearly every hop.

  • Onely: "Top 20 Bing results have 87% higher citation chance", hyperlinked to Triple Dart and not to any study.
  • Triple Dart: That same higher-chance wording on one page, then 87% of citations matching Bing's top 10 on a second, both linking Seer.
  • Metricus: An overlap stated correctly as a match against Bing's top organic results, attributed to "Seer Interactive (2025)" with no link.
  • Seer's Post: The measurement itself, an overlap share, under a section heading naming Bing's top 20 rather than its top 10.

An overlap share and a higher chance of citation are two separate claims...only the first of them was ever measured. Between two pages on one site the cutoff moves from top 20 down to top 10. Triple Dart states it wrong and links the study...the page that states it correctly links nothing.

Later measurements of the same relationship disagree with it...they disagree with each other too.

  • Ahrefs, August 2025: 8.1% overlap with Bing's top 10, against 8.0% for Google's top 10, across 15,000 long-tail queries.
  • Profound, August 2025: 240M citations across 1,000 prompts, Bing overlap falling from 26% to 8% while Google alignment rose 12% to 33%.
  • AI Plus Automation, April 2026: 27.4% across 400 queries, matched at domain level, on a nano model and not the consumer one.
  • Seer Interactive, February 2025: 87%+ of SearchGPT citations matching Bing's top organic results, 100 queries, matched at URL level.

But a range needs four teams measuring one thing. Four teams measured four slightly different objects on four different dates...averaging them produces a number that describes nothing. Bing's public Search APIs were also retired on 11 August 2025, so whatever arrangement exists between the two companies now is private and cannot be audited from outside.

Microsoft's Own Silence

Microsoft's silence is the more interesting evidence. When it shipped its AI Performance report inside Bing Webmaster Tools on 10 February 2026, it described the coverage as Microsoft Copilot, AI-generated summaries in Bing, and "select partner integrations". Every commercial incentive pointed toward writing the word ChatGPT there...the report leaves it out.

One Named Local Partner

One local data source is under contract and named out loud. Yelp began licensing reviews, photos and business information to OpenAI in a deal Axios reported on 23 July 2026, and OpenAI's developer docs describe a beta programme in which approved partners submit a paginated feed of business records that ChatGPT then indexes for search. Foursquare claimed a similar arrangement in a December 2024 LinkedIn post...nobody at OpenAI has ever confirmed it.

Bing Places And Local Facts

The chain everyone draws for a local business runs through Bing Places...the last link in it is disputed. Microsoft states that Bing Places is one of the primary sources behind a Bing local listing, then warns that other sources may take priority when their information is fresher, giving as its own example a Bing Places record left untouched for more than 6 months. Check yours.

  • OpenAI Side: Bing is named as one of exactly two third-party search providers, hedged with "sometimes".
  • Microsoft Side: Bing Places is named as a primary source behind a listing on Bing's own surfaces.
  • Unverified: Neither company has ever said ChatGPT reads that structured listing corpus.

Three checkable things point the other way. Search Engine Land, the highest-authority page in this topic, concludes that ChatGPT "does not see the local profile data managed by businesses via the Bing Places for Business platform". Bing's own robots.txt carries Disallow: /local under User-agent: * and contains no block for OAI-SearchBot, so the wildcard governs OpenAI's crawler, and that disallow closes the local surface to crawling. BrightLocal's July 2025 pass across four AI tools never mentions Bing Places, and puts business websites at 58% of the sources ChatGPT used.

Two Spanish consultants have inspected what ChatGPT actually calls, working days apart in May 2025, and neither of them found Bing Places. Both report a JSON payload carrying a provider field on every business record, and the only two values either of them observed were foursquare and serp, the second being a scraper reading the embedded Google Maps panel. Those are single unreplicated tests on models that have since been replaced three times...the published record still holds no other direct observation.

Vendor pages assert the missing link anyway...not one of them can source it. Claim the Bing Places listing because it is free and it governs what Bing and Copilot say about you, and expect nothing from ChatGPT on the strength of it.

Microsoft publishes three more rules for that listing, spread across two help pages...the audit blogs skip all of them.

  • Google Import: Link the accounts and Bing will "use Google data to overwrite the Bing Places for Business listing information".
  • Publish Cycle: Each change is validated against other sources first, so an edit can silently fail.
  • Renames And Moves: Microsoft advises creating a new listing and closing the old one, not editing.

The Worst Instrument You Own

The logged-in account a business owner uses daily is the worst available place to check any of this. Memory runs on by default, and OpenAI's Memory FAQ concedes the memory summary "will not include everything that ChatGPT remembers based on your chats"...the contamination stays hidden from inside the product.

Three mechanics make it worse than it sounds.

  • Query Rewriting: Memory changes the search that runs, upstream of the reply itself.
  • Temporary Chat: OpenAI says it "will still follow your custom instructions if they're enabled".
  • Cleared History: Saved memories sit separately and survive the chats that created them.

Query rewriting is the mechanism underneath that. OpenAI's worked example turns a restaurant question from a San Francisco address into the search "top restaurants San Francisco", and a vegan user's stored preferences turn the same prompt into "good vegan restaurants San Francisco". Memory rewrites the search before it runs...a different query retrieves different pages, and a different set of businesses is eligible to be named.

A Clean Test Procedure

Log out. It removes more distortion than any tool subscription will...that logged-out state is one OpenAI documents as supported.

  1. Use a logged-out browser instead of a Temporary Chat, because custom instructions survive the second one.
  2. Ask what a buyer with a problem would ask, and keep the business name out of the prompt.
  3. Log three columns per answer, named or not, cited or not, and which domains got cited.
  4. Run the same prompts again on other days, because a single reading is an anecdote.
  5. Check the Free plan and logged-out view, since OpenAI began testing ads on those tiers on 9 February 2026.

OpenAI puts that test in the United States, on the Free and Go plans, with a logged-out ad experience designed alongside it. By OpenAI's account the ads leave the answer itself alone, and a Plus subscriber auditing a category reads a different page from the buyer.

Repetition is the part people skip, and a Thinking Machines Lab experiment sampled 1,000 completions from one prompt at temperature zero and got 80 different outputs, tracing the cause to how many other requests were batched alongside each one. That is a property of inference, and the model tested was Qwen3-235B on the researchers' own stack.

Fluency reads like evidence...the Tow Center at Columbia ran 1,600 queries across eight AI search tools and found ChatGPT "incorrectly identified 134 articles, but signaled a lack of confidence just fifteen times out of its two hundred responses, and never declined to provide an answer".

The Documented Gaps

The gaps here are load-bearing...every competing page in the category fills them with confident invention.

  • Ranking Factors: OpenAI publishes "a number of factors" and declines to be more specific.
  • Recrawl Cadence: No guidance exists on how fast a site change reaches ChatGPT search.
  • Structured Data: No OpenAI documentation found mentions schema.org, JSON-LD or llms.txt.
  • Effect Size: Nobody has published a test showing that allowing OAI-SearchBot changes citations.

So vendor dashboards rush into that vacuum...they disagree by a factor of five. Wikipedia's share of ChatGPT citations is reported at 8.9% by Ahrefs, 13.15% by Similarweb and 47.9% by PikaSEO, while Contently claims the most-cited domain on any platform rarely exceeds 5 percent. Those four cannot all be true, and the spread is a property of the measuring tools. None of them publishes raw data. Read the method behind each one.

A Monthly Log

A trend line is the deliverable. Vendors document the mechanisms above...the magnitudes stay undocumented, which is why this check belongs on a calendar. vizmeta HQ is this site's own tool, and it runs the check monthly across ChatGPT, Claude and Google AI, free, recording what the machines get wrong.

The rest is unglamorous and well documented.

  1. Let OAI-SearchBot through, and check that no CDN rule is quietly blocking it.
  2. Keep the business listings OpenAI licenses accurate, and start with Yelp.
  3. Open the Cloudflare AI setting before 15 September, or Googlebot goes with it.
  4. Stop testing on an account that has spent two years learning what you like.