Nearly every vendor in this category sells one instruction, add schema and AI systems will cite you, while Google's own documentation says in writing that no special schema is needed. Both sentences describe the same features and only one can be true...the vendor instruction carries no study, and the denial comes from the company running the features. Almost every write-up prints the first and skips the second.
What Google Actually Wrote Down
Google removes the requirement in two sentences on its AI features and your website page. "You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add." A separate AI optimization guide, last updated 10 July 2026, repeats the point, and adds Markdown to the list nobody needs.
Schema keeps a documented job all the same. Google recommends structured data "as part of your overall SEO strategy, as it helps with being eligible for rich results on Google Search"...rich results, not AI citation.
Microsoft Says the Opposite
Fabrice Canel, Principal Product Manager at Microsoft Bing, told SMX Munich that "schema markup helps Microsoft's LLMs understand your content." Barry Schwartz reported the remark in Search Engine Land on 20 March 2025. No schema types were named, and no data was published.
Two companies operating these systems point opposite ways, neither published data, and neither statement settles the other...everyone else is quieter still.
- OpenAI: four crawlers documented, robots.txt controls explained, zero mentions of markup.
- Anthropic: three crawlers documented, robots.txt honored, zero mentions of rendering or markup.
- Perplexity: two crawlers documented, one of which generally ignores robots.txt because the fetch is user-triggered.
- Google: the only vendor addressing schema and AI together, and it calls the markup optional.
Controlled Evidence
Ahrefs ran a two-phase study published 11 May 2026 by Louise Linehan and Xibeijia Guan, whose causal phase tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against roughly 4,000 matched controls, measured 30 days either side. Linehan and Guan concluded that "the most consistent finding is that not much really changed," with schema showing no clear effect in either direction...one surface moved enough to register at all.
- AI Overviews: minus 4.6 percent, a small decline that reached statistical significance.
- AI Mode: plus 2.4 percent, which is indistinguishable from noise.
- ChatGPT: plus 2.2 percent, which is also indistinguishable from noise.
All four statistical approaches agreed, and the only change big enough to register was a decline...three caveats travel with the finding.
- Citation Floor: every page already held 100 or more AI Overview citations, so pages starting from zero went untested.
- Confounded Changes: pages that add schema often change other things at the same time.
- Type Blindness: the study pooled every schema type together, so type-level effects went unmeasured.
The Statistic That Does Not Exist
Daniel Cheung's systematic review of 6 July 2026 ran five arXiv queries, screened 71 papers, added a blind second-reviewer re-screen, and refused to count any claim until its primary source was fetched. It found that the "44 percent increase in AI search citations" attributed everywhere to BrightEdge has no study behind it. BrightEdge did publish a real 44 percent figure...that one says Google's AI Overviews are 44 percent more likely to criticize brands than ChatGPT. Somebody lifted that number off one claim and attached it to another.
His summary is blunt. "Every study that found schema helps is missing the controls." Cheung sells SEO tooling himself and has his own stake in the answer...the arithmetic is still checkable.
Backlinko's analysis of 11.8 million search results, last updated 14 April 2025, found that 72.6 percent of first-page results use schema and that structured data had no relationship with rankings...only the first half of that sentence ever travels.
Rank as the Confound
Cheung's formulation is the one to carry...any schema study that does not control for rank is mostly measuring rank. Kevin Indig, working with AirOps, reported ChatGPT citing pages with JSON-LD 38.5 percent of the time against 32.0 percent without. That post is paywalled, so the figures arrive secondhand through Cheung's review.
- With JSON-LD: cited by ChatGPT 38.5 percent of the time, against 32.0 percent without.
- Position 1: cited 58.4 percent of the time in that same ChatGPT dataset.
- Position 10: cited 14.2 percent of the time in the same dataset.
The 6.5 point schema gap sits inside a 44 point rank gap. Kurt Fischman's cross-platform study of 12 February 2026 cuts the same problem differently. Fischman's pooled analysis found a significant negative association between schema presence and citation, an odds ratio of 0.546...he attributes that to Google's ranking already over-selecting schema-bearing pages into the top 10. He is explaining away his own headline result.
Segmented by domain authority, the picture changes. Attribute-rich Product and Review markup "may confer modest citation advantages for lower-authority domains" at Domain Rating 60 or below. Generic Organization and BreadcrumbList types provide "none". Attribute-rich implementations were cited at 61.7 percent against 41.6 percent for generic ones. Those figures come from 730 citations across GPT-4o and Gemini 1.5 Pro, not the 353,799-page AirOps dataset. Fischman sells AI operations services, the paper is a preprint, and the models it tested were older generations...the domain-authority split still belongs to this paper alone. One paper points both ways depending on how it cuts the data.
Vendor Numbers, Traced Backwards
WordLift's arXiv preprint 2603.10700, submitted 11 March 2026, gets quoted everywhere for a 29.6 percent accuracy gain. Nobody peer reviewed it. Its own abstract says "while JSON-LD markup alone provides only modest improvements"...the gain came from an enhanced entity page format carrying agent instructions, breadcrumbs and neural search. Anyone quoting 29.6 percent as evidence for JSON-LD is quoting the paper against itself.
Search Atlas manages to hold both positions at once. One Search Atlas marketing page claims FAQPage schema produces 3.2 times more AI Overview appearances. Search Engine Land cites a different Search Atlas study...that one found no correlation between schema coverage and citation rates.
Where llms.txt Actually Stands
Almost nobody reads the file. A single month of Ahrefs Web Analytics data covering 137,210 domains found 28 percent publishing a valid llms.txt...97 percent of those files drew zero requests all month.
- Bot Share: 96 percent of the requests that did arrive came from bots, 4 percent from humans.
- Retrieval Bots: the AI systems capable of a citation accounted for 1.1 percent.
- Audit Tools: SEO auditing software formed the largest single group at 21.7 percent.
- Sample Bias: Ahrefs states its own users skew AI-aware...wider adoption likely runs lower.
SE Ranking studied roughly 300,000 domains, published on 7 November 2025 that 10.13 percent carried one, and concluded that AI systems see and cite content the same way today with or without one. Removing llms.txt from its predictive model improved that model's accuracy...the file carried no signal about citation behavior.
Notice the adoption gap. Two real studies on two different samples returned ten percent and 28 percent...anyone citing one study without the other is citing the number that suits them.
Not a Standard Anywhere
llms.txt is a proposal published by Jeremy Howard on 3 September 2024. Queries against the IETF Datatracker API for documents with llms.txt in the title return zero results, and the W3C URL returns 404. There is no spec. robots.txt is a published standard, RFC 9309...vendors calling llms.txt an IETF-spec discovery file are describing a document nobody has written.
Chrome shipped a Lighthouse audit for the file inside an Agentic Browsing category, and the trade press read that as endorsement. Lighthouse flags a server error while retrieving the file, and returns Not Applicable when the file is simply absent, because providing it is optional. It names no AI system that reads it, nor makes any claim about Google Search. An endorsement that names no consuming system endorses nothing.
Publishing a file for your own docs is a different decision from reading other people's. Anthropic, Perplexity and OpenAI all serve one on their own docs sites, and none has published a statement that its assistant consults third-party llms.txt files...the widely repeated claim that two of them have traces to no primary source.
One Rule Governs the Rest
One sentence in Google's structured data introduction decides more than any study above. "Don't create blank or empty pages just to hold structured data, and don't add structured data about information that is not visible to the user, even if the information is accurate." Read the last clause twice. Truth is no defense...the fact has to appear on the page in text.
That rule turns an abstract question into a mechanical one, because the visible page and the JSON-LD become one document written twice.
Client-Side Markup
Markup injected by a tag manager or a client-rendered framework does not reach the major American assistants. Andre Alpar's test design was simple.
- Decoy: the raw HTML carried a wrong reference number.
- Real Value: the true one arrived only after JavaScript ran.
- Distribution: twelve assistants got twelve secret, never-linked URLs.
Twelve assistants produced four distinct behaviors.
- Raw HTML only: ChatGPT, Claude, Gemini, Perplexity and Meta AI reported the decoy.
- Ran the JavaScript: DeepSeek, ERNIE, Qwen, Kimi and Mistral returned the real number.
- Fetched but never executed: Copilot downloaded the external JavaScript file and stopped.
- Ran it and ignored it: one Grok node executed the code and still returned the decoy.
Two of those rows are contested. Writesonic's controlled test of 30 March 2026 found Copilot executing JavaScript at zero and 500 milliseconds, and Grok executing out to roughly three seconds, the most patient of the six assistants it tried. Search Engine World's June test found the result unused in both cases. Nobody has reconciled them. Gemini carries the same problem from further back. Vercel wrote in December 2024 that "Gemini leverages Googlebot's infrastructure, enabling full JavaScript rendering"...two 2026 live-fetch tests found it reading raw HTML.
Alpar's own line is "not one major American assistant ran the JavaScript and used the result", and he scopes it as one site, twelve assistants, one prompt each, at one moment. Googlebot renders in a headless Chromium, and Search Console shows you that rendered view...the failure stays invisible from inside the tool you already check. A page can rank and still be nearly empty to ChatGPT, and it should still be eligible for AI Overviews, since Google's AI features draw on the same index and the index holds rendered content. That is reasoning from Google's documentation, not a tested result.
Common Crawl documents its own behavior, and its FAQ says JavaScript is not executed...checking your own page takes three steps.
- Fetch: curl the page under OpenAI's published GPTBot user agent.
- Strip: remove the script, style and noscript blocks.
- Count: the words that remain are all the assistant gets.
Run that against the React TodoMVC demo and 645 bytes come back carrying 15 words. Byte count lies, because a 350KB page can be almost entirely bundle code and inlined JSON. Google's recommended fix points away from prerendering...its documentation calls dynamic rendering "a workaround and not a long-term solution" and names server-side rendering, static rendering or hydration instead.
Fields Google Documents and Ignores
Google's Organization page opens by removing the checklist. "There are no required properties", it says. LocalBusiness asks for only two, name and address. Its worked JSON-LD example underneath runs to eleven keys with no @id, no founder and no knowsAbout. Organization also sits outside the visual rich results group in Google's structured data gallery...that gallery files it among the types that merely help Google understand a page.
- sameAs: schema.org reads it as an identity claim while Google asks only for a page carrying more information...one property does two jobs.
- areaServed: valid schema.org vocabulary inherited from Organization, and listed by Google as neither required nor recommended for LocalBusiness.
- Corporation: adds exactly one property over Organization, which is tickerSymbol...a private company gains nothing from the switch.
- ProfessionalService: carries schema.org's own note that the type was deprecated because of confusion with the Service type.
FAQPage is the cautionary one. Google's changelog records the rich result ending on 7 May 2026 and the documentation coming down on 15 June 2026, and yet FAQPage remains among the most heavily promoted types here...one vendor still sells it with that 3.2 times AI Overview claim attached.
A Page Worth Building Anyway
Documented reasons for a facts page are smaller and duller than the sales pitch. Instructions for correcting a Google knowledge panel tell you to "Provide links to web pages that can be used to verify this information." One page stating each fact plainly, in text, is the artifact that request needs.
- Legal Name: the registered name in a sentence a reader sees, and the identical string in legalName.
- Trading Name: whatever the signage says, matching Business Profile's rule that the name reflect real-world usage.
- Founding and Address: the year and the street written out, matching foundingDate and the PostalAddress fields.
- Service Area: named towns rather than a radius, since Business Profile stopped accepting a radius entirely.
- Review Date: lastReviewed and reviewedBy are real properties...the checklists in this category skip them.
No published study shows any of this moving an AI answer...saying so beats filling the gap with a number somebody invented. What is documented runs narrower. Structured data matching the visible text appears on Google's own list of things that do matter, and the correction workflow needs a URL to point at...one fact stated once in one phrasing is harder for anything to get wrong.