Playbook

Does an llms.txt File Get Your Agency Cited by AI?

A vendor sold you a file. Google says it reads it and throws it away.

Mike Moore, founder of Strategic AI Architects, at his desk looking at a laptop screen showing a plain text file labeled llms.txt next to a crossed-out AI citation checkmark, representing the gap between the file and actual AI search results
The short version

An llms.txt file, by itself, does not get an insurance agency cited by AI. A study of nearly 300,000 domains found no correlation between having the file and AI citation frequency, and only one of the top 50 most-cited domains had it at all3. Google states directly that it ignores the file1. What actually decides whether ChatGPT, Perplexity, or AI Overviews name your agency is the same unglamorous list it has always been: can a crawler reach your pages, are they indexed and snippet-eligible, and is there real sourced data on them worth repeating.

The file everyone told you to add

Somewhere in the last year, you probably heard about llms.txt. Maybe a web developer mentioned it on a call. Maybe you read a blog post, possibly one written by another agency's marketing vendor, telling you that adding this one file would get your agency "AI ready" and start showing up when people ask ChatGPT who to call about a Medicare plan. Maybe you already have one sitting on your site right now, dropped in by whoever built it, presented as a checkbox that's been ticked.

Here's the uncomfortable part. If you added it expecting it to move the needle on AI citations, the best available evidence says it didn't, and there's a real chance it never will. That's not a reason to panic about your AI visibility. It's a reason to stop chasing the wrong fix and look at what the actual gating factors are, because those are boring, well documented, and entirely within your control.

This guide walks through what llms.txt is, what Google itself says about it in its own published documentation, what a 300,000-domain study found when researchers actually tested whether it correlates with citation, why OpenAI and Anthropic still publish their own llms.txt files despite all of that, and then the part that matters more: what genuinely does move an insurance agency's odds of being named by an AI answer, sourced and checked this week.

What llms.txt actually is

llms.txt is a plain Markdown file served at the root of a domain, at yourdomain.com/llms.txt, a format proposed in 2024 as a way to hand AI systems a curated map of a site: a short summary, then a set of links to the pages the site owner considers most important. The idea, on paper, was reasonable. Search engines get a sitemap.xml full of every URL with no context. Maybe language models could use something shorter and more curated instead.

The problem isn't the idea. It's the assumption, repeated by a lot of vendors selling "AI SEO" packages, that any of the systems actually doing the citing treat the file as an input worth reading. That assumption is testable, and it has now been tested twice over: once by Google telling you directly what its own systems do with the file, and once by an independent study checking whether sites with the file get cited more often than sites without it. Both point the same direction.

Not the same as robots.txt

robots.txt is the file that actually controls whether a crawler, AI or otherwise, is allowed to access your pages at all. It is enforced by every major crawler and is where a real, fixable AI visibility problem is far more likely to be hiding. llms.txt is a separate, optional content summary that nothing is obligated to read. Confusing the two is one of the most common mistakes we see when an agency has been told its site is "AI blocked."

A real llms.txt file is short, plain, and unglamorous, which is part of why it's so easy for a vendor to add in an afternoon and count as a deliverable. A typical one opens with a single H1 line naming the business, a one-sentence summary underneath it, and then a handful of markdown sections grouping links: something like a "## Services" heading followed by a bulleted list of your service pages, a "## Guides" heading listing recent blog posts, maybe a "## Tools" heading if the site exposes any calculators or quoters an agent could call. That's the whole format. There's no field for keywords, no ranking signal to configure, nothing to optimize beyond writing an accurate one-line summary of each linked page. If a vendor quoted you a meaningful line item for building one, you were paying for maybe twenty minutes of work dressed up as a strategy.

What Google says about it, in Google's words

Google published an AI optimization guide, part of its official Search Central documentation, last updated July 10, 2026. It addresses llms.txt directly rather than leaving the question to speculation: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities)," and adds that doing so "will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them"1. That is about as unambiguous as a search engine ever gets about a specific tactic.

Google's companion AI features documentation says the same thing from a different angle. To appear as a supporting link in AI Overviews or AI Mode, "a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements," and "there are no additional technical requirements"2. The same page states plainly that "there's also no special schema.org structured data that you need to add"2, which is worth sitting with for a second: even structured data, the thing most AEO advice leans on hardest, is not a Google-stated requirement for AI Overviews eligibility. What Google does gate on is whether your page can be shown as a snippet at all, controlled by ordinary directives like nosnippet and max-snippet that predate the entire AI-search conversation2.

What Google's own documentation says gates AI Overviews eligibility, and what doesn't
Factor Does Google say it matters Source
Page is indexed by Google Yes, stated as the baseline requirement AI features documentation2
Page is snippet-eligible (no nosnippet or restrictive max-snippet) Yes, explicitly named AI features documentation2
Content is publicly accessible and crawlable Yes, part of "Search technical requirements" AI optimization guide1
A new AI-specific machine readable file or markup No, stated directly as unnecessary AI optimization guide1
Special schema.org structured data beyond normal SEO No, stated directly as unnecessary AI features documentation2
llms.txt specifically No. "Google Search ignores them" AI optimization guide1

Read that table again, because the second-to-last row surprises most people who've been sold an "AI schema package." We still recommend schema markup, and we still ship it on every site we build, but the honest reason is that FAQPage and Article schema help other engines parse your content cleanly and support classic rich results, not because Google has said it's a requirement for AI Overviews. Overselling schema as an AI Overviews requirement isn't something Google's own documentation backs up, and claiming otherwise to a client is the same category of overclaim as the llms.txt pitch this guide is about.

The 300,000-domain study

Google telling you it ignores a file is strong evidence on its own. It's still one company describing its own system. So it's worth looking at what happens when someone tests the claim empirically, across the wider AI-search ecosystem rather than just Google.

SE Ranking published a study in November 2025 that crawled close to 300,000 domains, checked which ones served an llms.txt file, and then tested whether having the file correlated with how often a domain actually got cited by AI systems. The adoption number alone is worth knowing: 10.13% of the domains studied had implemented llms.txt3, meaning roughly nine out of ten sites, across every industry in the sample, still don't have one. If you were told your competitors are all racing ahead of you on this, that's not what the data shows.

The correlation result is the part that actually matters for the sales pitch you may have heard. Using both Spearman correlation analysis and an XGBoost machine learning regression model, the researchers found no statistically significant relationship between llms.txt presence and AI citation frequency3. When they removed llms.txt as a variable from their predictive model entirely, the model's accuracy improved3, which is a researcher's way of saying the file was adding noise, not signal. Among the 50 most AI-cited domains in the whole dataset, only one had an llms.txt file3.

llms.txt adoption rate by site traffic tier, ~300,000 domains studied
Traffic tier llms.txt adoption
Low traffic (0 to 100 monthly visits) 9.88%
Mid traffic (1,001 to 5,000 monthly visits) 10.54%
High traffic (100,001+ monthly visits) 8.27%
Overall, all tiers 10.13%

Source: SE Ranking, "LLMs.txt: Why Brands Rely On It and Why It Doesn't Work," published November 7, 20253.

Notice that adoption doesn't even climb with traffic. High-traffic sites, the ones with the biggest marketing budgets and the most exposure to every new SEO trend, adopted llms.txt at a slightly lower rate than mid-traffic sites3. That's not the pattern you'd expect if the file were a genuine competitive lever. It's the pattern you'd expect from a tactic that spread through blog posts and vendor pitches faster than it spread through any measurable result.

Stat card titled What The Data Actually Shows, with three large figures: 300,000 domains studied, 10.13 percent had an llms.txt file, and 1 of the top 50 most AI-cited domains had one. Fourth figure reads 68.01 percent, the share of Google searches that ended without a click in early 2026. Sources noted as SE Ranking, November 2025, and SparkToro, 2026

This is where the free Audit earns its place in this guide rather than at the bottom as an afterthought.

Worth checking before you spend another hour on this

If you want to know whether your own site has the things that actually gate AI citation, indexing, snippet eligibility, crawlable static content, real schema, the free Audit checks it in about a minute rather than guessing from a blog post. Run a free Audit.

Then why do OpenAI and Anthropic use it?

This is the honest complication, and skipping it would make this guide as one-sided as the vendor pitch it's correcting. Both OpenAI and Anthropic maintain their own llms.txt files, and that's true. Anthropic's own engineering guidance on writing tools for AI agents states that "LLM-friendly documentation can commonly be found in flat llms.txt files on official documentation sites," pointing to its own API documentation's llms.txt as the example5. OpenAI publishes one too, at developers.openai.com/llms.txt, indexing its API guides, its Agents SDK documentation, and its developer cookbook6.

Look closely at what that file is actually for, though, and the picture changes. Anthropic's own text frames it as documentation a developer hands to a coding agent, like Claude Code, while building a tool prototype that calls a library or an API5. That's a reference-lookup job: an agent needs the shape of an API, and a clean Markdown index beats scraping a JavaScript-heavy docs site built for humans. It has nothing to do with whether ChatGPT cites your insurance agency's blog post when someone asks about Medicare Advantage plans in their county. Those are two different systems doing two different jobs. Conflating "AI companies use this file for developer docs" with "this file gets your business cited" is exactly the leap the vendor pitch depends on, and it's the leap the 300,000-domain study tested and found unsupported.

Documentation lookup

What llms.txt is actually built for

  • A coding agent needs an API's exact method names and parameters
  • The developer points the agent at a clean, curated Markdown index
  • The agent reads it once, mid-task, to write correct code
  • Used by Anthropic and OpenAI for their own developer docs56

RealA narrow, working use case

Consumer AI citation

What the vendor pitch promises instead

  • A prospect asks ChatGPT or an AI Overview who to call about a plan
  • The system decides which sources to name in its answer
  • No major consumer AI search product reads llms.txt for this decision1
  • No correlation found across nearly 300,000 domains tested3

UntestedSold as fact, not shown as fact

What actually gets a page cited

Infographic titled What Actually Gets a Page Cited. A faded dashed line shows an llms.txt file icon leading to a crossed-out checkmark labeled Ignored by Google. A bold indigo path below it shows five connected steps: Static HTML, Indexed, Snippet-eligible, Sourced data, Cited by AI, ending in a solid checkmark. Source noted as Google Search Central, AI optimization guide, 2026

None of this means AI citation is unwinnable or that the work doesn't matter. A large and growing share of your buyers are asking an assistant instead of typing a search query. SparkToro's analysis of Similarweb's US desktop and mobile panel data for January through April 2026 found that 68.01% of Google searches ended without a click, up from 60.45% just two years earlier in 20244. Being named in the answer, not just ranking a link nobody clicks, is increasingly the whole game. The question is which levers actually pull that outcome, and Google's own documentation and the SE Ranking data both point at the same unglamorous list.

Share of Google Searches Ending Without a Click 2024 60.45% Jan to Apr 2026 68.01% US Google searches, desktop and mobile combined. SparkToro / Similarweb panel data, 2026.
Zero-click Google searches rose 7.56 percentage points in two years4.

For Google's AI Overviews and AI Mode specifically, the gate is the same one that's governed classic search for years, with one AI-era addition. Your page has to be indexed. It has to be crawlable, meaning it renders without depending on JavaScript the crawler might not execute. And it has to be snippet-eligible, meaning nothing on the page or in its response headers carries a nosnippet or an overly restrictive max-snippet directive, because Google states outright that these directives "prevent or limit the content from being used as a direct input for AI Overviews and AI Mode"2. That last one is the sneaky one. A theme or a security plugin can ship a restrictive max-snippet by default, quietly disqualifying an otherwise well-built page from ever being pulled into an AI answer, with no visible error anywhere on the page itself.

For ChatGPT, Perplexity, and Claude, the mechanics differ because these are separate products with their own crawlers and their own retrieval systems, not extensions of Google's index. What they share with Google's version of the problem is more basic: can the crawler get in at all, and once it's in, is there something worth quoting. A crawler that hits a slow page, a JavaScript wall, or a robots.txt block never gets far enough to evaluate your content's quality in the first place. A crawler that gets in and finds a vague paragraph with no named source, no year, and no specific number has nothing citable to lift, even if it can technically read every word.

10.13%

Of ~300,000 domains had llms.txt; adoption doesn't track with AI citation3

1 of 50

Top AI-cited domains in the study that had llms.txt3

68.01%

Of Google searches ended without a click, Jan to Apr 20264

0

Additional AI-specific technical requirements Google states beyond snippet eligibility2

Why this hits page-builder sites harder

A GoHighLevel funnel page or a WordPress site running a heavy visual builder doesn't fail the AI citation test because it's missing a special file. It fails because of what has to happen before the actual content is even available to read. A page-builder platform loads a JavaScript runtime, a theme, and a stack of plugins in the browser before the visible content appears, which is slower for a human visitor and more expensive for any crawler to fully render and parse. We build on Astro, which ships zero JavaScript to the browser by default and renders pages to finished static HTML at build time, so a crawler, human or AI, downloads a complete page instead of downloading a framework, booting it, and waiting for the content to draw.

That architectural difference is the whole speed story, and it compounds directly into the AI citation conversation, because a slow or JavaScript-dependent page is exactly the kind that's expensive to crawl at scale and easy for a bot to deprioritize or skip. It's a separate problem from llms.txt entirely, but it's the one that actually shows up in the outcomes vendors promise the file will deliver.

The pattern worth remembering. Every real AI citation gate, indexing, crawlability, snippet eligibility, speed, sourced data, traces back to the same technical foundation classic SEO has asked for all along. llms.txt is the one item on the list that isn't actually part of that foundation, according to Google itself and according to the largest independent test of the claim so far.

Make it concrete with a single page. Say your agency runs a page comparing Medicare Advantage and Medicare Supplement plans, built on a GoHighLevel funnel template. A visitor's phone has to download the page-builder's JavaScript bundle, wait for it to parse, wait for the theme to render, and only then see the comparison table you wrote. On a decent mobile connection that can easily run past four or five seconds before the content anyone actually asked for shows up. An AI crawler hitting the same URL faces a version of the identical problem: it either has to execute that JavaScript to see the content at all, which many crawlers either can't or won't do at scale, or it gives up and moves to the next result. The exact same page rebuilt in static HTML delivers that comparison table in the initial response, no script required, readable by a human or a bot in the time it takes to receive the file. Nothing about that rebuild touches llms.txt. It touches the one thing that was actually broken.

How to check your own site

You don't need us, or anyone, to run this first pass. It takes about ten minutes with tools you already have open in a browser tab.

  1. Open your own robots.txt directly at yourdomain.com/robots.txt and read every User-agent block by name. Look specifically for GPTBot, OAI-SearchBot, and ChatGPT-User (OpenAI), ClaudeBot (Anthropic), PerplexityBot (Perplexity), Google-Extended, and Bingbot. If any of them appear under a Disallow rule, that's a real, fixable block, unlike llms.txt.
  2. Check for a Cloudflare Managed robots.txt block if your site sits behind Cloudflare. Cloudflare's own managed feature can prepend a block on several AI crawlers above your site's actual file, and it won't show up unless you check the zone settings directly rather than just reading the public robots.txt output.
  3. Search your page's HTML for nosnippet or a max-snippet directive in the meta robots tag, and check the X-Robots-Tag response header separately, since a header-level directive is invisible in the page source and just as disqualifying.
  4. Run your homepage and top service pages through PageSpeed Insights and look at the mobile score specifically, since mobile is scored separately from desktop and is the harder of the two to pass.
  5. Open your page's structured data in Google's Rich Results Test and confirm your Article, FAQPage, and Organization schema actually validate and match what's visibly on the page, not just present in the source.
  6. Ask three of your buyers' actual questions to ChatGPT, Perplexity, and a Google AI Overview and note whether your agency is named as a source, and if it is, which specific page gets the citation. That tells you which of your existing pages is already doing this job, so you know what to build more of.

Every step above is something you can run yourself

None of it requires buying anything. Plenty of agencies read a checklist like this and decide to run it on their own, and that's a fine outcome. If you'd rather have someone run all six checks and hand you the findings instead of doing it yourself, that's what the free Audit does in under a minute.

How we build the whole stack, not just the file

We still ship llms.txt on every site we build. It costs nothing to serve, it's genuinely useful for the narrower set of agentic and coding-assistant tools that do read it, and it's part of an emerging WebMCP tool-discovery pattern worth being ready for even where it isn't the deciding factor today. What we don't do is sell it to a client as the thing that gets them cited by AI, because the evidence in this guide says that claim doesn't hold up, and a claim we can't defend isn't one we're going to make to protect a sale.

What actually moves the needle is the full list this guide has walked through, and it's the list we build into every site by default rather than as an upsell. Every build ships on Astro, static HTML at the edge on Cloudflare, with Lighthouse performance scores of 90 or better on both mobile and desktop, and a number of pages that score 100. Every build ships llms.txt, sitemap.xml, and robots.txt that regenerate themselves the moment a page publishes, with the full absolute URL included automatically, and a robots.txt written to welcome every crawler by name rather than leaving the default alone. You can check ours directly: it explicitly allows GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, and Bingbot, and states outright that "every crawler is welcome, including AI answer engines. We WANT to be read, indexed, and cited"7.

Every page carries a real schema stack, Organization, Service, Article, FAQPage, and BreadcrumbList depending on the page type, matched to what's actually visible on the page rather than hidden markup contradicting the copy. Every blog post we publish cites the primary source for every number in it, the same discipline this guide has tried to model rather than just describe. And the quoters and calculators we build on a site are wired to live federal data instead of decorative forms with made-up example numbers, which is its own separate reason a real answer engine has something worth citing on your site beyond prose.

See the stack instead of one file

If you'd rather see what a full AI-citation build actually includes than evaluate each piece separately, Digital Foundation is the place to look, or book a call if your setup is more custom than a standard tier covers.

Digital Foundation includes this entire stack as part of building the site, not as a separate add-on you have to know to ask for, with a 14-day free trial so you can see it live on your own domain before billing starts8.

What you get

Skip the file, keep the fundamentals, and what changes is concrete rather than promotional. A site that loads fast enough for a crawler and a phone visitor to both get through it without waiting. A crawl file setup that doesn't quietly turn away the exact bots feeding the answers your buyers are reading. Pages with real, sourced, dated numbers instead of vague claims, which is the actual raw material an AI system needs to have something worth quoting. None of that is guaranteed to produce a specific citation count or a specific ranking, and we won't promise you one. What it does is put every lever Google and the independent research actually point to in place, instead of one that neither points to at all.

One honest caveat worth stating plainly: none of this is frozen. AI search is a young enough field that a file nobody reads today could plausibly matter more tomorrow, particularly as WebMCP and similar agent-tool-calling standards mature and browsing agents start looking for a structured list of what a site lets them do, not just what it says. That's a real, worth-watching possibility, and it's a different claim than "add this file and get cited now," which is the claim the study in this guide actually tested and didn't support. We'll keep shipping llms.txt on every build regardless, the same way we'd keep a low-cost door open rather than nail it shut, and we'll update this guide if the evidence changes. It just isn't the lever to pull today, and telling a client otherwise to close a sale isn't a claim we're willing to make.

Questions agents ask

Does adding an llms.txt file help my insurance agency get cited by ChatGPT or Google's AI Overviews?

The evidence says no, at least not on its own. An SE Ranking study of nearly 300,000 domains found no statistically significant correlation between having an llms.txt file and how often a domain gets cited by AI systems, and only one of the 50 most AI-cited domains in the dataset had the file at all (SE Ranking, November 2025). Google's own documentation states it plainly: Google Search ignores llms.txt entirely.

What does Google actually say about llms.txt?

Google's AI optimization guide, updated July 10, 2026, states that you don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search, including its generative AI features, and that doing so will neither harm nor help your site's visibility because Google Search ignores them. Google's AI features documentation adds that there are no additional technical requirements beyond being indexed and snippet-eligible.

If llms.txt doesn't help with AI citation, why do OpenAI and Anthropic use it?

Both companies publish llms.txt files for their own developer documentation, but for a different job: feeding coding agents like Claude Code clean API reference material while a developer is writing software. That is a documentation-lookup use case, not a consumer citation use case. Anthropic's own engineering guidance describes it as a place LLM-friendly documentation commonly lives on official docs sites, which is a narrower claim than getting a business cited in an AI answer.

What actually determines whether an AI engine cites my agency's website?

For Google's AI Overviews and AI Mode, it is the same foundation as classic search: the page has to be indexed, crawlable, and eligible to show with a snippet, with no nosnippet or restrictive max-snippet directive blocking it, per Google's own AI features documentation. For ChatGPT, Perplexity, and Claude, it comes down to whether their crawlers can reach fast, static, well-structured HTML with real sourced data and schema markup a machine can parse.

Is llms.txt worth adding to my insurance agency's website anyway?

Yes, as a small, low-cost addition to a real AEO strategy, not a substitute for one. It costs nothing to serve, it does not hurt anything, and it is genuinely useful for the smaller set of agentic tools that do read it, including coding-assistant workflows and emerging WebMCP tool discovery. Treat it the way you would treat a business card left on a counter: worth having, not a marketing plan.

Is this different for ChatGPT and Perplexity than it is for Google's AI Overviews?

Somewhat. Google's AI Overviews and AI Mode run on the same core Search index and ranking systems as classic search, so the gating factors are indexing, crawlability, and snippet eligibility, not any AI-specific file. ChatGPT and Perplexity operate their own separate crawlers, GPTBot, OAI-SearchBot, and PerplexityBot among them, and rely more heavily on whether those crawlers are allowed in and can parse the page cleanly and quickly.

How do I check whether AI crawlers can even reach my site?

Fetch your own robots.txt file directly in a browser and read every User-agent block, watching specifically for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, and Bingbot. If Cloudflare fronts your site, check its Bot Fight Mode and Managed robots.txt setting separately, since that layer can block a crawler even when your own file allows it.

Does a slow, page-builder insurance website hurt AI citation the same way it hurts Google rankings?

It hurts it more directly in some ways. A page-builder platform loads a JavaScript runtime, a theme, and a stack of plugins before the actual content appears, which makes the page slower and more expensive for any crawler, human or AI, to process. Static HTML built on a framework like Astro arrives as a finished page with nothing to boot, which is the architectural reason it is both faster for visitors and cheaper for a crawler to fully read.

Can having an llms.txt file actively hurt my agency's site?

No. Every source in this guide, including Google's own documentation, describes it as neutral at worst, not harmful. The real cost isn't the file itself, it's the opportunity cost of a vendor charging real money for it, or an owner spending their limited attention on it, while the fundamentals that actually gate citation, indexing, crawlability, snippet eligibility, and sourced content, go unchecked.

Sources

  1. Google Search Central. "AI Optimization Guide," on machine readable files and llms.txt, updated July 10, 2026. developers.google.com.
  2. Google Search Central. "AI Features and Your Website," on snippet eligibility, nosnippet, max-snippet, and technical requirements for AI Overviews and AI Mode. developers.google.com.
  3. SE Ranking. "LLMs.txt: Why Brands Rely On It and Why It Doesn't Work," published November 7, 2025, analysis of nearly 300,000 domains. seranking.com.
  4. SparkToro. "In 2026, Less Than One Third of Google Searches Still Send a Click," using Similarweb US desktop and mobile panel data, January through April 2026. sparktoro.com.
  5. Anthropic. "Writing Effective Tools for AI Agents, With Agents," engineering guidance mentioning llms.txt for API documentation. anthropic.com.
  6. OpenAI Developers. llms.txt documentation index for the OpenAI API, Agents SDK, and developer cookbook. developers.openai.com.
  7. Strategic AI Architects. robots.txt, live crawler-access policy for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, and Bingbot. strategicaiarchitects.com.
  8. Strategic AI Architects. "Digital Foundation," service pricing and trial terms. strategicaiarchitects.com.

Talk it through

Want a second pair of eyes on it?

Free 30 minutes. Bring what you found, or bring nothing and we will look together at how AI engines read your site and which fixes move first.

Find out what your site is actually missing

Run the free Audit, a live AEO Audit plus a HIPAA tracking scan of your site, in under a minute.

← All guides