llms.txt
A practical guide to the optional llms.txt convention: format, source governance, implementation, validation, differences from robots.txt and sitemaps, Google guidance, Lighthouse, and current limits.
- #LLM SEO
- #AI SEO
- #SEO Glossary
- #Technical SEO
In Plain English
llms.txt is a proposed, optional Markdown file at /llms.txt that gives inference-time agents a curated orientation to a website and its authoritative public sources.
Key Takeaways
- llms.txt is an informal optional convention, not a Google Search requirement or established ranking factor
- The file orients agents to public sources but does not control crawling, privacy, training, indexing, or citation
- A useful llms.txt should be generated from canonical content, deliberately curated, validated, and owned by a maintenance process
Deep dive
Quick Definition
llms.txt is a proposed, optional Markdown file published at /llms.txt. Its purpose is to give an inference-time agent a concise map of a website, explain what the project is, and point to a curated set of authoritative public resources.
It is an informal proposal, not a formal web standard. It does not tell a crawler what it may fetch, opt content out of model training, add URLs to a search index, act as a Google ranking factor, or guarantee that an answer system will read, cite, or follow the file.
What Problem the Proposal Tries to Solve
A documentation or knowledge site may contain hundreds of useful pages wrapped in navigation, scripts, layout, filters, changelogs, archives, and duplicated routes. An agent working with limited context may need a shorter route to the canonical introduction, current docs, primary definitions, and important policies.
llms.txt proposes a human-readable source map for that moment. The file can say, in effect: this is the project, these are the public sources that explain it, and these optional resources can be skipped when context is tight.
That is orientation, not authority by itself. A link in the file does not make a weak page trustworthy. The linked page still has to be current, accessible, specific, and supported.
Status of the Convention
Jeremy Howard published the proposal in September 2024, and the proposal page has been revised since. It defines a Markdown format and also suggests offering clean Markdown versions of individual pages at the same URL with .md appended. It does not prescribe how every model, agent, search product, or browser must process the file. Adoption and behavior remain product-specific.
Chrome Lighthouse describes llms.txt as an emerging convention in its Agentic Browsing guidance. Its audit treats a missing file as not applicable because the file is optional; a server error is reported because a broken advertised route is different from choosing not to publish one.
This Lighthouse check does not turn llms.txt into a Google Search requirement. Browser-agent readiness and Google Search ranking are different contexts.
The Proposed Format
The proposal uses Markdown because it is easy for people to review and for software to parse. Only one element is required: an H1 containing the project or site name. The remaining elements are optional and appear in this order:
- a blockquote with a short project summary;
- additional notes or context without headings;
- H2 sections containing Markdown lists of links;
- an H2 section named
Optionalfor resources that may be omitted when a shorter context is needed.
Each list item can contain a link and a short description after a colon. The description should tell the agent why the source matters, not repeat the page title.
markdown# Example Project
> A concise description of the project, audience, and source scope.
Use the current documentation for product behavior. Pricing and policy pages carry their own effective dates.
## Documentation
- [Getting started](https://example.com/docs/start): Current setup path and supported prerequisites.
- [API reference](https://example.com/docs/api): Canonical endpoint and schema documentation.
## Policies
- [Security](https://example.com/security): Published security scope and reporting contact.
## Optional
- [Changelog](https://example.com/changelog): Dated product changes and migration notes.
The proposal also discusses an optional llms-full.txt containing a more extensive Markdown compilation. That larger file increases maintenance and context costs. It should not be created merely because the filename exists in the proposal.
What llms.txt Is Not
Not robots.txt
robots.txt expresses crawl rules for named user agents and paths. llms.txt contains descriptions and links. Publishing a URL in llms.txt does not override a block, authentication, a firewall, a paywall, a noindex instruction, or a provider's policy.
Conversely, omitting a page from llms.txt does not stop a crawler or a user-requested agent from finding it elsewhere. Use real access, privacy, indexing, and bot controls for those decisions.
Not a sitemap
An XML sitemap helps search engines discover canonical URLs at scale. It can contain many URLs and technical metadata. llms.txt is meant to be selective and explanatory.
A sitemap answers “which URLs should a search engine know about?” A curated llms.txt answers “which public sources best explain this project?” A site may use both, one, or neither according to its needs.
Not a training opt-out
Crawler policy remains provider-specific. OpenAI, for example, documents OAI-SearchBot for ChatGPT search and GPTBot for content that may be used in training, and it treats those controls independently. Anthropic likewise names separate bots: ClaudeBot for content that may contribute to training, Claude-SearchBot for search quality, and Claude-User for retrieval a user asks for.
An llms.txt file does not replace these directives. It also cannot remove knowledge already present in a released model.
Not a privacy or security boundary
Never list private, staging, preview, customer, tokenized, internal, legally restricted, or confidential URLs. The file is public. A note asking agents not to use a link is not access control.
Not a citation or ranking guarantee
An application may ignore the file, read it but choose another source, or use the linked material without showing a citation. No reliable public formula converts inclusion in llms.txt into a probability of visibility.
What Google Says
Google's guidance for AI Overviews and AI Mode says that normal Search fundamentals apply. Site owners do not need new machine-readable files, special AI text files, or special schema markup to appear in these features. Supporting pages must still be indexed, eligible for a snippet, accessible, and compliant with Search policies.
Therefore, llms.txt is not required for Google Search, AI Overviews, or AI Mode. The presence of a Lighthouse Agentic Browsing audit does not make the file a ranking factor or a Site Audit defect when absent.
When the File May Be Useful
The convention is most plausible where a stable public knowledge space has more useful sources than an agent can sensibly inspect at once. Examples include:
- developer documentation with versioned guides and API references;
- a product with several maintained features, policies, and integration pages;
- a large glossary or research library with canonical definitions;
- a help center with authoritative troubleshooting routes;
- a standards body, university, government, or open-source project with clear source ownership;
- a multilingual site that can label the authoritative resource for each language.
The benefit is operational: one reviewed map can expose source priorities more clearly. That benefit still depends on an application choosing to use it.
When to Leave It Out
Do not publish llms.txt merely to satisfy a checklist. Leave it out when:
- no application or user need justifies the maintenance work;
- the team cannot name an owner or review schedule;
- product, pricing, policy, or documentation sources contradict one another;
- the proposed list would expose non-public routes;
- the file would become an automated dump of hundreds of URLs;
- the project expects a Google ranking or citation guarantee.
A clean 404 is an acceptable outcome for an optional file. A stale file that promotes retired features or old policy can be worse than no file because it publishes a convenient map of the wrong facts.
Build It from the Source of Truth
For a site with structured content, generate the file from the same canonical records that drive navigation, documentation, product metadata, locales, and public status. Keep a small explicit allowlist of resource groups rather than exporting every published route.
A maintainable process should define:
- the owner who approves descriptions and source priority;
- which content states are eligible, such as public and production only;
- how canonical and language-specific URLs are selected;
- which page types are included and how many links each may contribute;
- when retired, redirected, or materially changed pages are removed;
- how changes are reviewed before deployment;
- a scheduled or event-driven regeneration path.
Manual files are reasonable for a small stable site. For a large product, manual duplication invites drift.
Editorial Selection Rules
Choose links because they answer a distinct high-value question. Prefer a primary current source over a summary that merely repeats it. Keep descriptions factual and short. Include effective dates in the linked page when the information changes, rather than packing volatile details into the map.
Avoid:
- five pages that answer the same question;
- campaign URLs, tracking parameters, search results, filters, or pagination;
- raw inventories with 500 links;
- unsupported marketing superlatives;
- descriptions that contradict the linked page;
- obsolete language variants or redirected paths;
- linking to a homepage when a precise source exists.
The Optional section is useful for secondary examples, archives, and changelogs. It should not become a place to hide everything that failed the main curation decision.
Technical Validation
After deployment, verify the public route rather than only the source file:
- request the exact root path
/llms.txtwithout authentication; - expect a successful HTTP response and UTF-8 text;
- serve readable plain text or Markdown with a sensible content type;
- confirm the H1, section order, link syntax, and absolute public URLs;
- reject private hosts, localhost, preview deployments, tracking URLs, fragments, and duplicate canonicals;
- follow every selected link and flag server errors, soft 404s, blocked pages, or unexpected redirects;
- compare descriptions with current page titles and source facts;
- rerun the test after relevant content or routing releases.
Do not infer usage from one request in server logs. A bot can fetch the file without using it in an answer, and caching or proxy requests can obscure the caller.
Where Crawl Foundry Fits
Crawl Foundry's checks sit next to the file, on the question of access. The free AI Crawler Checker tests whether search and AI crawlers can fetch, index, and extract a public page. The free AI robots.txt Checker shows which AI crawlers each rule in a robots.txt actually allows or blocks.
If you decide to publish an llms.txt, the groundwork is ordinary SEO evidence: which pages carry your important topics, and whether those URLs return 200, are indexable, and are internally linked. The keyword database and Site Audit help with that. None of these checks shows whether an agent reads llms.txt, and Crawl Foundry promises no citation or ranking effect from the file.
Common Failure Modes
- Calling a proposal a standard or protocol supported by every LLM.
- Presenting absence as a technical SEO error.
- Confusing source orientation with bot permission.
- Claiming the file opts content out of training.
- Publishing private or staging URLs.
- Generating a full sitemap-shaped URL dump.
- Copying volatile prices and product claims into a file with no owner.
- Leaving links and descriptions stale after a migration.
- Treating a Lighthouse check as proof of Google Search impact.
- Reporting a citation change as caused by
llms.txtwithout controlled evidence.
FAQ
Is llms.txt required for Google SEO?
No. Google says no new AI text file is required for AI Overviews or AI Mode. The ordinary Search requirements remain relevant.
Does llms.txt replace robots.txt or a sitemap?
No. robots.txt expresses crawl policy, and a sitemap supports URL discovery. llms.txt is an optional, curated orientation file.
Does a 404 mean the site has an SEO defect?
No. The file is optional. Current Lighthouse Agentic Browsing guidance treats a missing file as not applicable, while a server error is a different implementation problem.
Does llms.txt block training or guarantee citation?
No. Training controls and search crawlers are governed separately by providers. Citation remains a product-level retrieval and answer decision.
Should every documentation site publish one?
Only when the site has a useful curated source set and a credible maintenance owner. The decision should follow the use case, not a generic readiness score.
Related Terms
- LLM SEO
- LLM Visibility
- Generative Engine Optimization
- Answer Engine Optimization
- robots.txt
- XML Sitemap
- AI Overviews
- Google AI Mode
Sources and Further Reading
Why It Matters for SEO
Teams need to distinguish a potentially useful orientation file from access controls and proven search requirements. Otherwise an optional experiment can create stale facts, false expectations, and unnecessary technical work.
Common questions
What is llms.txt?
llms.txt is a proposed, optional Markdown file at /llms.txt that gives inference-time agents a curated orientation to a website and its authoritative public sources.
Why does llms.txt matter for SEO?
Teams need to distinguish a potentially useful orientation file from access controls and proven search requirements. Otherwise an optional experiment can create stale facts, false expectations, and unnecessary technical work.
Deep blog guides
When you are ready to move from definition to workflow, continue with these practical guides.
Reviewed by
Crawl Foundry Team
Editorial TeamThe Crawl Foundry editorial team.
Give your content team one shared evidence base
Crawl Foundry gives editors, strategists, and leads one shared keyword database with lists, tags, SERP context, and statuses. Whether you publish an llms.txt stays your team's decision.