llms.txt and Machine-Readable Mirrors
llms.txt and machine-readable mirrors are distinct ways to organize public source material for machine use. Publishing either one does not prove that an agent will discover, retrieve, cite, or rank the material.
Not legal advice. This page is research, not compliance guidance.
What makes a source ready for machine use?
Machine-use readiness comes from stable canonical URLs, evidence layers, methodology, source references, freshness metadata, and stated limitations; those signals make a source inspectable but do not prove that a system will retrieve, cite, or rank it.
A machine-readable file is only one part of source readiness. A stable canonical URL gives a claim a durable location. An evidence layer and source references let a reader trace it. Methodology explains how the material was assembled. Freshness metadata dates the evidence, while stated limitations prevent a narrow fact from being stretched into a broader conclusion. How Private Pierce Makes Sources AI-Citable explains these structural signals in their own context.
Canonical Source Readiness by Artifact tracks which pages carry the signals. Its per-artifact rows are still being finalized; live pages already publish canonical URLs, JSON-LD, and scope notes. The inventory is therefore a status view of structural signals, not proof that a particular system uses them.
Path placement is another structural choice. Under the llms.txt proposal, a file may sit at a site root or another path and covers pages beneath that path. The proposal also separates llms.txt's on-demand contextual role from robots.txt rules about acceptable automated access. That distinction does not turn llms.txt into access control, a training opt-out, or a retrieval or ranking guarantee.
- Readiness signals: a stable canonical URL, evidence layer, methodology, source references, freshness metadata, and stated limitations.
- Inventory status: per-artifact rows are still being finalized; live pages already publish canonical URLs, JSON-LD, and scope notes.
- Interpretation limit: structural readiness does not establish discovery, retrieval, citation, or ranking by any system.
What does llms.txt propose?
The llms.txt proposal describes a small Markdown file that points agents toward relevant, machine-friendly content on demand; it does not establish that any crawler or agent actually follows the file.
Under the proposal, an llms.txt file may be placed at a site root or another path and covers pages beneath that path. The proposed flow is for an agent to view or search the file, follow a relevant link, and fetch more detailed, machine-friendly content only when needed. These are proposed mechanics, not evidence of behavior by any particular agent.
This site's root llms.txt lists citation-ready URLs aligned with priority pages in its sitemap.xml. That statement describes the site's publication choice. It does not establish whether a crawler or agent reads the file, follows a listed URL, or uses the resulting page. The evidence here also does not establish whether Google Search uses or does not use llms.txt.
The proposal assigns llms.txt a different role from robots.txt. robots.txt communicates what automated access a site considers acceptable; the proposal describes llms.txt as contextual information used on demand. An llms.txt file is therefore not an access-control rule, a training opt-out, or evidence that inclusion in search or generated answers will follow.
- Proposal mechanics: place llms.txt at the site root or another path to cover pages beneath that path.
- Intended flow: an agent views or searches the file, follows a relevant link, and fetches detailed content only when needed.
- Site implementation: the root llms.txt lists citation-ready URLs aligned with priority pages in sitemap.xml.
- Unknown: whether any particular crawler or agent follows the file, including any claim about Google Search.
- Not its role: access control, a training opt-out, or a discovery, retrieval, citation, or ranking guarantee.
How do machine-readable mirrors differ from llms.txt?
Machine-readable mirrors carry structured or simplified versions of source material, while llms.txt is a proposed index that can point toward relevant material; neither format proves that an agent will discover or use it.
The Dataset Catalog identifies this site's JSON and CSV mirrors and pairs them with the related human-readable research pages. Those mirrors expose material in structured formats. The linked research page remains the fact-owning source that supplies context, methodology, and limitations.
The llms.txt proposal describes a separate Markdown convention. It asks relevant pages to expose clean Markdown versions beside the original page URL, either by appending .md or replacing the existing extension with .md. In the proposal's intended flow, an agent finds a relevant link through llms.txt and fetches the detailed machine-friendly content only when needed. The convention does not establish that a specific agent supports, discovers, or uses a Markdown mirror.
JSON and CSV mirrors, proposed adjacent Markdown pages, and llms.txt therefore serve different structural jobs. A catalog inventories available data formats. A mirror carries another representation of source material. An llms.txt file points toward selected resources under the proposal. None of the three replaces the Methodology or the human-readable page that owns the underlying claim.
- JSON and CSV mirrors: structured representations inventoried in the Dataset Catalog and paired with research pages.
- Proposed Markdown mirrors: clean page versions at a neighboring URL formed by appending or replacing the extension with .md.
- llms.txt: a proposed compact index that points toward relevant machine-friendly resources.
- Fact ownership: the linked human-readable research page retains the context, methodology, and limitations for its claims.
- Evidence limit: publishing a mirror or listing it does not prove discovery, support, retrieval, citation, or ranking.