llms.txt — What It Is, What to Put In It, and What It Will Not Do For You
It takes an hour to write, costs nothing, and no engine guarantees to read it. Here is the honest version, including why robots.txt matters considerably more.
In short
llms.txt is a plain-text file at your site root listing your most useful pages for AI systems. It takes about 1 hour to write and no engine guarantees to read it. The 4 robots.txt tokens — GPTBot, ClaudeBot, PerplexityBot, Google-Extended — matter far more, and we find one blocked on half the sites we audit.
Reviewed 20 September 2026 · reviewed quarterly; adoption claims will be revised the moment an engine states a position publicly
What it is, plainly
A markdown-formatted text file placed at the root of your domain, listing the pages you most want an AI system to use, each with a short description of what it contains. The idea is simple: rather than making a model infer your site's structure from navigation and internal links, you hand it an index. It is an offer of help, not an instruction, and that distinction is the whole of the honest assessment.
Because it is an offer, nobody is obliged to accept it. Adoption is not universal, no major engine publishes a commitment to read it, and there is no reporting to tell you whether yours was used. What it does have in its favour is that it costs an hour, breaks nothing, and is trivially easy to keep current. Cheap insurance is a perfectly respectable category, and it should be sold as that rather than as a lever.
What belongs in yours
Keep it short. A file listing two hundred URLs is not an index, it is a sitemap with prose attached.
| Section | What to list | Why |
|---|---|---|
| A one-line description of the business | Who you are, where, what you sell | Use the same canonical description you use on every profile — entity consistency is the point |
| Core pages | Your 5 to 15 most useful pages, each with a sentence | If a model reads one thing, this is what you want it to be |
| Reference and data assets | Pricing pages, benchmarks, checklists, specifications | These are what get cited. Commercial pages average 3.1 AI citations against 5.6 for definitional ones |
| Contact and identity | Phone, email, address, hours | Cheap to include, and it reduces the chance of a stale directory record being used instead |
| What you do not do | A short line naming services you do not offer | Stops you being cited for work you would have to decline |
| A last-updated date | The date, in plain text | Same reason every page on this site carries one |
robots.txt is the file that actually decides
If you only have an hour, spend it here instead. Four user agents have their own tokens and each has to be checked by name: GPTBot, ClaudeBot, PerplexityBot and Google-Extended. Google-Extended is separate from Googlebot, so it is entirely possible to rank normally in Google search while being excluded from some AI surfaces, and almost nobody who is in that position chose it.
We find at least one of the four blocked on roughly half the sites we audit — including sites paying another agency ₹40,000 a month. Sometimes it is a platform default, sometimes a security plugin, sometimes a line someone copied from a blog post in 2023 when the advice was different. The other half of the check is the part people miss: test it from outside your own network, because a CDN, WAF or bot-management rule can block a crawler that robots.txt cheerfully allows. That is check 40 on our audit checklist and it is the commonest false pass in the whole list.
If you genuinely want to block AI crawlers
That is a legitimate choice — a publisher whose content is the product has a real argument for it. Make it deliberately, block by name, and accept the visibility cost: business websites earn 59.9% of local AI citations, so opting out means ceding those answers to directories and competitors. What is not defensible is being blocked by accident, which is the situation most sites are actually in.
llms.txt questions
Do I need an llms.txt file?
Need is too strong. It takes about an hour, costs nothing and cannot hurt, so the sensible answer for most businesses is yes, do it once and keep it current. But treat it as cheap insurance rather than as a lever — no engine guarantees to read it and there is no reporting to tell you whether yours was used.
Is llms.txt the same as robots.txt?
No. robots.txt controls access and is honoured; llms.txt offers guidance and is optional. If you have an hour, spend it on robots.txt — checking GPTBot, ClaudeBot, PerplexityBot and Google-Extended by name, and then testing from outside your network in case a CDN rule blocks what robots.txt allows.
What should go in llms.txt?
A one-line description of the business using the same canonical wording you use everywhere else, your five to fifteen most useful pages each with a sentence, your reference and data assets, contact details, a short note on what you do not do, and a last-updated date. Keep it short — a 200-URL file is not an index.
Will llms.txt get me cited by ChatGPT?
Not on its own. Citation comes from being in the retrievable pool with content that answers the question, plus a consistent identity across the profiles those systems read. llms.txt may make your good pages easier to use; it cannot make thin pages worth using.
How do I check if AI crawlers can reach my site?
Read robots.txt for each of the four tokens by name rather than assuming a blanket allow covers them, then request a page as those user agents from outside your network to catch CDN and firewall rules. Roughly half the sites we audit have at least one of the four blocked, almost always unintentionally.
Should a small Chennai business bother with any of this?
With robots.txt, absolutely — being invisible to AI systems by accident costs you answers you would otherwise appear in. With llms.txt, only if it takes an hour and stays current. Neither is a substitute for a complete Google Business Profile, which supplies 67% of local citations inside AI Overviews.
Related pages
- AEO checklist Section 5 covers crawler access
- 42-point checklist Checks 36–42
- AEO & GEO The full programme
- AI scoreboard What measurement looks like
- ₹12,500 audit Crawler access tested from outside
- Schema markup The 2.3× lever
Send the URL. Get a tier, a price and a start date.
One working day, on WhatsApp, from a person who has already opened your site in Search Console-shaped eyes. No discovery call required before you get a number.