LosenAgent-readiness for commercial websites

All articles

What llms.txt does not do

97% of published llms.txt files received zero requests. Publishing one changes nothing on its own — and some of the other files in the same genre are actively harmful to publish wrongly.

Published

The short answer: llms.txt is a file you publish, not a file anyone fetches. In a measurement from May 2026, 97% of published llms.txt files received zero requests. Google's John Mueller has been equally blunt: it is "purely speculative … none of the AI systems use it".

A score that rewards a file existing rewards the cult around the file rather than the work.

This is not only about llms.txt

The whole genre looks the same once you measure it instead of reading about it:

  • /.well-known/api-catalog (RFC 9727): of 74 domains probed, 68 answered with HTML and status 200 — a soft 404, not a catalogue.
  • A2A Agent Cards: 65 of 22,341 hosts served one (0.29%), and only 10 of those were conformant.
  • /.well-known/mcp.json is folklore. No specification defines that path.
  • When we probed a set of Norwegian sites for ai-plugin.json and mcp.json, every 200 we got back was an SPA shell — not one valid body.

Of all the agentic-commerce protocols that are detectable from outside at all, we found one valid manifest in the entire corpus: boozt.com, version 2026-04-08. That is not an argument for hurrying.

The part that can actually hurt you

Here is the trap. If your site answers 200 with HTML to any path — which every single-page application does by default — then you are claiming endpoints you do not have.

This is not theoretical, and we got it wrong ourselves: an early run of ours reported a batch of sites as advertising agentic-commerce endpoints when they were simply answering 200 to everything, an insurer and a clinic among them. The fix was to request a deliberately nonsensical path first, and believe an endpoint only if its body parses as the type the path promises.

But tools that do not do that still exist. And an agent that tries an endpoint you "have" and gets HTML back does not come again.

How to check

Ask for a path that certainly does not exist:

curl -s -o /dev/null -w "%{http_code}\n" https://yourshop.no/.well-known/does-not-exist-12345

The answer should be 404. If it is 200, your site is lying about everything under /.well-known/, and that is worth fixing regardless of what you think about agents.

Then look at what you actually publish:

for p in llms.txt robots.txt .well-known/security.txt; do
  printf "%s " "$p"
  curl -s -o /dev/null -w "%{http_code} %{content_type} %{size_download}\n" "https://yourshop.no/$p"
done

What to do instead

The order is not arbitrary. Everything here is fetched, unlike the files above.

  1. Make sure robots.txt does not shut out the assistants answering questions right now. Blocking crawlers that collect training data is a separate and legitimate decision. Sites block search and live assistants far more often than they intend to. Remember that robots.txt groups replace each other: a dedicated group for one bot name overrides the whole * group for that bot (RFC 9309 §2.2.1).
  2. Keep sitemap.xml real and current. It does get fetched — though note that an agent following links, as ours does, never reads it at all.
  3. Render listings, prices and availability on the server. That is where the other three articles here point, and that is where the return is.
  4. Publish llms.txt anyway if you like. It costs nothing and harms nothing. Just do not believe you have done something.

One last warning about tooling: Cloudflare's own ai-rules check currently demands robots blocks for Claude-Web and anthropic-ai — two retired tokens — while omitting every current Anthropic token and PerplexityBot. So a site can pass that check and still be invisible to Claude. Worth remembering every time a checklist goes green.

Where these numbers come from

The 97% figure for llms.txt, the api-catalog and A2A Agent Card numbers, and the Mueller quote are not our measurements. They come from external standards research from May and August 2026. The well-known probing that produced our own observations here was removed from the scanner in August 2026 — it measured what a site publishes, and we measure what an agent can do with it.

We measure whether an assistant that reaches your site can use it. We do not measure whether assistants reach you at all — that is largely decided off the website, and a tool claiming otherwise is guessing.

All articles · Scan a site