This week
Run of 2026-09-26, GitHub-hosted runner, US data center, 92 storefronts. Change against 2026-09-20 in brackets.
The full picture
| robots.txt withheld (refused or no answer) | 28 / 92 | 30% |
| robots.txt readable | 64 / 92 | 70% |
| of readable: names any AI crawler or agent | 16 / 64 | 25% |
| of readable: blocks a training crawler outright | 5 / 64 | 8% |
| of readable: blocks an AI search crawler outright | 3 / 64 | 5% |
| of readable: blocks a user-triggered agent outright | 2 / 64 | 3% |
| of readable: carries Content-Signal lines | 1 / 64 | 2% |
| homepage served, readable by a non-browser client | 49 / 92 | 53% |
| homepage refused or challenged | 30 / 92 | 33% |
| homepage an empty shell (script challenge or JavaScript-only) | 7 / 92 | 8% |
| homepage no answer | 4 / 92 | 4% |
| homepage other | 2 / 92 | 2% |
| of served homepages: JSON-LD structured data | 30 / 49 | 61% |
| llms.txt published | 16 / 92 | 17% |
| Universal Commerce Protocol profile at /.well-known/ucp | 3 / 92 | 3% |
| A2A agent card | 0 / 92 | 0% |
Week on week
| Run | robots.txt withheld | names AI agent | homepage served | llms.txt | UCP |
|---|---|---|---|---|---|
| 2026-09-26 | 28 | 16 | 49 | 16 | 3 |
| 2026-09-20 | 29 | 16 | 50 | 16 | 3 |
Method
Sample: NRF Top 100 Retailers 2026, one consumer storefront each; 92 of 100 have one. Where a company runs several banners, its own storefront is used if it has one, otherwise its largest US brand. Six plain requests per site, 1.5 seconds apart, under a user agent that names the research project and links to it: robots.txt, /llms.txt, /.well-known/ucp, two A2A agent-card paths, and the homepage. No login, no forms, no cart, no browser impersonation, no impersonating another company's bot, no retry after a refusal. A homepage counts as served only if it returns a title and at least ten links; a 200 response carrying a script challenge or an empty JavaScript shell is not a page an agent can read. Timeouts are reported as no answer, not as refusals, because a slow site and a deliberate stall look the same from outside.
The scan runs every Saturday from a GitHub-hosted runner in a US data center. Bot-management decisions are not fully deterministic, so a handful of sites answer differently between runs; read the trend, not a single week. Only aggregates are published here. Per-merchant results stay private; nothing on this page rates or ranks any retailer.
The launch piece: Amazon Blocked Meta's Shopping Agent in Public. A Third of Big Retailers Do It Quietly. For the identity layer that would change these numbers, see the Major Labs Agent Identity Tracker.
Machine-readable: /trackers/merchant-readiness/feed. Every figure is reproducible from the method above. Spot an error? Reply to any edition. See also all trackers.