This study covers 2,699 sampled rows across five national cohorts, representing 2,652 distinct domains. Whole-site crawler restriction is uncommon everywhere it was measured. Restriction reaching primary public content is higher, but still a minority position in every market. This page publishes the frozen comparative evidence base.
The upstream layer that most discussions of AI and search skip over.
AI systems increasingly depend on web-accessible information to discover, interpret and retrieve information about organisations. Yet crawler-access policies vary materially between websites and markets. C01 examines that upstream layer directly: whether and how websites permit or restrict crawler access.
The findings are relevant to researchers studying AI-mediated discovery and to businesses deciding how their websites should interact with search engines, AI crawlers and other automated systems.
Crawler access is a precondition question, not an outcome question. A site that excludes a retrieval crawler cannot be retrieved by it. A site that permits access has not thereby earned visibility, citation or recommendation — it has only remained eligible to be read. This study measures the first thing; it does not measure the second.
Descriptive findings from five independently analysed market cohorts.
C01 analysed crawler-access policy for organisational domains in five markets. For each market, two denominators are reported: the raw sample, and the policy-observed subset for which access policy could be analytically determined.
Differences between markets are reported as observed. This study does not establish why they differ, and does not attribute them to geography, jurisdiction or national policy.
The boundary of the instrument, stated explicitly.
C01 measures crawler access: whether a domain’s published access policy permits or restricts automated retrieval, and which functional class of content any restriction reaches.
Each of those is a separate empirical question requiring a different instrument. No result on this page should be read as evidence about any of them.
Measurement lens. The four research dimensions are the measurement lens of this instrument; they are distinct from, and not a relabelling of, the commercial Smart Card taxonomy.
The four research dimensions belong to a subsequent research series and were not the measurement basis of C01. C01 measured crawler access only.
Every percentage on this page is shown with the counts that produced it.
| Market | Raw N | Policy observed | Strict restriction | Meaningful restriction |
|---|---|---|---|---|
| Australia | 409 | 317/409 = 77.5% | 16/317 = 5.0% | 37/317 = 11.7% |
| United States | 808 | 624/808 = 77.2% | 24/624 = 3.8% | 67/624 = 10.7% |
| Great Britain | 754 | 540/754 = 71.6% | 22/540 = 4.1% | 52/540 = 9.6% |
| Singapore | 268 | 176/268 = 65.7% | 7/176 = 4.0% | 14/176 = 8.0% |
| India | 460 | 371/460 = 80.7% | 13/371 = 3.5% | 25/371 = 6.7% |
Policy observed = policy-observed count / raw market N.
Strict, Meaningful and Unknown-path = count / policy-observed denominator.
Strict, Meaningful and Unknown-path rates are not calculated against the raw sample. Mixing the two denominators is the most common way these figures are misread, which is why both counts are printed inline throughout this page.
Meaningful restriction is inclusive of Strict (Whole-site); it is not additional to it.
In other words, a domain counted as Strict is also counted within Meaningful. The two figures must never be added together to produce a combined restriction total.
The share of each raw sample for which crawler-access policy could be analytically determined.
India 371/460, Australia 317/409, United States 624/808, Great Britain 540/754, Singapore 176/268.
Coverage is a property of what could be observed, not a restriction finding. A domain outside the policy-observed subset is not thereby “open” or “closed” — its policy was not determinable within the instrument’s bounded protocol. All restriction measures below are calculated within each market’s policy-observed subset for exactly this reason.
Domains whose access policy restricts the whole site.
Australia 16/317, Great Britain 22/540, Singapore 7/176, United States 24/624, India 13/371. Bars are scaled to a 12% axis so that Strict and Meaningful can be compared visually on the same scale.
The narrow band across five independently sampled markets is the most stable result in the series: in none of them does whole-site restriction reach one domain in twenty.
Restriction reaching the primary public content a retrieval system would read.
Australia 37/317, United States 67/624, Great Britain 52/540, Singapore 14/176, India 25/371.
Meaningful restriction is inclusive of Strict (Whole-site); it is not additional to it.
Meaningful restriction is between roughly 1.9× and 2.8× the strict figure in each market. The gap between the two is the population of domains that restrict primary content without excluding the whole site — a position invisible to any binary “blocks AI / does not block AI” count.
Reported for transparency. Not a principal comparative measure.
| Market | Unknown-path domains | Policy-observed denominator | Rate |
|---|---|---|---|
| Australia | 146 | 317 | 46.1% |
| United States | 279 | 624 | 44.7% |
| Great Britain | 210 | 540 | 38.9% |
| Singapore | 68 | 176 | 38.6% |
| India | 131 | 371 | 35.3% |
An unknown path is a path referenced by an access policy whose functional content class could not be determined within the bounded protocol. Unknown paths are conservatively assigned to the secondary-content category, which means they contribute to the expanded measure only — not to strict or meaningful.
Expanded restriction is deliberately not used as a headline comparative measure in this publication. Doing so would make unresolved unknown-path treatment carry more interpretive weight than the frozen evidence supports. The principal comparison published here is Coverage → Strict → Meaningful.
What these numbers cannot be used to argue.
For organisations designing websites, access policies and AI-readiness strategies.
These implications follow from the observed distribution of access policy. They are offered as practical reading of the evidence, not as findings in their own right, and none of them is a claim about visibility or commercial outcome.
Each market was analysed and published independently.
Frozen sources, recorded hashes, and the records that govern them.
The five principal market results reproduce from frozen source files with zero discrepancies. Six frozen CSV artefacts passed a three-representation byte-portability check on 6 September 2026: worktree SHA256, Git blob SHA256 and git archive SHA256 each agreed with the recorded frozen hash, across 18 of 18 comparisons.
India access-analysis finality is established. India registry identity-verification finality is not asserted. The frozen India registry v1.0 retains non-final metadata fields relating to identity verification; those fields do not alter the R006-IN crawler-access measurements published here, and the India volume must not be described as a completed final identity-verification dataset.
Research reproducibility is intact. Historical publication-layout reproducibility is not established. PROVENANCE-DEFECT-002 records this as a publication-build archival defect — it is not a numerical or measurement defect, and it does not affect any figure on this page.