C01 · Global Crawler Access Series · Five-Market Comparative Findings

C01 — Global Crawler Access Series: Five-Market Comparative Findings — Australia, United States, Great Britain, Singapore and India

This study covers 2,699 sampled rows across five national cohorts, representing 2,652 distinct domains. Whole-site crawler restriction is uncommon everywhere it was measured. Restriction reaching primary public content is higher, but still a minority position in every market. This page publishes the frozen comparative evidence base.

Research instrumentPTODA C01 Crawler v1.5 / v1.5.1
Publication typeWeb findings publication
Published6 September 2026
Orientation

Why this research matters

The upstream layer that most discussions of AI and search skip over.

AI systems increasingly depend on web-accessible information to discover, interpret and retrieve information about organisations. Yet crawler-access policies vary materially between websites and markets. C01 examines that upstream layer directly: whether and how websites permit or restrict crawler access.

The findings are relevant to researchers studying AI-mediated discovery and to businesses deciding how their websites should interact with search engines, AI crawlers and other automated systems.

Crawler access is a precondition question, not an outcome question. A site that excludes a retrieval crawler cannot be retrieved by it. A site that permits access has not thereby earned visibility, citation or recommendation — it has only remained eligible to be read. This study measures the first thing; it does not measure the second.

3.5–5.0%
Whole-site restriction across all five markets, measured against each market’s policy-observed denominator.
Restriction reaching primary public content — the meaningful measure — ranges from 6.7% (India) to 11.7% (Australia). Policy-observed coverage ranges from 65.7% (Singapore) to 80.7% (India).
Executive summary

What the five markets show

Descriptive findings from five independently analysed market cohorts.

C01 analysed crawler-access policy for organisational domains in five markets. For each market, two denominators are reported: the raw sample, and the policy-observed subset for which access policy could be analytically determined.

Differences between markets are reported as observed. This study does not establish why they differ, and does not attribute them to geography, jurisdiction or national policy.

Scope

What C01 measures — and what it does not

The boundary of the instrument, stated explicitly.

C01 measures crawler access: whether a domain’s published access policy permits or restricts automated retrieval, and which functional class of content any restriction reaches.

C01 does not measure

Each of those is a separate empirical question requiring a different instrument. No result on this page should be read as evidence about any of them.

Measurement lens. The four research dimensions are the measurement lens of this instrument; they are distinct from, and not a relabelling of, the commercial Smart Card taxonomy.

The four research dimensions belong to a subsequent research series and were not the measurement basis of C01. C01 measured crawler access only.

Sample and denominators

Five-market sample and denominator table

Every percentage on this page is shown with the counts that produced it.

Table 1 — Sample, policy-observed coverage and restriction measures, five markets
Market Raw N Policy observed Strict restriction Meaningful restriction
Australia409317/409 = 77.5%16/317 = 5.0%37/317 = 11.7%
United States808624/808 = 77.2%24/624 = 3.8%67/624 = 10.7%
Great Britain754540/754 = 71.6%22/540 = 4.1%52/540 = 9.6%
Singapore268176/268 = 65.7%7/176 = 4.0%14/176 = 8.0%
India460371/460 = 80.7%13/371 = 3.5%25/371 = 6.7%

Denominator convention

Policy observed = policy-observed count / raw market N.

Strict, Meaningful and Unknown-path = count / policy-observed denominator.

Strict, Meaningful and Unknown-path rates are not calculated against the raw sample. Mixing the two denominators is the most common way these figures are misread, which is why both counts are printed inline throughout this page.

Meaningful restriction is inclusive of Strict (Whole-site); it is not additional to it.

In other words, a domain counted as Strict is also counted within Meaningful. The two figures must never be added together to produce a combined restriction total.

Measure 1 of 3

Policy-observed coverage

The share of each raw sample for which crawler-access policy could be analytically determined.

Policy-observed coverage — count / raw N
India80.7%
Australia77.5%
United States77.2%
Great Britain71.6%
Singapore65.7%

India 371/460, Australia 317/409, United States 624/808, Great Britain 540/754, Singapore 176/268.

Coverage is a property of what could be observed, not a restriction finding. A domain outside the policy-observed subset is not thereby “open” or “closed” — its policy was not determinable within the instrument’s bounded protocol. All restriction measures below are calculated within each market’s policy-observed subset for exactly this reason.

Measure 2 of 3

Strict restriction — whole-site

Domains whose access policy restricts the whole site.

Strict (whole-site) restriction — count / policy-observed denominator
Australia5.0%
Great Britain4.1%
Singapore4.0%
United States3.8%
India3.5%

Australia 16/317, Great Britain 22/540, Singapore 7/176, United States 24/624, India 13/371. Bars are scaled to a 12% axis so that Strict and Meaningful can be compared visually on the same scale.

The narrow band across five independently sampled markets is the most stable result in the series: in none of them does whole-site restriction reach one domain in twenty.

Measure 3 of 3

Meaningful restriction — primary content and whole-site

Restriction reaching the primary public content a retrieval system would read.

Meaningful restriction — count / policy-observed denominator
Australia11.7%
United States10.7%
Great Britain9.6%
Singapore8.0%
India6.7%

Australia 37/317, United States 67/624, Great Britain 52/540, Singapore 14/176, India 25/371.

Meaningful restriction is inclusive of Strict (Whole-site); it is not additional to it.

Meaningful restriction is between roughly 1.9× and 2.8× the strict figure in each market. The gap between the two is the population of domains that restrict primary content without excluding the whole site — a position invisible to any binary “blocks AI / does not block AI” count.

Methodological context

Unknown-path prevalence

Reported for transparency. Not a principal comparative measure.

Table 2 — Unknown-path prevalence against the policy-observed denominator
MarketUnknown-path domainsPolicy-observed denominatorRate
Australia14631746.1%
United States27962444.7%
Great Britain21054038.9%
Singapore6817638.6%
India13137135.3%

An unknown path is a path referenced by an access policy whose functional content class could not be determined within the bounded protocol. Unknown paths are conservatively assigned to the secondary-content category, which means they contribute to the expanded measure only — not to strict or meaningful.

Expanded restriction is deliberately not used as a headline comparative measure in this publication. Doing so would make unresolved unknown-path treatment carry more interpretive weight than the frozen evidence supports. The principal comparison published here is Coverage → Strict → Meaningful.

Limitations

Interpretation limitations

What these numbers cannot be used to argue.

Applied relevance

What this may mean in practice

For organisations designing websites, access policies and AI-readiness strategies.

These implications follow from the observed distribution of access policy. They are offered as practical reading of the evidence, not as findings in their own right, and none of them is a claim about visibility or commercial outcome.

The five country volumes

C01 — Global Crawler Access Series

Each market was analysed and published independently.

N = 409 · policy observed 317/409 = 77.5%
strict 16/317 = 5.0% · meaningful 37/317 = 11.7%
N = 808 · policy observed 624/808 = 77.2%
strict 24/624 = 3.8% · meaningful 67/624 = 10.7%
N = 754 · policy observed 540/754 = 71.6%
strict 22/540 = 4.1% · meaningful 52/540 = 9.6%
N = 268 · policy observed 176/268 = 65.7%
strict 7/176 = 4.0% · meaningful 14/176 = 8.0%
N = 460 · policy observed 371/460 = 80.7%
strict 13/371 = 3.5% · meaningful 25/371 = 6.7%
Web findings publication
The original AU / US / GB / SG cross-market study, retained as published.
Governance and reproducibility

How this evidence base is held

Frozen sources, recorded hashes, and the records that govern them.

The five principal market results reproduce from frozen source files with zero discrepancies. Six frozen CSV artefacts passed a three-representation byte-portability check on 6 September 2026: worktree SHA256, Git blob SHA256 and git archive SHA256 each agreed with the recorded frozen hash, across 18 of 18 comparisons.

Governance records

Integrity & freeze record
C01 five-market integrity and freeze record — frozen artefact hashes, byte-portability result, and the reproducibility boundary.
US correction record
GDR-004 — United States unknown-path review status correction. Documentary only; no numerical change.
Earlier correction records

India qualification

India access-analysis finality is established. India registry identity-verification finality is not asserted. The frozen India registry v1.0 retains non-final metadata fields relating to identity verification; those fields do not alter the R006-IN crawler-access measurements published here, and the India volume must not be described as a completed final identity-verification dataset.

Publication-render boundary

Research reproducibility is intact. Historical publication-layout reproducibility is not established. PROVENANCE-DEFECT-002 records this as a publication-build archival defect — it is not a numerical or measurement defect, and it does not affect any figure on this page.