C01 — Global Crawler Access Series · India

AI crawler access in India 2026

460 Indian business domains across ten sectors, read by functional purpose on a frozen instrument. The fifth and final market of the C01 series.

Status
Verified and frozen
Sample
460 domains · 10 sectors
Run
R006-IN · 460/460
Instrument
C01 v1.5
3.5%
of policy-observed Indian domains exclude AI retrieval crawlers from the whole site
13 of 371 domains. Meaningful restriction — primary content or whole site — is 25 of 371, or 6.7%. Both are the lowest of the five C01 markets. This is a property of this sample on its scan date, not an estimate for Indian business generally.

What this volume measures

Declared crawler access, and nothing beyond it.

This volume records what 460 Indian business domains declared about AI retrieval crawler access: their published robots.txt policy, meta-robots directives and observed response behaviour, captured in a single frozen run.

It does not measure AI visibility, AI ranking, AI citation, AI recommendation, digital authority or commercial performance. Whether a crawler may fetch a page is a different question from whether an AI system subsequently cites it.

Sample and denominators

Two denominator levels, both stated explicitly.

The India sample is 460 organisations across ten sectors, 46 per sector: accounting and finance, building trades, education and training, healthcare, hospitality and tourism, legal, professional services, real estate, retail and e-commerce, and technology and SaaS.

Crawler-access policy was analytically determined for 371 of the 460 domains. The remaining 89 returned no readable policy and are excluded from every restriction denominator rather than counted as open.

Meaningful restriction is inclusive of Strict (Whole-site); it is not additional to it.

80.7%
Policy observed
371 of 460 domains — the highest coverage of the five markets
3.5%
Strict restriction
13 of 371 policy-observed domains exclude the whole site
6.7%
Meaningful restriction
25 of 371 — primary content or whole site, inclusive of strict

Access classification

The five-class functional-purpose taxonomy, applied to all 460 rows.

India — access class distribution, run R006-IN
Access classDomainsShare of sampleShare of policy observed
Fully open20043.5% (200/460)53.9% (200/371)
Secondary paths restricted12126.3% (121/460)32.6% (121/371)
Functionally open255.4% (25/460)6.7% (25/371)
Primary content restricted122.6% (12/460)3.2% (12/371)
Whole site excluded132.8% (13/460)3.5% (13/371)
No determinable policy8919.3% (89/460)

Strict restriction is the whole-site class alone: 13/371 = 3.5%. Meaningful restriction is primary content plus whole site: (12 + 13)/371 = 25/371 = 6.7%.

Most restriction observed in India is neither strict nor meaningful: it is secondary and functional — administrative paths, search endpoints and checkout flows — which leaves primary editorial content retrievable.

Unknown-path context

Methodological context, not a headline measure.

131 of 371 policy-observed Indian domains (35.3%) declare at least one path whose functional purpose cannot be determined from the policy text alone. Under the frozen taxonomy those paths default to SECONDARY — the conservative treatment — and are logged for review.

India records the lowest unknown-path prevalence of the five markets. Unknown-path treatment affects only the expanded measure; strict and meaningful restriction are unaffected.

Where India sits

The same three measures, across all five C01 markets.

Five-market comparison, denominators inline
MarketPolicy observedStrictMeaningful
Australia77.5% (317/409)5.0% (16/317)11.7% (37/317)
United States77.2% (624/808)3.8% (24/624)10.7% (67/624)
Great Britain71.6% (540/754)4.1% (22/540)9.6% (52/540)
Singapore65.7% (176/268)4.0% (7/176)8.0% (14/176)
India80.7% (371/460)3.5% (13/371)6.7% (25/371)

India shows the highest policy-observed coverage and the lowest restriction on both measures. The series describes this ordering; it does not attribute it to geography. Sample construction, sector mix, hosting patterns and platform defaults differ between markets and are not controlled for.

Read the five-market comparative findings →

Method and provenance

What was run, and against what.

Run R006-IN

Sample
460 organisations, ten sectors, 46 per sector. Sample construction was documented before collection began.
Completion
460 of 460 rows collected under the locked protocol. No row was dropped.
Instrument
C01 crawler v1.5, unmodified before and after the run. Methodology →
Taxonomy
Five-class functional-purpose access taxonomy, frozen at v1.5.
Crawl input
The frozen India registry v1.0, byte-identical to the crawl input at run time.
Verification
Principal values independently recomputed from the frozen dataset with zero discrepancies.
Scope boundary — registry finality

India access-analysis finality is established. India registry identity-verification finality is not asserted.

The frozen India registry retains non-final identity-verification metadata: root-domain confirmation was not adjudicated, a robots-reachability field was deliberately left unwritten pending a definition fix, and rows remain marked provisional for inclusion purposes.

Those are registry-verification objects. They do not alter the R006-IN crawler-access measurements reported here, because verification status was never an analytical inclusion criterion for the access study. This volume does not describe the India registry as a completed final identity-verification dataset.