A structured benchmark of 409 publicly identifiable Australian business websites across ten industry groups, measuring AI-crawler access by the functional purpose of what is restricted — not by a single block/open count.
A binary count records 42.6% of policy-observed Australian domains as “blocking AI.” Read by the functional purpose of the restricted path, whole-site exclusion is just 5.0% and restriction reaching primary public content is 11.7%.
Most restriction falls on operational and secondary paths — carts, search endpoints, APIs, admin routes — not on the primary content an AI system retrieves to answer a question.
The same robots.txt policy data, read at three levels of stringency. Each measure includes the one above it and adds a less-substantive category. Strict and meaningful are final; the expanded measure is a conservative upper bound pending the unknown-path review.
The central finding: whole-site exclusion of AI retrieval crawlers is uncommon (5.0%), and meaningful restriction of public content is modest (11.7%). Most robots.txt activity classified as restrictive concerns functional, secondary, or infrastructure-layer paths rather than outright exclusion of AI crawlers from substantive content.
Every policy-observed domain is classified by its highest substantive restriction, from fully open to whole-site blocked. Of 317 policy-observed Australian domains:
Just under half of policy-observed domains (48.9%) are fully or functionally open — they restrict, at most, operational infrastructure. A further 39.4% restrict only secondary content. Primary-content restriction (6.6%) and whole-site blocking (5.0%) together account for the 11.7% meaningful-restriction figure.
Crawl-based measurement is subject to transient failure. The study used a bounded three-attempt recovery protocol with randomised order, and the Australian measurement was crawled twice on independent connections.
Independently replicated. Two crawls on separate network connections produced an identical policy denominator (317) with identical strict and meaningful restriction membership — the same 16 and same 37 domains. The wired run is canonical (c01_au_409_v15_wired); the mobile run is retained as replicated verification evidence. This establishes the result as reproducible rather than connection-dependent.
Declared robots.txt policy is separated from infrastructure non-response. A 403, timeout or non-resolving response is an access outcome, not evidence of crawler policy. These 92 domains are reported here and excluded from every policy-layer figure.
Why this matters: a domain that denies the crawler at the infrastructure layer has not expressed an AI-crawler policy — it has prevented one from being read. Reporting these 92 domains separately keeps the policy-layer figures based only on directly observed robots.txt behaviour. Across the two independent crawls the excluded set was identical in membership.
Of the 135 domains disallowing at least one retrieval crawler under the broad binary indicator, the source of the directive was classified into three categories.
Where Australian sites do restrict, the restriction is mostly a deliberate choice rather than a platform default — though a substantial share is indeterminate as to origin and is reported as such.
Share of policy-observed domains disallowing at least one retrieval crawler from at least one path — the broad binary indicator, reported for sector comparison. Rates rest on small per-sector denominators (21–44) and should be read with that in mind.
Real Estate (57.1%) and Accounting & Finance (51.6%) show the highest broad-disallow rates; Professional Services (22.6%) and Building & Trades (31.2%) the lowest. These are the broad binary indicator only — necessarily higher than the strict and meaningful measures because they count functional-path disallows.
Block rates across the 21-crawler panel, of 317 policy-observed domains. Retrieval crawlers (Group A) determine whether AI systems can access content to answer queries; training crawlers (Group B) gather data for model training. The two groups are reported separately and never combined.
Group A retrieval-crawler block rates fall in a narrow band (38.2–41.0%): restriction is largely applied across retrieval crawlers as a group rather than targeted at individual operators.
Roles. This study is published by the Periodic Table of Digital Authority (PTODA), the publisher and steward of the PTODA research methodology. It was conducted using the PTODA C01 Crawler v1.5.1, a deterministic robots.txt reference instrument, under PTODA C01 Crawler Methodology v1.5. The sample was constructed from named public sources using the published sampling standard. Commercial relationships played no role in domain selection, inclusion, exclusion, analysis, or interpretation. The methodology is fully documented and designed to support independent reproduction using the published specification and frozen datasets. This study publishes aggregate, anonymised findings only. No named individual site results are published.
Attribution chain: Douglas Lord (researcher and author) · Periodic Table of Digital Authority (publisher and methodology steward) · PTODA C01 Crawler v1.5.1 (research instrument) · Digital Dominator Pty Ltd ABN 28 616 931 116 (operating entity).
Intellectual property notice: This study, its methodology, findings, data, and all associated content are the original work of Douglas Lord and the property of Digital Dominator Pty Ltd (ABN 28 616 931 116). The Periodic Table of Digital Authority™ is a coined framework and trade mark pending (TM 2644497). AUTHORITY44™ is a trade mark pending (TM 2643932). All rights reserved.