An AI crawler access log study should answer a simple question with defensible evidence: when recognized AI crawlers request public pages, how often do they receive usable content, and what prevents access when they do not? This July 2026 research specification explains how Visible Pilot will measure that question across websites without turning user-agent strings into misleading statistics.
Evidence status: this page is a pre-registered study method, not a completed benchmark. No access rates or sample findings are claimed before verified log records are collected.
The distinction matters. A server log can show that a request reached an origin or edge, but it cannot by itself prove that an AI platform indexed, understood, cited, or recommended the page. Likewise, a request that says it is “GPTBot” or “ClaudeBot” is not automatically genuine. A credible report must verify identity, preserve denominators, and separate access, delivery, and downstream visibility.
Key findings: what this edition can and cannot report
The current finding is methodological: there is no supplied cross-site log dataset from which honest percentages can be calculated. Publishing invented pass rates would make the article look authoritative while weakening the very dataset Visible Pilot wants others to cite. Therefore, this edition establishes the measurement rules and the fields future releases will use.
- No fabricated denominators: every future percentage will state the number of eligible requests, sites, and URLs behind it.
- Identity before classification: recognizable user-agent text is a candidate signal, not proof of crawler ownership.
- Access is layered: robots directives, CDN/WAF decisions, HTTP delivery, redirects, rendering, and content quality are recorded separately.
- Purpose matters: search crawlers, training crawlers, and user-triggered agents should not be merged into one “AI bot” bucket.
- Absence is inconclusive: no log entry may mean no crawl demand, insufficient retention, edge-only logging, or a collection gap—not necessarily successful blocking.
- Citation is downstream: a successful 200 response is necessary for many crawl paths, but it does not guarantee discovery or a citation.
Methodology for the AI crawler access log study
The study unit will be a verified crawler request to an eligible public HTML URL during a fixed observation window. Participating sites must provide logs that include timestamp, host, request path, method, status, user agent, source IP or verified-bot signal, bytes sent, response time, and the security action applied. Personally identifying query strings and cookies should be excluded or removed before analysis.
A site enters the sample only when its collection window, timezone, retention policy, CDN or proxy layer, and known logging gaps are documented. The study will avoid convenience sampling claims: early Visible Pilot users may not represent the wider web, so results will be labeled by acquisition source, platform, size band, and site type.
Pass condition: a verified crawler requests an eligible canonical HTML URL and receives the intended final response with a 2xx status, non-empty content, no challenge page, and no conflicting index-control signal.
Inclusion and exclusion rules
- Include public, canonical HTML pages requested with GET or HEAD during the stated period.
- Deduplicate retries using crawler, verified identity, URL, response status, and a declared time window.
- Keep robots.txt requests as a separate diagnostic dataset rather than treating them as content-page successes.
- Exclude spoofed or unverifiable user agents from operator-level statistics; report them separately as unverified claims.
- Exclude admin, account, checkout, preview, staging, and intentionally private URLs from the accessibility denominator.
- Record 3xx, 4xx, 5xx, 402, challenge responses, connection failures, and origin timeouts as distinct outcomes.
Overall results: the reporting template
When the first dataset is available, the headline table will show four denominators: participating sites, eligible URLs, verified crawler requests, and unique crawler–URL pairs. Results will report successful delivery, redirect loops, robots exclusions, WAF or bot-management blocks, rate limits, authentication challenges, server errors, soft blocks returning 200, and inconclusive records.
The report will publish counts beside percentages. For example, a statement must read in the form “x of n verified requests,” not “most crawlers passed.” Confidence intervals may be added where sampling supports them, but statistical precision will not repair biased site recruitment or missing edge logs.
How results will be broken down
Aggregated access rates can hide useful differences. Future editions will segment results only when each group has enough data to avoid disclosing a participant or producing unstable comparisons.
- Site type: SaaS, ecommerce, publishing, local business, marketplace, and documentation.
- Platform: WordPress, Shopify, Wix, Webflow, custom applications, and headless stacks.
- Delivery layer: direct origin, major CDN/WAF, managed hosting security, and multi-proxy configurations.
- Crawler purpose: AI search, model training, user-triggered fetch, and mixed or unknown behavior.
- URL class: homepage, product or service page, article, documentation, category, and JavaScript-heavy application route.
Failure-pattern analysis
The most valuable output is not a league table of crawler brands. It is a map of failures that occur together. A robots.txt allow rule may coexist with a Cloudflare block. A 200 response may contain a managed challenge instead of the article. A crawler may reach a redirected URL but never receive the canonical page. A successful HTML response may still contain a noindex directive or almost no server-rendered text.
The analysis will therefore use multi-label failure clusters rather than assigning every request one oversimplified cause. Each record can carry separate flags for robots policy, identity confidence, edge action, HTTP outcome, redirect integrity, HTML payload, index controls, and render dependence.
Common interpretation mistake: matching a user-agent substring and calling it a verified visit. Spoofed strings, shared infrastructure, proxies, and changing IP ranges make identity verification essential.
Research workflow: from raw logs to defensible evidence
The workflow below starts with raw server or edge records, verifies the request identity, normalizes consistent fields, assigns pass/warning/fail outcomes, and only then calculates charts. This order prevents attractive dashboards from hiding weak inputs.

What website owners should do now
You do not need to wait for a benchmark to improve your own evidence. Export a short log window, list the AI-related user agents you observe, and verify them with the operator’s published method where available. Then sample representative URLs and compare robots.txt policy with the actual edge or origin outcome.
- Confirm that the crawler is allowed by the robots policy you intended to publish.
- Check CDN, firewall, bot-management, rate-limit, and challenge events for the same timestamp.
- Follow redirects to the final canonical URL and inspect the returned content type and body.
- Record status codes and response bodies; do not treat every 200 as usable HTML.
- Retest after the smallest safe change and preserve before-and-after evidence.
OpenAI publishes crawler identities and IP information for its bots; Anthropic documents its bot names and crawler IP guidance; Perplexity publishes separate identities for automated search crawling and user-triggered fetching. Cloudflare also distinguishes Search, Agent, and Training behavior and notes that stronger verification can rely on published IPs, reverse DNS, or cryptographic Web Bot Auth.
Reproducibility: fields and calculation notes
The planned public data dictionary will include anonymized site ID, site segment, timestamp bucket, crawler operator, crawler purpose, user-agent token, identity method, identity confidence, normalized URL class, robots outcome, edge action, HTTP status, redirect count, content type, payload class, index-control flag, and final outcome. Raw IP addresses and sensitive URLs should not be published.
Each chart will link to a calculation note defining its numerator, denominator, exclusions, deduplication rule, timezone, and software version. Code or formulas used to derive published tables should be versioned. Any manual classifications will be reviewed against a written rubric, with disagreements recorded instead of silently resolved.
Known limitations
- Participating Visible Pilot users may be more technically engaged than the average website owner.
- Edge providers expose different log fields and verification signals.
- Short observation windows may miss low-frequency crawlers.
- Operator IP lists, user agents, and policies change over time.
- Access logs cannot prove indexing, retrieval in every session, citation, referral quality, or model training.
- Privacy-safe aggregation may limit detailed subgroup reporting.
Update policy
The research date, collection window, schema version, and crawler-reference date will appear at the top of every edition. Future editions will preserve the original outcome definitions so year-over-year comparisons remain possible. If a definition changes, the report will either recalculate earlier data or clearly mark a series break.
Corrections will be logged with the affected chart, reason, date, and impact. Crawler identities will be checked against current operator documentation at publication time. The benchmark will never silently replace preliminary values with final ones.
Downloadable charts and dataset
The planned release package will contain reusable WEBP or PNG charts, an anonymized CSV, a field dictionary, inclusion rules, calculation notes, and a changelog. Until real records satisfy the methodology above, there is no “full benchmark dataset” to download. That is a deliberate quality control, not a missing statistic.
Want to contribute to the first defensible edition? Preserve a privacy-safe sample of edge or server logs, document the observation window and stack, and join the Visible Pilot early-access list.
Frequently asked questions
What is an AI crawler access log study?
It is an analysis of verified crawler requests recorded by web servers, CDNs, or security layers. A reliable study measures real delivery outcomes and failure causes while keeping crawler access separate from indexing and citation.
Can a user-agent string prove a crawler is genuine?
No. It is easy to copy a user-agent string. Use the operator’s published IP ranges, reverse-DNS process, Web Bot Auth signature, or a trusted verified-bot signal where available.
Does a 200 response mean the crawler read the page?
Not necessarily. A 200 response can contain a challenge, empty shell, error message, or thin client-rendered document. Inspect the content type, payload, canonical target, index controls, and body.
Why not publish illustrative percentages?
Illustrative numbers are often copied as facts after their label disappears. A benchmark intended to earn citations should wait for real denominators and publish calculation notes.
Where should I start?
Read the AI Crawler Access and Blocking Guide, review the AI crawler user agent list, and prepare for the forthcoming free AI crawler access tester.

Leave a Reply