Website Health Issues Benchmark Report

Website health issues benchmark report dashboard showing crawl, performance and content checks

Website health issues benchmark report data shows a web that is improving in some areas while carrying major technical debt in others. In the 2025 HTTP Archive evidence set, fewer than half of mobile websites achieved good Core Web Vitals, only about two-thirds of mobile pages exposed a canonical tag, and just 31% of mobile sites met minimum text-contrast requirements. These are not cosmetic details: they affect whether people and machines can reliably access, interpret, and trust a page.

Benchmark at a glance: 48% of mobile websites passed Core Web Vitals; the median mobile home page reached 2.6 MB; canonical tags appeared on 67% of mobile pages; meta descriptions appeared on 67.2%; structured data appeared on half of home pages; and only 31% of mobile sites passed minimum text contrast. The figures are sourced from the 2025 Web Almanac and are not presented as proprietary Visible Pilot scan results.

Key findings from the website health issues benchmark report

  • Performance is still a coin toss. Only 48% of mobile websites and 56% of desktop websites achieved good Core Web Vitals in 2025.
  • Pages are getting heavier. The median mobile home page reached 2.6 MB, up 8.4% year over year; the median desktop home page reached 2.9 MB.
  • Canonical coverage remains incomplete. Canonical tags appeared on 67% of mobile pages, while 9% pointed to another canonical URL.
  • Descriptions are missing at scale. Title elements appeared on 98.5% of mobile pages, but meta descriptions appeared on only 67.2%.
  • Structured data is not universal. It appeared on 50% of home pages across desktop and mobile; only 2% was added through JavaScript.
  • Accessibility failures are widespread. Only 31% of mobile sites met minimum text-contrast requirements in the automated test.
  • AI-crawler governance is emerging. GPTBot was named in 4.2% of mobile robots.txt files, ClaudeBot in 3.4%, and PerplexityBot in 2.7%.

Taken together, the figures show why a single “website health score” can be misleading. A site may use HTTPS and have a title tag yet still fail on mobile performance, unclear canonicalization, weak metadata, inaccessible presentation, or inconsistent crawler controls. Health is a chain: a serious failure at one stage can reduce the value of strengths elsewhere.

Methodology and research scope

This edition is an evidence synthesis based on the 2025 Web Almanac SEO chapter, its performance analysis, the accessibility chapter, and the page-weight chapter. HTTP Archive reports crawling roughly 17 million sites each month across home and secondary pages. Performance findings also use Chrome User Experience Report data where an origin has sufficient real-user observations.

The research period is the 2025 Web Almanac cycle, published in January 2026 and reviewed for this Visible Pilot report on 6 August 2026. Each percentage retains the source chapter’s own eligible-page or eligible-site denominator. We have not averaged unrelated percentages, converted them into a proprietary score, or claimed that every metric used all 17 million sites.

Important limitation: Automated benchmarks identify patterns, not the full quality of an individual site. They cannot reliably judge whether copy is accurate, whether a page fully answers its audience, or whether every accessibility criterion is met. Use the numbers to prioritize tests, then inspect real pages and server evidence.

Website benchmark methodology sampling pages, running technical checks and grouping health findings
Transparent benchmarks preserve the source, test conditions, field definitions and denominators.

Overall results

Website health signal2025 benchmarkWhat it means
Good Core Web Vitals48% mobile; 56% desktopAbout half of measured sites still miss the complete good threshold.
Median home-page weight2.6 MB mobile; 2.9 MB desktopPage weight continues to rise and can compound performance and rendering risk.
Canonical tag present67% of mobile pagesRoughly one-third lacked this consolidation hint in the measured set.
Meta description present67.2% of mobile pagesBasic descriptive metadata remains incomplete across many pages.
Structured data on home pages50% mobile and desktopHalf of home pages used no detected structured-data format.
Minimum text contrast passed31% of mobile sitesAccessibility and readability failures remain widespread.
GPTBot named in robots.txt4.2% of mobile sitesExplicit AI-crawler rules are growing, but are not yet common.

Breakdown by website health layer

1. Delivery and performance

Core Web Vitals improved year over year, but the overall pass rate still leaves more than half of measured mobile sites outside the complete “good” category. Rising page weight adds pressure: more bytes can mean more network work, image decoding, script execution, and rendering. Weight alone does not prove slowness, but it is an actionable risk indicator—especially on lower-powered mobile devices.

2. Crawl and index signals

Canonical adoption rose to 68% of desktop pages and 67% of mobile pages. That improvement still leaves substantial room for missing, conflicting, or unintended consolidation signals. The SEO chapter also found HTTPS on 91.5% of mobile pages, showing that encryption is now a baseline rather than a differentiator. Health audits should therefore move beyond “does HTTPS exist?” to response codes, redirect consistency, robots directives, rendered canonicals, and internal discovery paths.

3. Machine-readable meaning

Metadata and structured data reveal an uneven pattern. Titles are nearly universal, but one in three mobile pages lacks a detected meta description. Structured data appears on only half of home pages. Neither feature guarantees rankings or AI citations, but both can reduce ambiguity when they accurately describe visible content. A useful audit checks whether markup matches the page, not merely whether a script outputs schema.

4. Accessibility and readability

The 31% mobile contrast pass rate is the sharpest warning in this benchmark. Poor contrast directly harms people, and other accessibility problems can also weaken content clarity, navigation, and media understanding. Automated tools detect only part of the problem, so a passing score is a starting point—not proof of full accessibility.

5. AI crawler controls

Named AI agents are appearing more often in robots.txt files. GPTBot mentions grew by about 55% year over year, while ClaudeBot mentions nearly doubled. That does not mean those sites blocked the bots: a named rule may allow, restrict, or selectively govern access. The correct test is to read the matching rule, verify the live response, and confirm that the intended content is actually retrievable.

Which website health issues cluster together?

The source reports measure separate signals, so the following relationships are practical inferences rather than claims of statistical causation. In real audits, three clusters deserve attention:

  • Heavy delivery cluster: large images, third-party scripts, slow rendering, and weak mobile Core Web Vitals often share the same templates.
  • Ambiguous discovery cluster: missing canonicals, conflicting redirects, incomplete internal links, and unclear index directives can make the preferred URL difficult to establish.
  • Weak interpretation cluster: missing descriptions, thin headings, inaccurate structured data, inaccessible images, and poor contrast make a page harder to understand and use.

Trace representative URLs through discovery, crawl access, index eligibility, rendering, comprehension, and citation readiness. This sequence makes the root failure visible instead of producing a long list of disconnected warnings. For a practical implementation method, use our AI search readiness checklist for business websites.

Website health audit path from discovery and crawling to indexing, comprehension and citation readiness
A healthy page must pass through discovery, access, indexing, comprehension and citation-readiness checks.

What the benchmark means for website owners

  • Fix blockers before polish. A blocked, redirected, noindexed, or broken page needs attention before wording tweaks.
  • Test templates, not only the homepage. Compare the homepage, a commercial page, and a knowledge article because their delivery patterns may differ.
  • Keep raw evidence. Save response headers, rendered HTML, screenshots, directives, and test dates so changes can be retested.
  • Separate eligibility from visibility. A technically healthy page can still lack relevance, authority, or useful evidence.
  • Prioritize by business impact. Start with issues affecting revenue pages, frequently linked resources, and pages intended to earn search or AI citations.

Recommended first pass: Check HTTP status, robots access, meta robots, canonical URL, rendered main content, internal links, sitemap inclusion, Core Web Vitals, page weight, heading clarity, image alternatives, contrast, and structured-data validity. Record pass, warning, fail, evidence, owner, and retest date for every check.

Reproducibility and calculation notes

To reproduce the benchmark figures, use the linked 2025 Web Almanac chapters and open the data or query attached to each figure. HTTP Archive publishes its analysis process and SQL resources publicly. For your own site, keep the test device, user agent, location, URL set, and audit version stable. Report both the numerator and denominator—for example, “24 of 40 tested URLs passed”—rather than publishing only a percentage.

Do not combine crawlability, speed, accessibility, schema, and AI mentions into one unexplained grade. If you create a score, publish the rule for every component and retain the underlying findings. Our guide to AI visibility metrics that actually matter explains how to separate technical eligibility, mentions, citations, and business outcomes.

Update policy

Visible Pilot will review this page when a new Web Almanac edition is released or when a source publishes a material correction. Future editions should retain the same metric definitions where possible, show year-over-year changes, and flag any methodology break. A future proprietary Visible Pilot benchmark will be labeled separately and will publish its sample rules, anonymization method, field definitions, and limitations before presenting results.

Explore the benchmark data

The complete public charts, source notes, and reproducible queries are available through the linked HTTP Archive chapters. Start with the SEO chapter for canonicalization, metadata, structured data, and crawler controls; use the performance and page-weight chapters for delivery signals; and use the accessibility chapter for automated readability findings.

Next step: Use this website health issues benchmark report as a prioritization baseline, then test your own important URLs. Benchmark percentages describe the web; your server responses, rendered pages, and real-user experience determine what your team should fix first.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *