Zyte analysis finds 75% of top sites use robots.txt, but few name specific crawlers
Zyte published a study showing that three-quarters of the world's top websites publish a robots.txt file, yet most do not name individual crawlers.
Briefing
- Zyte: Zyte published a study showing that three-quarters of the world's top websites publish a robots.txt file, yet most do not name individual crawlers.
Why it matters
Robots.txt remains the web's primary opt-out mechanism, but its limited granularity means many sites cannot selectively block specific bots. This gap drives reliance on more aggressive anti-bot measures and complicates compliance for legitimate scrapers. The finding underscores the tension between the advisory nature of robots.txt and the growing need for precise crawler management.
Sources
- Primary source · Zyte · 2026-09-02 — 75% of the web uses robots.txt - here's how
Watch next
Will major sites begin adopting more granular robots.txt directives or move to alternative access-control mechanisms?