guides
Robots.txt for AI Crawlers: Search, Training and Bot Access
How should robots.txt handle AI crawlers without accidentally blocking search discovery?
Read guide →Knowledge base
Practical guides for crawler rules, indexing controls, host scope, parameter URLs, deployment behavior, and robots.txt troubleshooting.
guides
How should robots.txt handle AI crawlers without accidentally blocking search discovery?
Read guide →guides
Can robots.txt prevent a page from being indexed, and when should noindex be used instead?
Read guide →guides
Does crawl-delay work in robots.txt, especially for Googlebot?
Read guide →guides
How do robots.txt wildcards and conflicting Allow or Disallow rules actually match?
Read guide →guides
What happens when robots.txt returns 404, 403, 429, 500, or another HTTP status?
Read guide →guides
Where should we add a Sitemap directive in robots.txt, and how do we verify multiple sitemap references?
Read guide →guides
Does the root-domain robots.txt apply to subdomains, HTTP, HTTPS, or another port?
Read guide →guides
How should we control query-parameter and faceted URLs without blocking valuable pages by accident?
Read guide →guides
What happens when robots.txt is too large or not encoded as UTF-8, and how do we verify the live response?
Read guide →guides
Should we use robots.txt, a meta robots tag, X-Robots-Tag, canonical, or authentication for this URL?
Read guide →guides
How long can robots.txt stay cached, and how do we tell a crawler cache from a failed production deployment?
Read guide →guides
How should we keep a staging or preview site out of search without treating robots.txt as a security control?
Read guide →