Skip to main content

Robots.txt

concept

Robots.txt is a site-root text file that communicates crawler access rules under the Robots Exclusion Protocol.

Status: published
Last reviewed: 2026-09-12

Technical explanation

Rules identify user agents and paths that are allowed or disallowed for crawling and may declare sitemap locations. The file is publicly accessible, path matching is sensitive to syntax, and support for directives varies among crawlers.

Business relevance

A correct robots.txt helps manage crawler access and protect crawl capacity from low-value paths while pointing crawlers to sitemaps.

Implementation example

A site allows public pages, disallows an internal API path, and declares the canonical XML sitemap while keeping CSS and JavaScript required for rendering accessible.

Limitations and common misconceptions

Robots.txt is not an access-control or confidentiality mechanism and does not itself prevent indexing. Blocked URLs may still appear without content, and sensitive resources require authentication or removal.

Discuss your systems

Need help implementing or evaluating this concept? Keenfunnel designs connected AI, automation, and data systems.

Book a discovery session