Skip to main content

Robots.txt Generator

Build a robots.txt, then test any URL against it to see exactly which rule applies and why. Covers AI crawlers, wildcards and the precedence rules.

Rules

Editable — paste an existing file here to test it.

Test a URL

    This tool runs entirely in your browser. Nothing you enter is sent to our servers, so there is nothing for us to store or see.

    About the Robots.txt Generator

    Build a robots.txt from a form, or paste one you already have — then test a URL against it and see exactly which rule decides the answer.

    The tester is the part that matters. Generating the file is string concatenation; knowing whether it does what you meant is not, because the precedence rules are not what most people assume:

    • The longest matching pattern wins, not the first one in the file. Allow: /admin/public/ beats Disallow: /admin/ wherever it appears.
    • Only the first matching user-agent group applies. If Googlebot has its own group, it never reads the * group at all — including the Disallow lines you assumed covered everyone. This is how a site ends up blocking Bing and not Google while believing it blocked both.
    • Disallow: allows everything; Disallow: / blocks everything. One character apart, opposite meanings.

    The AI crawlers are in the list of known agents, since that is what most people are here to block: GPTBot, ClaudeBot, PerplexityBot and CCBot alongside the search engines. And the review panel says the thing that gets missed most often — robots.txt stops crawling, not indexing. A blocked page can still appear in results if other sites link to it, because the crawler cannot read a noindex tag on a page it was told not to fetch.

    How to use the Robots.txt Generator

    1. Add your rules

      Pick a crawler, then list the paths it may and may not fetch. Add more groups for crawlers that need different rules.

    2. Add your sitemap

      A full URL, including https://. This is the cheapest way to tell every crawler where your sitemap is, including ones you have never registered with.

    3. Test a URL against it

      Enter a path and a crawler name. The tester names the exact rule that decides the answer and the group it came from.

    4. Save it to your web root

      Download the file and put it at the root of your domain — it must be at /robots.txt and nowhere else to have any effect.

    Frequently asked questions

    Does robots.txt stop a page appearing in Google?

    No. It stops crawling, not indexing. If another site links to a blocked page, Google can still list it — usually with no description, because it was not allowed to fetch the content. To keep a page out of results, let it be crawled and serve a noindex meta tag or X-Robots-Tag header. Blocking it in robots.txt actually prevents the noindex from ever being seen.

    Which rule wins when Allow and Disallow both match?

    The one with the longer path pattern, regardless of the order they appear in the file. Allow: /admin/public/ beats Disallow: /admin/ because it is more specific. If two matching patterns are exactly the same length, Allow wins. This is Google's documented behaviour and the other major crawlers follow it.

    How do I block AI crawlers?

    Add a group for each one by name — GPTBot for OpenAI, ClaudeBot for Anthropic, PerplexityBot, CCBot for Common Crawl — with Disallow: /. They are all in the crawler list here. Be aware that robots.txt is voluntary: it works because reputable crawlers choose to honour it, and it stops nothing that does not.

    Why is my Disallow being ignored by one crawler?

    Almost certainly because that crawler has its own group somewhere in the file. A crawler uses only the first group whose user-agent matches its name, and a name match takes precedence over the wildcard. Once Googlebot has a group of its own, nothing in the * group applies to it at all — which is what the tester is for.

    Is anything sent to a server?

    No. The file is built and tested entirely in your browser. Nothing about your site structure, including the paths you are choosing to block, is transmitted or stored.