Skip to main content

Sitemap Generator

Crawl a site from one URL and build a valid XML sitemap. Honours robots.txt, follows redirects to the final URL, and leaves noindex pages out.

This tool needs our server to process your file. It is sent over an encrypted connection, stored only while it is being processed, and deleted automatically within 60 minutes. We never read the contents or keep a copy.

About the Sitemap Generator

Give a starting URL and the crawler follows the links on the site, up to fifty pages, and writes a valid XML sitemap of what it finds.

Several decisions are made for you, and they are the ones that make a sitemap useful rather than merely valid. Redirects are followed and the final URL is listed — a sitemap full of URLs that redirect is a sitemap that wastes crawl budget. Pages marked noindex are left out, because a sitemap says "please index this" and the two together are a contradiction that Search Console reports as an error. robots.txt is fetched first and honoured, so a page you have already told crawlers to skip does not reappear here.

The limits are deliberate and hard: fifty pages, twenty seconds, one request at a time with a quarter-second pause between them. A crawler without ceilings is a load generator pointed at somebody else's server. When a crawl stops early, the tool says so and reports how many URLs were still queued, rather than presenting a partial result as a complete one.

For a site larger than that, use the crawl to check the structure and generate the full sitemap from your CMS or framework, which knows every URL without having to discover them.

How to use the Sitemap Generator

  1. Enter a starting URL

    Usually the home page. Only pages on the same host are followed — subdomains are treated as different sites, because to a search engine they are.

  2. Set a page limit

    Up to fifty. A lower number is faster and is often enough to check that the structure is what you expect.

  3. Run the crawl

    One page at a time with a pause between them. If the budget runs out, the tool tells you how many URLs were still queued.

  4. Download sitemap.xml

    Put it in your web root and add a Sitemap line to robots.txt pointing at it. Submitting it in Search Console is optional but speeds up discovery.

Frequently asked questions

Why does the crawl stop at fifty pages?

Because each page is a request to somebody else's server, and an uncapped crawler run from a public form is a load generator. Fifty is enough to check a site's structure and to build a sitemap for a small site. For anything larger, generate the sitemap from the system that already knows every URL rather than by discovering them.

Do I need lastmod, changefreq and priority?

Google ignores priority entirely and has said changefreq is not used either. It does use lastmod, but only when it is consistently accurate — a file where every date is today's is treated as noise and disregarded. That is why lastmod is optional here and blank by default: a wrong date is worse than no date.

Why are some pages missing from the sitemap?

Four common reasons: the page is not linked from anywhere the crawl reached, it is blocked by robots.txt, it is marked noindex, or the crawl hit its budget first. The skipped list gives the reason for each URL that was seen and left out.

Should the sitemap list the URL before or after a redirect?

After — the URL that actually serves the page. A sitemap of redirecting URLs makes every crawl two requests instead of one, and search engines treat a sitemap full of redirects as a signal that it is not maintained. That is why the final URL is recorded rather than the one that was linked.

What does the crawl send to my site?

A plain GET request per page with our user agent, no cookies and no referrer. It reads only enough of each page to find its links and title. We record that a crawl happened, and not the URL or who asked.