Sitemaps Explained: How Search Engines Actually Use Your sitemap.xml
A sitemap is one of the easiest SEO wins available and one of the most commonly broken. New sites forget to create one; established sites break it with dead URLs, bad dates, and formatting errors they never notice. This is the complete, practical picture: what a sitemap is, how crawlers use it, the format rules that matter, and how to build a correct one in two minutes.
What a Sitemap Is (and Isn't)
A sitemap is an XML file that lists your site's important URLs and, for each one, tells the crawler when it was last modified, how often it typically changes, and how important you consider it relative to the rest of the site. It's a table of contents for robots, not for humans.
It is not a ranking button. A sitemap doesn't make Google index more of your pages or rank them higher — a crawler that wants your content will find it through links. What a sitemap does is remove friction: it guarantees the crawler knows your full URL inventory quickly, finds new pages faster, and spends its crawl budget on your real pages instead of guessing. On a small or new site, that "faster discovery" is the entire game — which is why brand-new sites especially need one submitted on day one.
The Format, Field by Field
<loc>— the URL. Absolute URL (https://), exactly as it should be indexed. No trailing-slash surprises, no parameters you don't want crawled, no URLs that 404.<lastmod>— last modified date. ISO format (2026-09-20). The #1 field sites get wrong: a lastmod that claims every page changed today teaches the crawler your dates are noise. Only list a lastmod you actually mean, and update it when content genuinely changes.<changefreq>— how often you expect changes. "weekly," "monthly," "yearly." It's a hint, not a promise — but "daily" on a static page is a hint you're not reading your own site.<priority>— relative importance. 0.0 to 1.0. Homepage 1.0, key pages 0.8–0.9, legal pages 0.4. Again a hint, but internally consistent priorities read as intentional.
What Belongs in It (and What Doesn't)
Include: every real page you want indexed — tools, articles, about, contact. On a site like ours that's 30+ URLs, all of them earn their place.
Exclude: pages blocked by robots.txt (listing them confuses the crawler), tag/archive pages with thin duplicates, login pages, PDFs-as-pages you don't want in results, and — the classic error — anchor fragments. A URL like site.com/#section isn't a separate page to a crawler; the page is site.com. Sitemaps of hash-fragments are how "I have a sitemap" becomes "my sitemap is worthless."
The Three Checks Before You Submit
- Every URL returns 200. One dead URL in a sitemap is a small wound; fifty is a credibility problem with the crawler. Test them, or build the list from your live pages.
- Every URL is canonical. If your page has a canonical tag pointing elsewhere, the sitemap should list the canonical, not the variant. www and non-www should not both appear.
- The file is referenced in robots.txt. One line —
Sitemap: https://yoursite.com/sitemap.xml— and most crawlers will find it even without a manual submission. Manual submission in Search Console/Bing is still the faster on-ramp for new URLs.
Large Sites: Sitemap Indexes
One sitemap file has hard limits: 50,000 URLs and 50 MB uncompressed. Beyond that, use a sitemap index — a small XML file that lists other sitemap files, each of which obeys the limit. The index file is the one referenced in robots.txt, not the individual sitemaps. Most sites never hit this; if you're running a product catalog or a news archive, you probably will.
Mistakes That Make a Sitemap Worthless
- Listing URLs that 404. A sitemap is a promise to the crawler. Broken promises in it teach it to discount the whole file.
- Hash fragments.
/page#sectionis not a page. List/pageonce, and let on-page navigation handle the section. - Tracking parameters and variants.
?ref=twitterand the plain URL are the same page to the crawler. Canonicalize first, then list the canonical. - Never updating. A sitemap from launch day, eight months stale, is worse than none — it claims an inventory your site no longer has.
Build Yours in Two Minutes
Our Sitemap Generator does exactly this: paste your URLs one per line, set the changefreq and priority for the batch, and it outputs valid, properly-escaped XML — ready to download and deploy, or copy straight into your editor. Then submit the URL in Google Search Console and Bing Webmaster Tools, and note the date. Two minutes now, and your next ten pages get discovered days faster instead of weeks.