Web & SEO

Sitemap

An organised list of site pages. An HTML sitemap can help visitors navigate; an XML sitemap supplies URLs for crawlers. Neither replaces useful internal links, and listing a page does not guarantee that a search engine will index it.

How it works

A sitemap is an inventory of the URLs you want search engines to discover. It complements normal navigation and internal links. For a multilingual site, each complete language page has its own canonical URL. A sitemap index can point to several smaller files, for example separate inventories for articles and glossary terms.

A realistic example

Imagine a glossary with 337 terms in four languages. The complete inventory could contain 1,348 URLs, but drafts, redirects and pages marked noindex must be excluded. A count in Search Console can lag behind a newly uploaded file; compare the actual XML first rather than assuming every missing count is a missing page.

What to check

Open the sitemap itself and check that it returns XML with a successful response. Count distinct loc values, compare them with the approved page inventory, and sample the URLs for status, robots directives and self-referencing canonical tags. Check the child files too when the submitted file is a sitemap index.

Limits and next steps

Submission does not guarantee crawling or indexing. If a listed URL is discovered but not indexed, investigate its content, canonical and internal links. Resubmitting an unchanged file repeatedly does not improve a thin page. Use real modification dates and remove duplicate URL forms instead of inflating the inventory.

Additional examples and checks

XML is the machine-readable format of this sitemap. A urlset contains page URLs; a sitemapindex contains links to sitemap files. Every loc value should be an absolute canonical address. Special characters must be XML-escaped. A file that looks readable in a browser can still contain invalid nesting or an incorrect namespace.

A site may split its inventory into pages-sitemap.xml and glossary-sitemap.xml. The index points to both files. If the second file returns an HTML error screen with a 200 status, the index itself may load normally while the glossary inventory cannot be processed. Inspect each response, not only the top-level address.

Technical sources

← All glossary terms