Google's Search Relations team dedicated a full episode of its Search Off the Record podcast to sitemaps, with Martin Splitt asking John Mueller the questions SEOs keep sending in. Most of it confirms what technical SEOs already suspect, but a few details are useful to have in one place. Do you need one?
Mueller said smaller sites can be crawled without a sitemap, but his advice is to leave it on anyway. Most CMS platforms and static hosts generate one by default, there is no harm in having it, and it is already in place when the site grows. News sites and large ecommerce sites benefit the most, since a sitemap tells Google a page changed before normal crawling would have found it.
What Google still reads
The priority and change frequency fields were dropped because every SEO marked every URL as top priority and always fresh.
The URL and the last modified date are what count, and the date should reflect a real, major edit. If every URL carries today's date Google will simply stop trusting the field. Mueller was clear this is not treated as spam.
List only the canonical version of each URL, the way you want it indexed. Adding date stamps to URLs for cache busting does not help, and a URL listed in the sitemap is a little more likely to be picked as canonical.
RSS, HTML and llms.txt
An RSS feed can be submitted in Search Console as a sitemap. It only lists recent URLs, so it helps a system find fresh changes without opening hundreds of sitemap files. An HTML sitemap is built for visitors and does not replace the XML file, and llms.txt cannot be processed as a sitemap at all. Mueller described the hope around llms.txt as bigger than the reality and said not to make it a priority, though it is fine if your CMS creates it.
The file limits Mueller gave from memory are 50,000 URLs and 50 MB uncompressed per file, with index files nesting only once. Check the current documentation before relying on those numbers.
Naming and AI crawlers
A sitemap can be named anything if you submit it in Search Console, but then Bing and other systems will not find it. AI crawlers have no console to submit to, so sitemap.xml or an RSS feed is the safer choice. Mueller said he has seen AI crawlers fetch both in his own server logs.
Why Search Console says "couldn't fetch"
When the sitemap is valid, public and listed in robots.txt, there are two usual causes. Google may be too busy with other crawling on the host, or it may see little reason to crawl the site because of how it judges overall quality. In the second case the fix is the content, not the file.
Hreflang, image and video sitemaps are still supported, although Mueller said image and video ones matter less now that Google recognises them on the page.
For a quick check this week, open your sitemap and strip out any priority and change frequency tags, then confirm the last modified dates only change when the page does.
Find the entire Podcast