Key takeaways
- Crawl budget is mainly a concern for large sites: 1 million+ pages, or 10,000+ pages that change daily.
- Filter combinations are the most common way big sites waste crawling.
- Google's preferred controls: robots.txt disallow or URL fragments for filters you don't need indexed, and 404s for empty results.
- Canonicals and nofollow help less over time.
Which sites need to worry about crawl budget?
Google says crawl budget guidance is for large sites with over a million unique pages that change about weekly, medium or larger sites with over 10,000 pages that change daily, and sites with many URLs reported as Discovered but not indexed. Most small business sites don't need to think about it.
The thresholds are from Google's crawl budget guide, which also warns that if Google spends too long on URLs it shouldn't crawl, it may not explore the rest of the site.
How should faceted navigation URLs be handled?
If filtered pages don't need to appear in Search, block them from crawling with robots.txt or move the filters into URL fragments, which Google generally ignores. Return a 404 when a filter combination has no results. Canonical tags and nofollow can reduce crawling of filter URLs, but Google calls them less effective in the long term.
All of this is in Google's guide to managing faceted navigation. A practical split:
- Keep indexable: the few filters people actually search for, such as a category plus a brand or a key attribute, with their own optimized page.
- Block or fragment: sorting, price ranges, multiple selections and combinations nobody searches.
- 404: any combination that returns no products.
- Consistent order: always the same parameter order, so one filter set has one URL.
Where canonical tags fit
A canonical URL tells Google which version to show, and over time can reduce how often duplicates are crawled. It does not stop crawling. Use it for small variations, such as tracking parameters or sort order, and robots.txt for large families of filter URLs.
How do you know if it's working?
Use the Crawl Stats report in Search Console. It shows how many requests Google made, when, the server responses and any availability problems. Compare requests to filter URLs before and after the fix, and watch whether important pages get crawled sooner. Google says sites under a thousand pages don't need this report.
See the Crawl Stats report help. Pair it with the checks in our technical SEO audit checklist; for stores on Shopify, the Shopify checklist covers the platform's own filter URLs, and variant structured data covers product pages. Delivery is part of our white label technical SEO and the wider white label SEO services.
Frequently asked questions
Does a 500-page site have a crawl budget problem?
Almost never. Google's crawl budget guidance is aimed at much larger sites.
Is noindex enough for filter pages?
Noindex keeps pages out of results but Google still has to crawl them to see it, so it doesn't save crawling.
Should filter links be nofollow?
It can help, but every link to the URL needs it, and Google calls it less effective long term than robots.txt.
What should an empty filter page return?
Google recommends an HTTP 404 when a filter combination returns no results.
Which filters should be indexable?
Only combinations people search for, given their own title, copy and internal links.