robots.txt
A text file in the site's root directory that tells search engine crawlers which parts of the site they should not crawl.
- 01
Why it matters
robots.txt blocks crawling of areas useless in search, such as admin panels, basket and internal search result pages, so crawlers spend their time on important pages. It does not, however, reliably keep a page out of the index; noindex is used for that. The file is public.
- 02
Example
A shop's addresses that generate endless filter combinations are eating up the crawler's time. Once rules blocking these parameters are added to robots.txt, crawling concentrates on product and category pages.
- 03
Common mistake
Blocking a page you want removed from the index with robots.txt. Because the blocked page cannot be crawled, its noindex tag is never seen and the page can stay in the index.
- 04
Let's talk about your project.
Tell us what you need; we will define the scope together.