Aller au contenu

Crawling & indexing: the technical SEO guide

Before rankings even come into play, Google has to clear two steps that are too often confused: crawling your pages (exploring them), then indexing them (storing them so they can be shown). This is where many technical SEO problems live: crawl budget wasted on useless URLs, pages blocked by mistake, duplicate content diluting signals, or deindexing that never takes effect.

This guide brings together around fifteen case studies and field reports, from the robots.txt file to migrations, to master what Googlebot really sees of your site. The linked articles are in French.

1. Mastering the robots.txt file

robots.txt controls crawling, not indexing: a URL blocked from crawling can still appear in the results. Its syntax looks simple but hides traps (grouping of User-agent lines, wildcards, case sensitivity, encoded URLs) that can block an entire site by accident.

robots.txt blocks crawling, not indexing Googlebot is stopped by a Disallow rule and does not crawl the page, but the URL can still be shown in search results, without a description. Googlebot crawls crawl robots.txt Disallow: /page Page not read (content unknown) but the URL stays indexable your-site.com/page No description available, blocked by robots.txt
Blocking crawling does not remove a page from the index: to deindex, you need a tag or a header, not a Disallow.

2. Indexing & deindexing

Getting a page into the index, or cleanly out of it, is not decided in the same place as crawling. Between meta robots, the X-Robots-Tag header and server directives, each method has its use cases. And robots.txt does not deindex.

Where deindexing happens The meta robots tag and the X-Robots-Tag header deindex a page; robots.txt does not. Meta tag in the HTML <head> content = noindex Deindexes X-Robots-Tag HTTP header noindex (PDF, images) Deindexes robots.txt blocks crawling Disallow: /page Does not deindex Three levers not to confuse: two deindex, one does not.

3. Crawl budget & log analysis

On a large site, Google only crawls part of the URLs on each visit: that is the crawl budget. To optimize it, nothing beats server log analysis, the only reliable source of what Googlebot actually does, provided you know how to isolate and interpret it.

Crawl budget and log analysis Server logs reveal what Googlebot actually crawls: keep the useful URLs and limit waste on useless ones. Googlebot Server access.log real requests Useful URLs to favor Wasted URLs parameters, facets
Logs are the only reliable source of what Googlebot really crawls: the basis for optimizing crawl budget.

4. Avoiding duplicate content

Duplicate content spreads signals across several URLs instead of concentrating them. The sources are many and often invisible: HTTP vs HTTPS, tracking parameters, indexed staging environments, or URL bugs specific to some CMSs.

Duplicate content spreads signals The same content served on several URLs spreads SEO signals; the canonical tag concentrates them on a single version. http://site.com/p https://site.com/p site.com/p?utm=x Duplicate content signals spread canonical 1 canonical URL signals concentrated
HTTP/HTTPS, tracking parameters, staging…: so many duplicate URLs to consolidate with the canonical.

5. Migrations & maintenance

A migration is when indexing is lost the fastest: forget a single URL and traffic evaporates. Likewise, badly handled maintenance mode (a 200 page or a redirect instead of a 503) sends the wrong signal to Googlebot.

A migration without losing indexing Every old URL must be 301-redirected to its new one; a single forgotten URL returns a 404 and loses its traffic. Old site New site /old-url-1 /new-url-1 301 /old-url-2 /new-url-2 301 /old-url-3 404 lost traffic
An exhaustive redirect map: every URL gets its 301, or it loses the rankings it has earned.

A crawling or indexing project to secure?

A migration to prepare, a crawl budget to optimize, an indexing problem that won’t go away? Tell me about your project.