Before rankings even come into play, Google has to clear two steps that are too often confused: crawling your pages (exploring them), then indexing them (storing them so they can be shown). This is where many technical SEO problems live: crawl budget wasted on useless URLs, pages blocked by mistake, duplicate content diluting signals, or deindexing that never takes effect.
This guide brings together around fifteen case studies and field reports, from the robots.txt file to migrations, to master what Googlebot really sees of your site. The linked articles are in French.
1. Mastering the robots.txt file
robots.txt controls crawling, not indexing: a URL blocked from crawling can still appear in the results. Its syntax looks simple but hides traps (grouping of User-agent lines, wildcards, case sensitivity, encoded URLs) that can block an entire site by accident.
- robots.txt: the small subtleties that can hurt
- robots.txt: testing encoded URLs
- robots.txt: leave Google what belongs to Google
- The Drupal robots.txt trap that blocks Google Images
2. Indexing & deindexing
Getting a page into the index, or cleanly out of it, is not decided in the same place as crawling. Between meta robots, the X-Robots-Tag header and server directives, each method has its use cases. And robots.txt does not deindex.
- X-Robots-Tag: controlling indexing through HTTP headers
- SEO duel: meta robots tags vs the X-Robots-Tag header
- Testing 5 methods to deindex a page from Google
- Removing a list of URLs from the Google index
- “Index of”: disabling directory listing on your web server
3. Crawl budget & log analysis
On a large site, Google only crawls part of the URLs on each visit: that is the crawl budget. To optimize it, nothing beats server log analysis, the only reliable source of what Googlebot actually does, provided you know how to isolate and interpret it.
- Isolating Googlebot logs with Apache
- Tracking the crawl errors Googlebot runs into
- Analyzing Googlebot’s crawl with the Watussi Box
- URL parameter combinations: the log script
- Why bots speak no language (Accept-Language)
4. Avoiding duplicate content
Duplicate content spreads signals across several URLs instead of concentrating them. The sources are many and often invisible: HTTP vs HTTPS, tracking parameters, indexed staging environments, or URL bugs specific to some CMSs.
- WordPress: duplicate content on category URLs
- Stop duplicate content: indexed staging sites
- Duplicate content: HTTP vs HTTPS
- Duplicate content & tracking parameters
- Negative SEO: protecting yourself from malicious duplicate content
- Anti-duplicate content script for subdomains
5. Migrations & maintenance
A migration is when indexing is lost the fastest: forget a single URL and traffic evaporates. Likewise, badly handled maintenance mode (a 200 page or a redirect instead of a 503) sends the wrong signal to Googlebot.
- During an SEO migration, forget no URL
- Site under maintenance: the HTTP 503 code to protect your SEO
A crawling or indexing project to secure?
A migration to prepare, a crawl budget to optimize, an indexing problem that won’t go away? Tell me about your project.