🕷️ Crawling & Indexing Guide

Master how search engines discover and store your pages

🔄 How Search Engines Work

1

Discover

Find URLs via links, sitemaps

→
2

Crawl

Fetch page content

→
3

Render

Execute JavaScript

→
4

Index

Store in database

→
5

Rank

Show in results

🤖 robots.txt Configuration

Control what search engines can crawl

Basic Example

# Allow all crawlers
User-agent: *
Allow: /

# Block admin areas
Disallow: /admin/
Disallow: /private/

# Sitemap location
Sitemap: https://site.com/sitemap.xml

Psilobase Example

User-agent: *
Allow: /

# Block utility pages
Disallow: /templates/
Disallow: /TODO/
Disallow: /scripts/

# Allow CSS/JS for rendering
Allow: /css/
Allow: /js/

📋 Robots.txt Directives

User-agent

Specify which crawler the rules apply to

User-agent: Googlebot

Disallow

Block crawling of specific paths

Disallow: /private/

Allow

Explicitly allow crawling (overrides Disallow)

Allow: /public/

Sitemap

Tell crawlers where your sitemap is

Sitemap: /sitemap.xml

Crawl-delay

Slow down crawl rate (not honored by Google)

Crawl-delay: 10

* (Wildcard)

Match any sequence of characters

Disallow: /*.pdf$

🗺️ XML Sitemaps

Sitemap Structure

<?xml version="1.0"?>
<urlset xmlns="...">
  <url>
    <loc>https://site.com/page</loc>
    <lastmod>2024-01-15</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.8</priority>
  </url>
</urlset>

Best Practices

  • ✅ Keep under 50,000 URLs per sitemap
  • ✅ File size under 50MB uncompressed
  • ✅ Use sitemap index for large sites
  • ✅ Include canonical URLs only
  • ✅ Update lastmod when content changes
  • ✅ Submit to Google Search Console
  • ✅ Don't include noindex pages
  • ✅ Use HTTPS URLs if site uses HTTPS

💰 Crawl Budget Optimization

✅

Fast server response

Under 200ms server response time lets Googlebot crawl more pages

✅

Flat site architecture

Important pages within 3 clicks from homepage

✅

Updated XML sitemap

Help Googlebot find new and updated content

✅

Strong internal linking

Pass link equity and help discovery of deep pages

❌

Duplicate content

Wastes crawl budget on identical pages

❌

Broken links (404s)

Wasted crawls on non-existent pages

❌

Redirect chains

Multiple hops waste crawl resources

❌

Infinite URL spaces

Calendar pages, faceted navigation traps

🏷️ Meta Robots Tags

Directive Crawling Indexing Use Case
index, follow ✅ Allowed ✅ Allowed Default - all pages
noindex, follow ✅ Allowed ❌ Blocked Thank you pages, internal search
index, nofollow ❌ Links not followed ✅ Allowed User-generated content pages
noindex, nofollow ❌ Blocked ❌ Blocked Admin, staging pages
noarchive ✅ Allowed ✅ Allowed No cached version shown
nosnippet ✅ Allowed ✅ Allowed No description in SERPs

🔧 Crawling Tools

🔍

Google Search Console

Coverage reports, indexing issues, sitemap submission

Access →
🕷️

Screaming Frog

Desktop crawler for technical SEO audits

Download →
🤖

robots.txt Tester

Test your robots.txt in Search Console

In GSC →
📋

Sitemap Validator

Validate XML sitemap format

Validate →

✅ Crawling Checklist

🤖 robots.txt

  • robots.txt exists
  • No important pages blocked
  • CSS/JS allowed
  • Sitemap referenced
  • Tested in GSC

🗺️ Sitemap

  • XML sitemap created
  • All pages included
  • No noindex pages
  • Submitted to GSC
  • Updated regularly

📊 Indexing

  • No unwanted noindex
  • Canonicals correct
  • No duplicate content
  • Coverage report clean
  • 404s fixed