跳到主要内容

robots.txt 要不要让 AI 爬虫进?

GEO1000问 / 2026-09-30 19:53 / uETyKCDsZeql / 查看原文

robots.txt 是 GEO 优化的”开关”——挡住 AI 爬虫 = 失去 GEO 引用的基础。

一、robots.txt 基础

robots.txt 放在网站根目录,告诉爬虫哪些页面可以抓取。

“`

# 全部允许

User-agent: *

Allow: /

# 禁止某爬虫

User-agent: BadBot

Disallow: /

“`

**主流 AI 爬虫 User-Agent**:

| AI 平台 | User-Agent | 备注 |

|—|—|—|

| OpenAI | GPTBot / ChatGPT-User | ChatGPT 训练 + 浏览 |

| Anthropic | ClaudeBot / Claude-User | Claude 训练 + 浏览 |

| Google | Google-Extended | Gemini / SGE 训练 |

| Meta | Meta-ExternalAgent | Llama 训练 |

| Baidu | Baiduspider | 文心 / 百度搜索 |

| 字节 | Bytespider | 豆包 / 抖音搜索 |

| Moonshot | MoonshotBot | Kimi 训练 + 浏览 |

| 阿里 | TongyiBot / TiktokBot | 通义 / 夸克 |

**GEO 友好的 robots.txt 范本**:

“`

User-agent: *

Allow: /

# 允许主要 AI 爬虫

User-agent: GPTBot

Allow: /

User-agent: ChatGPT-User

Allow: /

User-agent: ClaudeBot

Allow: /

User-agent: Google-Extended

Allow: /

User-agent: Baiduspider

Allow: /

User-agent: Bytespider

Allow: /

User-agent: MoonshotBot

Allow: /

User-agent: TongyiBot

Allow: /

# 仅禁止后台

Disallow: /wp-admin/

Disallow: /wp-login.php

Disallow: /admin/

“`

二、GEO robots.txt 常见错误

① **全禁 AI 爬虫**

– 错误:`User-agent: GPTBot / Disallow: /`

– 影响:完全失去 ChatGPT 引用

② **不知有 AI 爬虫**

– 错误:robots.txt 没写 AI 爬虫

– 影响:默认行为不一致(有些默认禁,有些默认允许)

③ **禁后台路径导致死链**

– 错误:`Disallow: /wp-content/`

– 影响:媒体 / 图片被禁,AI 无法引用

④ **sitemap.xml 没指定**

– 错误:没告诉爬虫 sitemap 在哪

– 影响:爬虫抓取效率低

三、GEO robots.txt 检查清单

– [ ] 至少包含 GPTBot / Claude / 百度 / 字节 / Kimi / 通义

– [ ] 不禁媒体 / 内容 / 静态资源

– [ ] sitemap.xml 路径明确

– [ ] 后台 / 登录页禁止

– [ ] 没有过度限制(避免反向影响)

**WordPress 实践**:

– Yoast / Rank Math 插件可以自动管理

– 自定义主机可手动编辑 robots.txt

一句话:robots.txt = GEO 的”开门开关”——挡住 AI 爬虫等于主动放弃 GEO。

相关标签 #robots.txt