自适应 Web 爬虫框架,从单次请求到全规模并发爬取一站式解决。60.3k stars,Python,MIT 协议。核心亮点:元素自适应追踪(网站改版后自动重定位)、开箱即用绕过 Cloudflare Turnstile、内置 MCP Server 供 AI Agent 调用。
from scrapling.fetchers import StealthyFetcher
p = StealthyFetcher.fetch('https://example.com', headless=True)
products = p.css('.product', auto_save=True) # 自动保存选择器位置
products = p.css('.product', adaptive=True) # 网站改版后自适应重定位核心能力
1. Spider — Scrapy 风格爬虫框架
from scrapling.spiders import Spider, Response
class MySpider(Spider):
name = "demo"
start_urls = ["https://example.com/"]
async def parse(self, response: Response):
for item in response.css('.product'):
yield {"title": item.css('h2::text').get()}
MySpider().start()| 特性 | 说明 |
|---|---|
| 并发爬取 | 可配置并发限制、per-domain 限速、下载延迟 |
| 多 Session | HTTP 请求与隐身浏览器在同一 Spider 中混用,按 Session ID 路由 |
| 暂停/恢复 | Checkpoint 持久化,Ctrl+C 优雅退出,重启续跑 |
| 流式模式 | async for item in spider.stream() 实时消费抓取结果,带实时统计 |
| 开发模式 | 首次运行缓存响应到磁盘,后续重放 — 无需反复请求目标服务器 |
| 导出 | 内置 JSON/JSONL:result.items.to_json() / .to_jsonl() |
2. Fetcher — 多模式请求器
| 类 | 场景 |
|---|---|
Fetcher / FetcherSession | 快速 HTTP 请求,可模拟浏览器 TLS 指纹、HTTP/3 |
StealthyFetcher / StealthySession | 高级反检测,自动绕过 Cloudflare Turnstile/Interstitial |
DynamicFetcher / DynamicSession | Playwright Chromium / Google Chrome 完整浏览器自动化 |
共同特性:代理轮换(ProxyRotator)、DNS-over-HTTPS 防泄漏、异步支持、域名 & 广告拦截(内置 3500+ 跟踪域名库)。
3. 自适应解析
- 智能元素追踪:
auto_save=True保存元素位置,adaptive=True在网站改版后自动重定位 - 多种选择器:CSS、XPath、过滤器、文本搜索、正则搜索、相似元素查找
- 自动生成选择器:为任意元素生成稳健的 CSS/XPath
4. MCP Server — AI 辅助爬取
内置 MCP Server,Claude/Cursor 等 AI Agent 可直接调用 Scrapling 抓取网页内容:
- 先用 Scrapling 精准提取目标内容再传给 AI,减少 token 消耗、加速处理
- 支持 Agent Skill:
github.com/D4Vinci/Scrapling/tree/main/agent-skill - OpenClaw Skill:
clawhub.ai/D4Vinci/scrapling-official
安装
pip install scrapling # 基础
pip install 'scrapling[all]' # 全部可选依赖(含浏览器)Docker 镜像(含所有浏览器)随每次 Release 自动构建推送。
CLI 直接使用
无需写代码,直接从终端抓取:
scrapling fetch https://example.com内置交互式 IPython Shell,支持 curl → Scrapling 命令转换。
架构亮点
- 92% 测试覆盖率 + 全类型注解(PyRight + MyPy 自动扫描)
- 10x 更快的 JSON 序列化(vs 标准库)
- 惰性加载 + 优化数据结构,内存占用极低
- 已被数百名爬虫工程师日常生产使用超一年
关联页面
- cloakbrowser — 同样面向反检测场景的隐身 Chromium,可与 Scrapling 的 DynamicFetcher 互补
- markitdown — 文件/网页内容转 Markdown,与 Scrapling 抓取的原始 HTML 配合可完整清洗内容