Engineering
Web Scraping
Intermediate5 uses

Scrapling โ€” Adaptive Web Scraping

An adaptive Python web scraping framework that handles everything from a single request to full-scale crawls. Automatically re-locates elements when pages update, bypasses Cloudflare Turnstile and other anti-bot systems out of the box, and supports concurrent multi-session crawls with pause/resume and proxy rotation.

web-scrapingpythonanti-botcloudflare-bypasscrawlingmcpautomationadaptive
๐Ÿ“‹

Spec

Scrapling โ€” Adaptive Web Scraping

Overview

Scrapling is an adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl. Its parser learns from website changes and automatically relocates your elements when pages update. Its fetchers bypass anti-bot systems like Cloudflare Turnstile out of the box. Its spider framework scales to concurrent, multi-session crawls with pause/resume and automatic proxy rotation โ€” all in a few lines of Python.

Key Features

  • Adaptive parsing: Automatically finds moved elements after site redesigns using adaptive=True
  • Anti-bot bypass: Bypasses Cloudflare Turnstile, stealth fetcher for fingerprint evasion
  • Multiple fetchers: Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher
  • Spider framework: Full crawl support with concurrent sessions, proxy rotation, pause/resume
  • MCP server: Use Scrapling as an MCP tool from Claude, OpenClaw, and other agents
  • CLI: Built-in command-line interface
  • 44k+ GitHub stars, production-proven

Quick Start

from scrapling.fetchers import StealthyFetcher

StealthyFetcher.adaptive = True
page = StealthyFetcher.fetch('https://example.com', headless=True, network_idle=True)
products = page.css('.product', auto_save=True)  # Survives site redesigns!

Full Crawl (Spider)

from scrapling.spiders import Spider, Response

class MySpider(Spider):
    name = "demo"
    start_urls = ["https://example.com/"]

    async def parse(self, response: Response):
        for item in response.css('.product'):
            yield {"title": item.css('h2::text').get()}

MySpider().start()

Fetcher Types

| Fetcher | Use Case | |---|---| | Fetcher | Simple HTTP requests (no JS needed) | | AsyncFetcher | Async HTTP requests | | StealthyFetcher | Stealth browser โ€” evades fingerprinting | | DynamicFetcher | Full browser automation with JS rendering |

Adaptive Parsing

# First scrape โ€” auto-save element fingerprints
elements = page.css('.price', auto_save=True)

# Later โ€” even if site redesigns, adaptive mode re-locates them
elements = page.css('.price', adaptive=True)

Installation

pip install scrapling
scrapling install  # Install browser dependencies

Agent Skill

Scrapling ships an official AgentSkill spec file. Install via Clawhub:

clawhub install scrapling-official

Or use the direct ZIP: https://github.com/D4Vinci/Scrapling/raw/refs/heads/main/agent-skill/Scrapling-Skill.zip

MCP Server

Scrapling exposes an MCP server for use with Claude Code, Cursor, and other AI tools. See: https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html

Docs

https://scrapling.readthedocs.io/

๐Ÿ”’

Constraints

  • Python only (not available for other languages)
  • Browser-based fetchers (StealthyFetcher, DynamicFetcher) require Playwright installation via scrapling install
  • Adaptive mode requires an initial auto_save=True scrape to build element fingerprints
  • Anti-bot bypass is not 100% guaranteed against all protection systems
โš ๏ธ

Anti-patterns

  • Don't use DynamicFetcher for simple pages โ€” use Fetcher or AsyncFetcher for speed
  • Don't skip scrapling install when using browser fetchers
  • Don't use adaptive=True without a prior auto_save=True run
  • Don't hardcode CSS selectors without considering auto_save for long-running scrapers
๐Ÿงช

Tests

  • Verify Fetcher.fetch() returns a parsed page object
  • Test adaptive=True recovers elements after mocking a page structure change
  • Confirm StealthyFetcher returns 200 on a Cloudflare-protected URL
  • Run spider with start_urls and assert yielded items are non-empty
โ–ถ๏ธ

Run Instructions

  1. Install: pip install scrapling
  2. Install browser deps: scrapling install
  3. Import desired fetcher: from scrapling.fetchers import StealthyFetcher
  4. Fetch a page: page = StealthyFetcher.fetch('https://example.com', headless=True)
  5. Select elements: items = page.css('.selector')
  6. For crawls, subclass Spider and define parse() method
  7. For MCP: run scrapling mcp and configure in your AI tool