ScrapingEvalsBrowse the catalog →

Scraping tools, tested against real protected sites.

No affiliate listicles, no vendor-supplied numbers. We run each tool against 21 real public sites — Amazon, Cloudflare-fronted stores, travel and real-estate portals — and save the request, the response, and the screenshot for every attempt. Then we write down what actually happened.

10 tools reviewed with real runs26 in the catalog21 live targets each

Tools tested · public-tough-v1 · proxy: none · 2026-05-24

ToolTypeVerdictPassedNotes
Camoufoxv0.6.0Anti-detect browserSharp edges11/21The highest ceiling we measured — if you pay the babysitting tax.
Botasaurusv4.0.97Anti-detect browserUsable6/21Best untuned score in the test — from its HTTP layer alone.
curl_cffiv0.15.0TLS-impersonation HTTPUsable5/21Chrome's TLS handshake without the browser.
Scraplingv0.4.8TLS-impersonation HTTPUsable5/21curl_cffi-class fetching with a scraping-first API.
SeleniumBasev4.49.2Browser automationSharp edges4/21Beats Playwright on Amazon, fails murkier everywhere else.
HTTPXv0.28.1HTTP clientUsable3/21Modern sync/async HTTP for Python, same wall as Requests.
Patchrightv1.60.0Browser automationSharp edges3/21Measured against stock Playwright: no difference (in this mode).
Playwrightv1.60.0Browser automationUsable3/21A rendering baseline, not a stealth tool.
Python Requestsv2.34.2HTTP clientUsable2/21The Python HTTP baseline — and our negative control.
wreqv0.11.3TLS-impersonation HTTPSharp edges2/21Real fingerprint, young edges.

Pass counts are out of 21 and measure the tool and its network position together. With no proxy, a low score often reflects a single residential IP's reputation as much as the tool. That is why these are reviews, not a leaderboard — read the per-tool page for what each number means. Camoufox's row shows its strongest of three configurations.

How we test

  1. Run the tool in its natural shape against 21 real public pages.
  2. Judge each attempt with deterministic validators: expected status code plus required and forbidden content markers — not an LLM guess.
  3. Save the request config, response metadata, HTML, and screenshots as artifacts.
  4. Write an honest verdict: what worked, what failed, and when you'd reach for it.

What this isn't (yet)

These are open-source tools you run yourself. We have not published numbers for hosted scraping APIs (Zyte, Scrapfly, Bright Data, and the like): the adapters exist, but no real paid runs have been made, so there is nothing honest to show. Proxy-mode results are also still to come. We'd rather ship a small true map than a big fake one.