Scraping tools, tested against real protected sites.
No affiliate listicles, no vendor-supplied numbers. We run each tool against 21 real public sites — Amazon, Cloudflare-fronted stores, travel and real-estate portals — and save the request, the response, and the screenshot for every attempt. Then we write down what actually happened.
Tools tested · public-tough-v1 · proxy: none · 2026-05-24
| Tool | Type | Verdict | Passed | Notes |
|---|---|---|---|---|
| Camoufoxv0.6.0 | Anti-detect browser | Sharp edges | 11/21 | The highest ceiling we measured — if you pay the babysitting tax. |
| Botasaurusv4.0.97 | Anti-detect browser | Usable | 6/21 | Best untuned score in the test — from its HTTP layer alone. |
| curl_cffiv0.15.0 | TLS-impersonation HTTP | Usable | 5/21 | Chrome's TLS handshake without the browser. |
| Scraplingv0.4.8 | TLS-impersonation HTTP | Usable | 5/21 | curl_cffi-class fetching with a scraping-first API. |
| SeleniumBasev4.49.2 | Browser automation | Sharp edges | 4/21 | Beats Playwright on Amazon, fails murkier everywhere else. |
| HTTPXv0.28.1 | HTTP client | Usable | 3/21 | Modern sync/async HTTP for Python, same wall as Requests. |
| Patchrightv1.60.0 | Browser automation | Sharp edges | 3/21 | Measured against stock Playwright: no difference (in this mode). |
| Playwrightv1.60.0 | Browser automation | Usable | 3/21 | A rendering baseline, not a stealth tool. |
| Python Requestsv2.34.2 | HTTP client | Usable | 2/21 | The Python HTTP baseline — and our negative control. |
| wreqv0.11.3 | TLS-impersonation HTTP | Sharp edges | 2/21 | Real fingerprint, young edges. |
Pass counts are out of 21 and measure the tool and its network position together. With no proxy, a low score often reflects a single residential IP's reputation as much as the tool. That is why these are reviews, not a leaderboard — read the per-tool page for what each number means. Camoufox's row shows its strongest of three configurations.
How we test
- Run the tool in its natural shape against 21 real public pages.
- Judge each attempt with deterministic validators: expected status code plus required and forbidden content markers — not an LLM guess.
- Save the request config, response metadata, HTML, and screenshots as artifacts.
- Write an honest verdict: what worked, what failed, and when you'd reach for it.
What this isn't (yet)
These are open-source tools you run yourself. We have not published numbers for hosted scraping APIs (Zyte, Scrapfly, Bright Data, and the like): the adapters exist, but no real paid runs have been made, so there is nothing honest to show. Proxy-mode results are also still to come. We'd rather ship a small true map than a big fake one.