When a site blocks you
How escalation works
Frankensurf starts with the cheapest tool and climbs a six-rung ladder only when a site pushes back.
Every way to fetch a page is a rung. Low rungs are fast and free; high rungs get through more. Each read starts low and climbs only as far as it has to.
| Rung | What it does | Tools |
|---|---|---|
| T0 Markdown | Asks the site for text/markdown |
HTTP |
| T1 Light fetch | Plain HTTP or a reader, no browser | HTTP, Scrapling, Jina Reader |
| T2 Browser | A real browser renders the page | Chromium, Steel, Crawl4AI, Browserbase, Hyperbrowser, Browserless, Kernel, Anchor, Cloudflare |
| T3 Stealth and signed | Looks like a person, or signs as a verified agent | Camoufox, Scrapling, Patchright, nodriver, signed requests |
| T4 Unblocker | Paid services built for the hardest walls | Firecrawl, ZenRows, fastCRW, Scrapfly, Bright Data, Zyte, Apify |
| T5 You | A person clears the wall | Human handoff |
What counts as a failure
Section titled “What counts as a failure”A tool “succeeding” isn’t enough. Frankensurf checks what came back and climbs when it sees:
- an app shell or unrendered template (
VISUAL_REQUIRED); - a rendered page with almost no text (
EMPTY_PAGE); - a challenge or “Just a moment…” page (
CAPTCHA,BLOCKED); - a sign-in page where content should be (
AUTH_REQUIRED).
Sites also fake walls for bots: a sign-in redirect, or a 404, that a real browser never sees. So on a public read:
- a
NOT_FOUNDfrom a plain fetch, or after another wall, gets one more try; - a sign-in wall stops the climb only after
auth_wall_confirmations(default 5) tools in a row hit one; - after a suspected fake wall, the tools that look most different from the one
that was fooled go next (
fake_wall_confirmers: Jina Reader, Camoufox, then the paid unblockers).
A page that is plainly a site’s bot page (a final URL such as /help/bots.html
or /captcha, or text like “thinks you are a bot”) counts as CAPTCHA, and a
short page that only asks you to sign in, in English, Spanish, Portuguese,
French, German, Italian or Dutch, counts as AUTH_REQUIRED.
Signed-in reads (identities and profiles) keep the strict rule: a sign-in wall there means the session needs renewing, and the climb stops.
Second opinion
Section titled “Second opinion”A plain HTTP page can pass every check and still lack its content: a site that
loads results with JavaScript serves only its header and footer as text. So when
an automatic read lands on plain HTTP with scripts and under 5,000 characters of
text (second_opinion_text_chars), Frankensurf also asks a browser and keeps
whichever page has clearly more content (at least 1.5 times as much, and 1,000
characters more). receipt.second_opinion shows both. If the render fails or
isn’t better, the HTTP page stands; the only cost is time.
Completeness
Section titled “Completeness”Then every automatic read gets a structural check (completeness.py).
It looks at the page’s shape:
- Search and category pages (a query parameter, or a path such as
/search) need at least 10 distinct same-site item links (product, listing or detail URLs; never assets or service and legal pages) or 4 prices. - Item pages need a price or 1,500 characters of text.
- Other pages need 1,500 characters, or 400 without scripts.
An incomplete page is re-read along completeness_ladder: Scrapling, Camoufox,
then Firecrawl, Zyte, Scrapfly and ZenRows when paid tools are allowed, then
Patchright and Bright Data. Each step is an ordinary read pinned to one tool,
with the usual checks and receipt. Frankensurf stops at the first complete page
and keeps the most complete one it saw. A search page that passes only narrowly
(fewer than completeness_borderline_items, 20, item links and few prices) also
gets one read from the strongest allowed tool, a paid one when allowed. completeness_max_extra_reads (default
4) and completeness_deadline_seconds (default 150) bound the extra work, and
receipt.completeness lists every step.
Free tools on the ladder race in pairs (completeness_parallel, default 2,
up to 4). The first complete page wins and the other read is cancelled.
Starts are half a second apart, so a site sees at most two extra reads at
once. Paid tools always run one at a time, so racing never pays twice. On the
139-site held-out set this cut the free tier’s p90 from 50 s to 37 s and its
median from 5.6 s to 4.5 s.
Placeholders and menus
Section titled “Placeholders and menus”Two more shapes of a page that came back without its content are caught on every kind of page, not only search results:
- Placeholders. Prices of
$0with no real price, or template values that leaked into the text ({{ price }},NaN,undefined), mean the page had not loaded.completeness.placeholderis set and the read escalates. - Menus only. When more than 80% of the text is link text and there is no
price, the read got the site’s navigation, not the page.
completeness.link_text_sharesays how much.
Some pages are complete but hide their prices behind a choice (“select guests
to see prices”). No stronger tool would show more, so these don’t escalate;
completeness.needs_interaction quotes the page so your agent knows why there
is no price.
Pages a stronger tool can’t fix
Section titled “Pages a stronger tool can’t fix”Some wrong pages look right to every tool, so Frankensurf names them instead of climbing:
- Off-query results. A site that ignores an unknown search parameter shows
its default feed, which passes the structure check. When the query is known
(from parameters such as
q,kworkeywords, or from yourexpect_terms), a results page whose items and text never mention it comes back observed but markedcompleteness.off_query.receipt.next_stepsays the search URL is probably wrong. A page still loading its results escalates as before. Once you know the right URL, a site module template keeps it. - Pages a site module vouches for. When a saved module’s assertions pass,
the page is complete even if it has few links, because the module knows its
results live in JSON. A module’s
invalidmarkers turn the site’s own error or empty-feed page intoNOT_FOUND. - Site error pages. A redirect to an
erroror404path, an error title such as “Page not found”, or a short page saying the page doesn’t exist fails asNOT_FOUND. A plain fetch’s error page gets one confirming read from a different tool, since some sites serve fake 404s to bots; then the read stops.
Hedging a slow read
Section titled “Hedging a slow read”A blocked site can take a dozen tools in turn before one gets through. Once an
automatic read has run for 10 seconds (hedge_after_seconds, 0 turns it off),
one more read starts beside it, pinned to the first free tool on the
completeness ladder (Jina Reader by default). If that read brings back a
complete page first, it wins and the climb is cancelled. A hedge page that is
incomplete or failed never replaces the main read, and a hedge never uses a
paid tool. When the hedge wins, receipt.hedge says so and the site’s route
hint learns it, so the next read starts there. On the 139-site held-out set,
reads that used to take about 50 seconds (Michaels, Naukri, Anthropologie,
Urban Outfitters) came back in 8 to 17.
Skipping ahead
Section titled “Skipping ahead”Three things reorder the ladder so repeat visits are fast:
| What it does | Setting | |
|---|---|---|
| Escalation | After 4 walls in one read, allowed paid tools move ahead of the free ones left. | escalate_after_walls |
| Site hints | When a read succeeds only after earlier tools failed, the tool that got through goes first on that site for 24 hours. | origin_route_hint_ttl_seconds |
| Route memory | After 3 clean reads in a row on the same path pattern, that tool goes first for that path. | route_memory_min_samples |
Measured on g2.com: 63.5 s before hints, 8.5 s on the first read with escalation, 2.4 s with the hint.
A few sites need a particular tool from the first visit. Those live as data in
bundled_route_seeds.json, never as site code, and your own settings always
win.
Pacing
Section titled “Pacing”Frankensurf is polite by default.
- Reads to one site are at least 2 seconds apart, across processes
(
origin_min_interval_seconds). - A
BLOCKED,CAPTCHAorRATE_LIMITEDanswer pauses direct reads of that site for 15 minutes (origin_cooldown_seconds). - Paid unblockers and handoff aren’t held by the pause: a different service or a person is handling the wall.
Choosing yourself
Section titled “Choosing yourself”| You want | Set |
|---|---|
| One specific tool, no fallback | provider="zenrows" |
| An exact order | provider_candidates=["http", "camoufox"] |
| Paid tools allowed, within a budget | allow_paid_fallbacks=True, max_cost_usd=0.05 |
| Nothing on your machine | allow_local_browser=False |