Skip to content

Read

PDFs

Read PDFs as text, with no extra setup.

When a URL returns a PDF, Frankensurf extracts its text with pypdf. Your agent gets text, not bytes.

page = await web.read("https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf")
Response
{
"url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
"text": "Dummy PDF file",
"content_type": "application/pdf; qs=0.001",
"structured": { "format": "pdf", "pages": 1 },
"receipt": { "method": "http", "latency_ms": 414 }
}
Option Default Description
pdf_max_pages 50 Most pages to extract.
max_bytes 40 MB Largest file accepted.