Read
PDFs
Read PDFs as text, with no extra setup.
When a URL returns a PDF, Frankensurf extracts its text with pypdf. Your agent gets text, not bytes.
page = await web.read("https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf")frankensurf read https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf{ "tool": "read", "arguments": { "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf" } }{ "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf", "text": "Dummy PDF file", "content_type": "application/pdf; qs=0.001", "structured": { "format": "pdf", "pages": 1 }, "receipt": { "method": "http", "latency_ms": 414 }}Options
Section titled “Options”| Option | Default | Description |
|---|---|---|
pdf_max_pages |
50 |
Most pages to extract. |
max_bytes |
40 MB | Largest file accepted. |