Tools / pdf.extract

PDF text extractor

pdf.extract

Page-by-page text from a text-based PDF, given as a URL or base64, up to 10 MB.

MediaLaunch set2 units per callnot cachedMCP: pdf_extractv1.0.0

When to use it

Use to read the text of a PDF such as a report, contract or paper. Works on text-based PDFs; for scanned pages with no text layer, use image.ocr on a rendered page image instead.

Input

FieldTypeDescription
urlstring (≤ 2048 chars)Absolute http or https URL of the file. Use instead of base64.
base64string (≤ 14000000 chars)Base64-encoded file bytes, up to 10 MB once decoded. Use instead of url.
max_pagesinteger (≥ 1, ≤ 200)Pages to read from the start of the document. Default: 50.

Output

Returns data with: page_count, pages_extracted, truncated, pages: [{page, text}].

Request

curl -X POST https://agentops.tools/v1/pdf.extract \
  -H 'content-type: application/json' \
  -d '{"url":"https://arxiv.org/pdf/1706.03762","max_pages":2}'

Over MCP the tool is named pdf_extract. Over A2A its skill id is pdf.extract. Same input, same envelope.

Try it

Counts against your rate limit.

Sources