Choose a PDF with an embedded text layer.
Extracted text
How it works
The accepted local PDF.js library reads each selected page’s text content. Items are grouped into lines using available text positions, ordered from top to bottom and left to right within each line, and separated by clear page markers. Complex layouts may not reproduce in exact visual reading order.
Limits and interpretation
Input is limited to 50 MB and 100 selected pages per operation. If a PDF contains only scanned images and no embedded text layer, this tool may return little or no text. Multi-column pages, tables, rotated text, and unusual font encodings can require manual cleanup. Encrypted PDFs are unsupported.
