Tools150

A useful PDF task, made simple

Extract Data from PDF.

Find the data you need. Keep the source page with every result.

Advertisements
Your PDF workspace
On your device
or drop your files here
Your PDF is processed in this browser. PDF
Advertisements

How to extract data from pdf

  1. 01

    Choose one PDF

    Load a document with selectable text, or try the sample.

  2. 02

    Extract and review

    Select Extract. Review values, counts and source pages; filter or sort the list.

  3. 03

    Copy or download

    Copy the values, or download CSV with page references or a simple TXT list.

BlogHow to extract structured data from a PDFRead the guide

Extract structured values from a PDF into CSV or text

Choose the data types you need: emails, phone numbers, URLs, numeric dates, IPv4 addresses, domains, DOI, ISBN, hashtags, and MAC addresses. Review the extracted values with counts and page references.

Recognized values become a list with source-page references.
QUICK VIDEO TUTORIAL How to extract structured data from a PDF Collect contacts and research identifiers in one organized list. English narration Watch the video

How to extract structured data from a PDF

When is this useful?

Research

Collect identifiers

Gather DOI and ISBN values with the pages where they appear.

Reports

Pull out numeric details

Find dated entries and IPv4 addresses.

Directories

Build a contact list

Combine emails, phones and web links in one structured export.

Features and controls

Select useful data types

Start with emails, phones and URLs, then open More data types for specialized values.

What the recognizers support

Dates must use a numeric format with a four-digit year; ambiguous day/month values retain their source order. IPv4 octets are validated; IPv6 is not included.

Domains are derived from recognized emails and web links, not arbitrary filenames. DOI extraction covers common modern DOI forms, not every historical form. ISBN-13 values need a valid checksum; ISBN-10 also needs an ISBN label. Hashtags can contain Unicode letters, including Arabic.

MAC addresses use six colon- or hyphen-separated pairs. The tool never contacts extracted addresses or identifiers.

Review values with their source pages

Check counts, jump to a PDF page, and filter the result list.

Review and organize

Each result records the physical page position in the PDF, which can differ from printed page numbers. Select a page chip to open the preview with an approximate highlight.

Remove duplicates starts enabled. Counts combine recognized repetitions of the same value; switch it off to see separate occurrences. Sort alphabetically changes the list order. Filtering applies to copying and downloads as well as the visible rows.

Copy values or download a structured list

Choose CSV for counts and pages, or TXT for one value per line.

What each download contains

CSV contains Type, Value, Count and Pages columns. UTF-8 encoding supports Arabic and other scripts. Values that could be interpreted as spreadsheet formulas are prefixed with an apostrophe for safer importing.

TXT and Copy all contain values only, one per line. All rows matching the filter are included, even when the on-screen table spans several result pages. No new PDF is created and the source PDF is unchanged.

Keep large extractions manageable

Process one PDF in your browser, with explicit limits and partial-result notices.

Files, limits and scanned pages

Load one PDF up to 50 MiB and 500 pages. Extraction checks up to 250,000 characters and 2,000 matches per page, 10 million characters and 20,000 matches overall, and 2,000 annotations per page. A 90-second extraction budget limits lengthy runs. Very complex documents may work better when split into smaller files.

When a limit is reached or a page cannot be fully read, the results show a notice and limited or failed runs export with partial in the filename. Pages without selectable text are counted separately. Use PDF OCR for image-only scans, then return to extract the recognized text.

PDF content is processed locally in this browser. Uploaded documents are not sent to an extraction server. Links, email addresses and phone numbers found inside the PDF are never contacted automatically.

Supported files & processing

Input
PDF
Output
CSV / TXT

Open one PDF up to 50 MB and 750 pages. Additional output, memory and processing limits can apply to complex documents.

Processing takes place on your device without uploading the document. Download your result before closing the page; keep the original separately.

Advertisements
FAQ

A few useful answers

Know what to expect before you start.

On your device
What can this tool recognize?

Choose the data types you need: emails, phone numbers, URLs, numeric dates, IPv4 addresses, domains, DOI, ISBN, hashtags, and MAC addresses. Review the extracted values with counts and page references. This extracts recognizable values, not tables, invoices or semantic fields. Patterns and checksums reduce noise but do not guarantee completeness or prove that an identifier is genuine.

Can I extract data from scanned PDFs?

Only selectable text and recognized link targets are read. Image-only pages need PDF OCR first. Run OCR, download the searchable PDF, then return to this extractor.

How are duplicates and counts handled?

Remove duplicates combines recognized equivalent values. Count shows detected occurrences and Pages lists physical PDF page positions. Disable the option to keep occurrences separate. A printed target and matching clickable link on the same page are not double-counted.

What is the difference between CSV and TXT?

CSV includes type, value, occurrence count and page references. TXT and Copy all provide values only, one per line. All rows matching the current filter are included, not only the visible table page.

Is the PDF uploaded for extraction?

No. PDF parsing, matching and exports run in this browser. The original PDF is unchanged. No extracted link, email or phone number is contacted automatically.

What are the limits?

One PDF up to 50 MiB and 500 pages. Extraction is limited to 250,000 text characters and 2,000 matches per page, 10 million characters and 20,000 matches overall, and 2,000 annotations per page. Results carry a notice if limits or unreadable pages make them partial.

Keep going

What would you like to do next?

Choose a tool. Your finished PDF goes with you.

PDF

Advertisements

Language38

English العربية Français Italiano Deutsch Español Português Nederlands Русский Türkçe 日本語 中文 हिन्दी Bahasa Indonesia Bahasa Melayu 한국어 Tiếng Việt ไทย Polski Svenska Українська Norsk বাংলা Ελληνικά فارسی اردو Čeština Dansk Magyar Română Suomi Български Filipino ქართული Slovenčina Azərbaycanca עברית Slovenščina