Tools150

A PRACTICAL PDF GUIDE

How to convert PDF to XML with layout data

Get structured page data for your own workflow.

Export PDF text, page dimensions and drawing data as XML. Keep referenced images together and inspect the structure before using it in code.

Follow the guide
A REAL TOOL EXAMPLE
BeforeFirst page of the four-page PDF source.
  • 4 PDF pages
AfterPretty-printed opening excerpt of the actual document.xml export.
  • document.xml · 4 page records
  • 1 referenced PNG asset
The result is structured XML data, shown as an excerpt. The complete export includes all four pages and the referenced image asset.
Advertisements
QUICK VIDEO TUTORIAL How to convert PDF to XML with layout data Get structured page data for your own workflow. English narration Watch the video

How to convert PDF to XML with layout data

Advertisements

Open PDF to XML to export page records, text blocks and drawing data. Our four-page PDF produces document.xml plus one referenced PNG image inside ZIP.

01

Choose the PDF pages

Load the PDF and leave Pages blank for the whole document, or enter the subset you need. Our example includes all four pages.

02

Export and keep the assets together

Choose Start. An export with image assets downloads as ZIP. Extract the whole archive and keep the images directory beside document.xml.

The sample includes images/page-001-image-005.png. Image references point to asset paths instead of embedding binary data in the XML. An export without image assets can download directly as .xml.

03

Read the document structure

The sample declares version 1, units pt and a top-left origin. Page records include source page number, dimensions, rotation, blocks and drawings. Coordinates describe the unrotated page; account for rotation when overlaying the data.

Parse the XML with an XML parser. Repeated values use item elements; some geometric tuples are text values. Inspect the actual structure instead of assuming a publishing schema such as DocBook.

Check reading order and image references against the original PDF before using the data in your processing workflow.

A FEW USEFUL DETAILS

Common questions

Does the XML contain semantic chapters?

It describes extracted page data. Heading meaning, reading order and document semantics may need your own interpretation.

Why did I receive a ZIP?

The sample contains an image. ZIP keeps document.xml and the referenced asset together so the image can be located.

YOUR TURN

Try it with your document

Review the result, then save the version you need.

Open PDF to XML

Image detail

Advertisements

Language38

English العربية Français Italiano Deutsch Español Português Nederlands Русский Türkçe 日本語 中文 हिन्दी Bahasa Indonesia Bahasa Melayu 한국어 Tiếng Việt ไทย Polski Svenska Українська Norsk বাংলা Ελληνικά فارسی اردو Čeština Dansk Magyar Română Suomi Български Filipino ქართული Slovenčina Azərbaycanca עברית Slovenščina