Open PDF to XML to export page records, text blocks and drawing data. Our four-page PDF produces document.xml plus one referenced PNG image inside ZIP.
Choose the PDF pages
Load the PDF and leave Pages blank for the whole document, or enter the subset you need. Our example includes all four pages.
Export and keep the assets together
Choose Start. An export with image assets downloads as ZIP. Extract the whole archive and keep the images directory beside document.xml.
The sample includes images/page-001-image-005.png. Image references point to asset paths instead of embedding binary data in the XML. An export without image assets can download directly as .xml.
Read the document structure
The sample declares version 1, units pt and a top-left origin. Page records include source page number, dimensions, rotation, blocks and drawings. Coordinates describe the unrotated page; account for rotation when overlaying the data.
Parse the XML with an XML parser. Repeated values use item elements; some geometric tuples are text values. Inspect the actual structure instead of assuming a publishing schema such as DocBook.
Check reading order and image references against the original PDF before using the data in your processing workflow.
Common questions
Does the XML contain semantic chapters?
It describes extracted page data. Heading meaning, reading order and document semantics may need your own interpretation.
Why did I receive a ZIP?
The sample contains an image. ZIP keeps document.xml and the referenced asset together so the image can be located.
Try it with your document
Review the result, then save the version you need.

