PDF Merge and Split

Merge, split, reorder and rotate PDF pages in this tab. The bytes never leave your machine: there is no upload, no server and no network call of any kind in this page. Every file this tool cannot handle is refused by name, and every structure it does not carry across is listed per file before you download anything.

Files stack in the order you add them. Merging is selecting all pages of two or more files. Splitting is selecting a subset of one.

Before you add a file

This tool refuses rather than guesses. On a corpus of 302 real PDF files collected from one working machine, deduplicated by SHA-256, and filtered to exclude every file this tool had written itself, 276 built, which is 91.4 percent of the corpus and 98.9 percent of the 279 that were not encrypted. Of the rest, 23 were refused as encrypted, 2 as not a PDF, and 1 because a compressed stream could not be inflated. Two further refusal classes, damaged or absent cross reference data and damaged page tree, are implemented and did not occur in that corpus. The full table is further down this page.

It is a page level tool, so it does not carry everything. Bookmarks and form fields are not carried over. Neither are named destinations, page labels, tagged structure, optional content groups or document metadata. Every one of those that your file actually contains is named in that file's entry below, before you build anything. Page level annotations are carried.

Files (none added)

Output page order (0 pages selected)

Results

Self test

Runs the parser and the writer against fixtures built into this page. It includes positive controls: assertions that are deliberately wrong and must be reported as correctly rejected. A checker that has never produced a failure has not been tested.

What this tool refuses, and why

Damaged cross reference data is refused, not reconstructed. This tool never scans a file for obj patterns to guess a cross reference table back into existence. If a file is refused here, qpdf and pdftk are strictly more capable and will very likely handle it: both do encryption, damaged cross reference reconstruction and linearization, none of which this tool attempts. The reason to use this page instead is that the bytes never leave the machine, not that the feature set is larger.

The five refusal classes, with counts from a corpus of 302 real PDF files collected from one working machine, deduplicated by SHA-256, and filtered to exclude every file this tool had written itself, run through this exact engine. 276 of the 302 built, which is 91.4 percent of the corpus and 98.9 percent of the 279 that were not encrypted. Two classes did not occur in that corpus and are listed with a count of zero rather than left out.
Refusal classFiles in that corpusWhat triggers it
encrypted23 of 302/Encrypt present in the trailer. No empty password decrypt is attempted.
not a PDF2 of 302The first 1024 bytes do not contain %PDF-, the file is zero bytes, or the file is larger than the 536,870,912 bytes this page will hold.
stream could not be inflated1 of 302A compressed stream still throws after the raw deflate and endstream trim fallbacks. A truncated stream is refused, never decoded partly. This covers the payload of a cross reference stream too, so a hybrid revision whose /XRefStm sits at a readable offset but will not inflate is refused under this class rather than the damaged cross reference one.
damaged or absent cross reference data0 of 302No startxref found, an offset outside the file, an offset pointing at neither the xref keyword nor a cross reference stream object, a hybrid revision whose /XRefStm is not an offset inside the file or does not point at a cross reference stream, or an object that is not at the offset the cross reference data claims.
damaged page tree0 of 302The catalog has no /Pages, or the page walk yields zero pages.

All 276 outputs were read back with pypdf 6.12.2, which did not write any of them: page counts matched on 276 of 276 with no warnings raised. pypdf is a recovery tolerant reader and so is more forgiving than a strict structural checker. qpdf, pdftk, mutool and Ghostscript are not installed on the machine these numbers came from and none of them was run.

What is not carried across

This is a page level tool. It copies the transitive closure of the page objects you select and builds a fresh catalog and page tree around them. Anything that hangs off the source catalog rather than off a page is not carried, and every one present in your file is named in that file's entry above before you build:

Page level annotations (/Annots) are carried. A link annotation whose destination points at a page you did not include has that destination replaced with null, and the build result counts how many were dropped.