Blog

Two ways to save a PDF, and why redaction uses the other one

A PDF is not one long run of text. It is a set of numbered objects — pages, fonts, the content streams that paint each page, annotations — with a table at the end saying where each one starts. That structure allows two very different ways of saving a change, and OpenPdfEdit uses both on purpose.

Appending: what most edits do

The format lets a file change without touching what is already in it. The new and changed objects are written after the end of the file, followed by a new table that points at them. This is called an incremental update. A reader follows the newest table, so it shows the edited document. The older objects are still there, earlier in the file.

OpenPdfEdit saves this way for almost every edit. A highlight, a sticky note, a replaced line of text, a rotated or deleted page: each one compares the document with how it stood at the last save and appends only the objects that changed. The bytes before that point are never rewritten.

That has two practical results. A digital signature already in the file survives an annotation, because the bytes it covers are never touched. And each save is a revision on top of the last one, not a new file that quietly replaces it.

Why that is wrong for redaction

The property that protects a signature defeats a redaction. Redacting in OpenPdfEdit removes the text from the page. The page's content stream is interpreted, and whatever paints inside the box you drew is taken out. Text is measured with the font's real widths where they are known, so a box over the address in Email: someone@example.com removes the address and keeps the label. The black box drawn on top is a second line of defence, not the mechanism.

Now picture that change appended. The page would point at a new content stream, while the old one — the one with the text in it — stayed in the file, one revision back. Anything that reads the raw file could still find it. The removal would be real, and the file would still carry a copy of what was removed.

So redaction rewrites the whole file

When you redact, OpenPdfEdit writes the document from scratch instead. Before it does, it drops every object that nothing in the document points to any more. That step is the one that matters. Redaction leaves the old content stream orphaned, and a plain rewrite would still write the orphan out. Pruning it is what makes the removal stick.

Find & redact follows the same rule. Every match you keep after reviewing the list is removed in one step: one rewrite of the file, one entry in the undo history. Then the app reads the saved file back. If any redacted area still has readable text in it, it names the page and asks you to deal with it, rather than counting the job as done.

What it costs

A full rewrite breaks an existing digital signature. That is not a side effect of rewriting. A signature exists to show that the signed bytes have not changed, and redaction changes them. Redacting a signed document was always going to invalidate the signature, whichever way the file was written.

Compress is the other operation that writes a file from scratch. It saves a new copy and leaves the open document alone. Even at its lossless level, that copy drops the chain of revisions appended edits build up, along with any orphaned objects. Signatures do not carry over to that copy either.

Checking it yourself

All of this happens on your own machine, in the desktop app or in the browser. The document is never uploaded. To check a redaction, open the saved file in any PDF reader, search for the words you removed, and try copying the text around where they were. Neither should find them.

The code that chooses between the two save paths is public, in the openpdfedit-doc and openpdfedit-session crates of the source on GitHub, under AGPL-3.0-or-later. How to redact a PDF walks through the tool itself.

Open the app All posts