What Your PDF Knows About You: Metadata and How to Clear It

Beyond its visible content a PDF carries bookkeeping fields: who authored it, which program created it, when it was modified. Sometimes that includes a full name, an internal project codename, or a path like C:\Users\jsmith\Desktop.

The main store is the Info dictionary: Title, Author, Subject, Keywords, Creator (the program the document was authored in) and Producer (whatever turned it into a PDF), plus creation and modification dates. Alongside it there may be an XMP block – an extended metadata section where editors write their own fields. Separately, the file can retain comments, page thumbnails, and traces of earlier saves.

Metadata leaks rarely look dramatic, but they happen constantly. An anonymously published document with a surname in the Author field. Tender paperwork whose Creator names the client's internal system. A report 'from the company' that, judging by Producer, was made in a free home-use tool. None of this shows on the page – which is exactly why it is worth checking before publication rather than after.

The workflow is simple: look at what the file actually contains, then clear or rewrite the fields that matter. For real anonymisation, metadata alone is not enough – it is worth also running the document through sanitisation (which strips scripts and attachments) and a hidden-layer check. If personal data appears in the body text, remove it by redaction, not by editing metadata.

Inspect PDF metadata