What a PDF says in its properties
Open any PDF, look at its properties, and there is usually a name. Sometimes it is the author. Sometimes it is whoever installed the software, or a colleague whose template was reused, or a company that has nothing to do with the document. Along with the name come two dates, the program that made the file and the program that last saved it. All of it is a few clicks away for anyone you send the file to.
What a PDF carries
- Document properties. Title, author, subject, keywords, the application that created the content and the one that produced the PDF, creation date and modification date.
- An XMP packet. The same fields again in a second format, plus editing history, document identifiers and whatever the authoring tool wanted to record.
- Identifiers. A document ID that stays the same across versions, which links a draft to the final and a leaked copy to the original.
None of this is in the page. You can read a PDF from cover to cover and never see the author field. That is what makes it easy to forget.
How to remove it
Open the tool, stay on the Files tab, drop the PDF in and press Clean. The properties
and the XMP packet are cleared and the file is written back with its pages intact. You get
name.cleaned.pdf. Then drop the cleaned file into the Metadata tab to read it
back and confirm the fields are empty. That round trip is the only proof that matters.
What cleaning does not do
Metadata is the label on the box. It is not the contents. A name in the header of every page, a signature block, a watermark drawn across the page, a comment left in the margin, the text of a tracked change that was flattened into the page: all of that is content, and it stays. If a document needs a name taken out of the page itself, that is redaction, which is a different job and needs a tool that rewrites the page.
Be careful with scanned PDFs too. The scan is an image, and the image can carry its own tags from the scanner or the phone that photographed the page. Cleaning the PDF properties does not reach inside an embedded image. If the scan came from a phone, clean the photo before you make the PDF.
When it matters
- Sending a CV or a cover letter. The author field may name a friend who helped, or an old employer whose template was used.
- Sending a quote or invoice. The producer field can name a product the client does not expect you to be using.
- Publishing anything anonymously. The document ID and the author field will identify you long after the name is gone from the page.
- Sharing a draft. The dates say how long it took, and the modification history says how many times.
Other documents
Word files and ODT files carry the same kind of properties in a different container. They go through the same Files tab and have their own guide on removing metadata from Word documents. Markdown, HTML, SVG and plain text files are also accepted, and are cleaned of hidden tag blocks and invisible characters as well as properties.
Common questions
Does removing PDF metadata change the pages?
No. The document properties and the XMP packet are cleared and the pages are written back as they were. Text, images and layout are untouched.
Will it remove my name from the page itself?
No. A name printed in the page, a header, a signature or a comment is content, not metadata. Removing it is redaction and needs a tool that rewrites the page.
How do I check the metadata is gone?
Drop the cleaned PDF into the Metadata tab. The author, creator, producer and date fields should come back empty.