Document Metadata: When Your Own Files Talk About You

Nobody sent them that. It was inside the file.

Law Firm Cybersecurity · Chapter 4~7 min read

Your firm publishes a PDF — a practice guide, a form, a brief. Someone downloads it, and without opening a single page of the actual document, learns the name of the associate who drafted it, the name of the partner who revised it, the internal file path it was saved to (which happens to reveal how your matters are organized and named), the software version your office runs, and the fact that it was last edited at 11:47pm on a Sunday.

Nobody sent them that. It was inside the file.

What metadata is, in one sentence

Metadata is data about data — information a file quietly carries about itself.1

Every document, spreadsheet, and image can carry it. It's generated automatically, it's usually invisible in normal use, and it survives being emailed, uploaded, and downloaded unless something specifically strips it out. Most of the time nobody notices, because nobody looks.

The adversary looks.

What it leaks, and why a firm should care

Names. Author, last-modified-by, reviewers, comment authors. String together the metadata from a dozen documents on your website and you have a substantial chunk of your staff list — including people who don't appear on your team page.

Internal structure. File paths (\\FIRM-DC01\Matters\2023\Client_Name\...) reveal your server naming, your matter numbering, and sometimes a client name that was never meant to be public.

Software and versions. Which suite, which version, which PDF generator. To an adversary, that's a shopping list: it tells them precisely what's worth trying against you.

Edit history and residue. Tracked changes. Deleted text that isn't as deleted as it looks. Comments. Earlier drafts embedded in the same file. This is the one that has ended careers — the settlement draft that reveals your client's actual walk-away number in the tracked changes, the brief whose comment bubbles contain candid assessments of your own case.

Interactive — see what's inside
SETTLEMENT AGREEMENT AND MUTUAL RELEASE — DRAFT

This Settlement Agreement ("Agreement") is entered into between Halbrook Manufacturing Ltd. ("Halbrook") and Vandermeer Logistics Inc. ("Vandermeer"), collectively "the Parties," who agree as follows…
This is a sample document — every name, firm, path, and matter invented. Nothing you have is being read. We'd never ask you to upload a real one, and you should be suspicious of any site that does.

The text was fine. Everything damaging was in the parts nobody thinks to check.

The photograph problem

Same principle, different file, and it's worth a moment because it's where people are most surprised.

Photographs have historically carried EXIF metadata — the timestamp, the device, the camera settings, and in many cases the GPS coordinates where the photo was taken. Free tools read it instantly.

There's genuine good news: most major social platforms now strip this data on upload, specifically to protect users.2 But that protection is not general. It applies where the platform chose to apply it. A photo emailed directly, or uploaded straight to your own website's media library, frequently retains everything it was born with.

So: the team photo on your About page. The candid from the firm retreat. The photo of a document that someone took with their phone rather than scanning. Any of those may be carrying coordinates and a timestamp.

That's the bridge to Chapter 12, where we deal with the social-media and geolocation side properly. For now, one rule: a photo is a file, and files carry metadata.

What this means for your professional obligations

In the United States

Rule 1.6(c)'s reasonable-efforts standard covers inadvertent disclosure — not just the dramatic breach, but the quiet leak.3 Transmitting a document with confidential metadata still embedded is textbook inadvertent disclosure, and it's one of the very few where the fix is a button rather than a budget.

In Canada

The confidentiality duty covers all information concerning a client's business and affairs acquired in the professional relationship — a broader concept than privilege, and it does not carve out the parts you disclosed by accident.4 The technological-competence commentary points directly at this kind of thing: understanding the risks of the tools you use, where the tool is Microsoft Word and the risk is that it remembers.5

The practical takeaway for both countries is the same, and it's unusually cheerful for a security post: this is a solved problem. The tools exist, they're built into software you already own, and the fix is a step in a workflow rather than a purchase.

What actually fixes this

1. Scrub before it leaves. Office's Document Inspector and Acrobat's equivalent remove metadata, comments, and revision history in a few clicks. The step is trivial. The discipline is the hard part.

2. Make it a policy, not a memory. "Everyone remembers to scrub" is not a control — it's a hope with a good record until the one Friday it isn't. Bake it into the workflow: nothing leaves the firm without going through the scrub step, the same way nothing gets filed without a check.

3. Send PDFs, and flatten them. Converting to PDF is not scrubbing — a PDF happily carries its own metadata, and "print to PDF" from a Word document can bring the residue along. Convert, then inspect, then send.

4. Strip EXIF before anything goes on the site. Any image published to your website should be stripped as part of the upload workflow. This is straightforwardly automatable — it should never be a human's job to remember.

5. Audit what's already published. Everything on your site right now was published under whatever your habits were at the time. Nobody has looked. It's worth looking — the documents on your website have been sitting there, downloadable, for years.

Michael Bazzell's Extreme Privacy is the practical reference for the scrubbing side of this if you want to go deeper than a policy and into technique.6

How we can help

We'll audit every document currently published on your website and tell you what's inside them — the names, the paths, the residue. Then we put a scrub-on-publish step into your site's workflow so it stops accumulating, and strip EXIF from images automatically on upload.

It's the kind of problem that's invisible until someone looks, and cheap to solve once they have.

Request a published-document metadata audit

We crawl the documents already on your site and report what's inside them. No access required — we only read what's already downloadable by the public. That's the point.

Your firm's metadata checklist

Spencer McLennan
Spencer McLennan is the founder and lead webmaster of LegalWebmasters, and has handled websites, hosting, and security for law firms and professional practices across the United States and Canada since 2007. He holds an MBA, a PCM from the American Marketing Association, and is a graduate student in intelligence studies. He writes the Law Firm Cybersecurity and Law Firm SEO & Design guides. Connect on LinkedIn.

Footnotes

  1. Rae Baker, Deep Dive: Exploring the Real-World Value of Open Source Intelligence 156 (Wiley 2023).
  2. Id. at 157.
  3. Model Rules of Pro. Conduct r. 1.6(c) (Am. Bar Ass'n 2023).
  4. Model Code of Pro. Conduct r. 3.3-1 (Fed'n of L. Soc'ys of Can.).
  5. Id. r. 3.1-2 cmts. [4A]–[4B].
  6. Michael Bazzell, Extreme Privacy: What It Takes to Disappear (2d ed. 2020).

Trusted by professionals like you.

Crease Harman LLP Borders Law Group David Aujla, Immigration Lawyer