Somewhere on your company's website there's a PDF nobody has opened in months. A whitepaper from some product launch. An old press release. Maybe an RFP response you sent to a client two years ago that they, weirdly, posted publicly on their own site. It looks fine. Clean, on-brand, exactly what you meant to say to the world.
Now open the properties panel. Or run it through one of those free metadata inspector sites. You might find the real name of whoever wrote it. An internal file path. A paragraph someone deleted that, technically, is still sitting inside the file. A comment from a reviewer arguing about a number that never even made it into the final version. Sometimes just the software version used to build it, which tells a stranger more about your internal systems than you'd like.
This isn't a hypothetical scenario I'm making up to scare you. It's one of the more common and least understood gaps in how businesses handle information. We spend real money locking down networks, training staff to spot phishing emails, encrypting sensitive files, and then someone on the marketing team publishes a document with tracked changes still baked into it. Kind of like shredding your bank statements and then leaving the shredder bin open on the curb, with the strips more or less in order.
All the Stuff You've Already Published
Most companies never think about their public documents as one connected pile of information. They think of them one at a time. This press release. That annual report. This job posting from last spring. But to someone doing competitive research, or due diligence before an acquisition, or just poking around with less friendly intentions, none of that is separate. It's a dataset. A pretty rich one, honestly.
Go through what's probably sitting out there already: financial filings and investor decks if you're public or fundraising. Press releases about new hires, partnerships, product launches. Case studies and whitepapers written to make the company look smart. RFP responses, some of which get reposted by the client without anyone at your company knowing. Patent filings, which tend to be detailed by nature. Job postings, and honestly these are underrated, since they'll casually describe internal tools and team structure in more depth than a press release ever would. Old conference slides that show up on a speaker's personal site a year later. Support documentation. Sustainability reports. Even a page torn out of the employee handbook that got shared during a hiring process because someone thought it was a nice touch.
None of these look dangerous on their own. String enough of them together and you've got a surprisingly detailed picture: how the company runs, who works there, what software they use, how teams are organized, sometimes even what's coming next. And that's all before you've looked at a single byte of metadata.
What Metadata Actually Is, and Why It's Not Just an IT Thing
People hear "metadata" and picture something technical, the kind of thing forensic analysts worry about, not something a comms team should care about. That's the mistake. Metadata is just data about the data. Every file carries it whether anyone in the room thought about that or not.
Right click a Word doc, a PDF, a spreadsheet, check the properties. You'll often see the author's actual name, whoever last touched the file, sometimes the organization the software license belongs to, timestamps for creation and every save after that, and frequently the exact software version used to build it. PDFs will sometimes carry details about the printer or converter that made them, which can leak internal network info nobody meant to share. Photos pasted into a slide deck can still have GPS coordinates baked in, so that friendly "team offsite" photo might tell a total stranger the exact address of your office.
None of this needs hacking. No exploit, no special access required. Usually it's one click. Properties on Windows, Get Info on a Mac, File then Info inside any Office app. There are also free websites that'll strip an uploaded file down to its metadata in about ten seconds flat.
Here's the part that actually matters: metadata often tells a more honest story than the polished content sitting on top of it. A press release gets reviewed by five people before it goes out, sure, but the metadata underneath can still say "Draft_v3_INTERNAL_DO_NOT_SHARE," or list an author whose job title reveals more about internal reporting lines than anyone intended.
The Stuff Left Behind
Metadata's only half the problem. The other half is content that's technically still in the file but was never supposed to be seen, leftovers from the editing process that nobody bothered to clean up before hitting publish.
Anyone who's used Track Changes in Word knows how easy it is to accept the visible edits and forget the revision history is still sitting there underneath. Depending on how the file gets exported to PDF, some of that history can still be pulled out later, or at least the deleted text recovered by someone who knows where to poke around inside the file structure. Comments are an even more common leak. Reviewers leave blunt notes in the margins. "This number's wrong, fix before sending." "Legal hasn't signed off on this yet." "Don't mention the delay publicly." If the document goes out without a real final export pass, those notes go right along with it.
Then there's hidden or white text. Sometimes left there on purpose for shady SEO reasons, more often just a leftover from someone changing the font color to leave themselves a private note and forgetting to delete it before the thing shipped. Spreadsheets have their own version of this. Hidden rows, hidden columns, entire hidden tabs that don't show up unless someone thinks to unhide them, or unless they open the file in a different tool that just shows everything by default.
PowerPoint decks are their own special mess. Speaker notes are a classic offender. Teams write unfiltered, honest talking points assuming only the presenter will ever see them, then send the deck around without stripping that layer out first. Deleted slides sometimes remain technically recoverable inside the file. And embedded objects, like a spreadsheet pasted straight into a slide, can carry the whole original file behind it. Every tab. Every column. Nothing cleaned up.
What All This Actually Reveals
Worth getting concrete here, because the risk isn't abstract at all.
Real names and titles show up in metadata even when a document is presented as coming from "the team" instead of a person. That matters more than it sounds like it should. Phishing works a lot better when there's a real name attached, a real title, a real project to reference, instead of a generic "Dear Sir/Madam." A file that says "prepared by Jane Doe, Senior Security Engineer, for internal review 3/14" hands an attacker a specific target and a believable excuse to reach out.
Internal file paths, sometimes visible in metadata or in embedded object links, can hint at network structure or leak project codenames that were never meant to leave the building. Something like \corp-fileserver\Projects\ProjectPhoenix_M&A\ tells a competitor that something called Project Phoenix exists and is apparently important, even if the visible document says nothing about it at all.
Software version info tells a quieter, slower story. One file built on an outdated version of some office suite is a tiny clue on its own. Enough tiny clues start to add up though, especially paired with job postings naming specific internal tools, or an engineer casually mentioning your stack at a conference. A patient enough researcher can piece together a decent map of your infrastructure that way, which happens to be exactly the groundwork that comes before a targeted technical attack.
Draft content and internal disagreements, preserved in comments and version history, can be embarrassing even when there's nothing actually wrong or unethical in them. Seeing that legal pushed back on a claim that shipped anyway, or that two reviewers argued over a statistic before it went public, chips away at trust in a way that has nothing to do with wrongdoing and everything to do with just looking sloppy.
And organizational structure becomes visible almost by accident, just from accumulation. Enough author names and titles across enough documents, and you've basically built an org chart for anyone patient enough to sit there and compile it. Useful for a competitor trying to poach talent, useful for a private equity firm doing diligence, useful for pretty much anyone trying to figure out who actually calls the shots, regardless of what the About Us page wants you to believe.
This Has Happened Before
None of this is new, which is part of why it's a little surprising it keeps happening.
One of the more famous examples in security circles goes back to a government document from the early 2000s, published to back up a policy position. Journalists dug into its metadata and found a detailed editing history showing how a document presented as independent analysis had actually been reshaped by communications staff along the way. The fallout had less to do with what the document said out loud and everything to do with what its hidden history quietly gave away. It's still used in security training today because it showed, pretty cleanly, how the invisible layer of a document can end up being the whole story.
Legal cases have produced their own version of this more than once. Litigation involving major tech companies has hinged, at least partly, on metadata pulled from documents submitted as evidence. A timestamp that contradicted an official timeline. Authorship info revealing who actually wrote something attributed to somebody else. In cases like that the metadata isn't a footnote. It becomes evidence, mostly because it's a lot harder to argue with than the words on the page.
Outside of the headline cases there are plenty of smaller, quieter incidents nobody hears about. A job listing that accidentally revealed a product name months before launch. A "sanitized" case study PDF that still had a client's real internal project name sitting in its file properties, which is an awkward call to have to make. These rarely blow up into anything huge. But they wear trust down bit by bit, and every so often they hand exactly the wrong person exactly what they needed.
Why This Keeps Happening
If everyone kind of knows this is a risk, why does it keep repeating? Mostly because document hygiene falls into a gap between departments that nobody's job description quite covers. Marketing cares about messaging, not metadata. Legal reviews for liability, not file properties. IT and security are staring at the network perimeter, not wondering what PDF export settings the comms team happened to use last Tuesday. Nobody owns the whole problem, so it never gets consistently fixed.
There's also just a human habit at work here. People treat the visible layer of a document as the entire document. If it looks clean on screen, it feels done. Almost nobody checks File then Info before hitting publish, and plenty of people don't realize that saving as a PDF doesn't automatically wipe out everything a file's picked up over its editing life. The tools we use to make these documents were built for collaboration. That history is a feature while you're drafting something and a liability the moment it goes public.
And then there's just scale. A company might have published thousands of documents over the years. Old press releases, archived case studies, job postings still indexed somewhere nobody remembers exists. Getting the hygiene right on one file isn't hard. Getting it right consistently, across years of output and a dozen different people with a dozen different habits, is a much harder thing to pull off by willpower alone.
What Actually Helps
None of this needs a specialized security team or a huge budget. Mostly it just needs a few habits to become standard instead of optional.
Build a final export step into the process for anything headed out to a public audience. Word and Excel both have an Inspect Document feature buried in there (sometimes labeled Check for Issues) that flags and strips metadata, comments, tracked changes, hidden rows and columns, and speaker notes before export. Takes maybe fifteen seconds. Should be as automatic as running spellcheck.
Treat the PDF export step as a security move, not just a formatting one. A genuinely flattened PDF, printed to PDF rather than saved straight out of a live editing file, strips out a lot of what causes trouble later. There are cheap or free tools built specifically to scrub author info and embedded data out of a finished PDF too, and running things through one before publishing is basically free insurance against an embarrassing find later.
Set a naming convention for anything that might get referenced or embedded in a public file, so internal codenames or sensitive folder paths don't slip through by accident. Less about technology, more about a simple rule: nothing built directly off an internal working file should ever be what actually ships. Duplicate it, rename it, work off the clean copy from there.
Go back and audit what's already out there. Most companies have never once pulled their own library of public PDFs and run them through a basic inspection tool. Do it once and you'll probably find a few things you didn't expect. Better you find them than a journalist or a competitor does.
Train the people who actually make this stuff, not just IT or security. Marketing writers, sales engineers answering RFPs, HR staff posting job listings, executives contributing slides to an investor deck. These are the people generating the risk day to day, and they're usually the last people anyone's ever bothered to tell that metadata is a thing worth thinking about. A short, practical session with real examples tends to land better than any policy document that's going to sit unread in a shared drive.
And build it into the templates themselves, so it doesn't depend on someone remembering under deadline pressure at 6pm the night before launch. If your standard templates disable automatic change tracking by default, or your export process automatically routes through a cleaning step, you've taken the dependency on any one person's memory out of the equation entirely. Systems that work by default beat policies that rely on people remembering, pretty much every time.
The Bigger Picture
This isn't really about paranoia. Most of what a company publishes is genuinely harmless, and most of the metadata sitting quietly in your archive will never be looked at by anyone with bad intentions. But the math here is lopsided. Cleaning a document up takes seconds. The cost of a leaked internal comment, or an employee's name showing up in a targeted phishing attempt, or a technical detail that helps someone map out your systems, can be wildly out of proportion to how little effort it would've taken to prevent in the first place.
Public and intentional aren't the same thing. A company controls what it means to say out loud. It doesn't automatically control everything a file happens to be carrying once it's left the building. Closing that gap isn't about getting secretive or locking everything down. It's really just about making sure the version of your company the world sees is the one you actually meant to show them. Nothing more, nothing accidentally extra. That's a small, achievable goal, and it starts with something as unglamorous as checking a document's properties before you hit publish.