What Is Dark Data and Why It Matters for Data Management and Analytics

Dark data is the digital information your organization collects, stores, and then never uses for analytics, decision-making, or monetization. Most companies are sitting on a lot of it.
Splunk’s global research survey of more than 1,300 business and IT decision-makers found that 60% of respondents reported that half or more of their organization’s data qualifies as dark, and other estimates put the figure closer to 55% of everything a large business holds (source). That gap between what gets collected and what gets examined is where risk, wasted spend, and unclaimed value all live at once.
What Is Dark Data?
Working definition
Think of it as data that exists, costs money to keep, and answers no questions. Emails from departed employees. Sensor readings from a production line nobody plotted. Call recordings that satisfied a compliance box and then went quiet. It was generated during ordinary business activities, stored because storing felt safer than deciding, and then forgotten by everyone except the invoice from your storage provider.
Gartner’s framing
Gartner’s definition is the one most analysts still lean on: information assets an organization collects, processes, and stores during regular business activities but generally fails to use for analytics, business relationships, or direct monetization. Note the word “assets.” That framing is deliberate. The data has latent worth, and the failure is one of activation, not collection.
Not the same as unstructured data
This is where people muddle things. Unstructured data is a format problem, text documents, video, audio, images without a schema. Dark data is a usage problem. A perfectly tidy relational table sitting in a decommissioned reporting database is dark. A well-tagged video library feeding a machine learning model is not. Plenty of dark data is structured data, pulled from structured data sources, and still nobody touches it. Also worth clarifying: dark data has nothing to do with the dark web or the deep web, despite the unfortunate name collision.
| Concept | Defining trait | Example |
|---|---|---|
| Dark data | Collected but unused | Archived server logs from 2019 |
| Unstructured data | No predefined schema | Customer support call audio |
| ROT data | Redundant, obsolete, trivial | Seventeen copies of one contract |
| Active data | In regular analytical use | Live sales dashboard tables |
Where Dark Data Hides
Documents and PDFs
PDFs are the worst offenders, and I say that with feeling. Scanned contracts, signed forms, vendor invoices, exported reports. They carry sensitive data inside images that no index can read without OCR, and they multiply across shared drives, email attachments, and document platforms with version histories nobody prunes. You will meet final_FINAL_revised_v7.pdf. You will meet its cousins.
Logs and machine output
As much as 90% of sensor and raw machine output never gets examined at all, per IBM’s research. Application logs, IoT telemetry, network flow records, clickstreams. These are the datasets with the highest analytical ceiling and the lowest inspection rate, which is a strange thing to keep being true year after year.
Archives and legacy systems
Cold storage, backup tapes, that acquired subsidiary’s CRM nobody migrated. Cyberhaven’s security researchers make the sharpest point here: abandoned cloud buckets and legacy systems form an unseen attack surface, because you cannot encrypt, monitor, or patch what you do not know exists.
Why Does Dark Data Keep Piling Up?
Cheap storage defaults
Per-terabyte costs kept falling, so “keep everything” became the path of least resistance. Western Digital’s storage economics work notes high-capacity HDDs underpin roughly 80% of data center capacity precisely because the math favors retention over judgment (source). Cheap is not free, though. One analysis puts enterprise waste at up to $2.5 million a year in storing data nobody uses.
No clear ownership
Ask who owns the folder. Watch the silence. Unowned governance, not the storage tier itself, is the actual defect. When nobody can say what data is there, who owns it, how long it must be kept, and what can safely go, you don’t have a data management problem so much as an accountability vacuum.
Reactive governance
Most organizations don’t have a dark data strategy. They have a panic button, pressed during an audit, a customer security questionnaire, or a breach postmortem.
What Risks and Costs Does It Create?
Security exposure
Unmonitored, uncategorized data is a soft target, and Komprise’s governance guidance flags exactly that combination as the compliance and breach risk (source). Mimecast’s research (source) indicates most security incidents trace back to human error, frequently mishandled files, and the picture gets uncomfortable.
Privacy and compliance
GDPR gives individuals deletion rights that assume you know where their information lives. When a subject access request arrives and the relevant records exist in a PDF, its OCR text index, three email inboxes, a reporting export, and last quarter’s backup, “delete the file” stops being a meaningful instruction.
Wasted spend and lost insight
There’s an environmental line item too. Storing unused data generates over 5.8 million tonnes of CO2 annually, roughly the emissions of 1.2 million cars, and researchers at the University of Pretoria warn that data centers already account for about 2% of global greenhouse gases with that share projected to roughly double by 2030 (source). Much of it is redundant, obsolete, trivial.
Discover and Classify What You Already Hold
Start narrow, not heroic. Alation’s data intelligence team frames dark data as an insight problem before a cleanup problem, and that ordering matters (source). A workable sequence:
- Inventory the repositories, including the ones people insist aren’t repositories.
- Scan and classify for sensitive data, applying metadata as you go.
- Score each store by risk exposure and analytical potential.
- Assign a named owner. Not a team. A person.
Modern data catalogs make step two survivable, and academic work on AI-driven mitigation suggests automated discovery and classification is now the only realistic way to cover an entire data estate.
Govern With Retention, Ownership, and Access Rules
Prevention beats archaeology. Classify inputs at the moment of capture, enforce retention at the document management layer, and gate uploads so anything containing sensitive information gets flagged immediately rather than discovered three years later by an auditor.
Why Total Deletion Rarely Works
Deletion isn’t an act. It’s a chain reaction across file copies, caches, search indexes, version history, downstream exports, and backups, and legal holds routinely override your intentions anyway. I’ve watched purge projects grind teams down chasing OCR artifacts and obsolete versions. Treat dark data like a fire hazard instead: contain and reduce exposure, don’t try to sandblast every ember out of the building.
Activate the Data Worth Keeping
Some of this stuff is genuinely a digital treasure trove. Retrieval-Augmented Generation lets AI pipelines query archived documents without restructuring them first, which changes the economics of keeping them. Predictive maintenance from unexamined sensor histories, personalized marketing from dormant behavioral records, NLP over support transcripts. The one caution, and reporting on AI companies harvesting training material makes it plain, is that curation quality decides whether machine learning models learn something useful or just inherit your data quality issues at scale.
FAQ
Is dark data the same as big data?
No. Big data describes volume and velocity. Dark data describes neglect.
Does deleting it always reduce risk?
Usually, but not if retention obligations apply. Check the legal hold first.
Can small businesses have dark data?
Absolutely. Shared drives and old email accounts are enough.
Conclusion
The organizations handling this well aren’t the ones with the cleanest archives. They’re the ones who can answer four questions quickly: what’s there, who owns it, how long it stays, and what can go. Everything else, the tooling, the classification, the analytics, follows from those answers.
Would you like to receive similar articles by email?


