Business

What Is Dark Data and Why It Matters for Data Management and Analytics

What Is Dark Data

Dark data is the digital information your organization collects, stores, and then never uses for analytics, decision-making, or monetization. Most companies are sitting on a lot of it.

Splunk’s global research survey of more than 1,300 business and IT decision-makers found that 60% of respondents reported that half or more of their organization’s data qualifies as dark, and other estimates put the figure closer to 55% of everything a large business holds (source). That gap between what gets collected and what gets examined is where risk, wasted spend, and unclaimed value all live at once.

What Is Dark Data?

Working definition

Think of it as data that exists, costs money to keep, and answers no questions. Emails from departed employees. Sensor readings from a production line nobody plotted. Call recordings that satisfied a compliance box and then went quiet. It was generated during ordinary business activities, stored because storing felt safer than deciding, and then forgotten by everyone except the invoice from your storage provider.

Gartner’s framing

Gartner’s definition is the one most analysts still lean on: information assets an organization collects, processes, and stores during regular business activities but generally fails to use for analytics, business relationships, or direct monetization. Note the word “assets.” That framing is deliberate. The data has latent worth, and the failure is one of activation, not collection.

Not the same as unstructured data

This is where people muddle things. Unstructured data is a format problem, text documents, video, audio, images without a schema. Dark data is a usage problem. A perfectly tidy relational table sitting in a decommissioned reporting database is dark. A well-tagged video library feeding a machine learning model is not. Plenty of dark data is structured data, pulled from structured data sources, and still nobody touches it. Also worth clarifying: dark data has nothing to do with the dark web or the deep web, despite the unfortunate name collision.

ConceptDefining traitExample
Dark dataCollected but unusedArchived server logs from 2019
Unstructured dataNo predefined schemaCustomer support call audio
ROT dataRedundant, obsolete, trivialSeventeen copies of one contract
Active dataIn regular analytical useLive sales dashboard tables

Where Dark Data Hides

Documents and PDFs

PDFs are the worst offenders, and I say that with feeling. Scanned contracts, signed forms, vendor invoices, exported reports. They carry sensitive data inside images that no index can read without OCR, and they multiply across shared drives, email attachments, and document platforms with version histories nobody prunes. You will meet final_FINAL_revised_v7.pdf. You will meet its cousins.

Logs and machine output

As much as 90% of sensor and raw machine output never gets examined at all, per IBM’s research. Application logs, IoT telemetry, network flow records, clickstreams. These are the datasets with the highest analytical ceiling and the lowest inspection rate, which is a strange thing to keep being true year after year.

Archives and legacy systems

Cold storage, backup tapes, that acquired subsidiary’s CRM nobody migrated. Cyberhaven’s security researchers make the sharpest point here: abandoned cloud buckets and legacy systems form an unseen attack surface, because you cannot encrypt, monitor, or patch what you do not know exists.

Why Does Dark Data Keep Piling Up?

Cheap storage defaults

Per-terabyte costs kept falling, so “keep everything” became the path of least resistance. Western Digital’s storage economics work notes high-capacity HDDs underpin roughly 80% of data center capacity precisely because the math favors retention over judgment (source). Cheap is not free, though. One analysis puts enterprise waste at up to $2.5 million a year in storing data nobody uses.

No clear ownership

Ask who owns the folder. Watch the silence. Unowned governance, not the storage tier itself, is the actual defect. When nobody can say what data is there, who owns it, how long it must be kept, and what can safely go, you don’t have a data management problem so much as an accountability vacuum.

Reactive governance

Most organizations don’t have a dark data strategy. They have a panic button, pressed during an audit, a customer security questionnaire, or a breach postmortem.

What Risks and Costs Does It Create?

Security exposure

Unmonitored, uncategorized data is a soft target, and Komprise’s governance guidance flags exactly that combination as the compliance and breach risk (source). Mimecast’s research (source) indicates most security incidents trace back to human error, frequently mishandled files, and the picture gets uncomfortable.

Privacy and compliance

GDPR gives individuals deletion rights that assume you know where their information lives. When a subject access request arrives and the relevant records exist in a PDF, its OCR text index, three email inboxes, a reporting export, and last quarter’s backup, “delete the file” stops being a meaningful instruction.

Wasted spend and lost insight

There’s an environmental line item too. Storing unused data generates over 5.8 million tonnes of CO2 annually, roughly the emissions of 1.2 million cars, and researchers at the University of Pretoria warn that data centers already account for about 2% of global greenhouse gases with that share projected to roughly double by 2030 (source). Much of it is redundant, obsolete, trivial.

Discover and Classify What You Already Hold

Start narrow, not heroic. Alation’s data intelligence team frames dark data as an insight problem before a cleanup problem, and that ordering matters (source). A workable sequence:

  1. Inventory the repositories, including the ones people insist aren’t repositories.
  2. Scan and classify for sensitive data, applying metadata as you go.
  3. Score each store by risk exposure and analytical potential.
  4. Assign a named owner. Not a team. A person.

Modern data catalogs make step two survivable, and academic work on AI-driven mitigation suggests automated discovery and classification is now the only realistic way to cover an entire data estate.

Govern With Retention, Ownership, and Access Rules

Prevention beats archaeology. Classify inputs at the moment of capture, enforce retention at the document management layer, and gate uploads so anything containing sensitive information gets flagged immediately rather than discovered three years later by an auditor.

Why Total Deletion Rarely Works

Deletion isn’t an act. It’s a chain reaction across file copies, caches, search indexes, version history, downstream exports, and backups, and legal holds routinely override your intentions anyway. I’ve watched purge projects grind teams down chasing OCR artifacts and obsolete versions. Treat dark data like a fire hazard instead: contain and reduce exposure, don’t try to sandblast every ember out of the building.

Activate the Data Worth Keeping

Some of this stuff is genuinely a digital treasure trove. Retrieval-Augmented Generation lets AI pipelines query archived documents without restructuring them first, which changes the economics of keeping them. Predictive maintenance from unexamined sensor histories, personalized marketing from dormant behavioral records, NLP over support transcripts. The one caution, and reporting on AI companies harvesting training material makes it plain, is that curation quality decides whether machine learning models learn something useful or just inherit your data quality issues at scale.

FAQ

Is dark data the same as big data?

No. Big data describes volume and velocity. Dark data describes neglect.

Does deleting it always reduce risk?

Usually, but not if retention obligations apply. Check the legal hold first.

Can small businesses have dark data?

Absolutely. Shared drives and old email accounts are enough.

Conclusion

The organizations handling this well aren’t the ones with the cleanest archives. They’re the ones who can answer four questions quickly: what’s there, who owns it, how long it stays, and what can go. Everything else, the tooling, the classification, the analytics, follows from those answers.

Would you like to receive similar articles by email?

Paul Tomaszewski is a science & tech writer as well as a programmer and entrepreneur. He is the founder and editor-in-chief of CosmoBC. He has a degree in computer science from John Abbott College, a bachelor's degree in technology from the Memorial University of Newfoundland, and completed some business and economics classes at Concordia University in Montreal. While in college he was the vice-president of the Astronomy Club. In his spare time he is an amateur astronomer and enjoys reading or watching science-fiction. You can follow him on LinkedIn and Twitter.

Leave a Reply

Your email address will not be published. Required fields are marked *