Data Discovery and Classification Solutions
Find the sensitive data scattered across file servers, databases and user devices with Veriket and DECE, and label it by sensitivity.
You Cannot Protect Data You Do Not Know About
Organizations produce millions of files, documents, e-mails and data records over the years.
This data can be scattered across;
- File servers,
- User computers,
- Network Share areas,
- Databases,
- Microsoft 365 environments,
- SharePoint,
- Cloud storage systems,
- Archive systems
these locations.
The real security problem is mostly not the existence of the data but the organization not knowing which data it holds and where that data is located.
Data Discovery & Classification is the data security approach providing for corporate data sources to be scanned so sensitive data is discovered, identified, classified and managed according to security policy.
Within SecureSys Data Discovery & Classification Solutions we carry out projects for making organizations' structured and unstructured data assets visible with the Veriket and DECE technologies.
What Is Data Discovery?
Data discovery is the process of scanning the data held on an organization's different systems through automated or controlled methods so sensitive content is identified.
The aim is to answer these questions:
Which data do we hold? Where is the data located? Does it contain sensitive data? Is it personal data? Is it special category personal data? Is it financial data? Is it corporate confidential information? Who holds access?
Without this visibility, building an effective data security policy is rather difficult.
What Is Data Classification?
Data classification is the separation of data into different categories according to sensitivity and business value.
Inside an organization, for example, a classification model can be built in the form of;
Public → Internal → Confidential → Strictly Confidential
this scale.
Another organization, in turn, may prefer the;
Public Internal Confidential Restricted
model.
What matters is that the classification model is compatible with the organization's information security policy.
Structured and Unstructured Data
Two core data types are present in data discovery projects.
Structured Data
Data held within defined data models.
For example:
- Oracle
- Microsoft SQL Server
- PostgreSQL
- MySQL
- ERP databases
- CRM databases
Unstructured Data
File and document based content.
For example:
- Word
- Excel
- Text
- CSV
- Shared folders
- User documents
- Archive files
In a modern data security strategy both data types need making visible.
Sensitive Data Discovery
One of the core duties of data discovery systems is identifying sensitive information automatically.
For example, data types such as;
Turkish ID Number Telephone Number E-Mail Address Postal Address IBAN Credit Card Customer Number Personnel Information Financial Information
can be searched for.
Sources holding sensitive data can thereby be identified without needing to examine millions of files manually.
Personal Data Discovery
From a KVKK perspective one of the most fundamental questions is:
"Which personal data does the organization hold, and where?"
this question.
A Turkish ID number, for example, may be present on;
CRM Database + Excel File + File Server + User Laptop + Old Backup + Shared Folder
these locations.
Data discovery technologies help make this scattered data estate visible.
Special Category Personal Data Discovery
Under KVKK some personal data may require higher protection.
Data discovery should therefore not be limited to finding general personal data alone.
Identifying the documents and systems holding special category personal data is separately important, according to the organization and the use case.
Regex and Pattern-Based Discovery
Data discovery solutions can make use of different analysis methods in identifying sensitive data.
One of these is Regular Expression – Regex based matching.
For example, data holding particular formats such as;
- Turkish ID Number
- IBAN
- Credit card
- Telephone
- Tax number
can be searched for by pattern.
Using regex alone can create false positives, however.
In advanced data discovery projects, additional validation and contextual analysis mechanisms are therefore important.
Keyword-Based Classification
Particular words and phrases can also be used to determine a document's sensitivity level.
For example, phrases such as;
CONFIDENTIAL STAFF ONLY CUSTOMER INFORMATION TRADE SECRET CONTRACT TENDER
can be part of classification policy.
Context-Aware Classification
The presence of a string of digits does not always mean sensitive data is present.
Not every 11-digit number is a Turkish ID number, for example.
In modern data discovery solutions, therefore, the;
Pattern + Context + Validation
approach produces more accurate results.
Data Fingerprinting
In some situations the organization's particular data sets need recognizing directly.
Taking the organization's customer list or personnel database as a reference, for example, whether similar data is present in different files and environments can be analyzed.
This approach is valuable particularly in detecting uncontrolled copies of sensitive data.
Data Labelling
Once discovered data has been classified, an appropriate security label can be applied to documents.
For example:
- PUBLIC
- INTERNAL
- CONFIDENTIAL
- RESTRICTED
Labels can show the user the document's security level and trigger the policies other security systems will apply.
Automated and Manual Classification
Data classification can be applied in two ways.
Manual Classification
The user selects the security level while creating the document.
Automated Classification
The system analyzes the document content and carries out the classification automatically, or suggests it to the user.
For example, the;
500 Turkish ID numbers found in the document → Personal Data → CONFIDENTIAL
policy can be applied.
The Difference Between Data Discovery and DLP
These two technologies particularly need separating from one another.
Data Discovery & Classification
Determines what the data is and where it is located.
DLP – Data Loss Prevention
Controls where the data goes and whether it can be taken outside.
For example:
Discovery → There is personal data in this Excel file.
Classification → The file was classified as CONFIDENTIAL.
DLP → Block this file being sent to a personal Gmail account.
Therefore;
Discovery → Classification → Protection
is a data security chain in which each part completes the others.
Data Discovery and DDR
The DDR – Data Detection and Response approach focuses on analyzing data usage behavior and the risks to data more dynamically.
Data discovery, in turn, brings out the organization's data surface first.
The ideal architecture can be built in the form of;
Data Discovery ↓ Classification ↓ DLP ↓ DDR ↓ SIEM / SOC
this sequence.
Data Discovery and DAM
DAM monitors database activity.
Data discovery, in turn, helps understand which sensitive data is present inside the database.
For example;
Discovery → A Turkish ID number was found in the CUSTOMER table.
Then;
DAM → Who reached the CUSTOMER table?
answers this question.
The two technologies together can therefore form a more meaningful data security visibility.
Data Discovery and KVKK
Data discovery is one of the extremely valuable technical tools in KVKK projects.
In KVKK work the questions;
Which personal data is processed? Where is it held? Who can reach it? How many different copies are there? Is unnecessary data present?
need answering.
Data discovery technologies can help support these processes with technical data.
It should not be forgotten, however, that data discovery software does not deliver KVKK compliance on its own.
Data Minimization
One of the core principles in data security is not holding data that is not needed.
For example;
An Excel file created 10 years ago + An old customer list + Turkish ID numbers + A File Share everyone can reach
can create serious risk.
Thanks to data discovery, ROT – Redundant, Obsolete, Trivial Data sources of this kind can be detected and brought into data lifecycle policy.
What Is Dark Data?
Dark data can be assessed as data stored by the organization whose existence, content or business value is not sufficiently known.
Millions of old documents may be present on a file server unused for years, for example.
Sensitive data being present inside those documents creates security and regulatory risk.
Data discovery solutions can help raise dark data visibility.
Data Security Posture Management – DSPM
In the modern data security approach, DSPM – Data Security Posture Management is becoming steadily more important.
DSPM's core questions are:
Where is the sensitive data? Who can reach it? Is it in the wrong place? Is there more privilege than necessary? Is it exposed in the cloud environment? What is the risk level?
these questions.
Data Discovery & Classification technologies can therefore be assessed as one of the core data sources of modern DSPM architecture.
Veriket
Turkish-Made Data Discovery and Classification Solution
Veriket can be positioned as a Turkish-made data security solution for discovering and classifying the sensitive data in corporate environments.
With Veriket the aim is to analyze the organization's different data sources so personal, financial or organization-specific sensitive data is made visible.
Sensitive Data Discovery
With Veriket, policies can be built for detecting different data categories inside the organization such as;
Personal Data Financial Data Customer Data Personnel Data Organization-Specific Data
and similar content.
Data Classification
Discovered content can be separated into categories according to the organization's information classification policy.
For example, organization-specific classification levels such as;
Public Service-Restricted Confidential Strictly Confidential
can be applied.
Veriket in KVKK Projects
Veriket can be assessed particularly in projects for determining where personal data is located inside the organization.
This visibility can help support processes such as;
KVKK Inventory + Data Minimization + Access Control + Retention Policy + DLP Policy
with technical data.
The Domestic Product Advantage
Veriket being a domestically developed solution;
- Public sector
- Defence industry
- Critical infrastructure
- Organizations with data sovereignty sensitivity
- Organizations wanting local technical support
can be an important assessment criterion for these estates.
DECE Data Discovery & Classification
Corporate Data Visibility and Sensitive Data Analysis
DECE is our second solution, one we can position for discovering and classifying the sensitive information held across organizations' different data sources.
With DECE the core aim is to build the;
Data Source → Discovery → Sensitive Data Detection → Classification → Risk Visibility
chain.
Making Scattered Data Visible
In organizations, sensitive data is mostly not held on a single system.
The same customer data, for example, may be present as different copies on;
Database + File Server + Excel + User Computer + Archive
these locations.
DECE can be used in projects for assessing data sources of this kind centrally.
Sensitive Data Analysis
Organization-specific data definitions can be built so particular content is detected.
This approach matters not only for standard personal data types but for the organization's own information assets such as;
Project Codes Customer Numbers Corporate Documents Commercial Information Organization-Specific Sensitive Data
these assets.
Data Classification Policies
Discovered data is classified by importance so a more meaningful data context that security systems can use is formed.
For example, the;
Public → Internal → Confidential → Restricted
model can be applied.
The Data Security Ecosystem
The data visibility DECE creates;
DLP DAM DDR SIEM SOC GRC
can be assessed as providing input to these processes.
Data discovery thereby stops being a system that merely reports on its own and becomes part of a wider data security architecture.
Veriket or DECE?
Rather than positioning these products directly as alternatives to one another, assessing them according to the organization's data sources and project scope is more accurate.
Veriket
Turkish-made Data Discovery + Classification + KVKK + Sensitive Data Visibility
can be assessed in projects holding these needs.
DECE
Corporate Data Discovery + Sensitive Data Analysis + Classification + Data Security Integrations
can be positioned in projects holding these requirements.
In the final product decision, a POC must be carried out on the organization's real data sources.
The Data Discovery POC Process
A controlled POC can be carried out in SecureSys data discovery projects.
- 1. Determining the Data Sources — file servers, databases, endpoints and other data sources are determined.
- 2. Defining the Sensitive Data Types — Turkish ID number, IBAN, telephone, e-mail and organization-specific data types are built.
- 3. Discovery Scan — the data sources determined are scanned.
- 4. False Positive Analysis — the accuracy of the data detected is checked.
- 5. Classification — sensitivity levels are applied.
- 6. Risk Analysis — the location and access state of the sensitive data is assessed.
- 7. DLP / DAM / DDR Integration — the use of discovered data in other data security technologies is planned.
- 8. Reporting — technical and management reports are produced.
SecureSys Data Discovery and Classification Solutions
At SecureSys, in Veriket and DECE projects we address the;
Data Source Analysis → Sensitive Data Inventory → Discovery → Classification → KVKK Data Mapping → Risk Analysis → POC → Deployment → Policy Tuning → DLP/DAM/DDR Integration → SIEM/SOC → Reporting
processes end to end.
Find the Data First, Then Protect It
An organization cannot effectively protect data it does not know it holds.
The starting point of data security is answering the questions;
Where is the data? What does it contain? How sensitive is it? Who can reach it? Does it genuinely need holding?
these questions.
With SecureSys Veriket and DECE Data Discovery & Classification Solutions, discover your corporate data, determine its sensitivity levels and build a strong data security foundation for your DLP, DAM, DDR and GRC processes.
Request a demo, POC and quote for Veriket and DECE
Want to learn more about this service?
Our expert team will reach out for a free consultation as soon as possible.