Cloud Snapshot Data Discovery for Unknown Data Stores
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cybersecurity solutions for cloud environments require manual identification of data sources, leading to inefficiencies, human errors, and incomplete protection due to incomplete data source information, and they can only protect known data stores, failing to identify unknown or inactive ones.
Innovation Solution
A method and system for data discovery that scans snapshots of disks to identify data stores automatically, creates engines to access and analyze data without manual permission, and classifies sensitivity, including inactive or unmanaged data stores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual identification of data sources is used, then deployment can be controlled, but the process becomes slower and less efficient
Solution Approach 1:
The system performs self-service by automatically scanning cloud environment snapshots to identify data stores without requiring manual input from security teams. The automated scanning process extracts data source information directly from cloud infrastructure snapshots, eliminating the need for human operators to manually identify and provide data source details.
Solution Approach 2:
The system performs preliminary action by pre-scanning and identifying data stores before security protection is deployed. By proactively discovering data sources in advance through automated scanning of cloud snapshots, the system prepares a comprehensive data source inventory that can be immediately used for security protection without waiting for manual identification.
2Reliability
If manual identification of data sources is used, then accuracy can be controlled, but human error introduces inaccuracies
Solution Approach 1:
The system replaces the mechanical human identification process with an automated computational scanning system. Instead of relying on human operators to manually identify and input data source information, the system uses automated scanning algorithms that parse cloud snapshots and extract data store locations, eliminating human error while maintaining high accuracy.
3Ease of operation
If permission is obtained for each data store manually, then access control is maintained, but the process becomes cumbersome
Solution Approach 1:
The system merges the permission granting process by obtaining a single unified permission to access cloud environment snapshots, rather than requiring individual permissions for each data store. This consolidated approach allows the scanning system to access multiple data stores through one authorization, significantly simplifying the operational process while maintaining security through controlled snapshot access.
4Reliability
If comprehensive data store listing is required, then complete protection is achieved, but technical expertise is required
Solution Approach 1:
The system performs self-service by automatically discovering and listing all data stores in the cloud environment through automated scanning of snapshots. This eliminates the need for operators to manually maintain comprehensive data store inventories or possess specialized technical expertise, as the system autonomously identifies and catalogs all data sources for protection.
Data Source
AI summary
A system and method for data discovery. A method includes performing a scan of a plurality of snapshots, each snapshot corresponding to a respective disk of a plurality of disks; identifying a plurality of data store files in the plurality of disks based on file metadata found during the scan; and detecting at least one data store based on the identified plurality of data store files, wherein each of the at least one data store is in a disk of the plurality of disks including one of the plurality of data store files.


