Data Classification via Statistical Sampling for Sensitive Information
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data loss prevention techniques require significant time and resources for categorizing large datasets, making them inefficient for identifying sensitive information.
Innovation Solution
A method involving selecting a sample set of files, classifying them using metadata and machine learning, and providing an estimate of sensitive information within the dataset, allowing for efficient data classification based on desired accuracy levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If automated methods are used to categorize and identify sensitive data in large datasets, then measurement precision of sensitive information is improved, but loss of time and resources increases significantly
Solution Approach 1:
The patent divides the large dataset into a sample subset for classification. Instead of scanning all files, the system selects a representative sample (e.g., 100 files from a million) and classifies only those, then extrapolates results to the entire dataset. This segmentation dramatically reduces time and resource consumption while maintaining identification accuracy through statistical representation.
Solution Approach 2:
The patent applies partial action by performing classification on only a portion (sample) of the total dataset rather than the complete set. The sample size is determined to be sufficient for providing an accurate estimate of sensitive information in the entire dataset, thus achieving the required measurement precision with significantly reduced time and resource investment.
2Productivity
If a sample set is used instead of analyzing all files, then productivity is improved, but measurement precision of sensitive information estimate may decrease
Solution Approach 1:
The patent changes the parameter of sample size to optimize the balance between productivity and measurement precision. By adjusting the sample size based on the desired confidence level and margin of error, the system can achieve high productivity with acceptable precision. The sample size is calculated using statistical formulas that ensure the estimate remains within an acceptable accuracy range.
Solution Approach 2:
The patent replaces the mechanical approach of analyzing every file with a statistical sampling methodology. Instead of exhaustive mechanical inspection, the system uses probabilistic sampling and statistical inference to estimate sensitive information in the entire dataset, achieving both high productivity and sufficient measurement precision through mathematical rather than mechanical means.
Data Source
AI summary
Techniques for data classification may be realized as a method including: selecting from a group of files a sample set representing fewer than all of the files; classifying each file in the sample set, wherein classifying each file includes identifying whether each file represents sensitive information; and providing an estimate for the group of files based on the classification of each file in the sample set, including an estimate of sensitive information within the group of files.


