Automated Data Classification via Rule Pattern Consolidation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage and security methods require substantial effort to accurately classify and tag sensitive information, leading to potential inappropriate dissemination or accessibility issues due to improper classification.
Innovation Solution
A data classification approach that captures and analyzes patterns in classification operations to derive rules for identifying sensitive data attributes, applies context information to label sensitivity, and anonymizes data to create general rules applicable across multiple datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual data classification operations are performed repeatedly, then classification accuracy is maintained, but time consumption and labor effort increase substantially
Solution Approach 1:
The system enables self-service by automatically capturing, analyzing, and applying classification rules without requiring repeated manual intervention. The automated rule generation and application process allows the system to classify data independently, reducing time consumption while maintaining accuracy through consistent rule application.
Solution Approach 2:
The system implements feedback mechanisms by capturing results from classification operations and using them to refine and update classification rules. This iterative feedback loop ensures that classification accuracy is maintained while reducing the need for repetitive manual operations, as the system learns from past performance and automatically adjusts its rules.
2Measurement precision
If data classification rules are created for each specific case, then classification precision is improved, but system complexity and rule management difficulty increase
Solution Approach 1:
The system merges similar classification rules into unified, generalized rules by consolidating duplicative conditions and eliminating inconsequential ones. This merging process reduces the number of individual rules while maintaining classification precision, as the consolidated rules capture the essential patterns across multiple specific cases without requiring separate management for each.
Solution Approach 2:
The system creates universal classification rules that can be applied across multiple data types and contexts. By identifying patterns that are common across different datasets and anonymizing specific details, the system generates multi-functional rules that maintain precision while simplifying management, as a single rule can handle various classification scenarios rather than requiring case-specific rules.
3Reliability
If detailed context information is applied to each data item, then classification accuracy is improved, but processing time and computational resources increase
Solution Approach 1:
The system performs preliminary actions by pre-processing and organizing context information into structured formats before classification operations. This pre-organization of contextual data allows for faster retrieval and application during classification, improving processing speed while maintaining accuracy. The preliminary structuring of context reduces the computational overhead of analyzing detailed information during the actual classification process.
Data Source
AI summary
Classification for data intake operations in an enterprise ensures that sensitive data is not disseminated inappropriately, but incurs substantial time, effort and expense. A method of classifying data in a large set of data repositories captures a set of raw rules resulting from inputs indicative of evaluations and conclusions of data classification operations, typically by logging data classification operations, and identifies patterns in the set of raw rules by consolidating duplicative conditions and eliminating inconsequential conditions. External conditions and observations may be referenced for applying a context to the rules based on a usage or domain of the data, and data sets of disparate entities may be examined for anonymizing the data and combining with other sets of anonymized data.


