Sensitive Data Classification via Weighted Rule Aggregation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems require manual effort to identify and manage sensitive data across various data sources, which is inefficient and prone to oversights, especially as data grows and becomes dispersed, lacking an automated solution to classify and secure sensitive information effectively.
Innovation Solution
A system comprising a sensitive data scanner with a data pre-processor, classifier, and reporting module that automatically identifies sensitive data across multiple data sources, using pattern matching, logical rules, contextual analysis, and machine learning to classify and apply security measures, thereby reducing the need for manual intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual methods are used to identify sensitive data, then users can determine sensitive information with human oversight, but the process is tedious, prone to oversights, and cannot scale with data growth
Solution Approach 1:
The system enables self-service automated classification of sensitive data using machine learning models and pattern matching algorithms that independently identify and classify sensitive information without requiring manual human review for each data element, thereby scaling efficiency while maintaining reliability through configurable confidence thresholds
Solution Approach 2:
The patent replaces the mechanical manual review process with an automated computational system using machine learning classifiers, pattern matching, and rule-based engines that can process vast amounts of data rapidly while maintaining high accuracy through multiple detection mechanisms and confidence scoring
2Productivity
If automated classification is implemented, then productivity and scalability improve, but system complexity increases
Solution Approach 1:
The system segments the classification task into distinct functional modules including pattern matching engines, machine learning classifiers, rule-based detection systems, and confidence scoring mechanisms, allowing each component to be independently developed, tuned, and maintained while working together to solve the overall classification problem
Solution Approach 2:
The patent introduces intermediary confidence scoring mechanisms and threshold-based filtering layers that mediate between the complex automated classification processes and the final decision-making, simplifying the interface between automated systems and human users by presenting only high-confidence classifications for review
3Reliability
If manual data transfer is performed to secure sensitive data, then data protection can be implemented, but the process creates additional problems and potential oversights
Solution Approach 1:
The system performs preliminary automated classification and identification of sensitive data before security measures are needed, continuously monitoring and categorizing data as it is created or moved, so that when security actions are required, the data is already identified and ready for rapid protection without time-consuming manual discovery
Solution Approach 2:
The patent replaces manual data transfer and security implementation processes with automated system-wide policies that can dynamically apply encryption, access controls, and protection measures across the entire data environment without requiring human intervention to move or secure individual data elements
Data Source
AI summary
A gateway device includes a network interface connected to data sources, and computer instructions, that when executed cause a processor to access data portions from the data sources. The processor accesses classification rules, which are configured to classify a data portion of the plurality of data portions as sensitive data in response to the data portion satisfying the rule. Each rule is associated with a significance factor representative of an accuracy of the classification rule. The processor applies each of the set of classification rules to a data portion to obtain an output of whether the data is sensitive data. The output are weighed by significance factors to produce a set of weighted outputs. The processor determines if the data portion is sensitive data by aggregating the set of weighted outputs, and presents the determination in a user interface. Security operations may also be performed on the data portion.


