Sensitive Data Classification via Weighted Rule Aggregation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems require manual effort to identify and manage sensitive data across various data sources, which is inefficient and prone to oversights, especially as data grows and becomes dispersed, lacking an automated solution to classify and secure sensitive information effectively.

Innovation Solution

A system comprising a sensitive data scanner with a data pre-processor, classifier, and reporting module that automatically identifies sensitive data across multiple data sources, using pattern matching, logical rules, contextual analysis, and machine learning to classify and apply security measures, thereby reducing the need for manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual methods are used to identify sensitive data, then users can determine sensitive information with human oversight, but the process is tedious, prone to oversights, and cannot scale with data growth

Engineering Contradiction:
Improveaccuracy of sensitive data identificationVSAvoidefficiency of sensitive data identification
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system enables self-service automated classification of sensitive data using machine learning models and pattern matching algorithms that independently identify and classify sensitive information without requiring manual human review for each data element, thereby scaling efficiency while maintaining reliability through configurable confidence thresholds

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual review process with an automated computational system using machine learning classifiers, pattern matching, and rule-based engines that can process vast amounts of data rapidly while maintaining high accuracy through multiple detection mechanisms and confidence scoring

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated classification is implemented, then productivity and scalability improve, but system complexity increases

Engineering Contradiction:
Improveautomation of sensitive data classificationVSAvoidcomplexity of classification system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the classification task into distinct functional modules including pattern matching engines, machine learning classifiers, rule-based detection systems, and confidence scoring mechanisms, allowing each component to be independently developed, tuned, and maintained while working together to solve the overall classification problem

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary confidence scoring mechanisms and threshold-based filtering layers that mediate between the complex automated classification processes and the final decision-making, simplifying the interface between automated systems and human users by presenting only high-confidence classifications for review

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If manual data transfer is performed to secure sensitive data, then data protection can be implemented, but the process creates additional problems and potential oversights

Engineering Contradiction:
Improvedata protection effectivenessVSAvoidtime required for data security operations
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary automated classification and identification of sensitive data before security measures are needed, continuously monitoring and categorizing data as it is created or moved, so that when security actions are required, the data is already identified and ready for rapid protection without time-consuming manual discovery

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual data transfer and security implementation processes with automated system-wide policies that can dynamically apply encryption, access controls, and protection measures across the entire data environment without requiring human intervention to move or secure individual data elements

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12153693B2Sensitive data classification
Publication Date: 2024.11.26 PROTEGRITY US HLDG LLC
  • US12153693B2 patent drawing
  • US12153693B2 patent drawing
  • US12153693B2 patent drawing

AI summary

A gateway device includes a network interface connected to data sources, and computer instructions, that when executed cause a processor to access data portions from the data sources. The processor accesses classification rules, which are configured to classify a data portion of the plurality of data portions as sensitive data in response to the data portion satisfying the rule. Each rule is associated with a significance factor representative of an accuracy of the classification rule. The processor applies each of the set of classification rules to a data portion to obtain an output of whether the data is sensitive data. The output are weighed by significance factors to produce a set of weighted outputs. The processor determines if the data portion is sensitive data by aggregating the set of weighted outputs, and presents the determination in a user interface. Security operations may also be performed on the data portion.