Machine Learning Classifiers for Sensitive Data Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional DLP systems fail to accurately identify and protect unstructured sensitive data and intellectual property, such as product formulas, source code, and sales reports, due to limitations in fingerprinting and describing technologies, which struggle with new or modified items.

Innovation Solution

The implementation of machine learning-based classifiers that can detect specific categories of sensitive information by training on customized data sets, identifying statistically significant features, and deploying within DLP systems to enforce policies for protection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If fingerprinting and describing technologies are used in conventional DLP systems, then known sensitive data can be protected, but new or modified sensitive data cannot be accurately identified

Engineering Contradiction:
Improvedetection accuracyVSAvoidcapability to detect new data
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary actions by collecting training data and training machine learning classifiers before actual DLP operations begin. The classifiers are pre-trained with positive examples (sensitive data) and negative examples (non-sensitive data) so they can automatically adapt to detect new and modified sensitive data without requiring updates to fingerprinting rules

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical fingerprinting and describing technologies with machine learning-based classifiers. Instead of using rigid pattern-matching algorithms that require manual rule configuration, the system uses trained classifiers that automatically learn detection patterns from training data, enabling them to generalize to new and modified sensitive data

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If machine learning-based classifiers are trained on customized data sets with statistically significant features, then detection precision is improved, but system complexity increases

Engineering Contradiction:
Improvedetection precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts only the statistically significant features from training data that are relevant for detection. By identifying and extracting only these key features rather than processing all possible data characteristics, the system achieves high detection precision while keeping the classifier models manageable in size and complexity

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes parameters by transforming raw training data into extracted features with specific statistical properties. The feature extraction process transforms unstructured data into structured feature vectors with defined dimensions and statistical significance thresholds, making the data suitable for machine learning classification while controlling model complexity

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8688601B2Systems and methods for generating machine learning-based classifiers for detecting specific categories of sensitive information
Publication Date: 2014.04.01 CA TECH INC
  • US8688601B2 patent drawing
  • US8688601B2 patent drawing
  • US8688601B2 patent drawing

AI summary

A computer-implemented method may include (1) identifying a plurality of specific categories of sensitive information to be protected by a DLP system, (2) obtaining a training data set for each specific category of sensitive information that includes a plurality of positive and a plurality of negative examples of the specific category of sensitive information, (3) using machine learning to train, based on an analysis of the training data sets, at least one machine learning-based classifier that is capable of detecting items of data that contain one or more of the plurality of specific categories of sensitive information, and then (4) deploying the machine learning-based classifier within the DLP system to enable the DLP system to detect and protect items of data that contain one or more of the plurality of specific categories of sensitive information in accordance with at least one DLP policy of the DLP system.