Transparent Data Loss Prevention Classifications

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional data loss prevention (DLP) systems rely on machine-learning classifiers that often function without transparency, leading to high rates of false positives and a lack of understanding among administrators about the basis for classifications, making it difficult to effectively manage sensitive data protection.

Innovation Solution

The system identifies passages within documents that contributed to classifications and displays these elements in context, allowing administrators to understand the basis for classifications and modify the classifiers to reduce false positives by highlighting specific mistakes and allowing user input for correction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning classifiers are used for DLP classifications, then classification accuracy is improved, but transparency and understandability of classification basis deteriorates

Engineering Contradiction:
Improveclassification accuracyVSAvoidtransparency of classification basis
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent introduces an intermediary component that acts as a bridge between the machine learning classifier and the administrator. This intermediary extracts and presents the classification basis in human-understandable form, allowing administrators to see why documents were classified without compromising the accuracy of the underlying ML model. The intermediary translates complex classifier decisions into transparent, explainable information.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If machine learning classifiers are used for DLP classifications, then classification capability is improved, but false positive rate increases

Engineering Contradiction:
Improveclassification capabilityVSAvoidfalse positive rate
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where administrators can review the classification basis presented by the intermediary and provide corrections or adjustments. This feedback loop allows the system to learn from administrator decisions and refine its classification behavior over time, reducing false positives while maintaining the adaptability and versatility of the machine learning classifier.

Inventive Principle:
Principle #23Feedback

3Loss of information

If traditional DLP systems with keywords and regular expressions are used, then transparency of classification basis is improved, but classification accuracy deteriorates

Engineering Contradiction:
Improvetransparency of classification basisVSAvoidclassification accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent merges the strengths of traditional DLP systems (transparency through keywords and rules) with the strengths of machine learning systems (accuracy and adaptability). The intermediary component combines explainable keyword-based reasoning with ML-based classification, providing both transparency and high classification accuracy simultaneously.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9235562B1Systems and methods for transparent data loss prevention classifications
Publication Date: 2016.01.12 CA TECH INC
  • US9235562B1 patent drawing
  • US9235562B1 patent drawing
  • US9235562B1 patent drawing

AI summary

A computer-implemented method for transparent data loss prevention classifications may include 1) identifying a document that received a classification by a machine learning classifier for data loss prevention, 2) identifying at least one linguistic constituent within the document that contributed to the classification, 3) identifying a relevant passage of the document that contextualizes the linguistic constituent, and 4) displaying a user interface including the linguistic constituent in context of the relevant passage. Various other methods, systems, and computer-readable media are also disclosed.