Transparent Data Loss Prevention Classifications
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional data loss prevention (DLP) systems rely on machine-learning classifiers that often function without transparency, leading to high rates of false positives and a lack of understanding among administrators about the basis for classifications, making it difficult to effectively manage sensitive data protection.
Innovation Solution
The system identifies passages within documents that contributed to classifications and displays these elements in context, allowing administrators to understand the basis for classifications and modify the classifiers to reduce false positives by highlighting specific mistakes and allowing user input for correction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine learning classifiers are used for DLP classifications, then classification accuracy is improved, but transparency and understandability of classification basis deteriorates
Solution Approach 1:
The patent introduces an intermediary component that acts as a bridge between the machine learning classifier and the administrator. This intermediary extracts and presents the classification basis in human-understandable form, allowing administrators to see why documents were classified without compromising the accuracy of the underlying ML model. The intermediary translates complex classifier decisions into transparent, explainable information.
2Adaptability or versatility
If machine learning classifiers are used for DLP classifications, then classification capability is improved, but false positive rate increases
Solution Approach 1:
The patent implements a feedback mechanism where administrators can review the classification basis presented by the intermediary and provide corrections or adjustments. This feedback loop allows the system to learn from administrator decisions and refine its classification behavior over time, reducing false positives while maintaining the adaptability and versatility of the machine learning classifier.
3Loss of information
If traditional DLP systems with keywords and regular expressions are used, then transparency of classification basis is improved, but classification accuracy deteriorates
Solution Approach 1:
The patent merges the strengths of traditional DLP systems (transparency through keywords and rules) with the strengths of machine learning systems (accuracy and adaptability). The intermediary component combines explainable keyword-based reasoning with ML-based classification, providing both transparency and high classification accuracy simultaneously.
Data Source
AI summary
A computer-implemented method for transparent data loss prevention classifications may include 1) identifying a document that received a classification by a machine learning classifier for data loss prevention, 2) identifying at least one linguistic constituent within the document that contributed to the classification, 3) identifying a relevant passage of the document that contextualizes the linguistic constituent, and 4) displaying a user interface including the linguistic constituent in context of the relevant passage. Various other methods, systems, and computer-readable media are also disclosed.


