ML Leak Detection for Sensitive Data Release Prevention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies struggle to optimally detect and prevent leaks of sensitive information outside secure computer networks, particularly in documents, emails, audio, and communication sessions, due to inefficiencies in identifying such leaks.
Innovation Solution
A trained machine learning algorithm, utilizing sensitive or insensitive training data, determines the presence of sensitive information and takes preventive actions, such as blocking the release, using data processors to capture and analyze input data across various communication channels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine learning algorithms are trained to identify sensitive information leaks, then detection accuracy is improved, but system complexity increases
Solution Approach 1:
The machine learning algorithms are trained in advance with extensive sensitive information data before deployment. This preliminary training enables the system to automatically identify and detect sensitive information patterns without requiring complex real-time analysis, thereby improving detection accuracy while managing system complexity through pre-computation
Solution Approach 2:
The patent introduces trained machine learning models as intermediary components between the input data and the detection decision. These models act as mediators that have already processed and learned sensitive information patterns during training, allowing the system to achieve high detection accuracy by delegating complex pattern recognition to pre-trained algorithms rather than implementing complex real-time analysis
2Reliability
If comprehensive training data is used to improve leak detection, then detection capability is improved, but training data requirements increase
Solution Approach 1:
Comprehensive training data is collected and processed in advance to train the machine learning algorithms before deployment. This preliminary action allows the system to learn from extensive examples of sensitive and non-sensitive information, improving detection capability while managing data requirements through pre-computation and offline training
Solution Approach 2:
The patent uses synthetic or representative training data that copies the essential characteristics of real sensitive information without requiring actual sensitive data. This approach improves detection capability by learning from realistic patterns while reducing the quantity of actual sensitive training data needed, as the model learns from replicated patterns rather than requiring every possible real-world example
Data Source
AI summary
A trained machine learning algorithm receives input data that may contain sensitive information. For example, the input data may be top secret military specifications that are sent as an attachment in an email that is being sent outside of a government computer network. The trained machine learning algorithm is trained with one of: sensitive training data or insensitive training data (or there may be two trained machine learning algorithms where one is trained with the sensitive training data and one is trained with the insensitive training data). The trained machine learning algorithm determines whether the input data contains the sensitive information. In response to determining that the input data contains the sensitive information, an action is taken to prevent release of the input data. For example, the action may be to block the sending of the email.


