ML-Based Sensitive Event Classification for Data Loss Prevention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data loss prevention solutions are manual and limited to identifying highly formatted sensitive information, failing to protect against the transmission of sensitive information in non-standardized formats and contextualizing the sensitivity of information entering or leaving a network.
Innovation Solution
A computer-implemented method using a machine-learning model to evaluate electronic communications for specific information, tagging and blocking transmissions based on identified patterns of words or phrases related to sensitive information, with the ability to dynamically update and apply controls to prevent sensitive information from leaving a secure network.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data loss prevention solutions are used to identify formatted sensitive information, then protection of highly formatted data (e.g., social security numbers) is achieved, but the ability to detect sensitive information in non-standardized formats is insufficient
Solution Approach 1:
The patent replaces manual, rule-based detection mechanisms with a machine-learning model that automatically evaluates electronic communications. The model uses natural language processing to identify sensitive information patterns, contextual phrases, and semantic meanings beyond simple format matching, thereby achieving both precision and versatility in detecting sensitive data across multiple formats including non-standardized ones
Solution Approach 2:
The system dynamically adjusts detection parameters by training the machine-learning model on organizational-specific data and contexts. The model learns to recognize sensitive information patterns tailored to the organization's terminology, document structures, and communication styles, enabling accurate detection across varied formats while adapting to changing information types and formats over time
2Adaptability or versatility
If machine-learning models are used to evaluate electronic communications for sensitive information, then detection of sensitive information in non-standardized formats is improved, but system complexity increases
Solution Approach 1:
The patent implements a multi-functional machine-learning model that performs multiple tasks within a single system: evaluating electronic communications for sensitive information, classifying types of sensitive events, determining contextual sensitivity, and generating appropriate responses. This consolidates what would otherwise require multiple separate systems into one unified platform, managing complexity while enhancing versatility
Solution Approach 2:
The system introduces an intermediary classification layer that processes electronic communications before final sensitivity determination. The machine-learning model acts as a mediator between raw communication data and policy enforcement decisions, breaking down the complex evaluation process into manageable stages: initial scanning, pattern recognition, contextual analysis, and classification, thereby reducing overall system complexity
3Productivity
If automated machine-learning evaluation is implemented to tag and block electronic communications, then productivity of data loss prevention is improved, but the system requires significant computational resources
Solution Approach 1:
The patent implements a two-stage evaluation process where the machine-learning model first performs a quick initial scan to identify potentially sensitive communications, then applies more intensive analysis only to those flagged as suspicious. This partial action approach processes the majority of communications with minimal resource consumption while maintaining high detection accuracy for sensitive information, thereby improving productivity without proportionally increasing computational resource requirements
Solution Approach 2:
The system performs preliminary classification and tagging of electronic communications as they enter the network, before full policy evaluation and enforcement actions are required. By pre-processing and categorizing communications upfront, the system reduces the computational burden during subsequent policy enforcement stages, improving overall productivity while managing resource consumption through staged processing
Data Source
AI summary
Identification of an electronic communication containing specific information is provided. Content of the electronic communication may be evaluated by a machine-learning model, and based on an evaluation of the content, it may be determined that the electronic communication contains the specific information. The electronic communication may be tagged with tag information indicating that the electronic communication contains the specific information, and transmission of the electronic communication may be blocked based on the tag information.


