Machine Learning Data Loss Prevention Profile Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data loss prevention (DLP) systems that use machine learning-based detection (MLD) profiles require expertise and do not provide users with tools to generate or modify profiles, limiting their configurability and adaptability, especially when dealing with unstructured data like product formulas and sales reports.
Innovation Solution
A computing device and method that allow users to modify training data sets by adding incorrectly classified documents as negative or positive examples, enabling the generation of updated MLD profiles through machine learning analysis, with quality ratings and periodic retraining, facilitating user-generated and improved MLD profiles without requiring vector machine learning expertise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If VML technology is used to protect sensitive unstructured data, then detection accuracy is improved, but system complexity increases and requires expert knowledge
Solution Approach 1:
The system enables end users to generate and modify MLD profiles independently through a user-friendly interface, eliminating the need for VML experts. Users can upload documents, select categories, and train profiles without technical expertise, making the complex VML technology self-service accessible.
Solution Approach 2:
The patent introduces an intermediary layer between the complex VML algorithms and the end users. This intermediary includes automated profile generation, template-based configurations, and guided workflows that translate user needs into accurate MLD profiles without exposing the underlying complexity.
2Ease of operation
If predefined MLD profiles are shipped with the DLP system, then ease of operation is improved, but adaptability deteriorates as customers cannot modify profiles
Solution Approach 1:
The system transitions from static predefined profiles to dynamic, user-modifiable profiles. Users can upload their own documents, select categories, and train custom profiles that adapt to their specific needs while maintaining the ease of operation through automated processing and user-friendly interfaces.
Solution Approach 2:
The patent segments the profile generation process into discrete, manageable steps: document upload, category selection, training configuration, and profile generation. This segmentation makes the previously complex and inaccessible profile modification process simple and user-friendly while maintaining full adaptability.
3Measurement precision
If manual training data curation is performed, then profile quality is improved, but time consumption increases
Solution Approach 1:
The system performs automated quality assessment and training data curation, eliminating the need for manual review. The automated process evaluates document quality, selects appropriate training data, and generates profiles efficiently, maintaining high profile quality while significantly reducing time consumption compared to manual curation.
Solution Approach 2:
The patent implements continuous automated training and quality improvement processes. The system continuously refines profiles based on feedback and performance metrics, ensuring ongoing quality improvement without requiring periodic manual intervention, thus maintaining high quality while minimizing time loss.
Data Source
AI summary
A computing device receives a document that was incorrectly classified as sensitive data based on a machine learning-based detection (MLD) profile. The computing device modifies a training data set that was used to generate the MLD profile by adding the document to the training data set as a negative example of sensitive data to generate a modified training data set. The computing device then analyzes the modified training data set using machine learning to generate an updated MLD profile.


