Multimodal AI DLP for Dictionary-Free Sensitive Data Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Data Loss Protection (DLP) systems require up-front dictionaries and are often overly restrictive or ineffective in detecting sensitive data across various file formats, particularly in multimodal contexts such as images and videos.
Innovation Solution
A multimodal DLP system utilizing artificial intelligence and machine learning, including Large Language Models (LLMs) and zero-shot classifiers, processes and integrates information from multiple data modalities like text, images, and video without the need for pre-defined dictionaries, enhancing detection accuracy and reducing false positives/negatives.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If predefined dictionaries and exact data Matching (EDM) are used for DLP, then specific sensitive data can be detected, but the system becomes overly restrictive and may miss crucial files in various formats
Solution Approach 1:
The patent replaces traditional mechanical DLP methods (predefined dictionaries, regex matching, EDM, IDM) with an AI-based multimodal system that can automatically understand and classify sensitive data across multiple formats including text, images, video, and audio without requiring format-specific rules
Solution Approach 2:
The multimodal AI system provides universal detection capability across diverse data formats (text, images, video, audio) through a single unified platform, eliminating the need for separate detection mechanisms for each file type while maintaining high detection accuracy
2Measurement precision
If up-front dictionaries are provided for DLP detection, then specific sensitive data formats can be identified, but the system requires significant up-front input and configuration
Solution Approach 1:
The AI-based DLP system performs self-service by automatically learning and adapting to sensitive data patterns without requiring manual dictionary creation or configuration. The system autonomously processes and classifies sensitive information across multiple formats through trained machine learning models, eliminating the need for IT administrators to manually configure detection rules
Solution Approach 2:
The system performs preliminary training actions during model development where AI models are pre-trained on diverse datasets to recognize sensitive data patterns across multiple formats. This preliminary training enables the system to immediately detect sensitive data without requiring organization-specific configuration, while still maintaining high detection accuracy
3Reliability
If traditional DLP methods are used, then data protection can be implemented, but the system cannot comprehend the combination of various file formats (images, text, video)
Solution Approach 1:
The patent employs a composite approach by integrating multiple AI specialized processors (NLP for text, computer vision for images and video, audio processing models) into a unified multimodal DLP system. Each specialized component processes its native format while the integrated system comprehends combinations of various formats, providing reliable data protection across diverse media types
Data Source
AI summary
Multimodal Data Loss Protection (DLP) includes receiving an input comprising data in any of a plurality of formats; processing the input to determine whether or not the data includes sensitive data; and responsive to the input including sensitive data, performing steps of: processing the input to classify the input into a category of a plurality of categories; and providing an indication of the category of the plurality of categories. Advantageously, the trained multimodal system can detect categories of data being accessed, transferred, etc., without the requirement of up-front dictionaries from corporate Information Technology (IT).


