Cloud DLP Indexed Document Matching for Unstructured Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Data Loss Prevention (DLP) systems struggle to effectively monitor and protect sensitive data, especially in unstructured documents and encrypted SSL/TLS traffic, leading to increased risks of data loss due to blind spots and high costs associated with inspection.
Innovation Solution
The implementation of Indexed Document Matching (IDM) technology, which identifies and protects content by matching whole or partial documents through cryptographic hashes and Context Triggered Piecewise Hashes, enabling similarity detection and fragment identification within a cloud-based system, allowing for real-time data leak protection across various file types and profiles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional DLP approaches using software agents and physical appliances are used, then data monitoring capability is provided, but blind spots are created when users access applications directly from the cloud bypassing security controls
Solution Approach 1:
The patent introduces a cloud-based DLP service as an intermediary between users and cloud applications. This service intercepts and inspects data in motion without requiring users to install local agents or change their access patterns. The intermediary performs SSL/TLS inspection to detect sensitive data in encrypted traffic, thereby maintaining both protection coverage and user flexibility.
2Loss of information
If SSL/TLS traffic inspection is performed to detect sensitive data in encrypted traffic, then data visibility is improved, but cost and processing capability requirements increase significantly
Solution Approach 1:
The patent extracts only the essential elements needed for DLP inspection from the encrypted traffic flow. Instead of fully decrypting and analyzing all SSL/TLS traffic, the system selectively inspects data based on predefined policies and data patterns. This extraction approach reduces processing overhead while maintaining effective detection of sensitive information.
Solution Approach 2:
The system performs partial inspection of encrypted traffic by applying DLP rules only to specific data patterns and contexts rather than analyzing every byte of SSL/TLS traffic. This partial action approach provides sufficient data visibility for security purposes while significantly reducing processing costs compared to comprehensive traffic analysis.
3Measurement precision
If DLP dictionaries and engines are used for Exact Data Matching, then structured data detection is improved, but detection of unstructured documents becomes difficult
Solution Approach 1:
The patent implements a universal DLP service that handles multiple data types and formats through a single cloud-based platform. The service can detect sensitive data in structured documents using traditional DLP dictionaries while also analyzing unstructured documents, emails, and data in motion through SSL/TLS inspection. This multi-functional approach eliminates the need for separate solutions for different document types.
Solution Approach 2:
The system dynamically adjusts detection parameters based on the type of data being analyzed. For structured documents, it uses precise dictionary-based Exact Data Matching, while for unstructured documents and encrypted traffic, it employs pattern recognition and contextual analysis. This parameter adaptation enables effective detection across diverse document types while maintaining high accuracy.
Data Source
AI summary
Cloud-based data loss prevention (DLP) systems and methods include monitoring a file to be checked for sensitive data from a user associated with a tenant; obtaining one or more dictionaries for the tenant; identifying a DLP match based on any of identifying exact document matches between the file and files in the one or more dictionaries, identifying same text in the file as in an indexed document in the one or more dictionaries, identifying content in the file that contains a subset of text in an indexed document in the one or more dictionaries, and identifying content that is similar but not exact as the text in an indexed document in the one or more dictionaries; and, responsive to the DLP match, blocking the file in the cloud-based system.


