Indexed Document Matching for Unstructured Data Leak Protection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Data Loss Prevention (DLP) approaches struggle to effectively monitor and protect sensitive data, especially in unstructured documents and encrypted traffic, leading to increased risks of data loss due to unintentional or malicious reasons.
Innovation Solution
The implementation of Indexed Document Matching (IDM) systems and methods, which enable the identification and protection of content matching whole or partial documents from a repository, providing data leak protection for unstructured documents through similarity detection and fragment identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional DLP approaches use DLP dictionaries and engines for Exact Data Matching, then structured documents with specific data formats can be detected, but unstructured documents cannot be effectively analyzed
Solution Approach 1:
The patent implements a universal document matching system that can handle multiple document types (structured and unstructured) through a common fingerprinting and hashing mechanism. The system extracts content fingerprints from various document formats and compares them against indexed documents, making the DLP solution adaptable to different document structures while maintaining detection accuracy
Solution Approach 2:
The system changes the detection parameter from format-specific pattern matching to content-based fingerprint comparison. By converting document content into comparable hash values and fingerprints, the system can detect similarities across different document formats without being constrained by specific structural parameters
2Reliability
If DLP systems inspect SSL/TLS encrypted traffic to detect sensitive data, then data loss risk is reduced, but processing cost and latency increase significantly
Solution Approach 1:
The patent extracts only the essential identifying features (fingerprints and hashes) from encrypted traffic for comparison, rather than attempting to fully decrypt and analyze the entire content. This extraction approach maintains detection effectiveness while significantly reducing processing overhead and latency
Solution Approach 2:
The system performs preliminary indexing of document fingerprints and creates lookup tables in advance. When inspecting encrypted traffic, it only needs to compare extracted fingerprints against the pre-built index, rather than performing full content analysis, thereby reducing real-time processing requirements
3Ease of operation
If DLP systems use conventional approaches with software agents and physical appliances, then on-network data can be monitored, but users bypass security controls when off-network creating blind spots
Solution Approach 1:
The patent introduces a cloud-based indexing service as an intermediary that maintains a centralized repository of document fingerprints. This mediator enables consistent DLP policy enforcement across both on-network and off-network scenarios, as the fingerprint comparison service can be accessed remotely, eliminating blind spots when users are away from the corporate network
Data Source
AI summary
Cloud-based data loss prevention (DLP) systems and methods include monitoring a file to be checked for sensitive data from a user associated with a tenant; obtaining one or more dictionaries for the tenant; identifying a DLP match based on any of identifying exact document matches between the file and files in the one or more dictionaries, identifying same text in the file as in an indexed document in the one or more dictionaries, identifying content in the file that contains a subset of text in an indexed document in the one or more dictionaries, and identifying content that is similar but not exact as the text in an indexed document in the one or more dictionaries; and, responsive to the DLP match, blocking the file in the cloud-based system.


