Security Parameter Index for Data Leak Prevention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data leak prevention technologies face challenges in effectively identifying and managing sensitive information across structured, semi-structured, and unstructured data, leading to high false positives and negatives, and require extensive manual effort for policy specification and enforcement.
Innovation Solution
A system that computes a Security Parameter Index (SPI) for data sets, determining a Security Quotient (Sq) based on the SPI, source, and destination, to dynamically define and enforce data leak prevention policies, allowing for real-time automatic policy application and minimizing false alerts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review of communications is performed by security staff, then data leak prevention accuracy is improved, but processing speed decreases and privacy compromise risk increases
Solution Approach 1:
The patent introduces an automated DLP system as an intermediary between security staff and communications. This system uses lexical matching techniques with pre-specified keywords to automatically screen communications, flagging only suspicious cases for manual review. This intermediary layer maintains high detection accuracy while dramatically increasing processing speed and reducing the privacy compromise risk associated with manual review of all communications.
Solution Approach 2:
The patent segments the data review process into two distinct stages: automated preliminary screening using keyword matching, and selective manual review of flagged communications. This segmentation allows the system to leverage the speed and consistency of automated processing for the majority of communications while reserving human judgment for complex or ambiguous cases, thereby resolving the contradiction between accuracy and speed.
2Measurement precision
If Regular Expressions based matching is used to discover sensitive information, then structured data detection is improved, but detection capability for semi-structured and unstructured data deteriorates
Solution Approach 1:
The patent transitions from Regular Expressions-based matching to lexical matching techniques that operate on different parameters. Instead of relying on strict syntactic patterns, the system uses semantic keyword matching that can identify sensitive information regardless of its structural context. This parameter change enables the system to effectively detect sensitive data across structured, semi-structured, and unstructured formats while maintaining detection accuracy.
3Speed
If lexical matching techniques with pre-specified keywords are used, then detection speed is improved, but false positive rate increases due to inability to discern context
Solution Approach 1:
The patent segments the detection process into automated keyword matching followed by contextual analysis. The lexical matching technique rapidly screens communications at high speed, but the system then segments suspicious cases into different categories based on contextual indicators. This segmentation allows the system to maintain high processing speed while reducing false positives by applying additional contextual filters to ambiguous cases.
Solution Approach 2:
The patent implements feedback mechanisms where the results of lexical matching are continuously refined based on contextual analysis and manual review outcomes. The system learns from false positives and adjusts its keyword matching and contextual interpretation accordingly, thereby maintaining high detection speed while progressively improving reliability and reducing false positive rates over time.
Data Source
AI summary
Some embodiments of high granularity reactive measures for selective pruning of information have been presented. The system and apparatus embody algorithms to automatically evaluate the security based significance (also referred to as “information enthalpy”) of a given set of structured, semi-structured or unstructured Data. This is also termed as security parameter index (SPI), represented by a numerical value, and is regarded as the intrinsic property of a given set of structured, semi-structured or unstructured Data. In one embodiment, a security parameter index (SPI) of a set of data is determined based on content of the set of data. If the SPI is above a predetermined threshold, then a security quotient (Sq) of the set of data is further determined based on the SPI and an action to be performed on the set of data in the current situation. Based on the value of the Sq, a data leak prevention policy is automatically defined and enforced on the set of data in the current situation. The system and apparatus also embody a Security Map that enumerates the security based inter-relationship between Agents, Data Set(s) and permissible Action(s) that can be invoked on the data. The Security Map enables automatic and dynamic generation and enforcement of security policies to prevent data leak.


