Sensitive Data Discovery Using Regex Metadata Relationships
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems fail to effectively detect and protect sensitive data in complex environments, leading to increased vulnerability to data breaches due to unexpected storage locations and inadequate security measures.
Innovation Solution
A system that automatically discovers sensitive data using regular expressions (regexes) and metadata, generating confidence scores based on multiple factors, and applies security operations based on these scores to ensure data protection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional security policies are used to protect sensitive data, then basic security coverage is provided, but false positives and negatives increase in complex environments
Solution Approach 1:
The system implements feedback loops where detection results and security operations are continuously analyzed to improve future detections. Confidence scores from multiple factors (regex matches, metadata, relationships) feed into a learning mechanism that refines detection accuracy over time, reducing both false positives and negatives while maintaining reliable security coverage.
Solution Approach 2:
The patent combines multiple detection methodologies (regular expressions, metadata analysis, relationship discovery) into a composite detection system. Each method contributes different strengths, and their results are integrated through confidence scoring to achieve more accurate and reliable sensitive data detection than any single method could provide alone.
2Reliability
If manual tracking of sensitive data locations is performed, then security control is maintained, but human error increases leading to unexpected storage locations
Solution Approach 1:
The system enables self-service automated detection and tracking of sensitive data locations. The automated discovery system continuously monitors and identifies sensitive data across the environment without requiring manual intervention, eliminating human error while maintaining reliable security control through machine-driven location tracking and security policy enforcement.
Solution Approach 2:
The patent replaces manual mechanical tracking methods with automated computational systems. Machine learning algorithms and automated discovery mechanisms substitute for human administrators in tracking sensitive data locations, providing more reliable and error-free monitoring while reducing the operational burden on human staff.
3Reliability
If security updates are applied to all storage servers, then security coverage is maximized, but resource consumption increases
Solution Approach 1:
The system applies security measures locally and selectively based on actual sensitive data locations and risk assessments. Rather than uniformly applying security updates to all storage servers, the system identifies specific locations containing sensitive data and applies appropriate security controls only where needed, maintaining maximum security coverage while reducing unnecessary resource consumption on systems without sensitive data.
4Productivity
If automated discovery systems are implemented, then detection speed improves, but false positives increase without relationship context
Solution Approach 1:
The patent merges multiple detection approaches (regex-based rapid scanning, metadata analysis, and relationship discovery) into a unified system. The relationship context from discovered connections between data objects is integrated with rapid automated detection results, filtering out false positives by verifying detections against known relationships and business logic, thus maintaining high detection speed while improving accuracy.
Data Source
AI summary
Techniques for automatically discovering and protecting sensitive data are disclosed. In some embodiments, a set of data objects is searched for data matching a first set of one or more regular expressions and for metadata matching a second set of one or more regular expressions. A confidence score is then generated for a particular data objects in the set of data objects as a function of regular expressions in the first set of one or more regular expressions that match data stored in the particular data object and regular expression in the second set of one or more regular expressions that match metadata associated with the particular data object. One or more operations may be performed to protect sensitive data stored in the particular data object based, at least in part, on the confidence score.


