Sensitive Data Discovery Using Regex Metadata Relationships

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems fail to effectively detect and protect sensitive data in complex environments, leading to increased vulnerability to data breaches due to unexpected storage locations and inadequate security measures.

Innovation Solution

A system that automatically discovers sensitive data using regular expressions (regexes) and metadata, generating confidence scores based on multiple factors, and applies security operations based on these scores to ensure data protection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional security policies are used to protect sensitive data, then basic security coverage is provided, but false positives and negatives increase in complex environments

Engineering Contradiction:
Improvedata security reliabilityVSAvoidsensitive data detection precision
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The system implements feedback loops where detection results and security operations are continuously analyzed to improve future detections. Confidence scores from multiple factors (regex matches, metadata, relationships) feed into a learning mechanism that refines detection accuracy over time, reducing both false positives and negatives while maintaining reliable security coverage.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent combines multiple detection methodologies (regular expressions, metadata analysis, relationship discovery) into a composite detection system. Each method contributes different strengths, and their results are integrated through confidence scoring to achieve more accurate and reliable sensitive data detection than any single method could provide alone.

Inventive Principle:
Principle #40Composite materials

2Reliability

If manual tracking of sensitive data locations is performed, then security control is maintained, but human error increases leading to unexpected storage locations

Engineering Contradiction:
Improvesecurity control reliabilityVSAvoiddata location tracking ease
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system enables self-service automated detection and tracking of sensitive data locations. The automated discovery system continuously monitors and identifies sensitive data across the environment without requiring manual intervention, eliminating human error while maintaining reliable security control through machine-driven location tracking and security policy enforcement.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical tracking methods with automated computational systems. Machine learning algorithms and automated discovery mechanisms substitute for human administrators in tracking sensitive data locations, providing more reliable and error-free monitoring while reducing the operational burden on human staff.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If security updates are applied to all storage servers, then security coverage is maximized, but resource consumption increases

Engineering Contradiction:
Improvesecurity coverageVSAvoidsystem resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system applies security measures locally and selectively based on actual sensitive data locations and risk assessments. Rather than uniformly applying security updates to all storage servers, the system identifies specific locations containing sensitive data and applies appropriate security controls only where needed, maintaining maximum security coverage while reducing unnecessary resource consumption on systems without sensitive data.

Inventive Principle:
Principle #3Local quality

4Productivity

If automated discovery systems are implemented, then detection speed improves, but false positives increase without relationship context

Engineering Contradiction:
Improvedetection speedVSAvoiddetection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent merges multiple detection approaches (regex-based rapid scanning, metadata analysis, and relationship discovery) into a unified system. The relationship context from discovered connections between data objects is integrated with rapid automated detection results, filtering out false positives by verifying detections against known relationships and business logic, thus maintaining high detection speed while improving accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12373584B2Systems and methods for securing data based on discovered relationships
Publication Date: 2025.07.29 ORACLE INT CORP
  • US12373584B2 patent drawing
  • US12373584B2 patent drawing
  • US12373584B2 patent drawing

AI summary

Techniques for automatically discovering and protecting sensitive data are disclosed. In some embodiments, a set of data objects is searched for data matching a first set of one or more regular expressions and for metadata matching a second set of one or more regular expressions. A confidence score is then generated for a particular data objects in the set of data objects as a function of regular expressions in the first set of one or more regular expressions that match data stored in the particular data object and regular expression in the second set of one or more regular expressions that match metadata associated with the particular data object. One or more operations may be performed to protect sensitive data stored in the particular data object based, at least in part, on the confidence score.