Automated Personal Data Discovery in Relational Databases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In large organizations with multiple disparate databases, locating and removing personal information is a time-consuming and human-driven process, making it difficult to comply with privacy laws and regulations effectively.

Innovation Solution

A system that processes personal data in relational databases by sampling data, analyzing it against known types of personal data, and marking relevant data tables and columns, allowing for the identification and processing of personal information across multiple data tables, including those that reference each other.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If manual methods are used to locate personal information in databases, then flexibility and adaptability are maintained, but time consumption and labor requirements increase significantly

Engineering Contradiction:
Improvetime to locate personal informationVSAvoidautomation of personal data discovery
Core Design Contradiction:
Loss of timeVSExtent of automation

Solution Approach 1:

The system performs self-service by automatically discovering and cataloging personal data across multiple databases without requiring manual intervention. The automated discovery process samples data, identifies personal information patterns, and builds a comprehensive catalog of data locations, eliminating the need for human operators to manually search through databases.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical search processes with automated computational systems. The system uses computer algorithms to sample data, recognize patterns associated with personal information, and automatically generate catalogs of data locations, substituting human manual operations with automated mechanical/electronic processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If comprehensive catalogs of personal data locations are built manually, then accuracy and completeness can be maintained, but the process becomes complex and time-consuming

Engineering Contradiction:
Improvecompleteness of personal data catalogVSAvoidcomplexity of data cataloging process
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by automatically sampling data from multiple databases and pre-identifying personal information patterns before a comprehensive catalog is needed. The automated discovery process continuously monitors and catalogs data locations in real-time, maintaining completeness without requiring complex manual procedures.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating a virtual catalog representation of personal data locations across multiple databases. Instead of manually tracking each data location, the system creates a digital copy/catalog that automatically reflects the current state of personal data distribution, simplifying the management of comprehensive information.

Inventive Principle:
Principle #26Copying

3Productivity

If automated data sampling and analysis is implemented, then productivity and speed of personal data discovery improve, but system complexity and computational requirements increase

Engineering Contradiction:
Improvespeed of personal data discoveryVSAvoidcomplexity of data processing system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system applies segmentation by dividing the data discovery process into distinct modular components: data sampling module, pattern recognition module, catalog generation module, and data location tracking module. This segmentation allows each component to be optimized independently and reduces overall system complexity while maintaining high productivity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements universality by designing a multi-functional automated discovery system that can handle multiple data types, database formats, and personal information patterns through a single unified platform. The system's ability to perform sampling, analysis, cataloging, and tracking functions in one integrated system reduces complexity compared to multiple separate tools.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11762833B2Data discovery of personal data in relational databases
Publication Date: 2023.09.19 RUBRIK INC
  • US11762833B2 patent drawing
  • US11762833B2 patent drawing
  • US11762833B2 patent drawing

AI summary

Described herein is a system that processes personal data in databases. The system samples data stored in columns of data tables and analyzes the sampled data to determine whether the sampled data includes personal data. Based on the analysis, the system marks which data tables and which columns of the data tables store personal data. The system receives a request to process personal data for a subject. From data tables that are marked as storing personal data, the system identifies records storing personal data for the subject. The system additionally identifies other data tables marked as storing personal data that reference or are referenced by the data tables including the records referencing the subject. The system processes the data stored in the columns that are marked as storing personal data.