Automated Personal Data Discovery in Relational Databases
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In large organizations with multiple disparate databases, locating and removing personal information is a time-consuming and human-driven process, making it difficult to comply with privacy laws and regulations effectively.
Innovation Solution
A system that processes personal data in relational databases by sampling data, analyzing it against known types of personal data, and marking relevant data tables and columns, allowing for the identification and processing of personal information across multiple data tables, including those that reference each other.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If manual methods are used to locate personal information in databases, then flexibility and adaptability are maintained, but time consumption and labor requirements increase significantly
Solution Approach 1:
The system performs self-service by automatically discovering and cataloging personal data across multiple databases without requiring manual intervention. The automated discovery process samples data, identifies personal information patterns, and builds a comprehensive catalog of data locations, eliminating the need for human operators to manually search through databases.
Solution Approach 2:
The patent replaces manual mechanical search processes with automated computational systems. The system uses computer algorithms to sample data, recognize patterns associated with personal information, and automatically generate catalogs of data locations, substituting human manual operations with automated mechanical/electronic processes.
2Loss of information
If comprehensive catalogs of personal data locations are built manually, then accuracy and completeness can be maintained, but the process becomes complex and time-consuming
Solution Approach 1:
The system performs preliminary actions by automatically sampling data from multiple databases and pre-identifying personal information patterns before a comprehensive catalog is needed. The automated discovery process continuously monitors and catalogs data locations in real-time, maintaining completeness without requiring complex manual procedures.
Solution Approach 2:
The patent uses copying by creating a virtual catalog representation of personal data locations across multiple databases. Instead of manually tracking each data location, the system creates a digital copy/catalog that automatically reflects the current state of personal data distribution, simplifying the management of comprehensive information.
3Productivity
If automated data sampling and analysis is implemented, then productivity and speed of personal data discovery improve, but system complexity and computational requirements increase
Solution Approach 1:
The system applies segmentation by dividing the data discovery process into distinct modular components: data sampling module, pattern recognition module, catalog generation module, and data location tracking module. This segmentation allows each component to be optimized independently and reduces overall system complexity while maintaining high productivity.
Solution Approach 2:
The patent implements universality by designing a multi-functional automated discovery system that can handle multiple data types, database formats, and personal information patterns through a single unified platform. The system's ability to perform sampling, analysis, cataloging, and tracking functions in one integrated system reduces complexity compared to multiple separate tools.
Data Source
AI summary
Described herein is a system that processes personal data in databases. The system samples data stored in columns of data tables and analyzes the sampled data to determine whether the sampled data includes personal data. Based on the analysis, the system marks which data tables and which columns of the data tables store personal data. The system receives a request to process personal data for a subject. From data tables that are marked as storing personal data, the system identifies records storing personal data for the subject. The system additionally identifies other data tables marked as storing personal data that reference or are referenced by the data tables including the records referencing the subject. The system processes the data stored in the columns that are marked as storing personal data.


