Metadata-Based Personal Data Discovery Without Database Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current information-security processes and machines are unable to efficiently identify, manage, and utilize personal data within structured data sources or software applications without risking data exposure, often requiring database access and encountering compliance issues, and are not proactive in handling changing data environments.
Innovation Solution
Utilizing artificial-intelligence processes and machines that predict personal data presence based on metadata fields through natural language processing, embedding characters into vectors, and employing bidirectional LSTM algorithms to generate contextualizations without accessing the data content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual processes are used to identify and classify personal data, then application owners can maintain information about their applications, but the process is laborious and error-prone
Solution Approach 1:
The system enables automated self-service identification and classification of personal data through machine learning models that automatically scan metadata, predict personal data presence, and classify data types without requiring manual application owner intervention
Solution Approach 2:
The patent replaces manual mechanical processes with automated computational systems including natural language processing, bidirectional LSTM algorithms, and neural networks that analyze metadata and predict personal data presence automatically
2Measurement precision
If database access is required to identify personal data, then complete data analysis can be performed, but permissible-use issues and security risks arise
Solution Approach 1:
The system extracts and analyzes only the metadata portion of database structures, separating the identification process from the actual data content. This allows personal data detection without accessing or exposing the sensitive data itself, eliminating security risks associated with data access
Solution Approach 2:
The patent uses metadata as an intermediary layer between the detection system and the actual personal data. The machine learning models analyze metadata fields to predict personal data presence without directly accessing the sensitive information, serving as a secure intermediary that prevents harmful exposure
3Reliability
If manual processes are used to keep application information up to date, then application owners can maintain accuracy, but the frequency of updates is limited due to labor requirements
Solution Approach 1:
The automated machine learning system enables continuous monitoring and updating of personal data identification as applications are developed, deployed, and modified. The system can operate continuously without interruption, automatically adapting to new data formats and regulatory requirements without the frequency limitations of manual processes
4Reliability
If current security processes are used to manage personal data, then data can be secured, but the processes are inefficient and cannot keep pace with constantly changing data
Solution Approach 1:
The patent implements a dynamic system using machine learning models that continuously adapt to changing data formats, new personal data types, and evolving regulatory requirements. The bidirectional LSTM and neural network architectures can learn from new data patterns, making the security process dynamic rather than static, enabling it to keep pace with constantly changing data
Data Source
AI summary
Artificial-intelligence computer-implemented processes and machines predict whether personal data may be present in structured software based on metadata field(s) contained therein. Natural language processing preprocesses input strings corresponding to the metadata field(s) into normalized input sequence(s). Individual characters in the sequence(s) are embedded into fixed-dimension vectors of real numbers. Bidirectional LSTM(s) or other machine-learning algorithm(s) are utilized to generate forward and backward contextualization(s). Neural network output(s) are provided based on element-wise averaging or feed forwarding based on the contextualization(s) in order to predict whether one or more value fields corresponding to the metadata field(s) may contain personal data.


