Anonymized Data Tables for Secure Cloud Processing with Referential Integrity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Financial institutions face challenges in transmitting confidential customer data to distributed or cloud-based computing clusters due to confidentiality and privacy restrictions, and existing cryptographic or tokenization processes fail to maintain referential integrity and data format across multiple, mutually incompatible database tables, rendering tokenized data unsuitable for machine-learning or artificial-intelligence processes.
Innovation Solution
Implementing de-risking processes that selectively tokenize or anonymize confidential data within source data tables using configuration data and mapping data to maintain referential integrity, allowing secure transmission to cloud-based systems while preserving data format and structure, enabling suitable data ingestion for machine-learning or artificial-intelligence processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If confidential data is transmitted to distributed or cloud-based computing clusters, then data processing capability is improved, but confidentiality and privacy restrictions are violated
Solution Approach 1:
The patent segments confidential data into multiple de-identified data sets by applying different de-identification techniques to different portions or aspects of the data. This allows the data to be processed externally while maintaining confidentiality, as no single data set contains the complete confidential information in its original form.
Solution Approach 2:
The patent introduces de-identification techniques as an intermediary process between the confidential data and the distributed computing clusters. This intermediary layer transforms the data into de-identified form that can be safely processed externally while preserving the ability to maintain referential integrity through mapping data.
2Object-affected harmful factors
If cryptographic or tokenization processes are applied to confidential data, then confidentiality is improved, but referential integrity and data format are lost
Solution Approach 1:
The patent changes the parameters of data de-identification by using configurable de-identification techniques that can be selected and adjusted based on specific requirements. This allows optimization of both confidentiality protection and referential integrity maintenance by selecting appropriate techniques such as generalization, suppression, or swapping with configurable parameters.
Solution Approach 2:
The patent introduces dynamic configurability to the de-identification process, allowing the techniques and parameters to be adjusted based on the specific data sets and processing requirements. This dynamic approach enables the system to maintain referential integrity when needed while providing strong confidentiality protection when required.
3Object-affected harmful factors
If data is anonymized for secure transmission, then confidentiality is improved, but data usability for machine-learning processes deteriorates
Solution Approach 1:
The patent segments the de-identification approach by applying different techniques to different data sets or data elements. This segmentation allows certain portions of data to be de-identified for confidentiality while other portions maintain their original format and usability for machine-learning processes.
Solution Approach 2:
The patent applies local quality by using different de-identification techniques for different portions of the data based on their specific requirements. Critical confidential data receives stronger de-identification while other data maintains higher usability, optimizing the balance between confidentiality and data usability for machine-learning applications.
Data Source
AI summary
In some examples, computer-implemented systems and processes deploy securely de-risked elements of confidential data within a distributed computing environment. For example, an apparatus may obtain configuration data associated with a source data table. The configuration data may specify an identifier of a column of the source data table that includes elements of confidential data, and based on the configuration data, the apparatus perform operations that anonymize the elements of confidential data within the column of the source data table and generate an anonymized column within the source data table. The apparatus may also perform operations that provision an anonymized data table that includes the anonymized column to at least one computing system, which may process the anonymized data table and generate an output data table that includes the anonymized column and maintains a referential integrity between the source data table and the output data table.


