Anonymized Data Tables for Secure Cloud Processing with Referential Integrity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Financial institutions face challenges in transmitting confidential customer data to distributed or cloud-based computing clusters due to confidentiality and privacy restrictions, and existing cryptographic or tokenization processes fail to maintain referential integrity and data format across multiple, mutually incompatible database tables, rendering tokenized data unsuitable for machine-learning or artificial-intelligence processes.

Innovation Solution

Implementing de-risking processes that selectively tokenize or anonymize confidential data within source data tables using configuration data and mapping data to maintain referential integrity, allowing secure transmission to cloud-based systems while preserving data format and structure, enabling suitable data ingestion for machine-learning or artificial-intelligence processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If confidential data is transmitted to distributed or cloud-based computing clusters, then data processing capability is improved, but confidentiality and privacy restrictions are violated

Engineering Contradiction:
Improvedata processing capabilityVSAvoidconfidentiality violation
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent segments confidential data into multiple de-identified data sets by applying different de-identification techniques to different portions or aspects of the data. This allows the data to be processed externally while maintaining confidentiality, as no single data set contains the complete confidential information in its original form.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces de-identification techniques as an intermediary process between the confidential data and the distributed computing clusters. This intermediary layer transforms the data into de-identified form that can be safely processed externally while preserving the ability to maintain referential integrity through mapping data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If cryptographic or tokenization processes are applied to confidential data, then confidentiality is improved, but referential integrity and data format are lost

Engineering Contradiction:
Improveconfidentiality protectionVSAvoidreferential integrity
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent changes the parameters of data de-identification by using configurable de-identification techniques that can be selected and adjusted based on specific requirements. This allows optimization of both confidentiality protection and referential integrity maintenance by selecting appropriate techniques such as generalization, suppression, or swapping with configurable parameters.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces dynamic configurability to the de-identification process, allowing the techniques and parameters to be adjusted based on the specific data sets and processing requirements. This dynamic approach enables the system to maintain referential integrity when needed while providing strong confidentiality protection when required.

Inventive Principle:
Principle #15Dynamics

3Object-affected harmful factors

If data is anonymized for secure transmission, then confidentiality is improved, but data usability for machine-learning processes deteriorates

Engineering Contradiction:
Improveconfidentiality protectionVSAvoiddata usability
Core Design Contradiction:
Object-affected harmful factorsVSAdaptability or versatility

Solution Approach 1:

The patent segments the de-identification approach by applying different techniques to different data sets or data elements. This segmentation allows certain portions of data to be de-identified for confidentiality while other portions maintain their original format and usability for machine-learning processes.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by using different de-identification techniques for different portions of the data based on their specific requirements. Critical confidential data receives stronger de-identification while other data maintains higher usability, optimizing the balance between confidentiality and data usability for machine-learning applications.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250238536A1Secure deployment of de-risked confidential data within a distributed computing environment
Publication Date: 2025.07.24 THE TORONTO DOMINION BANK
  • US20250238536A1 patent drawing
  • US20250238536A1 patent drawing
  • US20250238536A1 patent drawing

AI summary

In some examples, computer-implemented systems and processes deploy securely de-risked elements of confidential data within a distributed computing environment. For example, an apparatus may obtain configuration data associated with a source data table. The configuration data may specify an identifier of a column of the source data table that includes elements of confidential data, and based on the configuration data, the apparatus perform operations that anonymize the elements of confidential data within the column of the source data table and generate an anonymized column within the source data table. The apparatus may also perform operations that provision an anonymized data table that includes the anonymized column to at least one computing system, which may process the anonymized data table and generate an output data table that includes the anonymized column and maintains a referential integrity between the source data table and the output data table.