Non-linear Data Masking for Secure ML Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models require large datasets for training, which often include sensitive data, posing challenges in data security and privacy when using cloud computing resources, as existing encryption schemes may be insufficient due to computational power limitations and risks of data exposure.

Innovation Solution

A system and method for training machine learning models using encoded data sets that preserve interrelationships, employing autoencoders for non-linear masking, allowing secure storage and processing on cloud platforms without decrypting sensitive data, thus protecting it from unauthorized access.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is encrypted using traditional encryption schemes, then data security is improved, but computational power requirements increase beyond available local resources

Engineering Contradiction:
Improvedata securityVSAvoidcomputational power
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

The patent introduces cloud computing resources as an intermediary to perform the computationally intensive encoding and decoding operations. The encoding scheme is implemented through cloud-based machine learning models that transform sensitive data into encoded representations without requiring high computational power at the local edge devices, thus resolving the contradiction between security and computational constraints.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces traditional mechanical encryption systems with machine learning-based encoding models deployed in the cloud. These neural network models perform complex transformations on sensitive data that would be computationally prohibitive for local devices, substituting the mechanical encryption approach with a more powerful computational paradigm available in cloud infrastructure.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If data is decrypted for machine learning training, then model training capability is improved, but data exposure risks increase

Engineering Contradiction:
Improvemodel training capabilityVSAvoiddata exposure risk
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent applies preliminary encoding to sensitive data before it enters the machine learning training pipeline. The encoding models transform the data into a protected representation that preserves the statistical and relational properties needed for training, while removing sensitive information. This preliminary transformation allows model training to proceed without exposing the original sensitive data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameter representation of sensitive data through encoding transformations. The encoding models alter the data parameters into a new representation space that maintains the structural relationships necessary for machine learning training but eliminates the sensitivity of the original data, thus enabling training while reducing exposure risk.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If encoding schemes are implemented to protect sensitive data, then data privacy is improved, but data interrelationship preservation becomes challenging

Engineering Contradiction:
Improvedata privacyVSAvoiddata interrelationship
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The encoding models perform parameter transformations that change the representation of sensitive data while preserving the statistical relationships and patterns needed for machine learning. The transformations are designed to maintain data interrelationships such as correlations and distributions, ensuring that encoded data retains the structural properties necessary for effective model training.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20210256393A1System and method for distributed non-linear masking of sensitive data for machine learning training
Publication Date: 2021.08.19 ROYAL BANK OF CANADA
  • US20210256393A1 patent drawing
  • US20210256393A1 patent drawing
  • US20210256393A1 patent drawing

AI summary

Described in various embodiments herein is a technical solution directed to training downstream machine learning models. In particular, specific machines, computer-readable media, computer processes, and methods are described that are utilized to improve data security during training downstream machine learning models, including decreasing the risk of unauthorized access of training data, decreasing the risk of unauthorized use of training data by authorized users, increasing system systemic speed, and reduced overall computational resource requirements. Training data is manipulated prior to being provided for training machine learning models.