Stochastic Noise Layer for Secure Foundation Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models face challenges in obfuscating training data without access to the specific model architecture, especially when data needs to be shared with third parties or when the model has not been created yet, and there is a risk of data exposure during training, particularly in federated learning scenarios.

Innovation Solution

The implementation of a stochastic noise layer in machine learning models, specifically using autoencoders to learn parametric noise distributions and apply noise to training data, allowing for the creation of obfuscated datasets that can be used for training without revealing the original sensitive information, even in distributed training environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If training data is shared with third parties or used in federated learning scenarios, then model training capability is improved, but data privacy and confidentiality are compromised

Engineering Contradiction:
Improvemodel training capabilityVSAvoiddata exposure risk
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent applies preliminary action by obfuscating training data before it is shared with third parties or used in federated learning scenarios. The obfuscation process transforms the original training data into an obfuscated version that retains the necessary statistical properties for model training while removing sensitive information. This preliminary transformation ensures that data privacy is protected from the outset, preventing data exposure risks before they can materialize during data sharing or distributed training operations.

Inventive Principle:
Principle #10Preliminary action

2Object-affected harmful factors

If obfuscation techniques are applied to training data, then data privacy is improved, but model training accuracy may deteriorate

Engineering Contradiction:
Improvedata privacy protectionVSAvoidmodel training accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent employs parameter changes by carefully controlling the obfuscation parameters to maintain the statistical properties of the training data. The obfuscation process transforms data parameters in a way that preserves the distribution, correlations, and other statistical characteristics necessary for accurate model training. By adjusting and optimizing these parameters, the system achieves a balance where data privacy is enhanced through obfuscation while model training accuracy is maintained at acceptable levels.

Inventive Principle:
Principle #35Parameter changes

3Object-affected harmful factors

If stochastic noise layers are applied to training data, then data obfuscation is improved, but data processing complexity increases

Engineering Contradiction:
Improvedata obfuscation effectivenessVSAvoiddata processing complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The patent applies mechanics substitution by replacing complex, manual obfuscation processes with automated neural network-based stochastic noise layers. Instead of using intricate mechanical or manual methods to obfuscate data, the system employs learnable neural network components that automatically apply appropriate noise transformations. This substitution simplifies the overall data processing pipeline while maintaining effective obfuscation, as the neural network learns the optimal noise application strategy during training.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240185080A1Self-supervised data obfuscation in foundation models
Publication Date: 2024.06.06 PROTOPIA AI INC
  • US20240185080A1 patent drawing
  • US20240185080A1 patent drawing
  • US20240185080A1 patent drawing

AI summary

Provided are methods and system for obtaining, by a computer system, a machine learning/machine learning model; obtaining, by the computer system, a training data set; training, with the computer system, an obfuscation transform based on the machine learning/machine learning model and the training data set; and storing, with the computer system, the obfuscation transform in memory.