Machine Learning Model Training Using Segmented Privacy Data Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning model training processes often violate privacy laws due to the handling and processing of privacy-relevant data, which requires secure and efficient methods to protect sensitive information.

Innovation Solution

Separating privacy-relevant and non-relevant data portions and processing them in dedicated secure and efficient computing environments, respectively, using techniques like Trusted Execution Environment (TEE) and Homomorphic Encryption (HE), while applying transfer-learning for model retraining.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all training data processing is performed in a privacy-secure computing environment, then data security is improved, but computational efficiency deteriorates

Engineering Contradiction:
Improvedata securityVSAvoidcomputational efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments training data into privacy-relevant portions and privacy-nonrelevant portions, then processes these portions in different computing environments. Privacy-nonrelevant portions are processed in computationally-efficient environments while privacy-relevant portions are processed in privacy-secure environments, thereby resolving the contradiction between security and efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing qualities to different portions of data based on their privacy characteristics. Privacy-relevant data receives high-security processing in trusted environments, while privacy-nonrelevant data receives standard processing in efficient environments, optimizing the overall system by matching security level to data sensitivity.

Inventive Principle:
Principle #3Local quality

2Productivity

If privacy-relevant data is processed in computationally-efficient environments, then computational efficiency is improved, but data security deteriorates

Engineering Contradiction:
Improvecomputational efficiencyVSAvoiddata security
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The training data is segmented into privacy-relevant and privacy-nonrelevant portions. The privacy-nonrelevant portions can be safely processed in computationally-efficient environments without security risks, while privacy-relevant portions are directed to secure environments, thus achieving both efficiency and security where appropriate.

Inventive Principle:
Principle #1Segmentation

3Reliability

If privacy-relevant data is removed or masked from training data, then data security is improved, but model training quality deteriorates

Engineering Contradiction:
Improvedata securityVSAvoidmodel training quality
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

Instead of uniformly removing or masking all privacy-relevant data, the patent processes privacy-relevant portions in privacy-secure computing environments where the data can be utilized for model training while maintaining security. This allows the model to learn from complete information while the security constraint is satisfied through the secure environment.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240378311A1Method to improve training of classifiers when using data with personal identifiable information
Publication Date: 2024.11.14 ROBERT BOSCH GMBH
  • US20240378311A1 patent drawing
  • US20240378311A1 patent drawing
  • US20240378311A1 patent drawing

AI summary

An approach for managing privacy-relevant data. Disclosed embodiments significantly improve the computational efficiency of training machine-learning models while still protecting privacy-relevant data in the training data.