Machine Learning Model Training Using Segmented Privacy Data Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning model training processes often violate privacy laws due to the handling and processing of privacy-relevant data, which requires secure and efficient methods to protect sensitive information.
Innovation Solution
Separating privacy-relevant and non-relevant data portions and processing them in dedicated secure and efficient computing environments, respectively, using techniques like Trusted Execution Environment (TEE) and Homomorphic Encryption (HE), while applying transfer-learning for model retraining.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all training data processing is performed in a privacy-secure computing environment, then data security is improved, but computational efficiency deteriorates
Solution Approach 1:
The patent segments training data into privacy-relevant portions and privacy-nonrelevant portions, then processes these portions in different computing environments. Privacy-nonrelevant portions are processed in computationally-efficient environments while privacy-relevant portions are processed in privacy-secure environments, thereby resolving the contradiction between security and efficiency.
Solution Approach 2:
The patent applies different processing qualities to different portions of data based on their privacy characteristics. Privacy-relevant data receives high-security processing in trusted environments, while privacy-nonrelevant data receives standard processing in efficient environments, optimizing the overall system by matching security level to data sensitivity.
2Productivity
If privacy-relevant data is processed in computationally-efficient environments, then computational efficiency is improved, but data security deteriorates
Solution Approach 1:
The training data is segmented into privacy-relevant and privacy-nonrelevant portions. The privacy-nonrelevant portions can be safely processed in computationally-efficient environments without security risks, while privacy-relevant portions are directed to secure environments, thus achieving both efficiency and security where appropriate.
3Reliability
If privacy-relevant data is removed or masked from training data, then data security is improved, but model training quality deteriorates
Solution Approach 1:
Instead of uniformly removing or masking all privacy-relevant data, the patent processes privacy-relevant portions in privacy-secure computing environments where the data can be utilized for model training while maintaining security. This allows the model to learn from complete information while the security constraint is satisfied through the secure environment.
Data Source
AI summary
An approach for managing privacy-relevant data. Disclosed embodiments significantly improve the computational efficiency of training machine-learning models while still protecting privacy-relevant data in the training data.


