Flexible Data Security System for Secure Third-Party ML Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Entities are reluctant to share their data for machine learning due to data security concerns, limiting the availability and accuracy of training data and thus the utility of machine-learned models.
Innovation Solution
A flexible data security and machine learning system that allows third-party entities to select from various data security options for storage, accessibility, and mixing, ensuring secure processing and sharing of their data while enabling more accurate model training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If entities share their data sets for machine learning training, then model accuracy and utility are improved, but data security risks increase
Solution Approach 1:
The patent introduces a machine learning system as an intermediary that processes data from multiple entities. The system receives data sets from entities, processes them according to security policies, and generates models without requiring entities to directly share or access each other's raw data. This intermediary approach enables collaborative model training while maintaining data security and privacy for each entity.
Solution Approach 2:
The patent implements data security policies that allow entities to control parameters of data processing, such as specifying which data can be accessed, how it can be processed, and under what conditions. Entities can adjust these parameters to balance between data sharing for improved model accuracy and data security requirements, enabling flexible control over data utilization.
2Object-affected harmful factors
If entities do not share their data sets due to security concerns, then data security is maintained, but model accuracy and utility deteriorate
Solution Approach 1:
The machine learning system serves as a secure intermediary that processes data without requiring entities to share direct access to their data sets. The system implements security policies that prevent unauthorized access while still enabling the aggregation of information needed for accurate model training, thus maintaining both security and model quality.
Solution Approach 2:
The patent segments the data processing workflow into distinct stages where entities maintain control over their data. Each entity provides data according to their security policies, and the system processes it independently before generating models. This segmentation allows entities to retain security control while contributing to overall model accuracy through their segmented data contributions.
3Productivity
If more training data is used to improve model accuracy, then model utility increases, but data security concerns intensify
Solution Approach 1:
The machine learning system acts as a secure intermediary that aggregates data from multiple entities to create comprehensive training data sets. By processing data through this controlled intermediary system, the platform can accumulate sufficient training data for high model utility while maintaining security policies that prevent direct entity access to each other's data, thus reducing security concerns.
Solution Approach 2:
The system implements universal data processing capabilities that can handle data from multiple entities under various security policies. This multi-functional approach allows the system to aggregate diverse data types and sources for improved model utility while providing flexible security controls that address different entity requirements, thereby increasing productivity without proportionally increasing security risks.
Data Source
AI summary
Techniques for a flexible data security and machine learning system for merging third-party data are provided. In one technique, the system receives a data set from a third-party entity and receives selection data that indicates that the third-party entity selected a set of data security policies that includes an encryption option and a data mixing option from among multiple data mixing options. In response to receiving the selection data, the system stores data that associates the set of data security policies with the data set, encrypts the data set according to the encryption option, and persistently stores the encrypted data set. Later, the system decrypts the encrypted data set in volatile memory, generates, based on the data mixing option, training data based on the decrypted version of the data set, trains a machine-learned model based on the training data, and stores the machine-learned model in association with the data set.


