Prediction Model Training Using Averaged Samples for Data Privacy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The use of personal information in training data for machine learning models poses a risk of personal information leakage, leading to reduced security in prediction models.
Innovation Solution
A prediction model creation apparatus that utilizes averaged samples obtained by averaging multiple samples of explanatory and objective variables, estimating a pre-averaging distribution, and performing machine learning on a prediction model based on this distribution to predict the objective variable, thereby reducing the risk of personal information leakage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine learning is performed using case data containing personal information, then prediction accuracy is improved, but security deteriorates due to risk of personal information leakage
Solution Approach 1:
The patent extracts only the necessary statistical information (averaged samples and pre-averaging distribution) from the original case data, separating the prediction model training from the raw personal information. This allows the model to learn from data patterns while removing identifiable personal information, thus maintaining prediction accuracy while preventing information leakage.
Solution Approach 2:
The patent introduces an intermediary representation layer consisting of averaged samples and pre-averaging distribution. These intermediaries serve as a bridge between the raw case data and the prediction model, enabling the model to access statistical patterns without direct access to personal information, thereby resolving the contradiction between accuracy and security.
2Reliability
If raw case data is used for training, then model performance is improved, but data security deteriorates
Solution Approach 1:
The patent creates a simplified copy of the essential statistical characteristics from the raw data through averaging. Instead of using the original detailed case data, the model trains on averaged samples that preserve the statistical patterns while eliminating individual data points, thus maintaining model performance while preventing raw data leakage.
Solution Approach 2:
The patent transforms the data representation by changing from individual case-level parameters to aggregated statistical parameters (averaged samples and distribution). This parameter transformation maintains the predictive power needed for model performance while fundamentally altering the data structure to prevent raw information leakage.
Data Source
AI summary
A prediction model creation apparatus of the present disclosure includes: an acquiring unit acquiring training data including averaged samples each obtained by averaging a plurality of samples each composed of a pair of an explanatory variable and an objective variable; an estimating unit estimating a pre-averaging distribution that is a distribution of explanatory variables before averaging corresponding to explanatory variables composing the averaged samples of the training data; and a training unit performing machine learning on a prediction model that predicts an objective variable from an explanatory variable, based on the training data and the pre-averaging distribution for decision making support.


