Prediction Model Training Using Averaged Samples for Data Privacy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The use of personal information in training data for machine learning models poses a risk of personal information leakage, leading to reduced security in prediction models.

Innovation Solution

A prediction model creation apparatus that utilizes averaged samples obtained by averaging multiple samples of explanatory and objective variables, estimating a pre-averaging distribution, and performing machine learning on a prediction model based on this distribution to predict the objective variable, thereby reducing the risk of personal information leakage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning is performed using case data containing personal information, then prediction accuracy is improved, but security deteriorates due to risk of personal information leakage

Engineering Contradiction:
Improveprediction accuracyVSAvoidpersonal information leakage risk
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the necessary statistical information (averaged samples and pre-averaging distribution) from the original case data, separating the prediction model training from the raw personal information. This allows the model to learn from data patterns while removing identifiable personal information, thus maintaining prediction accuracy while preventing information leakage.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary representation layer consisting of averaged samples and pre-averaging distribution. These intermediaries serve as a bridge between the raw case data and the prediction model, enabling the model to access statistical patterns without direct access to personal information, thereby resolving the contradiction between accuracy and security.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If raw case data is used for training, then model performance is improved, but data security deteriorates

Engineering Contradiction:
Improvemodel performanceVSAvoidraw data leakage
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent creates a simplified copy of the essential statistical characteristics from the raw data through averaging. Instead of using the original detailed case data, the model trains on averaged samples that preserve the statistical patterns while eliminating individual data points, thus maintaining model performance while preventing raw data leakage.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the data representation by changing from individual case-level parameters to aggregated statistical parameters (averaged samples and distribution). This parameter transformation maintains the predictive power needed for model performance while fundamentally altering the data structure to prevent raw information leakage.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20260037697A1Prediction model creation apparatus
Publication Date: 2026.02.05 NEC CORP
  • US20260037697A1 patent drawing
  • US20260037697A1 patent drawing
  • US20260037697A1 patent drawing

AI summary

A prediction model creation apparatus of the present disclosure includes: an acquiring unit acquiring training data including averaged samples each obtained by averaging a plurality of samples each composed of a pair of an explanatory variable and an objective variable; an estimating unit estimating a pre-averaging distribution that is a distribution of explanatory variables before averaging corresponding to explanatory variables composing the averaged samples of the training data; and a training unit performing machine learning on a prediction model that predicts an objective variable from an explanatory variable, based on the training data and the pre-averaging distribution for decision making support.