Prediction Model Correction via Statistical Distribution Differences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data analysis methods face challenges in performing accurate prediction analysis when limited data is available, leading to over-fitting and requiring large-scale data sets for effective prediction, which can be time-consuming and costly to collect.

Innovation Solution

A data analysis apparatus that generates a first prediction model using a smaller number of samples and corrects the prediction results using correction information derived from comparing statistical models based on actual measurement results from different data sets, allowing for highly accurate prediction independent of data quantity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If prediction analysis is performed using limited data samples, then the data collection time and cost are reduced, but prediction accuracy deteriorates due to over-fitting

Engineering Contradiction:
Improvedata collection timeVSAvoidprediction accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent introduces correction information as an intermediary element that mediates between the limited facility data and the target prediction goals. This correction information, derived from distribution comparisons between actual data and pseudo data, acts as a bridge that enhances prediction accuracy without requiring extensive data collection, thus resolving the contradiction between data collection time and prediction accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter of data representation by transforming raw facility data into correction information through statistical distribution comparisons. By modifying how data is processed and represented (from raw samples to correction factors), the system achieves accurate predictions with limited data, resolving the contradiction between sample size and prediction reliability

Inventive Principle:
Principle #35Parameter changes

2Reliability

If a large-scale data set is collected to improve prediction accuracy, then prediction reliability is improved, but the time and cost required for data collection increase

Engineering Contradiction:
Improveprediction reliabilityVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates pseudo data sets that copy the essential statistical characteristics of large-scale data without actually collecting the full data set. By generating synthetic data with matching distribution properties and using it to derive correction information, the system achieves reliable predictions without the time-consuming process of collecting and processing massive amounts of actual data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent extracts only the essential statistical distribution characteristics from data sets rather than collecting and processing complete large-scale data. By taking out and utilizing only the distributional properties needed for correction information, the system achieves reliable predictions while avoiding the time and resource costs of comprehensive data collection

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If machine learning is performed with insufficient data samples, then the learning process is faster and requires fewer resources, but over-fitting occurs reducing generalization capability

Engineering Contradiction:
Improvelearning speedVSAvoidgeneralization capability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The correction information serves as an intermediary that compensates for the insufficient data used in machine learning. By introducing this correction layer derived from distribution comparisons, the system maintains fast learning with limited samples while preventing over-fitting, thus preserving generalization capability without sacrificing learning speed

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a composite approach by combining machine learning models trained on limited facility data with correction information derived from statistical distribution comparisons. This composite prediction system leverages both the speed of machine learning on small data sets and the reliability of statistical correction, achieving both fast processing and good generalization

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS20220343122A1Data analysis apparatus, data analysis method, and data analysis program
Publication Date: 2022.10.27 HITACHI LTD
  • US20220343122A1 patent drawing
  • US20220343122A1 patent drawing
  • US20220343122A1 patent drawing

AI summary

To implement a highly accurate prediction analysis that does not depend on data amount. A data analysis apparatus has a processor that is configured to execute: an acquisition processing of acquiring a first statistical model based on a distribution of actual measurement results of a group and a second statistical model based on a distribution of a first actual measurement result of first samples having a smaller number of samples than the number of samples of the group; a calculation processing of calculating correction information indicating a difference between the first statistical model and the second statistical model; a learning processing of generating a first prediction model by performing machine learning using the first actual measurement result and first feature amount data corresponding to the first actual measurement result; and a correction processing of correcting a first prediction result , and outputting a second prediction result.