Artificial Case Generation for Machine Learning Prediction Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating artificial cases for machine learning models do not effectively improve predictive performance, as they often produce redundant cases that do not contribute to accuracy and may distort the original data distribution.
Innovation Solution
An information processing device and method that selectively generates and adds artificial cases based on actual cases where the model's prediction is uncertain, focusing on cases close to the decision boundary and ensuring diversity to enhance prediction accuracy without over-generating redundant cases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If artificial cases are generated based on actual cases, then the number of training cases is increased, but the predictive performance of the machine learning model does not necessarily improve
Solution Approach 1:
The patent applies local quality by generating artificial cases with different degrees of similarity to actual cases. Specifically, it creates multiple artificial cases with varying similarity levels (high, medium, low) to different actual cases, rather than uniformly generating all artificial cases with the same similarity. This allows the system to selectively use artificial cases that locally match the characteristics of uncertain actual cases, thereby improving predictive performance in specific regions of the feature space where the model is uncertain.
Solution Approach 2:
The patent implements partial action by generating only the necessary number of artificial cases with appropriate similarity levels, rather than excessively generating all possible artificial cases. It selectively generates artificial cases based on the uncertainty of actual cases and their similarity relationships, avoiding redundant generation. This partial generation approach efficiently increases training data quantity while maintaining model reliability by focusing computational resources on the most beneficial artificial cases.
2Quantity of substance
If multiple artificial cases are generated from each actual case, then more training data is available, but redundant cases are created that do not contribute to accuracy improvement
Solution Approach 1:
The patent applies parameter changes by varying the similarity parameter when generating artificial cases. Instead of generating all artificial cases with a fixed similarity level to their source actual cases, it creates artificial cases with different similarity parameters (high, medium, low). This parameter variation ensures that the generated artificial cases have diverse characteristics and reduces redundancy by spreading the artificial cases across different similarity ranges, thereby improving the effectiveness of training data expansion.
3Adaptability or versatility
If artificial cases are added to training data, then the model can handle more cases, but the original data distribution may be distorted
Solution Approach 1:
The patent applies local quality by generating artificial cases with different similarity levels to preserve local data distribution characteristics. High-similarity artificial cases maintain the local structure near actual cases, medium-similarity cases extend to intermediate regions, and low-similarity cases explore broader areas. This multi-level similarity approach allows the model to handle more diverse cases while preserving the original data distribution at different scales, preventing distortion of the underlying data structure.
Data Source
AI summary
In an information processing device, an input means acquires each actual case formed by features. An artificial case generation means generates a plurality of artificial cases based on each acquired actual case. An artificial case selection means configured to select each artificial case in which a prediction of a machine learning model is to be uncertain, from the plurality of artificial cases. After that, an output means outputs each selected artificial case.


