Prediction Model Building with Missing-Data Augmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine-learning prediction models fail to generate accurate results due to incomplete training or source data with missing values, particularly in physiological data collection scenarios where personal information is partially missing, leading to reduced dataset accuracy and inability to predict states reliably.
Innovation Solution
A method to augment incomplete training and source data by generating augmented values for missing features, forming complete datasets for model training and prediction, using processors and machine-learning algorithms to enhance data completeness and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data collection includes personal information with missing values, then data completeness is improved, but prediction model accuracy deteriorates
Solution Approach 1:
The patent creates copies of incomplete data records by generating augmented data through machine learning algorithms. These augmented copies fill in missing personal information fields using learned patterns from complete records, thereby maintaining data quantity while improving data quality for model training.
Solution Approach 2:
The patent introduces an intermediary processing layer between raw data collection and prediction model training. This intermediary uses machine learning algorithms to preprocess and augment incomplete data records, transforming them into complete virtual records that can be used for training without losing original data diversity.
2Ease of operation
If training data has missing values, then data collection ease is improved, but model training reliability deteriorates
Solution Approach 1:
The patent performs preliminary data augmentation before model training begins. By pre-processing incomplete data records and generating augmented versions in advance, the system ensures that the training data is complete and reliable before the actual model training starts, maintaining both data collection ease and training reliability.
Solution Approach 2:
The patent creates virtual copies of incomplete data records by generating augmented data through machine learning algorithms. These augmented copies fill in missing personal information fields using learned patterns from complete records, thereby maintaining data quantity while improving data quality for model training.
3Area of stationary object
If radar data scanning is performed, then measurement coverage is improved, but data processing complexity increases
Solution Approach 1:
The patent extracts only the essential features from the radar data scanning process and separates them from the complex processing requirements. By identifying which features are critical for prediction and which can be simplified or imputed, the system maintains comprehensive measurement coverage while reducing processing complexity through selective feature extraction and augmentation.
Data Source
AI summary
A prediction-model-building method includes: generating a plurality of pieces of augmented training data after a piece of training data with a missing item is determined; constituting a training dataset comprising pieces of training data without any missing item and the plurality of pieces of augmented training data; and building a prediction model based on the training dataset by a machine-learning algorithm. Similarly, a state prediction method includes: generating a plurality of pieces of estimation augmented data after a piece of source data with a missing item is determined; inputting the said plurality of pieces of estimation augmented data into a prediction model; and outputting a predicted state based on the output of the prediction model. Devices for performing the above methods are also disclosed.


