Missing Value Correction via Iterative Machine Learning Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for correcting missing values in data analysis are limited by their reliance on small models, leading to low accuracy and reduced statistical power due to the deletion of data sets containing missing values.
Innovation Solution
A method and apparatus that use machine learning algorithms to select variables, learn data, and predict missing values through a prediction model configuration process, involving integrity data extraction, feature data extraction, and iterative correction using multiple prediction models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data sets containing missing values are deleted, then data analysis can proceed without missing values, but the total amount of data is reduced and statistical power is lowered
Solution Approach 1:
The patent extracts and separates the missing value problem from the complete data set. Instead of removing entire data sets containing missing values, the method extracts only the missing value components and replaces them with predicted values generated by machine learning models, thereby preserving the complete data structure while eliminating the harmful effect of missing values.
Solution Approach 2:
The patent creates copies of complete data patterns to fill in missing values. Machine learning models are trained on complete data sets and then used to generate predicted values that copy the statistical patterns and relationships observed in the complete data, allowing these predicted copies to replace missing values while maintaining data integrity.
2Reliability
If missing data is replaced with average data or most frequent data, then missing values are corrected, but the accuracy of the correction is not high due to limited models
Solution Approach 1:
The patent changes the parameters of the correction models from simple statistical measures (average, mode) to sophisticated machine learning models with multiple algorithms and tunable parameters. The system selects and configures appropriate machine learning models based on the specific characteristics of the data and missing value patterns, thereby significantly improving correction accuracy.
Solution Approach 2:
The patent implements a self-service mechanism where the system automatically selects and configures the most appropriate machine learning models for correcting missing values based on the data characteristics. The model selection and parameter optimization are performed automatically without requiring manual intervention, enabling the system to adapt to different data types and missing value patterns dynamically.
3Measurement precision
If machine learning algorithms are used to predict missing values, then prediction accuracy is improved, but computational complexity increases
Solution Approach 1:
The patent applies partial machine learning modeling by training models only on the specific features and relationships relevant to predicting missing values in each context. Rather than applying exhaustive or overly complex models to all data, the system uses targeted machine learning approaches that apply computational resources only where necessary and proportionate to the prediction accuracy requirements.
Data Source
AI summary
A method and apparatus for correcting missing values in data are provided. A method of correcting missing values in basic data according to an embodiment includes a data extraction step, a prediction model configuration step, a first correction step, and a second correction step. The method corrects missing values in data by repeating the steps of generating a prediction model for correcting the missing value and correcting the missing value with the use of the prediction model.


