Missing Data Imputation Prioritization for Machine Learning Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies for imputing missing data in machine learning do not consider the extent of imputation required to enhance total prediction accuracy, leading to inefficiencies.
Innovation Solution
A method and device that prioritize data imputation based on an imputation amount adjustment parameter, predicting missing data using machine learning, and determining an integrated imputation priority order to reduce the workload and enhance prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If regression imputation is performed on all missing data, then imputation accuracy is improved, but the workload and time required for imputation increases
Solution Approach 1:
The patent applies partial action by selectively imputing only the most important missing data rather than all missing data. The system calculates importance scores for each missing data point based on correlation with target variables and other features, then prioritizes imputation of high-importance data points. This approach achieves sufficient prediction accuracy while significantly reducing imputation time and computational resources.
Solution Approach 2:
The patent applies local quality by treating different missing data points differently based on their individual importance. Instead of uniform imputation treatment, the system assigns different priority levels to different missing values based on their impact on prediction accuracy. High-importance missing data receives full imputation attention while low-importance data may be left as-is or handled with simpler methods.
2Measurement precision
If regression imputation is performed on all missing data, then imputation accuracy is improved, but device complexity and computational resources increase
Solution Approach 1:
The system performs imputation only on a subset of missing data points that are deemed most important for prediction accuracy. By calculating importance scores and selecting only high-priority missing values for imputation, the system reduces computational complexity and resource requirements while maintaining adequate accuracy for the overall machine learning task.
Solution Approach 2:
The system changes the parameter of imputation completeness from 100% to a selective subset based on importance scoring. By introducing an importance threshold parameter, the system can adjust the balance between accuracy and complexity, imputing only those missing values that exceed the threshold importance level.
3Reliability
If all missing data is imputed, then prediction accuracy is improved, but the proportion of data requiring processing increases
Solution Approach 1:
The system processes only a partial subset of missing data points based on their calculated importance scores. By identifying and imputing only the most critical missing values that have the greatest impact on prediction accuracy, the system maintains high reliability while significantly improving processing efficiency and reducing the proportion of data that requires intensive imputation processing.
Data Source
AI summary
There are provided an imputation priority flagging unit that gives an imputation priority flag to at least one entry falling within a predetermined proportion defined by an imputation amount adjustment parameter in pieces of data of specific column data items, and a priority order determination unit that counts the number of entries each given the imputation priority flag in entries included in a data table, and that determines an integrated imputation priority order in a descending order of the number of the imputation priority flags.


