Missing Data Imputation Prioritization for Machine Learning Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies for imputing missing data in machine learning do not consider the extent of imputation required to enhance total prediction accuracy, leading to inefficiencies.

Innovation Solution

A method and device that prioritize data imputation based on an imputation amount adjustment parameter, predicting missing data using machine learning, and determining an integrated imputation priority order to reduce the workload and enhance prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If regression imputation is performed on all missing data, then imputation accuracy is improved, but the workload and time required for imputation increases

Engineering Contradiction:
Improveimputation accuracyVSAvoidimputation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by selectively imputing only the most important missing data rather than all missing data. The system calculates importance scores for each missing data point based on correlation with target variables and other features, then prioritizes imputation of high-importance data points. This approach achieves sufficient prediction accuracy while significantly reducing imputation time and computational resources.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent applies local quality by treating different missing data points differently based on their individual importance. Instead of uniform imputation treatment, the system assigns different priority levels to different missing values based on their impact on prediction accuracy. High-importance missing data receives full imputation attention while low-importance data may be left as-is or handled with simpler methods.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If regression imputation is performed on all missing data, then imputation accuracy is improved, but device complexity and computational resources increase

Engineering Contradiction:
Improveimputation accuracyVSAvoidimputation system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs imputation only on a subset of missing data points that are deemed most important for prediction accuracy. By calculating importance scores and selecting only high-priority missing values for imputation, the system reduces computational complexity and resource requirements while maintaining adequate accuracy for the overall machine learning task.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system changes the parameter of imputation completeness from 100% to a selective subset based on importance scoring. By introducing an importance threshold parameter, the system can adjust the balance between accuracy and complexity, imputing only those missing values that exceed the threshold importance level.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If all missing data is imputed, then prediction accuracy is improved, but the proportion of data requiring processing increases

Engineering Contradiction:
Improveprediction accuracyVSAvoiddata processing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system processes only a partial subset of missing data points based on their calculated importance scores. By identifying and imputing only the most critical missing values that have the greatest impact on prediction accuracy, the system maintains high reliability while significantly improving processing efficiency and reducing the proportion of data that requires intensive imputation processing.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250265139A1Missing data imputation device and missing data imputation method
Publication Date: 2025.08.21 HITACHI LTD
  • US20250265139A1 patent drawing
  • US20250265139A1 patent drawing
  • US20250265139A1 patent drawing

AI summary

There are provided an imputation priority flagging unit that gives an imputation priority flag to at least one entry falling within a predetermined proportion defined by an imputation amount adjustment parameter in pieces of data of specific column data items, and a priority order determination unit that counts the number of entries each given the imputation priority flag in entries included in a data table, and that determines an integrated imputation priority order in a descending order of the number of the imputation priority flags.