Latent Missing Feature Detection in Machine Learning Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning model training datasets often suffer from latent missing features, leading to degraded performance and increased bias, as traditional data augmentation techniques are either costly or inaccurate, and current methods fail to effectively prioritize feature collection.

Innovation Solution

The development of data augmentation techniques that predict feature impacts and leverage combinatoric optimization to prioritize data collection operations, generating a datapoint priority matrix that optimizes data collection based on entity-feature value pairs, thereby improving model performance while minimizing resource constraints.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If comprehensive data collection operations are conducted to improve model performance, then model accuracy is improved, but time consumption and cost increase significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by selectively collecting only the most impactful missing features rather than comprehensively collecting all possible features. The system identifies and prioritizes a subset of features that will yield the greatest model performance improvement, thereby reducing data collection time and cost while still achieving significant accuracy gains.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs self-service by automatically analyzing the training dataset to identify missing features and their impact on model performance. The patent uses the existing data and model to generate insights about what additional data would be most valuable, eliminating the need for external expert guidance in data collection strategy.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If comprehensive data collection operations are conducted to improve model performance, then model accuracy is improved, but computational resources are wasted on redundant features

Engineering Contradiction:
Improvemodel accuracyVSAvoidcomputational resource waste
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent applies partial action by collecting only the necessary subset of features rather than all possible features. By identifying and prioritizing features based on their impact on model performance, the system avoids collecting redundant features that would consume computational resources without providing proportional value.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system changes parameters by dynamically adjusting feature priority based on model performance metrics. The patent uses sensitivity analysis to determine which feature parameters have the greatest impact on model output, allowing selective collection of high-impact features while ignoring low-impact redundant features.

Inventive Principle:
Principle #35Parameter changes

3Loss of energy

If synthetic data augmentation techniques are used to replace missing features, then data collection cost is reduced, but data accuracy and real-world characteristics are compromised

Engineering Contradiction:
Improvedata collection costVSAvoiddata accuracy
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The patent applies partial action by using synthetic data augmentation only for low-priority features while collecting real data for high-priority features. This selective approach balances cost reduction with accuracy preservation, using synthetic data only where it will have minimal impact on model performance.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system applies local quality by differentiating between high-priority and low-priority features in terms of data collection strategy. High-priority features that significantly impact model performance receive real collected data, while low-priority features use synthetic augmented data, optimizing the balance between cost and accuracy locally for each feature.

Inventive Principle:
Principle #3Local quality

4Measurement precision

If unguided comprehensive data augmentation techniques are used, then all features are targeted, but resource efficiency decreases due to redundant feature collection

Engineering Contradiction:
Improvefeature coverageVSAvoidresource efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies partial action by selectively targeting only the most impactful missing features rather than attempting to collect all possible features. The system identifies a priority subset of features based on their contribution to model performance, collecting data for these high-priority features while skipping lower-priority redundant features.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system changes parameters by dynamically prioritizing features based on their impact on model performance. The patent uses sensitivity analysis and importance scoring to transform the feature set into a prioritized list, allowing the data collection process to focus on high-impact features and improve resource efficiency.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240265304A1Optimized latent missing feature detection for machine learning models
Publication Date: 2024.08.08 OPTUM INC
  • US20240265304A1 patent drawing
  • US20240265304A1 patent drawing
  • US20240265304A1 patent drawing

AI summary

Various embodiments of the present disclosure provide techniques for optimally augmenting a training dataset for a machine learning model based on multiple model-focused predictions. The techniques may include generating a datapoint priority matrix that corresponds to a plurality of entity-feature value pairs of a training dataset for a machine learning model, generating a plurality of impact predictions and feature sensitivity predictions for the plurality of entity-feature value pairs, generating a refined datapoint priority matrix by updating the datapoint priority matrix based on the plurality of impact predictions and sensitivity predictions, and providing a datapoint collection output for the training dataset based on the refined datapoint priority matrix and a data augmentation threshold.