Filtering Imbalanced Training Sets for Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing predictive data analysis systems face inefficiencies and resource wastage due to the need for numerous training iterations when using imbalanced training sets, which lead to ineffective classification models.

Innovation Solution

The system selects optimal imbalance adjustment conditions to filter training entries, maximizing the cumulative target score while keeping the cumulative non-target score within a threshold, allowing for effective training of machine learning models without extensive iterations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional machine learning training is performed with imbalanced training sets, then the model can be trained without additional filtering, but the training requires numerous iterations and results in ineffective classification

Engineering Contradiction:
Improveclassification effectivenessVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies preliminary action by filtering training entries before the actual training process begins. The system identifies and removes unrepresentative training entries based on comparison between training set characteristics and live data characteristics, ensuring that only high-quality training data is used from the start, thereby avoiding the need for numerous training iterations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies parameter changes by modifying the composition of the training set through selective filtering. The system changes the parameters of the training data by removing entries that do not match the characteristics of live data, thereby transforming an imbalanced and ineffective training set into a balanced and effective one.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If numerous training iterations are performed with imbalanced training sets, then the model attempts to learn from all available data, but computational resources are wasted and training time increases

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary filtering of training entries by comparing training set characteristics with live data characteristics. This preliminary action identifies and removes unrepresentative entries before training begins, ensuring that subsequent training iterations use only high-quality data, thereby reducing total training time while maintaining or improving model accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system employs self-service by automatically detecting and removing unrepresentative training entries through characteristic comparison. The training data selection process serves itself by identifying its own deficiencies and correcting them through automated filtering, eliminating the need for manual data curation and reducing overall training time.

Inventive Principle:
Principle #25Self-service

3Reliability

If traditional training approaches are used with imbalanced datasets, then all training entries are processed uniformly, but the resulting models are ineffective and resource wastage occurs

Engineering Contradiction:
Improvemodel effectivenessVSAvoidcomputational resource usage
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system applies preliminary action by filtering training entries based on characteristic comparison before training commences. By removing unrepresentative entries in advance, the system ensures that computational resources are dedicated only to processing high-quality, representative training data, thereby improving model effectiveness while reducing energy waste.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies the extraction principle by removing unrepresentative training entries from the training set. The system extracts and eliminates data points that do not conform to live data characteristics, concentrating computational resources on the subset of training data that actually contributes to model effectiveness, thereby reducing energy waste.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20230134348A1Training classification machine learning models with imbalanced training sets
Publication Date: 2023.05.04 OPTUM INC
  • US20230134348A1 patent drawing
  • US20230134348A1 patent drawing
  • US20230134348A1 patent drawing

AI summary

Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for performing predictive data analysis operations. For example, certain embodiments of the present invention utilize systems, methods, and computer program products that perform predictive data analysis operations by machine learning models that are trained using one or more filtered training entries that are selected from a plurality of candidate training entries in accordance with one or more optimal imbalance adjustment conditions, where the one or more optimal imbalance adjustment conditions that are selected from a plurality of candidate imbalance adjustment conditions in a manner that is configured to maximize a cumulative target score for the one or more optimal imbalance adjustment conditions while a cumulative non-target score for the one or more optimal imbalance adjustment conditions satisfies an upper cumulative non-target score threshold.