Filtering Imbalanced Training Sets for Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing predictive data analysis systems face inefficiencies and resource wastage due to the need for numerous training iterations when using imbalanced training sets, which lead to ineffective classification models.
Innovation Solution
The system selects optimal imbalance adjustment conditions to filter training entries, maximizing the cumulative target score while keeping the cumulative non-target score within a threshold, allowing for effective training of machine learning models without extensive iterations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional machine learning training is performed with imbalanced training sets, then the model can be trained without additional filtering, but the training requires numerous iterations and results in ineffective classification
Solution Approach 1:
The patent applies preliminary action by filtering training entries before the actual training process begins. The system identifies and removes unrepresentative training entries based on comparison between training set characteristics and live data characteristics, ensuring that only high-quality training data is used from the start, thereby avoiding the need for numerous training iterations.
Solution Approach 2:
The patent applies parameter changes by modifying the composition of the training set through selective filtering. The system changes the parameters of the training data by removing entries that do not match the characteristics of live data, thereby transforming an imbalanced and ineffective training set into a balanced and effective one.
2Measurement precision
If numerous training iterations are performed with imbalanced training sets, then the model attempts to learn from all available data, but computational resources are wasted and training time increases
Solution Approach 1:
The system performs preliminary filtering of training entries by comparing training set characteristics with live data characteristics. This preliminary action identifies and removes unrepresentative entries before training begins, ensuring that subsequent training iterations use only high-quality data, thereby reducing total training time while maintaining or improving model accuracy.
Solution Approach 2:
The system employs self-service by automatically detecting and removing unrepresentative training entries through characteristic comparison. The training data selection process serves itself by identifying its own deficiencies and correcting them through automated filtering, eliminating the need for manual data curation and reducing overall training time.
3Reliability
If traditional training approaches are used with imbalanced datasets, then all training entries are processed uniformly, but the resulting models are ineffective and resource wastage occurs
Solution Approach 1:
The system applies preliminary action by filtering training entries based on characteristic comparison before training commences. By removing unrepresentative entries in advance, the system ensures that computational resources are dedicated only to processing high-quality, representative training data, thereby improving model effectiveness while reducing energy waste.
Solution Approach 2:
The patent applies the extraction principle by removing unrepresentative training entries from the training set. The system extracts and eliminates data points that do not conform to live data characteristics, concentrating computational resources on the subset of training data that actually contributes to model effectiveness, thereby reducing energy waste.
Data Source
AI summary
Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for performing predictive data analysis operations. For example, certain embodiments of the present invention utilize systems, methods, and computer program products that perform predictive data analysis operations by machine learning models that are trained using one or more filtered training entries that are selected from a plurality of candidate training entries in accordance with one or more optimal imbalance adjustment conditions, where the one or more optimal imbalance adjustment conditions that are selected from a plurality of candidate imbalance adjustment conditions in a manner that is configured to maximize a cumulative target score for the one or more optimal imbalance adjustment conditions while a cumulative non-target score for the one or more optimal imbalance adjustment conditions satisfies an upper cumulative non-target score threshold.


