ML Training Data Filtering for Label Error Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Inaccurate training data leads to poor performance in machine learning models, resulting in incorrect predictions and inefficiencies across various applications, including consumer and enterprise scenarios.
Innovation Solution
Identify and remove suspect training data based on vector space variance and prediction confidence levels, using a system that includes a data item labeler, AI engine, and a portal to refine and retrain models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If training data is expanded to improve model coverage, then model versatility improves, but data accuracy deteriorates due to increased human error and labeling conflicts
Solution Approach 1:
The system performs preliminary detection and removal of inaccurate training data before model training. By identifying and eliminating erroneous samples in advance through automated analysis of labeling conflicts and data quality metrics, the system ensures that only high-quality data is used for training, thus maintaining data accuracy while still achieving comprehensive model coverage through selective data curation
Solution Approach 2:
The patent introduces an intermediary data quality assessment layer between data collection and model training. This intermediary system evaluates training data quality by detecting labeling conflicts, analyzing data consistency, and filtering erroneous samples, thereby mediating between the need for extensive training data and the requirement for high data accuracy
2Reliability
If manual data labeling is performed to improve data accuracy, then data quality improves, but processing time increases
Solution Approach 1:
The system implements self-service automated detection mechanisms that evaluate and filter training data quality without requiring extensive manual review. The automated system performs quality assessment, identifies labeling conflicts, and removes erroneous samples independently, thereby maintaining high data quality while eliminating the time-consuming manual processing step
Solution Approach 2:
The patent replaces the mechanical manual labeling and verification process with an automated computational system. The system uses algorithms to detect data quality issues, identify labeling conflicts, and filter erroneous samples automatically, substituting human manual work with machine-based processing that achieves comparable or superior quality at much faster speeds
3Reliability
If all training data is retained to improve model completeness, then model accuracy improves, but computational efficiency deteriorates due to processing inaccurate data
Solution Approach 1:
The system extracts and removes inaccurate training samples from the complete dataset before model training. By identifying and eliminating erroneous data points through automated quality assessment and conflict detection, the system retains only high-quality data for training, thereby maintaining model accuracy while improving computational efficiency by reducing the volume of data that needs to be processed
Solution Approach 2:
The patent changes the parameter of training data quality by implementing filtering criteria and quality thresholds. The system transforms the training dataset from containing all available data to containing only data that meets specified quality parameters, thereby optimizing the balance between model accuracy and computational efficiency through parameter-based data selection
Data Source
AI summary
Methods, systems and computer program products are described to improve machine learning (ML) model-based classification of data items by identifying and removing inaccurate training data. Inaccurate training samples may be identified, for example, based on excessive variance in vector space between a training sample and a mean of category training samples, and based on a variance between an assigned category and a predicted category for a training sample. Suspect or erroneous samples may be selectively removed based on, for example, vector space variance and/or prediction confidence level. As a result, ML model accuracy may be improved by training on a more accurate revised training set. ML model accuracy may (e.g., also) be improved, for example, by identifying and removing suspect categories with excessive (e.g., weighted) vector space variance. Suspect categories may be retained or revised. Users may (e.g., also) specify a prediction confidence level and/or coverage (e.g., to control accuracy).


