ML Training Data Filtering for Label Error Removal

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Inaccurate training data leads to poor performance in machine learning models, resulting in incorrect predictions and inefficiencies across various applications, including consumer and enterprise scenarios.

Innovation Solution

Identify and remove suspect training data based on vector space variance and prediction confidence levels, using a system that includes a data item labeler, AI engine, and a portal to refine and retrain models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If training data is expanded to improve model coverage, then model versatility improves, but data accuracy deteriorates due to increased human error and labeling conflicts

Engineering Contradiction:
Improvemodel coverageVSAvoiddata accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system performs preliminary detection and removal of inaccurate training data before model training. By identifying and eliminating erroneous samples in advance through automated analysis of labeling conflicts and data quality metrics, the system ensures that only high-quality data is used for training, thus maintaining data accuracy while still achieving comprehensive model coverage through selective data curation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary data quality assessment layer between data collection and model training. This intermediary system evaluates training data quality by detecting labeling conflicts, analyzing data consistency, and filtering erroneous samples, thereby mediating between the need for extensive training data and the requirement for high data accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If manual data labeling is performed to improve data accuracy, then data quality improves, but processing time increases

Engineering Contradiction:
Improvedata qualityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system implements self-service automated detection mechanisms that evaluate and filter training data quality without requiring extensive manual review. The automated system performs quality assessment, identifies labeling conflicts, and removes erroneous samples independently, thereby maintaining high data quality while eliminating the time-consuming manual processing step

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual labeling and verification process with an automated computational system. The system uses algorithms to detect data quality issues, identify labeling conflicts, and filter erroneous samples automatically, substituting human manual work with machine-based processing that achieves comparable or superior quality at much faster speeds

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If all training data is retained to improve model completeness, then model accuracy improves, but computational efficiency deteriorates due to processing inaccurate data

Engineering Contradiction:
Improvemodel accuracyVSAvoidcomputational efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system extracts and removes inaccurate training samples from the complete dataset before model training. By identifying and eliminating erroneous data points through automated quality assessment and conflict detection, the system retains only high-quality data for training, thereby maintaining model accuracy while improving computational efficiency by reducing the volume of data that needs to be processed

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of training data quality by implementing filtering criteria and quality thresholds. The system transforms the training dataset from containing all available data to containing only data that meets specified quality parameters, thereby optimizing the balance between model accuracy and computational efficiency through parameter-based data selection

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12524707B2System and method for improving machine learning models by detecting and removing inaccurate training data
Publication Date: 2026.01.13 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12524707B2 patent drawing
  • US12524707B2 patent drawing
  • US12524707B2 patent drawing

AI summary

Methods, systems and computer program products are described to improve machine learning (ML) model-based classification of data items by identifying and removing inaccurate training data. Inaccurate training samples may be identified, for example, based on excessive variance in vector space between a training sample and a mean of category training samples, and based on a variance between an assigned category and a predicted category for a training sample. Suspect or erroneous samples may be selectively removed based on, for example, vector space variance and/or prediction confidence level. As a result, ML model accuracy may be improved by training on a more accurate revised training set. ML model accuracy may (e.g., also) be improved, for example, by identifying and removing suspect categories with excessive (e.g., weighted) vector space variance. Suspect categories may be retained or revised. Users may (e.g., also) specify a prediction confidence level and/or coverage (e.g., to control accuracy).