Automated Ground Truth Curation for Classifier Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The quality of ground truth data used for classifier training is often compromised by redundant and conflicting records, leading to inefficiencies and reduced accuracy in classifier performance.

Innovation Solution

A method is implemented to identify and calculate similarity metrics between classifier training data elements, determining classifications and curating ground truth data to avoid duplication and contradiction, thereby enhancing data efficiency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If raw ground truth data sets are used for classifier training, then the quantity of training data is maximized, but the quality and accuracy of classifier training deteriorates due to redundant and conflicting records

Engineering Contradiction:
Improvequantity of training dataVSAvoidquality of training data
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system extracts and removes redundant and conflicting records from the raw ground truth data set through automated curation processes. By identifying and taking out harmful duplicate records and contradictory information, the system preserves only the high-quality, non-redundant training data, thus resolving the contradiction between maintaining large data quantity and ensuring data quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the parameters of the training data by transforming raw uncurated data into curated high-quality data through automated processes. This involves changing the purity parameter by removing contaminants (duplicates and conflicts) while maintaining the volume of useful training data, thereby improving reliability without sacrificing quantity.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If manual curation of ground truth data is performed to eliminate duplicates and conflicts, then the quality of training data improves, but the time and computational resources required increase significantly

Engineering Contradiction:
Improvequality of training dataVSAvoidcuration time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system implements self-service automated curation where the ground truth data curates itself without human intervention. The automated system independently identifies duplicates, detects conflicts, and removes problematic records from the training data set, eliminating the need for manual curation efforts while maintaining high data quality standards.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces manual mechanical curation processes with automated computational mechanisms. By substituting human reviewers and manual inspection methods with algorithmic duplicate detection and conflict identification systems, the patent dramatically reduces the time and labor required for data curation while maintaining or improving the quality outcomes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If all records in the ground truth data set are retained for training, then the productivity of data preparation is maximized, but the classifier accuracy decreases due to conflicting information

Engineering Contradiction:
Improvedata preparation efficiencyVSAvoidclassifier accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary automated curation actions on the ground truth data before it is used for classifier training. By pre-identifying and removing conflicting records and duplicates in advance, the system ensures that only high-quality data enters the training process, thereby maintaining both high productivity and high classifier accuracy without requiring post-processing corrections.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11200452B2Automatically curating ground truth data while avoiding duplication and contradiction
Publication Date: 2021.12.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11200452B2 patent drawing
  • US11200452B2 patent drawing
  • US11200452B2 patent drawing

AI summary

A computer-implemented method according to one embodiment includes identifying a first classifier training data element and a second classifier training data element, calculating a similarity metric between the first classifier training data element and the second classifier training data element, and determining a classification for the first classifier training data element and the second classifier training data element, utilizing the similarity metric between the first classifier training data element and the second classifier training data element.