Automated Training Data Selection via Error Likelihood Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating and validating ground truth data for cognitive question answering systems are inefficient and prone to errors due to the high cost and imperfections of hand-labeled data, leading to SME fatigue and inaccuracies.

Innovation Solution

An automated training data selection system that uses error classifications from machine-learning models to identify and select unlabeled data for annotation by SMEs, reducing specific error types through error type models and thresholds, thereby focusing annotation efforts on high-error data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If SMEs manually review and correct large volumes of machine-annotated ground truth data, then the training data quality can be improved, but SME accuracy deteriorates due to fatigue and sloppiness

Engineering Contradiction:
Improvetraining data qualityVSAvoidSME validation accuracy
Core Design Contradiction:
Manufacturing precisionVSMeasurement precision

Solution Approach 1:

The patent segments the large volume of training data review task into smaller batches by identifying and separating high-error cases from low-error cases. SMEs are presented with segmented portions of data that are most likely to contain errors, rather than requiring them to review all data uniformly. This segmentation reduces cognitive load and maintains accuracy by presenting manageable portions of work at each stage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary error classification model that acts as a filter between the machine-annotated data and the SME reviewer. This intermediary component predicts which cases are likely to contain errors and prioritizes them for SME review, thereby reducing the total volume of data SMEs must manually validate while maintaining overall data quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If all unlabeled training data is presented to SMEs for annotation, then complete ground truth coverage can be achieved, but the cost and time required increases significantly

Engineering Contradiction:
Improveground truth coverageVSAvoidannotation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by having SMEs annotate only a subset of unlabeled training data - specifically, the portions predicted by the error classification model to have the highest error likelihood. Rather than requiring complete annotation of all data, the system achieves sufficient ground truth coverage by focusing resources on the most critical cases, thereby reducing annotation time while maintaining reliability.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The error classification model serves as a self-service mechanism that automatically identifies and prioritizes cases needing SME annotation. The system autonomously determines which unlabeled data should be presented to SMEs based on error predictions, eliminating the need for manual prioritization and reducing the overall annotation workload while ensuring comprehensive coverage of error-prone cases.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If weak supervision is used to label large amounts of training data, then data quantity can be increased, but data quality deteriorates due to noise and errors in weak labels

Engineering Contradiction:
Improvetraining data quantityVSAvoidtraining data quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediary error classification model that acts as a quality filter between weakly supervised labeled data and the final training set. This intermediary component identifies cases where weak labels are likely to be erroneous and either flags them for SME review or excludes them from the training data, thereby maintaining data quantity while improving quality by removing noisy weak labels.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback by using the error classification model to continuously identify and flag problematic weak labels. The feedback loop allows the system to learn from error patterns in weakly supervised data and progressively improve the filtering of low-quality labels, thereby maintaining large data quantities while progressively improving data quality through iterative refinement.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11443209B2Method and system for unlabeled data selection using failed case analysis
Publication Date: 2022.09.13 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11443209B2 patent drawing
  • US11443209B2 patent drawing
  • US11443209B2 patent drawing

AI summary

A method, system, and a computer program product automatically select training data for updating a model by applying human-annotated training data to a model to generate results that are evaluated to identify correct case results and false case results that are categorized into error type categories for use in building error models corresponding to the error type categories, where each error model is built from at least failed case results belonging to a corresponding error type, and where unlabeled data samples are applied to each error model to compute an error likelihood for each unlabeled data sample with respect to each error type category, thereby enabling the selection and display of unlabeled data samples for annotation by a subject matter expert based on a computed error likelihood for the one or more unlabeled data samples in a specified error type category meeting or exceeding an error threshold requirement.