Automated Training Data Selection via Error Likelihood Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating and validating ground truth data for cognitive question answering systems are inefficient and prone to errors due to the high cost and imperfections of hand-labeled data, leading to SME fatigue and inaccuracies.
Innovation Solution
An automated training data selection system that uses error classifications from machine-learning models to identify and select unlabeled data for annotation by SMEs, reducing specific error types through error type models and thresholds, thereby focusing annotation efforts on high-error data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If SMEs manually review and correct large volumes of machine-annotated ground truth data, then the training data quality can be improved, but SME accuracy deteriorates due to fatigue and sloppiness
Solution Approach 1:
The patent segments the large volume of training data review task into smaller batches by identifying and separating high-error cases from low-error cases. SMEs are presented with segmented portions of data that are most likely to contain errors, rather than requiring them to review all data uniformly. This segmentation reduces cognitive load and maintains accuracy by presenting manageable portions of work at each stage.
Solution Approach 2:
The patent introduces an intermediary error classification model that acts as a filter between the machine-annotated data and the SME reviewer. This intermediary component predicts which cases are likely to contain errors and prioritizes them for SME review, thereby reducing the total volume of data SMEs must manually validate while maintaining overall data quality.
2Reliability
If all unlabeled training data is presented to SMEs for annotation, then complete ground truth coverage can be achieved, but the cost and time required increases significantly
Solution Approach 1:
The patent applies partial action by having SMEs annotate only a subset of unlabeled training data - specifically, the portions predicted by the error classification model to have the highest error likelihood. Rather than requiring complete annotation of all data, the system achieves sufficient ground truth coverage by focusing resources on the most critical cases, thereby reducing annotation time while maintaining reliability.
Solution Approach 2:
The error classification model serves as a self-service mechanism that automatically identifies and prioritizes cases needing SME annotation. The system autonomously determines which unlabeled data should be presented to SMEs based on error predictions, eliminating the need for manual prioritization and reducing the overall annotation workload while ensuring comprehensive coverage of error-prone cases.
3Quantity of substance
If weak supervision is used to label large amounts of training data, then data quantity can be increased, but data quality deteriorates due to noise and errors in weak labels
Solution Approach 1:
The patent introduces an intermediary error classification model that acts as a quality filter between weakly supervised labeled data and the final training set. This intermediary component identifies cases where weak labels are likely to be erroneous and either flags them for SME review or excludes them from the training data, thereby maintaining data quantity while improving quality by removing noisy weak labels.
Solution Approach 2:
The system implements feedback by using the error classification model to continuously identify and flag problematic weak labels. The feedback loop allows the system to learn from error patterns in weakly supervised data and progressively improve the filtering of low-quality labels, thereby maintaining large data quantities while progressively improving data quality through iterative refinement.
Data Source
AI summary
A method, system, and a computer program product automatically select training data for updating a model by applying human-annotated training data to a model to generate results that are evaluated to identify correct case results and false case results that are categorized into error type categories for use in building error models corresponding to the error type categories, where each error model is built from at least failed case results belonging to a corresponding error type, and where unlabeled data samples are applied to each error model to compute an error likelihood for each unlabeled data sample with respect to each error type category, thereby enabling the selection and display of unlabeled data samples for annotation by a subject matter expert based on a computed error likelihood for the one or more unlabeled data samples in a specified error type category meeting or exceeding an error threshold requirement.


