Automated Ground Truth Curation for Classifier Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The quality of ground truth data used for classifier training is often compromised by redundant and conflicting records, leading to inefficiencies and reduced accuracy in classifier performance.
Innovation Solution
A method is implemented to identify and calculate similarity metrics between classifier training data elements, determining classifications and curating ground truth data to avoid duplication and contradiction, thereby enhancing data efficiency and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If raw ground truth data sets are used for classifier training, then the quantity of training data is maximized, but the quality and accuracy of classifier training deteriorates due to redundant and conflicting records
Solution Approach 1:
The system extracts and removes redundant and conflicting records from the raw ground truth data set through automated curation processes. By identifying and taking out harmful duplicate records and contradictory information, the system preserves only the high-quality, non-redundant training data, thus resolving the contradiction between maintaining large data quantity and ensuring data quality.
Solution Approach 2:
The system changes the parameters of the training data by transforming raw uncurated data into curated high-quality data through automated processes. This involves changing the purity parameter by removing contaminants (duplicates and conflicts) while maintaining the volume of useful training data, thereby improving reliability without sacrificing quantity.
2Reliability
If manual curation of ground truth data is performed to eliminate duplicates and conflicts, then the quality of training data improves, but the time and computational resources required increase significantly
Solution Approach 1:
The system implements self-service automated curation where the ground truth data curates itself without human intervention. The automated system independently identifies duplicates, detects conflicts, and removes problematic records from the training data set, eliminating the need for manual curation efforts while maintaining high data quality standards.
Solution Approach 2:
The system replaces manual mechanical curation processes with automated computational mechanisms. By substituting human reviewers and manual inspection methods with algorithmic duplicate detection and conflict identification systems, the patent dramatically reduces the time and labor required for data curation while maintaining or improving the quality outcomes.
3Productivity
If all records in the ground truth data set are retained for training, then the productivity of data preparation is maximized, but the classifier accuracy decreases due to conflicting information
Solution Approach 1:
The system performs preliminary automated curation actions on the ground truth data before it is used for classifier training. By pre-identifying and removing conflicting records and duplicates in advance, the system ensures that only high-quality data enters the training process, thereby maintaining both high productivity and high classifier accuracy without requiring post-processing corrections.
Data Source
AI summary
A computer-implemented method according to one embodiment includes identifying a first classifier training data element and a second classifier training data element, calculating a similarity metric between the first classifier training data element and the second classifier training data element, and determining a classification for the first classifier training data element and the second classifier training data element, utilizing the similarity metric between the first classifier training data element and the second classifier training data element.


