Automated Data Quality Assessment for Training Sets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for creating and verifying training data for decision-making systems are time- and cost-intensive, requiring significant expert effort and producing excessive data, especially when dealing with large hierarchies of document types.
Innovation Solution
A computer-implemented method that selects training documents, estimates the quality of category organization, and determines if it meets a predetermined threshold, allowing for automatic decision systems to focus on one category at a time, reduce cognitive load, and diagnose suboptimal data quality for improved performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual verification and labeling of training data is performed, then data quality is ensured, but time consumption and cost increase significantly
Solution Approach 1:
The system automatically performs data quality assessment and verification tasks that were previously done manually. The quality assessment module autonomously evaluates training data without requiring expert intervention, thereby maintaining data quality standards while eliminating the time-consuming manual verification process.
Solution Approach 2:
Manual mechanical processes of expert verification are replaced with automated computational systems. The quality assessment module uses algorithmic approaches to evaluate data quality, substituting the mechanical process of human experts reviewing and verifying training data with an automated system that achieves the same goal more efficiently.
2Measurement precision
If experts manually label and verify training data, then accurate categorization is achieved, but cognitive effort and cost increase
Solution Approach 1:
The quality assessment module autonomously performs categorization verification without requiring expert cognitive effort. The system self-evaluates the accuracy of training data categorization using automated metrics and algorithms, maintaining high measurement precision while eliminating the need for experts to manually review and verify categorizations.
Solution Approach 2:
The system implements automated feedback mechanisms where the quality assessment module continuously evaluates categorization accuracy and provides feedback for improvement. This closed-loop approach maintains high categorization precision by automatically identifying and correcting errors without requiring ongoing expert cognitive intervention.
3Reliability
If manual processes are used to create training data, then data quality can be verified, but excessive training examples are produced
Solution Approach 1:
Manual processes that inevitably produce excessive training data are replaced with automated quality assessment mechanisms. The system uses computational methods to precisely determine when sufficient training data has been collected, stopping the data collection process automatically rather than relying on manual judgment that tends to over-collect.
Solution Approach 2:
The quality assessment module acts as an intermediary between data collection and the learning algorithm. It provides objective quality metrics that determine when training data is sufficient, mediating between the need for adequate data quantity and the risk of producing excessive examples that waste resources.
4Adaptability or versatility
If large hierarchies of document types are used, then comprehensive classification is achieved, but cognitive load increases significantly
Solution Approach 1:
The quality assessment module divides the evaluation of large document type hierarchies into manageable segments. Instead of requiring experts to evaluate all categories simultaneously, the system automatically assesses quality across the entire hierarchy by breaking down the complex classification structure into evaluable components, maintaining comprehensive classification coverage while eliminating excessive cognitive load.
Data Source
AI summary
According to one embodiment, a computer-implemented method for cleaning up a data set having a possible incorrect label includes: selecting a plurality of training documents; estimating a quality of an organization of a plurality of categories; and determining whether the quality of the organization is greater than a predetermined quality threshold. Corresponding system and computer program product embodiments are also presented. Other aspects and advantages of the present invention will become apparent from the following detailed description, which, when taken in conjunction with the drawings, illustrate by way of example the principles of the invention.


