Automated Data Quality Assessment for Training Sets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for creating and verifying training data for decision-making systems are time- and cost-intensive, requiring significant expert effort and producing excessive data, especially when dealing with large hierarchies of document types.

Innovation Solution

A computer-implemented method that selects training documents, estimates the quality of category organization, and determines if it meets a predetermined threshold, allowing for automatic decision systems to focus on one category at a time, reduce cognitive load, and diagnose suboptimal data quality for improved performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual verification and labeling of training data is performed, then data quality is ensured, but time consumption and cost increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system automatically performs data quality assessment and verification tasks that were previously done manually. The quality assessment module autonomously evaluates training data without requiring expert intervention, thereby maintaining data quality standards while eliminating the time-consuming manual verification process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical processes of expert verification are replaced with automated computational systems. The quality assessment module uses algorithmic approaches to evaluate data quality, substituting the mechanical process of human experts reviewing and verifying training data with an automated system that achieves the same goal more efficiently.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If experts manually label and verify training data, then accurate categorization is achieved, but cognitive effort and cost increase

Engineering Contradiction:
Improvecategorization accuracyVSAvoidcognitive effort
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The quality assessment module autonomously performs categorization verification without requiring expert cognitive effort. The system self-evaluates the accuracy of training data categorization using automated metrics and algorithms, maintaining high measurement precision while eliminating the need for experts to manually review and verify categorizations.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements automated feedback mechanisms where the quality assessment module continuously evaluates categorization accuracy and provides feedback for improvement. This closed-loop approach maintains high categorization precision by automatically identifying and correcting errors without requiring ongoing expert cognitive intervention.

Inventive Principle:
Principle #23Feedback

3Reliability

If manual processes are used to create training data, then data quality can be verified, but excessive training examples are produced

Engineering Contradiction:
Improvedata quality verificationVSAvoidnumber of training examples
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Manual processes that inevitably produce excessive training data are replaced with automated quality assessment mechanisms. The system uses computational methods to precisely determine when sufficient training data has been collected, stopping the data collection process automatically rather than relying on manual judgment that tends to over-collect.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The quality assessment module acts as an intermediary between data collection and the learning algorithm. It provides objective quality metrics that determine when training data is sufficient, mediating between the need for adequate data quantity and the risk of producing excessive examples that waste resources.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If large hierarchies of document types are used, then comprehensive classification is achieved, but cognitive load increases significantly

Engineering Contradiction:
Improveclassification coverageVSAvoidcognitive load
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The quality assessment module divides the evaluation of large document type hierarchies into manageable segments. Instead of requiring experts to evaluate all categories simultaneously, the system automatically assesses quality across the entire hierarchy by breaking down the complex classification structure into evaluable components, maintaining comprehensive classification coverage while eliminating excessive cognitive load.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10235446B2Systems and methods for organizing data sets
Publication Date: 2019.03.19 TUNGSTEN AUTOMATION CORPORATION
  • US10235446B2 patent drawing
  • US10235446B2 patent drawing
  • US10235446B2 patent drawing

AI summary

According to one embodiment, a computer-implemented method for cleaning up a data set having a possible incorrect label includes: selecting a plurality of training documents; estimating a quality of an organization of a plurality of categories; and determining whether the quality of the organization is greater than a predetermined quality threshold. Corresponding system and computer program product embodiments are also presented. Other aspects and advantages of the present invention will become apparent from the following detailed description, which, when taken in conjunction with the drawings, illustrate by way of example the principles of the invention.