Crowdsourced Data Labeling Validation via Consensus Criteria

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The quality of training data obtained from crowdsourcing platforms is unstable due to worker bias and variance, which affects the accuracy of machine learning classifiers, despite the cost-effectiveness and scalability of crowdsourcing for generating labeled data.

Innovation Solution

A hybrid solution that combines crowdsourced labeling with domain expert validation, using a multi-level worker platform and an integrated data labeling engine (IDLE) to dynamically manage tasks, assess worker quality, and apply consensus criteria to aggregation of worker decisions, ensuring high-quality training data is collected efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If crowdsourcing platforms are used to generate labeled training data, then cost-effectiveness and scalability are improved, but the quality and stability of the training data deteriorate due to worker bias and variance

Engineering Contradiction:
Improvescalability of data generationVSAvoidquality stability of training data
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces an intermediary validation layer between crowdsourced workers and the final training dataset. Domain experts or automated validation mechanisms act as mediators to review and verify worker annotations, filtering out low-quality labels while preserving the scalability benefits of crowdsourcing. This intermediary step resolves the contradiction by maintaining high throughput while improving reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically adjusts worker selection parameters, task assignment parameters, and validation thresholds based on observed worker performance metrics. By changing these parameters adaptively, the system maintains optimal balance between cost-effectiveness and data quality stability, resolving the contradiction through parameter optimization.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If more crowdsourced workers are engaged to improve scalability, then the volume of labeled data increases, but worker bias and variance increase leading to decreased data quality

Engineering Contradiction:
Improvevolume of labeled dataVSAvoidannotation accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the crowdsourcing workforce into specialized groups based on domain expertise, task type, and performance metrics. Instead of treating all workers uniformly, the system divides workers into segments that are optimally matched to specific annotation tasks, reducing overall variance while maintaining scalable data production volume.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies partial validation to all annotations and excessive validation (domain expert review) to a subset of challenging or ambiguous cases. This selective approach maintains measurement precision for critical annotations while preserving scalability through efficient resource allocation.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If domain expert validation is added to verify crowdsourced labels, then the quality and stability of training data are improved, but the time and cost of data generation increase

Engineering Contradiction:
Improvequality of training dataVSAvoiddata generation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements partial validation where domain experts review only a subset of annotations - specifically those from new workers, ambiguous cases, or high-impact categories. Most routine annotations are validated through automated consistency checks, reducing time loss while maintaining quality through targeted expert review.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system enables worker self-validation through automated feedback mechanisms that provide immediate quality assessments and guidance. Workers can self-correct annotations based on system feedback, reducing the need for time-consuming manual expert validation while improving overall data quality.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240193645A1Quality of labeled training data
Publication Date: 2024.06.13 NIELSEN CONSUMER LLC
  • US20240193645A1 patent drawing
  • US20240193645A1 patent drawing
  • US20240193645A1 patent drawing

AI summary

Methods, systems, apparatus, and tangible non-transitory carrier media encoded with one or more computer programs for classifying an item. In accordance with particular embodiments, a labeling task is issued to workers participating in a crowdsourcing system. The labeling task includes evaluating an inferred classification that includes one or more of the class labels in a hierarchical classification taxonomy based at least in part on a description of the item and the class labels in the classification. Evaluation decisions are received from the crowdsourcing system. The classification is validated based on the evaluation decisions to obtain a validation result. The validating includes applying at least one consensus criterion to an aggregation of the received evaluation decisions. Data corresponding to one or more of the class labels in the classification is routed to respective destinations based on the validation result.