Crowdsourced ML Training Data via Super Recognizer Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Acquiring high-quality training data for machine learning models is challenging due to reliance on unreliable synthetic data or manual merges, which are computationally flawed and costly, especially in fields like health where accuracy is critical.

Innovation Solution

The system identifies 'super recognizers' among crowdworkers through confidence metrics to generate high-volume, high-accuracy training data by providing inputs to a crowdsourced platform, where super recognizers classify data accurately, and their annotations are used to create training datasets for machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If synthetic data or manual merges are used to acquire training data, then the process is simpler to implement, but the data quality and reliability deteriorate

Engineering Contradiction:
Improveease of training data acquisitionVSAvoiddata quality
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The system automatically evaluates crowdworker performance through confidence metrics and super recognizer identification, enabling the training data generation process to self-optimize without manual intervention. The platform autonomously selects high-quality annotators and generates reliable training data through automated assessment mechanisms.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the parameter of crowdworker selection from random or manual selection to automated selection based on confidence metrics and super recognizer status. This parameter transformation enables the system to dynamically identify and utilize the most reliable annotators, significantly improving data quality while maintaining ease of implementation.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If manual merging of training data is performed, then data quality may be maintained, but the cost and time required increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoidtraining data generation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system replaces the mechanical process of manual data merging with an automated computational system that uses confidence metrics and super recognizer identification. This substitution eliminates manual labor while maintaining or improving data quality through algorithmic selection of high-quality annotations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces confidence metrics and super recognizer identification as intermediary mechanisms between raw crowdworker annotations and final training data. These intermediaries automatically filter and select high-quality annotations, replacing manual review processes and significantly improving productivity while maintaining data reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If a larger pool of crowdworkers is utilized, then the talent pool and potential accuracy improve, but the complexity of identifying reliable annotators increases

Engineering Contradiction:
Improveannotator qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system extracts the key characteristic of reliable annotators through confidence metrics and super recognizer identification, separating quality assessment from the overall crowdworker pool. This extraction enables the system to focus on identifying high-quality annotators without being overwhelmed by the complexity of evaluating the entire crowdworker population.

Inventive Principle:
Principle #2Taking out (Extraction)

4Measurement precision

If confidence metrics and super recognizer identification are implemented, then data accuracy improves, but the computational processing required increases

Engineering Contradiction:
Improveannotation accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system applies partial action by using confidence metrics and super recognizer identification only for the portion of crowdworker annotations that require quality assessment. Rather than processing all annotations equally, the system focuses computational resources on evaluating and selecting high-quality annotations, reducing overall computational overhead while maintaining measurement precision.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20230306303A1Systems and Methods for Crowdsourced Machine Learning
Publication Date: 2023.09.28 THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
  • US20230306303A1 patent drawing
  • US20230306303A1 patent drawing
  • US20230306303A1 patent drawing

AI summary

Systems and methods for crowdsourced machine learning in accordance with embodiments of the invention are illustrated. In many embodiments, particular crowdworkers from a plurality of crowdworkers who are able to perform with high accuracy and reliability (referred to herein as “super recognizers”) are identified and used to generate training data for machine learning models. In various embodiments, super recognizers are identified by providing a request to answer questions regarding a particular type of input to the plurality of crowdworkers and providing received answers to a machine learning model trained using expert-annotated inputs similar to the inputs provided to the crowdworkers.