Machine Learning Algorithm for Crowdsourced Label Noise Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Crowdsourced platforms generate noisy labels due to non-professional assessors, which affects the quality of training data for machine learning algorithms, leading to inefficiencies in data labeling processes.

Innovation Solution

A computer-implemented method and system that uses a machine learning algorithm (MLA) to generate digital task labels by training on a dataset of labels submitted by workers, incorporating their activity histories and latent features to predict accurate labels, with a majority vote mechanism for determining the final label.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If crowdsourced platforms are used to obtain labels, then the quantity of labels and speed of acquisition are improved, but the quality and reliability of labels deteriorate due to non-professional assessors

Engineering Contradiction:
Improvespeed of label acquisitionVSAvoidquality of labels
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces an intermediary system that mediates between crowdsourced assessors and the final labeled data. This system uses multiple assessment methods including consistency checks, confidence scoring, and aggregation algorithms to filter and validate labels from non-professional assessors, thereby maintaining high productivity while improving reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent combines multiple labels from different assessors through aggregation mechanisms. By merging assessments from multiple workers and applying consensus algorithms, the system produces higher quality final labels that compensate for individual assessor limitations while maintaining efficient throughput

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If expert assessors are used to obtain labels, then the quality and reliability of labels are improved, but the cost and time required increase significantly

Engineering Contradiction:
Improvequality of labelsVSAvoidcost and time efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent creates a virtual copy of expert assessment capability through algorithmic systems. By implementing automated validation, consistency checking, and aggregation algorithms, the system replicates expert-quality label production at much lower cost and higher speed, without requiring actual expert human assessors

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system enables self-service label validation through automated confidence scoring and consistency checks. The crowdsourced platform itself performs quality assurance functions through built-in validation mechanisms, eliminating the need for expensive external expert review while maintaining label quality

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240160935A1Method and a system for generating a digital task label by machine learning algorithm
Publication Date: 2024.05.16 Y E HUB ARMENIA LLC
  • US20240160935A1 patent drawing
  • US20240160935A1 patent drawing
  • US20240160935A1 patent drawing

AI summary

A method system for selecting a label for a task, the method including, at a training phase: acquiring, a digital training task; acquiring, by the server, a plurality of digital training task labels having been submitted by a plurality of workers; acquiring, a worker activity history associated with each of the worker; training the MLA, including: inputting, the digital training task into the MLA; inputting, the worker activity histories into the MLA; generating a triplet of training objects, the triplet of training object including: the task vector representation, a given worker vector representation and a given digital training task label associated with the given worker vector representation; using the triplet of training objects to train the MLA to predict a given digital task label for a given digital task's task vector representation and a given worker vector representation.