Crowd-Sourced Labeling System Using ML Example Subsets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Assessors on crowd-sourced platforms, being non-professional and varying in expertise, face difficulties in understanding tasks without actual examples, leading to inconsistent labeling quality.
Innovation Solution
A computer-implemented method and system that generates a subset of examples similar to the digital task, based on past tasks, to provide maximum benchmark coverage with minimal samples, and presents these examples to crowd-sourced workers to solicit labels, using machine learning algorithms to remove bias and cluster labels for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If crowd-sourced workers are used to perform tasks, then cost and time are reduced compared to professional assessors, but labeling quality becomes inconsistent due to varying expertise levels
Solution Approach 1:
The system performs preliminary actions by generating and presenting example labels before the worker performs the actual task. These examples are created using machine learning models trained on historical data, showing workers what correct labels look like for similar tasks, thereby preparing them to produce consistent quality work
Solution Approach 2:
The system introduces an intermediary mechanism - a machine learning model that generates example labels based on historical data. This intermediary bridges the gap between professional assessor quality and crowd worker capability by providing reference examples that guide workers toward consistent labeling standards
2Measurement precision
If more examples are provided to workers to improve understanding, then labeling accuracy improves, but the complexity and time required for task preparation increases
Solution Approach 1:
The system changes parameters by dynamically adjusting the number and type of examples provided based on task characteristics and worker performance. The machine learning model selects relevant examples from historical data, transforming the static example provision into a dynamic, adaptive process that optimizes for both accuracy and efficiency
Solution Approach 2:
The system creates copies of historical task examples and presents them to workers. Instead of requiring complex original content creation for each task, the system replicates and adapts proven examples from the past, reducing preparation complexity while maintaining labeling accuracy
Data Source
AI summary
There is disclosed a method and system for receiving a label for a digital task executed within a computer-implemented crowd-sourced environment, the method comprising: receiving, an indication of the digital task to be processed in the computer-implemented crowd-sourced environment; generating, a subset of examples, the subset of examples based on past digital tasks executed in the computer-implemented crowd-sourced environment, each of the subset of examples being similar to the digital task within a pre-determined similarity threshold; the subset of examples having a number of examples selected such that to provide maximum benchmark coverage with a minimum number of samples in the subset of examples; associating, the subset of examples to the digital task to be presented; causing the digital task to be presented on a computing device of at least one crowd-sourced worker in the computer-implemented crowd-sourced environment to solicit the label for the digital task.


