Crowdsourcing System with Community Learning for Sparse Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Crowdsourcing systems face challenges in aggregating answers from diverse workers due to uncertainty in trustworthiness and quality, especially with sparse data, and scalability issues in machine learning for training processes.
Innovation Solution
A machine learning system that jointly learns characteristics of individual crowd workers and their communities using probabilistic graphical models, enabling efficient training and aggregation of labels through message passing schedules, and facilitating active learning, targeted training, and reward mechanisms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional machine learning approaches are used to aggregate crowdsourced answers, then training can be performed, but accuracy is poor when observed data about worker behavior is sparse
Solution Approach 1:
The patent combines individual worker characteristics with community-level characteristics into a unified probabilistic model. By merging data at both individual and community levels, the system achieves accurate label aggregation even when individual worker data is sparse, as community-level patterns compensate for individual data deficiencies
Solution Approach 2:
The patent introduces community characteristics as an intermediary layer between individual workers and the aggregation process. This intermediary captures shared behaviors and biases within worker communities, enabling accurate predictions for individual workers even when their personal data is limited
2Measurement precision
If machine learning systems are trained on web-scale data, then model accuracy improves, but training becomes time-consuming and computationally resource intensive
Solution Approach 1:
The patent segments the training process into separate phases: community-level pattern learning and individual worker characteristic learning. This segmentation allows for more efficient computation by processing data at appropriate granularities, reducing overall training time while maintaining accuracy on web-scale datasets
Solution Approach 2:
The patent employs dynamic training strategies where the system adapts its learning process based on data availability and computational resources. The probabilistic graphical model allows flexible inference that can operate effectively with varying amounts of training data and computational budget
3Measurement precision
If individual worker characteristics are modeled independently, then worker-specific accuracy may improve, but the system cannot leverage community-level patterns
Solution Approach 1:
The patent merges individual worker characteristics with community-level characteristics in a unified probabilistic framework. This combination allows the system to simultaneously capture worker-specific behaviors and community-level patterns, achieving both individual accuracy and collective intelligence
Data Source
AI summary
Crowdsourcing systems with machine learning are described. Specifically, item-label inference methods and systems are presented, for example, to provide aggregated answers to a crowdsourced task in a manner achieving good accuracy even where observed data about past behavior of crowd members is sparse. In various examples, an item-label inference system infers variables describing characteristics of both individual crowd workers and communities of the workers. In various examples, an item-label inference system provides aggregated labels while considering the inferred worker characteristics and the inferred characteristics of the worker communities. In examples the item-label inference system provides uncertainty information associated with the inference results for selecting workers and generating future tasks.


