Crowdsourcing System with Community Learning for Sparse Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Crowdsourcing systems face challenges in aggregating answers from diverse workers due to uncertainty in trustworthiness and quality, especially with sparse data, and scalability issues in machine learning for training processes.

Innovation Solution

A machine learning system that jointly learns characteristics of individual crowd workers and their communities using probabilistic graphical models, enabling efficient training and aggregation of labels through message passing schedules, and facilitating active learning, targeted training, and reward mechanisms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional machine learning approaches are used to aggregate crowdsourced answers, then training can be performed, but accuracy is poor when observed data about worker behavior is sparse

Engineering Contradiction:
Improveaccuracy of aggregated labelsVSAvoidamount of observed data
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent combines individual worker characteristics with community-level characteristics into a unified probabilistic model. By merging data at both individual and community levels, the system achieves accurate label aggregation even when individual worker data is sparse, as community-level patterns compensate for individual data deficiencies

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces community characteristics as an intermediary layer between individual workers and the aggregation process. This intermediary captures shared behaviors and biases within worker communities, enabling accurate predictions for individual workers even when their personal data is limited

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If machine learning systems are trained on web-scale data, then model accuracy improves, but training becomes time-consuming and computationally resource intensive

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the training process into separate phases: community-level pattern learning and individual worker characteristic learning. This segmentation allows for more efficient computation by processing data at appropriate granularities, reducing overall training time while maintaining accuracy on web-scale datasets

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs dynamic training strategies where the system adapts its learning process based on data availability and computational resources. The probabilistic graphical model allows flexible inference that can operate effectively with varying amounts of training data and computational budget

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If individual worker characteristics are modeled independently, then worker-specific accuracy may improve, but the system cannot leverage community-level patterns

Engineering Contradiction:
Improveworker-specific prediction accuracyVSAvoidability to leverage community patterns
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent merges individual worker characteristics with community-level characteristics in a unified probabilistic framework. This combination allows the system to simultaneously capture worker-specific behaviors and community-level patterns, achieving both individual accuracy and collective intelligence

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10762443B2Crowdsourcing system with community learning
Publication Date: 2020.09.01 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10762443B2 patent drawing
  • US10762443B2 patent drawing
  • US10762443B2 patent drawing

AI summary

Crowdsourcing systems with machine learning are described. Specifically, item-label inference methods and systems are presented, for example, to provide aggregated answers to a crowdsourced task in a manner achieving good accuracy even where observed data about past behavior of crowd members is sparse. In various examples, an item-label inference system infers variables describing characteristics of both individual crowd workers and communities of the workers. In various examples, an item-label inference system provides aggregated labels while considering the inferred worker characteristics and the inferred characteristics of the worker communities. In examples the item-label inference system provides uncertainty information associated with the inference results for selecting workers and generating future tasks.