Distributed Data Categorization via Crowdsourced Pairwise Annotations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for categorizing large datasets are inefficient as they rely on a single person or machine, which is time and cost ineffective, and may not accurately identify categories due to differing annotation criteria among sources.

Innovation Solution

A distributed data categorization system that utilizes a combination of human and machine annotators to cluster and categorize data subsets, generating pairwise annotations and identifying categories based on metadata, with the ability to iteratively refine sub-categories and construct taxonomies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a single person or machine is used for categorizing large datasets, then the process is simple to manage, but it is time and cost ineffective and may not accurately identify categories

Engineering Contradiction:
Improvecategorization speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the large dataset into multiple subsets and distributes them to different human and machine annotators for parallel processing. This segmentation enables simultaneous categorization of multiple data portions, significantly increasing productivity while maintaining manageable task units for each annotator.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent combines the efforts of multiple human annotators and machine learning models into a unified categorization system. By merging diverse annotation results and reconciling differences through consensus algorithms, the system achieves both high productivity and improved accuracy while managing complexity through structured integration.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If multiple annotators are used to categorize data, then productivity increases, but annotation accuracy may decrease due to differing criteria among sources

Engineering Contradiction:
Improvecategorization throughputVSAvoidcategory identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where annotation results are continuously evaluated and used to refine categorization criteria. Discrepancies between annotators trigger review processes and feedback loops that adjust annotation guidelines, ensuring that productivity gains from multiple annotators do not compromise accuracy.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent dynamically adjusts annotation parameters and criteria based on performance data from multiple annotators. By changing thresholds, confidence levels, and categorization rules based on observed patterns, the system maintains high accuracy while leveraging the productivity of multiple sources.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If human intelligence is leveraged for categorization, then accuracy improves, but cost and time resources increase

Engineering Contradiction:
Improveannotation accuracyVSAvoidcategorization time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial human annotation combined with machine processing for routine cases. Human annotators focus on complex or ambiguous data portions requiring judgment, while machine learning models handle straightforward categorizations, optimizing the balance between accuracy and time efficiency.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent introduces machine learning models as intermediaries between raw data and final human annotation. These models pre-process data, suggest categorizations, and filter obvious cases, reducing the time burden on human annotators while maintaining high accuracy through human review of critical decisions.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Productivity

If machine learning is used for categorization, then speed increases, but accuracy may decrease due to inability to capture nuanced annotation criteria

Engineering Contradiction:
Improveprocessing speedVSAvoidnuanced category identification
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements a hybrid system where machine learning models continuously learn from human annotations and improve their own performance. The system serves itself by using human-annotated data to retrain and refine machine models, enabling the machine to capture nuanced criteria over time while maintaining high processing speed.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10157217B2Systems and methods for the distributed categorization of source data
Publication Date: 2018.12.18 CALIFORNIA INST OF TECH
  • US10157217B2 patent drawing
  • US10157217B2 patent drawing
  • US10157217B2 patent drawing

AI summary

Systems and methods for the crowdsourced clustering of data items in accordance embodiments of the invention are disclosed. In one embodiment of the invention, a method for determining categories for a set of source data includes obtaining a set of source data, determining a plurality of subsets of the source data, where a subset of the source data includes a plurality of pieces of source data in the set of source data, generating a set of pairwise annotations for the pieces of source data in each subset of source data, clustering the set of source data into related subsets of source data based on the sets of pairwise labels for each subset of source data, and identifying a category for each related subset of source data based on the clusterings of source data and the source data metadata for the pieces of source data in the group of source data.