Consensus Label Set Generation for Machine Learning Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The effectiveness of machine learning algorithm training is limited by the quality of data classification, as different individuals may introduce biases and inconsistencies in labeling data points, leading to varying label sets that do not effectively coordinate multiple perspectives.

Innovation Solution

A computer system generates a consensus label set by retrieving multiple label sets, creating a graph with edges between labels, determining weights based on data point consistency, grouping labels, identifying conflicts, and transmitting issues for further classification to entities, ultimately producing a new label set that reduces bias and improves data classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple individuals label data points independently, then diverse perspectives are captured, but biases and inconsistencies are introduced

Engineering Contradiction:
Improvemultiple perspectivesVSAvoidconsistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent merges multiple label sets from different individuals into a unified consensus label set. The system combines the diverse perspectives captured by multiple labelers with the consistency requirements through graph-based integration, where edges represent relationships between labels and weights reflect agreement levels, ultimately producing a harmonized label set that preserves diversity while ensuring reliability

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements feedback mechanisms where the system identifies conflicts between labels and returns them to individuals for resolution. The graph structure enables the system to detect inconsistencies and iterate with labelers to reach consensus, ensuring that the final label set reflects both diverse perspectives and reliable classification

Inventive Principle:
Principle #23Feedback

2Ease of manufacture

If label sets are created without coordination, then labeling process is simple, but conflicts and inconsistencies arise

Engineering Contradiction:
Improvelabeling process simplicityVSAvoidclassification accuracy
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent introduces a coordination system as an intermediary between independent labelers and the final label set. This mediator uses graph structures to represent label relationships and automatically resolves conflicts, maintaining the simplicity of independent labeling while ensuring high classification accuracy through coordinated consensus

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If a single label set is used, then consistency is maintained, but biases from single perspective are introduced

Engineering Contradiction:
ImproveconsistencyVSAvoidperspective diversity
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent merges multiple label sets from different individuals into a unified consensus label set. The system combines the diverse perspectives captured by multiple labelers with the consistency requirements through graph-based integration, where edges represent relationships between labels and weights reflect agreement levels, ultimately producing a harmonized label set that preserves diversity while ensuring reliability

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10902352B2Labeling of data for machine learning
Publication Date: 2021.01.26 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10902352B2 patent drawing
  • US10902352B2 patent drawing
  • US10902352B2 patent drawing

AI summary

A computer generates labels for machine learning algorithms by retrieving, from a data storage circuit, multiple label sets that contain labels that each classify data points in a corpus of data. A graph is generated that includes a plurality of edges, each edge between two respective labels from different label sets of the multiple label sets. Weights are determined for the plurality of edges based upon a consistency between data points classified by two labels connected by the edges. An algorithm is applied that groups labels from the multiple label sets based upon the weights for the plurality of edges. Data points are identified from the corpus of data that represent conflicts within the grouped labels. An electronic message is transmitted in order to present the identified data points to entities for further classification. A new label set is generated using the further classification received from the entities.