Consensus Label Set Generation for Machine Learning Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The effectiveness of machine learning algorithm training is limited by the quality of data classification, as different individuals may introduce biases and inconsistencies in labeling data points, leading to varying label sets that do not effectively coordinate multiple perspectives.
Innovation Solution
A computer system generates a consensus label set by retrieving multiple label sets, creating a graph with edges between labels, determining weights based on data point consistency, grouping labels, identifying conflicts, and transmitting issues for further classification to entities, ultimately producing a new label set that reduces bias and improves data classification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple individuals label data points independently, then diverse perspectives are captured, but biases and inconsistencies are introduced
Solution Approach 1:
The patent merges multiple label sets from different individuals into a unified consensus label set. The system combines the diverse perspectives captured by multiple labelers with the consistency requirements through graph-based integration, where edges represent relationships between labels and weights reflect agreement levels, ultimately producing a harmonized label set that preserves diversity while ensuring reliability
Solution Approach 2:
The patent implements feedback mechanisms where the system identifies conflicts between labels and returns them to individuals for resolution. The graph structure enables the system to detect inconsistencies and iterate with labelers to reach consensus, ensuring that the final label set reflects both diverse perspectives and reliable classification
2Ease of manufacture
If label sets are created without coordination, then labeling process is simple, but conflicts and inconsistencies arise
Solution Approach 1:
The patent introduces a coordination system as an intermediary between independent labelers and the final label set. This mediator uses graph structures to represent label relationships and automatically resolves conflicts, maintaining the simplicity of independent labeling while ensuring high classification accuracy through coordinated consensus
3Reliability
If a single label set is used, then consistency is maintained, but biases from single perspective are introduced
Solution Approach 1:
The patent merges multiple label sets from different individuals into a unified consensus label set. The system combines the diverse perspectives captured by multiple labelers with the consistency requirements through graph-based integration, where edges represent relationships between labels and weights reflect agreement levels, ultimately producing a harmonized label set that preserves diversity while ensuring reliability
Data Source
AI summary
A computer generates labels for machine learning algorithms by retrieving, from a data storage circuit, multiple label sets that contain labels that each classify data points in a corpus of data. A graph is generated that includes a plurality of edges, each edge between two respective labels from different label sets of the multiple label sets. Weights are determined for the plurality of edges based upon a consistency between data points classified by two labels connected by the edges. An algorithm is applied that groups labels from the multiple label sets based upon the weights for the plurality of edges. Data points are identified from the corpus of data that represent conflicts within the grouped labels. An electronic message is transmitted in order to present the identified data points to entities for further classification. A new label set is generated using the further classification received from the entities.


