Context-Aware Data Labeling for Machine Learning Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Inaccuracies in training datasets for machine learning models due to ambiguous data inputs lead to operational inefficiencies and reduced model accuracy, as errors in labeling and context clarification are necessary during the training process.

Innovation Solution

The provision of context sets with specific context objects to data labelers, which are evaluated for their impact on model accuracy and performance, allowing for the creation of optimal context sets that enhance data labeling efficiency and accuracy, thereby improving the training datasets and the resulting ML models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If data labelers are provided with extensive context information during labeling tasks, then labeling accuracy improves, but data transmission volume and processing complexity increase

Engineering Contradiction:
Improvelabeling accuracyVSAvoidcontext data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system dynamically adjusts context set parameters (size, composition, detail level) based on the specific data input being labeled. Rather than providing all possible context information uniformly, the system modifies context parameters to match the labeling requirements of each specific input, reducing unnecessary data transmission while maintaining accuracy where needed

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

Context information is divided into multiple hierarchical levels or categories. The system segments context data into essential vs. supplementary information, and selectively transmits only the necessary segments to data labelers based on the specific labeling task, reducing overall data volume while preserving critical context for accurate labeling

Inventive Principle:
Principle #1Segmentation

2Reliability

If context sets are customized and evaluated for each data labeler, then model accuracy improves, but system complexity and processing time increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

Context sets are pre-evaluated and pre-configured based on historical performance data and model requirements before actual labeling tasks begin. The system performs preliminary analysis to determine which context objects are most valuable for each labeler, so that during runtime, the system can simply retrieve and transmit pre-optimized context sets rather than performing complex real-time customization

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements a feedback loop where model accuracy results from different context set configurations are continuously monitored and fed back into the context selection process. This feedback mechanism allows the system to learn which context objects and configurations produce the best results, automatically refining the customization strategy without requiring manual intervention or increasing system complexity

Inventive Principle:
Principle #23Feedback

3Stability of the object's composition

If comprehensive context information is provided to all data labelers, then labeling consistency improves, but operational efficiency decreases

Engineering Contradiction:
Improvelabeling consistencyVSAvoidlabeling efficiency
Core Design Contradiction:
Stability of the object's compositionVSProductivity

Solution Approach 1:

Instead of providing uniform comprehensive context to all data labelers, the system applies local quality by tailoring context information to each labeler's specific needs, expertise level, and the particular data input being labeled. Each labeler receives a customized context set that contains precisely the right amount and type of information for their specific task, maintaining consistency through appropriate context rather than excessive context

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240232699A1Enhanced data labeling for machine learning training
Publication Date: 2024.07.11 CAPITAL ONE SERVICES LLC
  • US20240232699A1 patent drawing
  • US20240232699A1 patent drawing
  • US20240232699A1 patent drawing

AI summary

Embodiments relate to enhancing data labeling for machine learning model training using context data. The context data is provided to data labelers to improve accuracy of labels assigned by the data labelers, such as those assigned for ambiguous or unclear target objects. In an example, a system determines and provides context sets for target objects at user devices. The system generates a first training dataset and a second training dataset with labels obtained in connection with a first subset and a second subset of the context sets, respectively. The system trains a first and second instance of a machine learning model on the first training dataset and the second training dataset, respectively, and determines a respective accuracy score for the model instances. If the first instance is more accurate than the second instance, the system generates subsequent context sets based on characteristics of the first subset of context sets.