AI Training Data Labeling Interface

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing artificial intelligence systems require large volumes of labeled training data, which can introduce classification bias and are inefficient due to slow, error-prone, and inconsistent manual labeling processes by subject-matter experts.

Innovation Solution

A user interface is developed that allows subject-matter experts to quickly and accurately label training data by displaying a subset of training data items alongside relevant labels, enabling selection and storage for use in training AI systems, with features like label hierarchies and filtering to enhance efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling is performed by subject-matter experts viewing each training data item individually, then labeling accuracy can be maintained, but labeling speed and productivity are severely reduced

Engineering Contradiction:
Improvelabeling accuracyVSAvoidlabeling speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary actions by automatically generating candidate labels and pre-sorting training data items before presentation to the subject-matter expert. This reduces the cognitive load and time required for each labeling decision while maintaining accuracy through pre-computed suggestions.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the purely mechanical manual review process with an automated system that uses machine learning models to generate candidate labels and sort data items. This substitution accelerates the labeling process while human experts provide final validation to maintain accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If standard pre-labeled data sets are used for training, then training efficiency is improved, but classification bias is introduced

Engineering Contradiction:
Improvetraining efficiencyVSAvoidclassification bias
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system applies local quality by using domain-specific labeling expertise for particular problem areas while using automated candidate generation for other areas. This allows the system to maintain high reliability in critical classification areas while preserving overall training efficiency.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces asymmetry in the labeling approach by treating different data items and label types differently based on their importance and complexity. Critical labels receive more rigorous expert review while less critical ones use automated suggestions, reducing bias in important areas while maintaining efficiency overall.

Inventive Principle:
Principle #4Asymmetry

3Quantity of substance

If existing manual labeling techniques are used for new training data, then labeling completeness can be achieved, but consistency and accuracy deteriorate due to human error

Engineering Contradiction:
Improvelabeling completenessVSAvoidlabeling consistency
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The system implements feedback mechanisms where subject-matter experts review and correct automated candidate labels, and their corrections are fed back to improve the automated labeling system. This continuous feedback loop improves consistency and accuracy over time while maintaining labeling completeness.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary automated labeling system that sits between the training data and the subject-matter expert. This intermediary generates candidate labels that reduce human error while experts provide final oversight to ensure consistency and correct any systematic biases.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240046155A1Interface for artificial intelligence training
Publication Date: 2024.02.08 HRB INNOVATIONS
  • US20240046155A1 patent drawing
  • US20240046155A1 patent drawing
  • US20240046155A1 patent drawing

AI summary

Media and method for a user interface for training an artificial intelligence system. Many artificial intelligence systems require large volumes of labeled training data before they can accurately classify previously unseen data items. However, for some problem domains, no pre-labeled training data set may be available. Manually labeling training data sets by a subject-matter expert is a laborious process. An interface to enable such a subject-matter expert to accurately, consistently, and quickly label training data sets is disclosed herein. By allowing the subject-matter expert to easily navigate between training data items and select the applicable labels, operation of the computer is improved.