Adaptive Annotation Instructions for ML Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual annotation of training data for machine learning models is tedious, error-prone, and time-consuming, especially when using crowdsourced labels, due to low-quality instructions and the difficulty in handling edge cases, leading to inconsistent and biased results.

Innovation Solution

A machine learning-based adaptive instruction system that automatically identifies and selects data elements for annotation job instructions, including edge cases, to improve annotator performance by providing high-quality examples and soliciting feedback, thereby reducing the need for manual iteration and enhancing annotation quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation is used to obtain training data labels, then human annotators can provide labels for data, but the process becomes tedious, error-prone, and time-consuming

Engineering Contradiction:
Improveannotation qualityVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically generating annotation instructions, selecting representative data elements (including edge cases), and preparing training materials before annotators begin work. This upfront automation reduces the time annotators need to spend on manual label creation while maintaining quality standards.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary layer that automatically processes data elements and generates instructional content between the raw data and the annotators. This intermediary system selects representative examples and creates guidance materials, reducing the direct manual effort required while improving annotation consistency and quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If crowdsourced labels are used to speed up annotation, then more annotators can work in parallel, but the results become inconsistent and biased due to low-quality instructions

Engineering Contradiction:
Improveannotation throughputVSAvoidannotation consistency
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system applies local quality by selecting and presenting specific representative data elements and edge cases in the instructions to annotators. Rather than providing generic guidance, the system tailors the instructional content to highlight particular annotation challenges and examples that are most relevant to maintaining consistency across different annotators.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system incorporates feedback mechanisms where annotator responses and performance data are used to iteratively improve the instructional materials and data element selections. This feedback loop helps maintain consistency by identifying and addressing annotation patterns that deviate from quality standards.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If more data elements are provided in annotation instructions to cover edge cases, then annotator performance improves, but the instruction set becomes larger and more complex

Engineering Contradiction:
Improveannotation accuracyVSAvoidinstruction complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts and selects only the most representative and critical data elements for inclusion in instructions, rather than providing all possible examples. By carefully curating a focused set of representative cases including edge cases, the system maintains high annotation accuracy while keeping instruction sets manageable in size and complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies partial action by providing a selective subset of data elements that are most useful for annotation guidance. Rather than overwhelming annotators with exhaustive examples, the system identifies and presents the critical few examples that will have the greatest impact on annotation quality and consistency.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10977518B1Machine learning based adaptive instructions for annotation
Publication Date: 2021.04.13 AMAZON TECH INC
  • US10977518B1 patent drawing
  • US10977518B1 patent drawing
  • US10977518B1 patent drawing

AI summary

Techniques for generating and utilizing machine learning based adaptive instructions for annotation are described. An annotation service can use models to identify edge case data elements predicted to elicit differing annotations from annotators, “bad” data elements predicted to be difficult to annotate, and/or “good” data elements predicted to elicit matching or otherwise high-quality annotations from annotators. These sets of data elements can be automatically incorporated into annotation job instructions provided to annotators, resulting in improved overall annotation results via having efficiently and effectively “trained” the annotators how to perform the annotation task.