Digital Annotation System for Training Automatic Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional electronic document management systems face inaccuracies, inefficiencies, and inflexibility in annotating key portions of documents due to errors in individual annotation models, lack of accurate training data, and rigid reliance on expert annotators, leading to high computational resource usage and scalability limitations.

Innovation Solution

A digital document annotation system that collects and analyzes annotation performance data to generate accurate digital annotations by presenting documents to annotators based on topic preferences, using question-answering protocols, and monitoring annotator interactions to identify reliable annotations, which are then used to train or test automatic annotation models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional electronic document management systems use sequential labeling techniques to identify key portions of documents, then the systems can automatically annotate documents, but the annotation accuracy deteriorates due to high subjectivity levels amongst reviewers

Engineering Contradiction:
Improveautomatic annotationVSAvoidannotation accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system collects annotation performance data from multiple annotators and uses this feedback to iteratively improve annotation accuracy. By analyzing patterns in annotator behavior and comparing annotations across multiple reviewers, the system resolves subjective differences and converges on accurate ground-truth labels, thereby maintaining high automation while improving precision.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system employs multiple annotators with diverse expertise to perform the same annotation task, allowing the system to leverage collective knowledge. This multi-functional approach enables the system to handle subjective annotation challenges by aggregating perspectives from different reviewers, ultimately producing more reliable annotations through consensus or expert validation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If conventional electronic document management systems require a significant amount of training data to generate annotation models, then the models can be trained comprehensively, but the training time and computational resources increase significantly

Engineering Contradiction:
Improvemodel training completenessVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by collecting and curating high-quality annotation performance data from multiple annotators before model training begins. This pre-processing step creates a refined training dataset that is more efficient for model learning, reducing the overall training time and computational resources needed while maintaining comprehensive model training.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes key parameters by transforming raw annotation data into optimized training datasets through filtering, validation, and curation processes. By adjusting data quality parameters and selecting the most informative annotations, the system achieves comprehensive model training with reduced dataset size, thereby decreasing training time and computational requirements.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If conventional electronic document management systems use unreliable training data to test or train annotation models, then the systems can proceed with training, but the convergence time and computational resources increase

Engineering Contradiction:
Improvetraining progressionVSAvoidcomputational resources
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The system performs preliminary validation and curation of training data before model training begins. By pre-screening annotation quality and selecting reliable training examples, the system ensures that training progresses efficiently without wasting computational resources on poor-quality data, thereby maintaining productivity while reducing energy consumption.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms that evaluate training data quality during the data collection and preparation phases. This feedback loop allows the system to identify and exclude unreliable annotations before they consume computational resources during training, ensuring efficient convergence while maintaining training progression.

Inventive Principle:
Principle #23Feedback

4Measurement precision

If conventional electronic document management systems rigidly require expert annotators to generate training data sets, then the annotation quality can be maintained, but the scalability is significantly limited and expenses increase

Engineering Contradiction:
Improveannotation qualityVSAvoidscalability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system employs multiple annotators with varying levels of expertise to perform annotation tasks, making the system more scalable and adaptable. By designing annotation guidelines and validation processes that work across different skill levels, the system maintains annotation quality while expanding its capacity to handle larger volumes of documents and diverse annotation needs.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system implements feedback loops where annotations from less experienced annotators are validated and refined through comparison with expert annotations or through iterative review processes. This feedback mechanism enables the system to scale up annotation capacity while maintaining quality standards, as non-expert annotators can contribute to large-scale annotation projects with their outputs being quality-controlled through systematic feedback.

Inventive Principle:
Principle #23Feedback

5Loss of information

If conventional electronic document management systems require users to navigate through multiple user interfaces to analyze annotations, then comprehensive analysis is possible, but the time and resources required increase significantly

Engineering Contradiction:
Improveannotation analysis completenessVSAvoidanalysis time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system merges multiple annotation analysis functions into a unified interface that consolidates annotation viewing, validation, and evaluation capabilities. By integrating these previously separate functions into a single coherent interface, the system enables comprehensive annotation analysis without requiring users to navigate through multiple separate interfaces, thereby reducing analysis time while maintaining completeness.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11232255B2Generating digital annotations for evaluating and training automatic electronic document annotation models
Publication Date: 2022.01.25 ADOBE INC
  • US11232255B2 patent drawing
  • US11232255B2 patent drawing
  • US11232255B2 patent drawing

AI summary

Systems, methods, and non-transitory computer-readable media are disclosed that collect and analyze annotation performance data to generate digital annotations for evaluating and training automatic electronic document annotation models. In particular, in one or more embodiments, the disclosed systems provide electronic documents to annotators based on annotator topic preferences. The disclosed systems then identify digital annotations and annotation performance data such as a time period spent by an annotator in generating digital annotations and annotator responses to digital annotation questions. Furthermore, in one or more embodiments, the disclosed systems utilize the identified digital annotations and the annotation performance data to generate a final set of reliable digital annotations. Additionally, in one or more embodiments, the disclosed systems provide the final set of digital annotations for utilization in training a machine learning model to generate annotations for electronic documents.