Neural Network Figure Captioning via Attention Mechanisms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional computer vision and caption generation technologies struggle to generate meaningful and accurate captions for electronic figures, which depict quantified data, as they fail to effectively describe the nuanced and comparative complexities of these figures.

Innovation Solution

A reasoning and sequence-level training approach is employed using a recurrent neural network with attention models to generate captions for electronic figures, optimizing resource consumption and improving accuracy by calculating weights for specific characteristics such as labels, visual aspects, and relationships between them.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional computer vision-based caption generation techniques are used, then the system is simple and easy to implement, but it fails to accurately decipher complex characteristics of electronic figures such as labels, relative values, relationships, and trends

Engineering Contradiction:
Improveaccuracy of deciphering figure characteristicsVSAvoidcomplexity of caption generation system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the electronic figure analysis into distinct components: visual element detection, text extraction, relationship identification, and caption generation. Each component processes specific aspects separately before integrating results, enabling accurate handling of complex figure characteristics while maintaining manageable system architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces intermediary processing layers between raw figure input and final caption output, including intermediate representations for visual elements, extracted text data, and identified relationships. These intermediaries bridge the gap between simple computer vision techniques and complex analytical requirements

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If conventional caption generation is used for digital images, then the captions are relatively concise, but they lack the analytical and thoughtful depth needed for electronic figures with multiple sets of quantified data

Engineering Contradiction:
Improvecompleteness of figure informationVSAvoidcomplexity of analysis process
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis actions before caption generation, including pre-processing steps to extract visual elements, detect text, and identify relationships. This preliminary structuring of information enables comprehensive analysis without overwhelming the final caption generation process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system adds analytical dimensions to the caption generation process by incorporating multiple layers of interpretation: visual structure analysis, text content extraction, relationship mapping, and trend identification. This multi-dimensional approach ensures complete information capture while organizing complexity into manageable layers

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If computer vision-based systems are used to identify characteristics of electronic figures, then the system operation is simple, but the system has difficulty deciphering complex characteristics such as labels, relative values, and relationships

Engineering Contradiction:
Improveaccuracy of characteristic identificationVSAvoidease of system operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system employs self-service mechanisms where the caption generation model automatically adapts to different figure types and complexities through sequence-level training. The system self-adjusts its analysis depth and methodology based on the input figure characteristics, maintaining ease of operation while achieving high precision in characteristic identification

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11461638B2Figure captioning system and related methods
Publication Date: 2022.10.04 ADOBE INC
  • US11461638B2 patent drawing
  • US11461638B2 patent drawing
  • US11461638B2 patent drawing

AI summary

Embodiments of the present invention are generally directed to generating figure captions for electronic figures, generating a training dataset to train a set of neural networks for generating figure captions, and training a set of neural networks employable to generate figure captions. A set of neural networks is trained with a training dataset having electronic figures and corresponding captions. Sequence-level training with reinforced learning techniques are employed to train the set of neural networks configured in an encoder-decoder with attention configuration. Provided with an electronic figure, the set of neural networks can encode the electronic figure based on various aspects detected from the electronic figure, resulting in the generation of associated label map(s), feature map(s), and relation map(s). The trained set of neural networks employs a set of attention mechanisms that facilitate the generation of accurate and meaningful figure captions corresponding to visible aspects of the electronic figure.