Image Captioning Augmented with Surrounding Text Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image captioning systems fail to provide contextualization of images with respect to the surrounding text, leading to a lack of enriched and enhanced captions.

Innovation Solution

The system generates an augmented caption by creating a knowledge graph that combines an image caption graph and a contextual graph, using natural language understanding to recalculate node and edge importance, incorporating emotion, entity, and tone analysis to enhance the caption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional image captioning systems are used, then image description accuracy is improved, but contextualization with surrounding text is lost

Engineering Contradiction:
Improveimage description accuracyVSAvoidcontextual information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent merges the image caption graph (containing object detection and relationship information) with the contextual graph (containing surrounding text information) into a unified knowledge graph. This combination allows the system to preserve both accurate image description and contextual information from surrounding text, resolving the contradiction between description accuracy and contextualization.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a knowledge graph as an intermediary structure that bridges image content and surrounding text. The knowledge graph serves as a mediator that integrates visual information from the image caption graph with textual context from the contextual graph, enabling both accurate description and contextualization to coexist.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If manual caption creation is used, then caption quality is improved, but productivity is reduced

Engineering Contradiction:
Improvecaption qualityVSAvoidcaption generation speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent implements an automated system that generates enhanced captions without manual intervention. The knowledge graph automatically integrates image content and surrounding text, and the natural language generation system autonomously produces contextualized captions, eliminating the need for manual caption creation while maintaining high quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the parameters of automated caption generation by incorporating contextual information from surrounding text into the knowledge graph. This parameter change transforms the output from simple object detection to enriched contextualized descriptions, achieving manual-quality captions through automated processes.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated caption generation is used, then productivity is improved, but contextual understanding is reduced

Engineering Contradiction:
Improvecaption generation speedVSAvoidcontextual understanding
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent performs preliminary action by constructing the knowledge graph before generating captions. The system pre-integrates image content and surrounding text into the knowledge graph structure, establishing contextual relationships in advance. This preliminary contextualization enables fast automated generation of contextually understanding captions.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical caption generation (simple object detection) with a knowledge graph-based system that incorporates natural language understanding. This substitution enables automated systems to achieve contextual understanding by processing semantic relationships in the knowledge graph rather than relying on basic image recognition alone.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10915572B2Image captioning augmented with understanding of the surrounding text
Publication Date: 2021.02.09 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10915572B2 patent drawing
  • US10915572B2 patent drawing
  • US10915572B2 patent drawing

AI summary

To augment an image caption, a caption graph containing entity nodes corresponding to entities contained in the image and relationship edges between entity nodes corresponding to relationships between entities as illustrated in the image is generated. In addition, a contextual graph containing one or more of entity nodes corresponding to entities contained in the image and described in text associated with the image, textual entity nodes corresponding to textual entities described in text associated with the image and textual relationship edges between entity node pairs, textual entity node pairs and entity node and textual entity node pairs is generated. The textual relationship edges correspond to relationships described in the text associated with the image between entity pairs, textual entity pairs or entity and textual entity pairs. From the contextual graph, an augmented caption graph containing entity nodes, relationship edges, textual entities and textual relationship edges is generated.