Image Captioning Using Spatial Relationship Transformer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current deep learning approaches for automatic digital content captioning fail to effectively extract important features from images, leading to inaccurate captions due to the neglect of spatial relationships between objects.

Innovation Solution

The use of an Object Relation Transformer (ORT) with an encoder-decoder architecture that incorporates object spatial relationship information, specifically relative position and size, to generate improved image captions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep learning approaches (CNN+RNN) are used for automatic image captioning, then the system can generate natural language captions, but the captions lack accuracy due to failure to extract important spatial relationship features

Engineering Contradiction:
Improvecaption accuracyVSAvoidspatial relationship information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent transitions from traditional 2D image processing to 3D spatial relationship modeling by adding depth information about object positions, distances, and spatial arrangements. The system extracts not just what objects are present but also their spatial relationships, transforming the feature extraction from surface-level to deep spatial understanding, thereby improving caption accuracy while preserving spatial information.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent segments the image processing into distinct modules: object detection, spatial relationship extraction, and caption generation. By separating the extraction of spatial features from the language generation process, the system can accurately capture spatial relationships and then effectively translate them into natural language descriptions, resolving the contradiction between generating captions and maintaining spatial information.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If spatial relationship information is extracted and incorporated into the captioning system, then caption quality improves, but system complexity increases

Engineering Contradiction:
Improvecaption qualityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges the object detection module and spatial relationship extraction module into an integrated processing pipeline. By combining these functions in a unified architecture where spatial features are extracted and represented alongside object features, the system achieves improved caption quality while managing complexity through functional integration rather than separate independent components.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent changes the parameter representation by transforming raw image data into structured spatial relationship parameters (positions, distances, orientations). This parameter transformation allows the system to handle spatial information in a standardized format that can be efficiently processed by the language generation model, improving quality while controlling complexity through parameter standardization.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12271814B2Automatic digital content captioning using spatial relationships method and apparatus
Publication Date: 2025.04.08 YAHOO ASSETS LLC
  • US12271814B2 patent drawing
  • US12271814B2 patent drawing
  • US12271814B2 patent drawing

AI summary

Disclosed are systems and methods for improving interactions with and between computers in content hosting and/or providing systems supported by or configured with personal computing devices, servers and/or platforms. The systems interact to identify and retrieve data within or across platforms, which can be used to improve the quality of data used in processing interactions between or among processors in such systems. The disclosed systems and methods provide systems and methods for automatically creating a caption comprising a sequence of words in connection with digital content.