Cross-Modal Image-Audio Feature Mapping via Shared Latent Space

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image recognition technologies face challenges in accurately associating visual information with linguistic information, particularly in identifying audio segments that describe objects in images across different languages, due to variations in duration and word count.

Innovation Solution

A learning device that trains image and audio encoders to map images and speeches into a shared latent space using a neural network with a self-attention mechanism, ensuring that image features are similar to corresponding audio features, thereby enabling accurate association of visual and linguistic information across multiple languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional image recognition technologies are used to associate visual information with linguistic information, then the system can process images and audio captions, but the association accuracy deteriorates due to difficulties in identifying audio segments across different languages

Engineering Contradiction:
Improveassociation accuracyVSAvoidaudio segment identification difficulty
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent merges image encoding and audio encoding into a unified framework where both modalities are processed by encoders that map to a common latent space. The image encoder and audio encoder are trained jointly using paired data, allowing the system to capture cross-modal relationships and improve association accuracy between visual and linguistic information.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal encoding framework that handles multiple languages and modalities through a single shared latent space. The encoders are designed to process diverse input (images and audio captions in different languages) and transform them into a common representation space, enabling the system to generalize across languages and improve detection of audio segments regardless of language.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If the system processes audio captions in multiple languages with varying durations and word counts, then the system becomes more versatile, but the feature mapping consistency deteriorates

Engineering Contradiction:
Improvemulti-language capabilityVSAvoidfeature mapping consistency
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent employs parameter changes by dynamically adjusting encoder outputs based on input characteristics such as audio caption length and language. The encoders process audio captions with varying durations and word counts by adapting their internal parameters, while the shared latent space ensures that despite these variations, the feature mappings remain consistent and comparable across different languages and modalities.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11817081B2Learning device, learning method, learning program, retrieval device, retrieval method, and retrieval program
Publication Date: 2023.11.14 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11817081B2 patent drawing
  • US11817081B2 patent drawing
  • US11817081B2 patent drawing

AI summary

A learning device calculates an image feature using a model (image encoder) that receives an image and outputs the image feature obtained by mapping the image into a latent space. The learning device calculates an audio feature using a model (audio encoder) that receives a speech in a predetermined language and outputs the audio feature obtained by mapping the speech into the latent space, and that includes a neural network provided with a self-attention mechanism. The learning device updates parameters of the models used by an image feature calculation unit and an audio feature calculation unit such that the image feature of a first image is similar to the audio feature of a speech corresponding to the first image.