Image Text Description Generation via Semantic Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional multimodal large language models struggle with generating accurate text descriptions for images due to the presence of low-quality data in image-text datasets, leading to poor performance and inaccurate descriptions.

Innovation Solution

The approach involves semantically aligning a visual encoder with a text encoder, which is then used to train a conversion model and a language model. This alignment allows the visual encoder to convert image features into a space compatible with the language model, enabling the generation of accurate text descriptions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional multimodal large language models process image-text data directly, then the model can handle diverse input types, but low-quality data in the dataset leads to inaccurate text descriptions

Engineering Contradiction:
Improvemultimodal input capabilityVSAvoidtext description accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent extracts and removes low-quality data from the image-text dataset through data filtering and cleaning processes. Only high-quality data pairs are retained for training the conversion model and language model, thereby eliminating the harmful effect of low-quality data while preserving the multimodal processing capability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces a conversion model as an intermediary component between the visual encoder and the language model. This conversion model transforms visual features into a representation space compatible with the language model, enabling accurate text generation while filtering out low-quality data through the intermediary processing layer

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If the visual encoder and language model are trained independently, then the training process is simpler, but the semantic alignment between visual and text representations is insufficient

Engineering Contradiction:
Improvetraining process simplicityVSAvoidsemantic alignment accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent merges the training processes by using a shared conversion model that is trained on both visual features and text representations. This unified training approach simultaneously achieves semantic alignment between modalities while maintaining training simplicity through a single integrated model structure

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent changes the parameter space by transforming visual features through the conversion model into a representation space that is semantically aligned with the language model's text space. This parameter transformation enables direct semantic alignment while keeping the training process manageable through controlled feature space manipulation

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250124730A1Method and apparatus for generating text description for image, electronic device, and medium
Publication Date: 2025.04.17 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20250124730A1 patent drawing
  • US20250124730A1 patent drawing
  • US20250124730A1 patent drawing

AI summary

Embodiments of the present disclosure relate to a method and apparatus for generating a text description for an image, an electronic device, and a medium. The method includes generating a first feature of the image through a visual encoder, where a text encoder semantically aligned with the visual encoder is used to train a conversion model and a language model. The method further includes converting the first feature into a second feature through the conversion model, where the first feature and the second feature correspond to different feature spaces. In addition, the method also includes generating, by the language model, a text description for the image based on the second feature.