Image-Text Model Training for Accurate Image Description

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image description methods, such as question-answer and picture-based forms, often result in inaccurate descriptions due to simplification or interference from irrelevant content, leading to discrepancies between generated text and original image content.

Innovation Solution

A pre-trained image-text model is used to generate target text by constructing an image loss based on a first and second sample image, ensuring consistency in the conversion process to enhance accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If question-answer form or picture-based description form is used to generate image description text, then the generation process is simple, but the accuracy of the description text deteriorates due to excessive simplification or interference from irrelevant content

Engineering Contradiction:
Improvesimplicity of generation processVSAvoidaccuracy of description text
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent introduces an image-text model as an intermediary between the input image and the output description text. This model is pre-trained using a combination of image-to-text and text-to-image transformation losses, enabling it to accurately capture and translate image content into text while filtering out irrelevant information, thus resolving the contradiction between simple generation process and accurate description text

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies preliminary action by pre-training the image-text model using a specially constructed loss function that incorporates both image-to-text and text-to-image transformation losses. This pre-training process prepares the model to accurately understand and generate image descriptions before actual use, improving description accuracy without complicating the generation process during deployment

Inventive Principle:
Principle #10Preliminary action

2Productivity

If image-to-text transformation is performed directly without text-to-image conversion, then the process is efficient, but information loss occurs leading to reduced consistency

Engineering Contradiction:
Improveefficiency of transformation processVSAvoidinformation loss in conversion
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent implements feedback by incorporating text-to-image transformation loss into the pre-training process. The model generates text from images, then reconstructs images from the generated text, and compares the reconstructed images with originals. This feedback loop ensures the text accurately represents the original image content, minimizing information loss while maintaining transformation efficiency

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent adds another dimension to the transformation process by incorporating the reverse text-to-image transformation path. Instead of only transforming images to text in one direction, the model operates in both directions (image→text and text→image), creating a bidirectional transformation space that preserves information while maintaining efficiency through shared model parameters

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20260004475A1Image processing method and apparatus, device, medium, and program product
Publication Date: 2026.01.01 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US20260004475A1 patent drawing
  • US20260004475A1 patent drawing
  • US20260004475A1 patent drawing

AI summary

Embodiments of this application disclose an image processing method and apparatus, a device, and a medium. The method includes: obtaining a target image to be processed; inputting the target image to a pre-trained image-text model, a model loss of the image-text model including an image loss, and the image loss being constructed according to a first sample image and a second sample image that is obtained by converting a first sample text configured for describing the first sample image; and obtaining a target text configured for describing the target image and generated by the image-text model. In technical solutions of the embodiments of this application, the generated target text can describe the target image as accurately as possible, thereby ensuring accuracy of the target text.