Image-Text Model Training for Accurate Image Description
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image description methods, such as question-answer and picture-based forms, often result in inaccurate descriptions due to simplification or interference from irrelevant content, leading to discrepancies between generated text and original image content.
Innovation Solution
A pre-trained image-text model is used to generate target text by constructing an image loss based on a first and second sample image, ensuring consistency in the conversion process to enhance accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If question-answer form or picture-based description form is used to generate image description text, then the generation process is simple, but the accuracy of the description text deteriorates due to excessive simplification or interference from irrelevant content
Solution Approach 1:
The patent introduces an image-text model as an intermediary between the input image and the output description text. This model is pre-trained using a combination of image-to-text and text-to-image transformation losses, enabling it to accurately capture and translate image content into text while filtering out irrelevant information, thus resolving the contradiction between simple generation process and accurate description text
Solution Approach 2:
The patent applies preliminary action by pre-training the image-text model using a specially constructed loss function that incorporates both image-to-text and text-to-image transformation losses. This pre-training process prepares the model to accurately understand and generate image descriptions before actual use, improving description accuracy without complicating the generation process during deployment
2Productivity
If image-to-text transformation is performed directly without text-to-image conversion, then the process is efficient, but information loss occurs leading to reduced consistency
Solution Approach 1:
The patent implements feedback by incorporating text-to-image transformation loss into the pre-training process. The model generates text from images, then reconstructs images from the generated text, and compares the reconstructed images with originals. This feedback loop ensures the text accurately represents the original image content, minimizing information loss while maintaining transformation efficiency
Solution Approach 2:
The patent adds another dimension to the transformation process by incorporating the reverse text-to-image transformation path. Instead of only transforming images to text in one direction, the model operates in both directions (image→text and text→image), creating a bidirectional transformation space that preserves information while maintaining efficiency through shared model parameters
Data Source
AI summary
Embodiments of this application disclose an image processing method and apparatus, a device, and a medium. The method includes: obtaining a target image to be processed; inputting the target image to a pre-trained image-text model, a model loss of the image-text model including an image loss, and the image loss being constructed according to a first sample image and a second sample image that is obtained by converting a first sample text configured for describing the first sample image; and obtaining a target text configured for describing the target image and generated by the image-text model. In technical solutions of the embodiments of this application, the generated target text can describe the target image as accurately as possible, thereby ensuring accuracy of the target text.


