Image description method based on symmetric diffusion model and visual regeneration

By introducing symmetric diffusion models and visual regeneration methods in image description tasks, combined with visual regeneration loss function, the semantic alignment problem existing in image description generation of existing diffusion models is solved, and a higher quality image description generation is achieved, which transcends the traditional autoregression method.

CN120148033APending Publication Date: 2025-06-13NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510190500.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

There are significant gaps in the generation quality of existing image description methods based on diffusion models, especially due to the lack of consideration of temporal and spatial historical information, resulting in problems such as omissions, duplications and object illusions in the generated descriptions.

Method used

A method of image description based on symmetric diffusion model and visual regeneration is proposed. Through two diffusion models used for image-to-text and text-to-image generation, combined with a pre-trained image encoder and text encoder, the visual regeneration loss function is used to maximize the visual semantic consistency between the original image and the regenerated image, ensuring semantic alignment between the input image and the generated text description.

Benefits of technology

Effectively improve the image description generation performance based on diffusion model, significantly better than previous diffusion-based methods, and surpass most existing autoregressive methods, experiments on MS COCO data sets show that MirrorDiff achieves better performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148033A_ABST
    Figure CN120148033A_ABST
Patent Text Reader

Abstract

The invention discloses an image description method based on symmetric diffusion models and visual regeneration, comprising two diffusion models respectively used for image-to-text generation and text-to-image generation, and respectively obtaining image representation and text representation by using a pre-trained image encoder and a pre-trained text encoder, inputting the image representation and the noisy text representation into a de-noising device to generate a text description; and finally, designing a loss function named as visual regeneration loss, wherein the loss function can ensure semantic alignment between the input image and the generated text description by maximizing visual semantic consistency between the regenerated image and the original input image. Different from most existing image description methods, the method can further evaluate and optimize the generated sentences through the visual similarity between the input image and the regeneration image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to an image description method based on a symmetric diffusion model and visual regeneration. Background Art

[0002] Image Captioning is an important cross-task in the fields of computer vision and natural language processing, and its goal is to generate a natural language description that conforms to semantics based on the input image. This task has a wide range of applications, including visual assistive devices, content retrieval, social media automation, and robot navigation. In recent years, with the development of deep learning technology, this task has made significant progress. According to the way of generating natural language descriptions, existing image description methods can be divided into two categories: autoregressive and non-autoregressive.

[0003] Autoregressive image description generates words in sequence, and at the same time refers to the semantic features of the image and the words that have been generated to predict the next word until the last word or the end character is generated and then stops. Early autoregressive methods mainly used convolutional neural networks (CNNs) and recurrent neural networks (RNNs) as their backbone networks. The CNN is responsible for extracting high-level semantic representations from the image, and the RNN is responsible for decoding these representations and generating descriptions in an autoregressive manner. During this period, many studies tried to improve the generation results by introducing new designs in the CNN-RNN-based architecture, including means such as incorporating and improving various attention mechanisms, modeling visual objects, and improving image semantic representation methods.

[0004] Non-autoregressive image description methods generate the entire text description at once, rather than generating word by word, and use parallel processing to significantly speed up the inference speed. Early non-autoregressive models were mainly some variants based on the Transformer with multi-head attention. In recent years, inspired by the excellent performance of diffusion models in text-to-image generation, diffusion-based image description has emerged. Compared with traditional autoregressive methods, it has advantages such as faster inference speed, lower cumulative error, and increased diversity of generation results. Chen et al. proposed a new diffusion paradigm that converts discrete text into binary bits and applied it to image description generation and achieved performance improvement. Luo et al. proposed SCDNet, which added cross-modal retrieval on the basis of the theory of Chen et al. aiming to improve the generation quality. However, due to the lack of consideration of temporal and spatial historical information in non-autoregressive methods (such as the words that have been generated and the hidden states of multi-modal representations), there is a semantic inconsistency between the given image and the generated description. This leads to problems such as content omission, repetition, and object hallucination in the generated description, and ultimately results in a consistent quality difference between diffusion-based methods and autoregressive methods.

[0005] The inherent limitations of autoregressive methods lie in slow inference speed, large cumulative errors, and low generation diversity. Diffusion models can address the inherent limitations of autoregressive models, but there is a significant gap in description quality compared to autoregressive methods. Previous diffusion-based methods have obvious problems such as word omission, repetition, and object hallucination, which are the direct reasons for poor generation quality. This has prompted researchers to consider whether there are inherent limitations in diffusion models for image captioning tasks that prevent them from matching the performance of autoregressive models. LaDiC replaces the text itself with high-level text representations and proposes a diffusion-based image captioning method, further enhancing the capabilities of diffusion models in image-to-text generation. However, compared to autoregressive methods, due to the complex transformation from the visual space to the discrete text space, the noise addition and denoising processes, and the lack of necessary semantic alignment guidance during this process, the semantic gap between visual content and text descriptions still exists. Summary of the Invention

[0006] To overcome the deficiencies of the prior art, the present invention provides an image captioning method based on a symmetric diffusion model and visual regeneration, which includes two diffusion models respectively for image-to-text and text-to-image generation. The image representation and text representation are obtained by using a pre-trained image encoder and text encoder respectively, and then the image representation and the noisy text representation are input into a denoiser to generate a text description; finally, a loss function called visual regeneration loss is designed, which can ensure the semantic alignment between the input image and the generated text description by maximizing the visual semantic consistency between the regenerated image and the original input image. Different from most existing image captioning methods, the present invention can further evaluate and optimize the generated sentences through the visual similarity between the input image and the regenerated image.

[0007] The technical solutions adopted by the present invention to solve its technical problems are as follows:

[0008] Step 1: Image and text encoding;

[0009] Use BLIP and BERT to encode the image and text respectively to obtain the visual and text latent representations;

[0010] Step 2: Add noise;

[0011] Step 3: Remove noise;

[0012] Use Transformer to remove the noise in the noisy text latent representation, and finally obtain the noise-free text latent representation;

[0013] Step 4: Text decoding;

[0014] Decode the latent representation of the denoised text obtained by BERT to obtain the generated text description.

[0015] Step 5: Image regeneration;

[0016] Use ADD XL to input the generated text description into ADD XL to generate a new image;

[0017] Step 6: Calculate the visual regeneration loss function;

[0018] The generated new image and the original image are respectively encoded using CLIP to obtain the CLIP visual representation, calculate the cosine similarity of the two visual representations, calculate the gradient using the visual regeneration loss function, and update the model parameters using the backpropagation algorithm.

[0019] Preferably, step 1 is specifically:

[0020] Given an image i ∈ R H×W×C and the true label describing the image where H, W, and C are the height, width, and number of channels of the input image, and d c is the length of the description label; through the visual encoder E v encode the image i into the visual representation v i , and through the text encoder E T and the regularization module encode the label c into the text representation x 0 ; the whole process is expressed as:

[0021] v i = E V (i)

[0022]

[0023] Preferably, step 2 is specifically:

[0024] Gradually add Gaussian noise to the latent representation of the text and perform a forward diffusion process on x 0 to generate a series of noisy text representations x 1 , x 2 ,..., x T ;

[0025] Preferably, step 3 is specifically:

[0026] The visual representation v i and x T are input into the denoiser f θ for reverse denoising; the denoiser outputs the predicted text latent representation This process is expressed as:

[0027]

[0028] where \(t\) is the current time step.

[0029] Preferably, step 4 is specifically as follows:

[0030] Input into the text decoder \(D\) T and the language model head to generate a predicted description The formula is as follows:

[0031]

[0032] Preferably, step 5 is specifically as follows:

[0033] Input the output by the image-to-text generation module into ADD XL to generate an image

[0034] Preferably, step 6 is specifically as follows:

[0035] Represent the original image \(i\) and the regenerated image respectively through the CLIP image encoder to obtain \(m\) i and Define the loss function as:

[0036] \(m\) i \(=\) CLIP(\(i\)),

[0037]

[0038]

[0039] Calculate the loss function and the gradient, and update the model parameters.

[0040] The beneficial effects of the present invention are as follows:

[0041] The present invention proposes a new image captioning framework, namely MirrorDiff, which unifies image-to-text and text-to-image generation. This framework specifically addresses the semantic alignment problem between the input image and the generated caption, effectively improving the performance of image captioning generation based on the diffusion model. In addition, the present invention proposes a new loss function called visual regeneration loss, which realizes the semantic alignment between the input image and the generated text by maximizing the CLIP semantic similarity between the original image and the image regenerated according to the caption. Experiments on the MS COCO dataset show that MirrorDiff significantly outperforms previous diffusion-based methods and surpasses most existing autoregressive methods. Brief Description of the Drawings

[0042] Figure 1 is the structural diagram of the present invention. Detailed Embodiment

[0043] The present invention will be further described below in conjunction with the drawings and embodiments.

[0044] The object of the present invention is to propose a novel image description generation framework based on symmetric diffusion, providing a solution that can generate higher-quality text descriptions for the model training and inference of image description methods based on diffusion models.

[0045] Considering that image description and its inverse process text-to-image generation have the same goal, namely aligning image content with text semantics, the present invention proposes a new image description framework based on symmetric diffusion models, named MirrorDiff. Specifically, MirrorDiff is a framework with a symmetric structure that first generates text from an image and then generates an image from the text, including two diffusion models for image-to-text and text-to-image generation respectively. It uses pre-trained image encoders and text encoders to obtain image representations and text representations respectively, and then inputs the image representation and the noisy text representation into a denoiser to generate text descriptions. To align the text description with the input image semantically, MirrorDiff uses a diffusion-based visual regenerator to regenerate a new image conditioned on the generated text description, aiming to make the semantics of the new image have as high a similarity as possible with the original input image. Therefore, the present invention designs a loss function named visual regeneration loss, which can ensure the semantic alignment between the input image and the generated text description by maximizing the visual semantic consistency between the regenerated image and the original input image. Different from most existing image description methods, MirrorDiff can further evaluate and optimize the generated sentences through the visual similarity between the input image and the regenerated image. In addition, the proposed MirrorDiff is a plug-and-play framework that can be inserted into many previous image description methods to further improve performance.

[0046] The technical solutions adopted by the present invention to solve its technical problems are as follows:

[0047] Step 1: Image and text encoding;

[0048] Use BLIP and BERT to encode the image and text respectively to obtain visual and text latent representations. Specifically, given an image i ∈ R H×W×C and the true label describing the image where H and W are the height and width of the input image respectively, and d c is the length of the description label. Through the visual encoder E VEncode the image i into a visual representation v i , and encode the label c into a text representation x T through the text encoder E and the regularization module 0 . The whole process can be expressed as:

[0049] v i = E V (i)

[0050]

[0051] Step 2: Add noise;

[0052] Gradually add Gaussian noise to the latent representation of the text, and perform a forward diffusion process on x 0 to generate a series of noisy text representations x 1 , x 2 ,..., x T ;

[0053] Step 3: Remove noise (generation);

[0054] Use the Transformer to remove the noise in the latent representation of the noisy text, and finally obtain the latent representation of the text without noise. Specifically, the visual representation v i and x T are input into the denoiser f θ for reverse denoising. The denoiser outputs the predicted latent representation of the text This process can be expressed as:

[0055]

[0056] Step 4: Text decoding;

[0057] Use BERT to decode the latent representation of the text without noise obtained after denoising to obtain the generated text description. Specifically, input the text decoder D T and the language model head to generate the predicted description The formula is as follows:

[0058]

[0059] Step 5: Image regeneration;

[0060] Use ADD XL to input the generated text description into ADD XL to generate a new image. Specifically, input the output by the image-to-text generation module into ADD XL to generate the image

[0061] Step 6: Calculate the generation loss function visually;

[0062] Encode the generated new image and the original image using CLIP respectively to obtain the CLIP visual representation, calculate the cosine similarity of the two visual representations, calculate the gradient using the visual regeneration loss function, and update the model parameters using the backpropagation algorithm. Specifically, first, the original image i and the regenerated image are respectively represented through the CLIP image encoder to obtain m i and Next, define the loss function as:

[0063] m i = CLIP(i)

[0064]

[0065] Calculate the loss function and the gradient, and update the model parameters.

[0066] Example:

[0067] The present invention proposes a new image description method based on the symmetric diffusion model, called MirrorDiff. MirrorDiff can be divided into two main modules: the image-to-text generation module and the text-to-image generation module. These two modules cooperate with each other during the training and inference processes to achieve better image description generation.

[0068] In the image-to-text generation module, LaDiC achieves a good diffusion-based image description generation baseline, so a similar pipeline is used to complete the image-to-text generation. The difference is that in the present invention, the frozen text decoder in LaDiC is set to a trainable state to expand the influence range of the designed visual regeneration loss. Specifically, given an image i ∈ R H×W×C and the true label describing the image where H and W are the height and width of the input image respectively, and d c is the length of the description label. Encode the image i into the visual representation v v through the visual encoder E i , and encode the label c into the text representation x T through the text encoder E and the regularization module 0 . The whole process can be expressed as:

[0069] v i = E v (i)

[0070]

[0071] Next, for x o Perform a forward diffusion process to generate a series of noisy text representations x 1 , x 2 ,..., x T . The visual representation v i and x T are input into the denoiser f θ for reverse denoising. The denoiser outputs the predicted text latent representation This process can be expressed as:

[0072]

[0073] Finally, the input text decoder D T and the language model head are used to generate the predicted description The formula is as follows:

[0074]

[0075] A pre-trained visual reconstructor AdversarialDiffusionDistillation (ADD) XL model is selected for the text-to-image generation module. ADD XL uses score distillation and combines adversarial loss to ensure high-fidelity images in the low-step state. Specifically, the present invention inputs the output by the image-to-text generation module into ADD XL to generate an image

[0076] The present invention proposes a visual regeneration loss to correctly guide the underlying semantic alignment between the input image and the generated description, and calculates the similarity between the new image generated from the description and the original image, which is a new dimension. Previous work directly compared the similarity between the description and the original image, but there was a problem: there are often many descriptions that can correctly describe the content in an image. If one wants to obtain a description that can completely and correctly describe the image content, directly measuring the semantic similarity between the description and the image often makes it difficult to achieve correct guidance.

[0077] The present invention can solve this problem through image regeneration, and the visual regeneration loss is very sensitive to object hallucinations. Those objects that appear in the regenerated image but not in the original image can be easily detected, thus effectively alleviating this problem. Specifically, first, the original image i and the regenerated image are respectively represented by the CLIP image encoder to obtain m i and Next, the loss function is defined as:

[0078] m i =CLIP(i)

[0079]

[0080] Finally, calculate the loss function and the gradient, and optimize all trainable parameters in the model until the loss function converges.

[0081] To verify the quantization performance of this patent in the image captioning task, the present invention conducted experiments on Microsoft COCO. Microsoft COCO (MS COCO) is one of the most classical datasets in the field of image captioning, containing 82,783 training images, 40,504 validation images, and 40,775 test images. Each image is annotated with 5 descriptions.

[0082] Table 1 summarizes the performance of MirrorDiff on the MS COCO data, including two groups of baselines (autoregressive and non-autoregressive). The blank parts in Table 1 indicate that the data is not given in the corresponding literature. The horizontal line in the middle differentiates autoregressive and non-autoregressive methods. The bold numbers represent the best results among all non-autoregressive methods.

[0083] Table 1 Comparative experiments on image captioning in the MS COCO dataset

[0084]

[0085]

[0086] Generally speaking, there is still a certain performance gap between the current state-of-the-art autoregressive methods and non-autoregressive methods. However, MirrorDiff has surpassed some autoregressive methods trained on datasets much larger than MS COCO and achieved results comparable to the state-of-the-art autoregressive methods. From the internal comparison of non-autoregressive methods, diffusion-based methods have led traditional non-autoregressive methods, especially in terms of the CIDEr score, leading by more than 10%. From the internal comparison of diffusion-based methods, MirrorDiff has achieved better performance compared to the state-of-the-art methods.

Claims

1. An image description method based on a symmetric diffusion model and visual reproduction, characterized in that: The steps include: Step 1: Image and text encoding; Use BLIP and BERT to encode images and texts respectively to obtain latent representations of vision and text; Step 2: Add noise; Step 3: Remove noise; Use Transformer to remove noise from the noisy text latent representation, and finally obtain the noise-free text latent representation; Step 4: Text decoding; Use BERT to decode the noise-free text latent representation obtained after denoising to obtain the generated text description; Step 5: Image regeneration; Using ADD XL, input the generated text description into ADD XL to generate a new image; Step 6: Visualization generates loss function calculation; The generated new image and the original image are encoded using CLIP respectively to obtain the CLIP visual representation, the cosine similarity of the two visual representations is calculated, the gradient is calculated using the visual reproduction loss function, and the model parameters are updated using the back-propagation algorithm.

2. The image description method based on symmetric diffusion model and visual reproduction according to claim 1, characterized in that: The step 1 is specifically as follows: Given an image i∈R H×W×C and the true label describing the image Where H, W, and C are the height, width, and number of channels of the input image, respectively. c is the length of the description label; through the visual encoder E V Encode image i into a visual representation v i , and through the text encoder E T and regularization module Encode the label c into a text representation x0; the whole process is expressed as: v i =E V (i) 3. The image description method based on symmetric diffusion model and visual reproduction according to claim 2, characterized in that: The step 2 is specifically as follows: Gaussian noise is gradually added to the latent representation of the text, and a forward diffusion process is performed on x0 to generate a series of noisy text representations x1, x2, ..., x T .

4. The image description method based on symmetric diffusion model and visual reproduction according to claim 3, characterized in that: The step 3 is specifically as follows: Visual Representation i and x T is input to the denoiser f θ In the process, reverse denoising is performed; Textual latent representation of denoiser output prediction This process is expressed as: where t is the current time step.

5. The image description method based on symmetric diffusion model and visual reproduction according to claim 4, characterized in that: The step 4 is specifically as follows: Will Input text decoder D T and language model header to generate predicted descriptions The formula is as follows:

6. The image description method based on symmetric diffusion model and visual reproduction according to claim 5, characterized in that: The step 5 is specifically as follows: The output of the image-to-text generation module Type ADD XL to generate the image 7. The image description method based on symmetric diffusion model and visual reproduction according to claim 6, characterized in that: The step 6 is specifically as follows: The original image i and the regenerated image Respectively represented by CLIP image encoder, we get m i and The loss function is defined as: m i =CLIP(i), Calculate the loss function and gradient and update the model parameters.