Multi-scale multi-modal image condition generation method based on de-noising diffusion model

By designing a multi-modal image denoising module and a multi-scale denoising strategy, combined with the introduction of infrared images during training, the problems of traditional denoising diffusion models in multi-scale image generation, local image blurring and lack of details are solved, and high-precision multi-scale multi-modal image generation and conditional generation are achieved.

CN119963673APending Publication Date: 2025-05-0910TH RES INST OF CETC
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411948048.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Traditional single denoising and diffusion models are difficult to handle multi-scale image generation tasks simultaneously, generating images are blurred locally, lacking details, and there is room for improvement in multimodal image generation and conditional generation.

Method used

A multi-modal image denoising module is designed, and through a multi-scale denoising strategy, combined with the introduction of infrared images during training, the image structure is fusion with complementary denoising modules and multi-scale diffusion forward noise addition modules to achieve high-precision denoising and generation of multi-modal images.

Benefits of technology

It effectively solves the problems of multi-scale image generation, local image blur and lack of details in a single denoising diffusion model, improves the effect of multi-modal image generation, and enhances the generalization and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963673A_ABST
    Figure CN119963673A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale multi-modal image condition generation method based on a denoising diffusion model, and relates to the field of image generation. According to the method, the space and information complementarity of the multi-modal image is effectively utilized, and the generation is guided based on the text embedding features generated by the fine-tuning CLIP text encoder, so that the problems that the de-noising diffusion model cannot effectively generate the high-fidelity multi-modal image and the types of generated targets and environmental conditions are limited are effectively solved, and the multi-modal image with high fidelity is obtained. And meanwhile, a multi-scale denoising strategy is introduced in the denoising generation process, so that the capability of generating different-scale images by a single denoising diffusion model is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image generation, and in particular to a multi-scale multi-modal image conditional generation method based on a denoising diffusion model. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and may not constitute prior art.

[0003] The emergence of Generative Artificial Intelligence (GenAI) marks a revolutionary breakthrough in image generation technology. The image generation task aims to learn from existing data and create new images. Conditional Image Generation refers to the use of additional conditional information to guide and control the generation results during the generation process. By adding conditional variables (text descriptions, input images, etc.) to the generation process, the generative model can generate images with specific features, attributes or styles according to specific needs. With the development of technology, the most advanced generative models are now able to generate high-quality and realistic images. Conditional image generation has shown broad application potential in many fields, including data augmentation, computer graphics, artistic creation, virtual reality content generation, medical image processing, etc.

[0004] Traditional image generation methods include generative adversarial networks (GANs), variational autoencoders (VAEs), and autoregressive models. These models have achieved remarkable results in the field of image generation, but they also have certain limitations. Although GANs can generate highly realistic images, their training process is highly unstable and prone to mode collapse. In addition, the generator has limited generation types and lacks diversity. VAEs have a more stable training process and can better capture the relationship between latent variables. However, the images generated by VAEs are usually blurry, and it is difficult to accurately control the details of the generated results. Although autoregressive models can generate high-quality images, the generation process is very slow and inefficient.

[0005] Recently, denoising diffusion models have demonstrated excellent performance in image generation tasks. The model simulates the gradual denoising process of an image from noise to a clear image, reversely infers the true distribution of the image, and generates high-quality, detailed images. However, denoising diffusion models still have certain defects. A single denoising diffusion model can usually only learn and generate images of a fixed scale, and it is difficult to handle multi-scale image generation tasks at the same time. In addition, these models still have room for improvement in multimodal image generation and conditional generation. Summary of the invention

[0006] The purpose of the present invention is to address the deficiencies in the prior art and provide a multi-scale multi-modal image conditional generation method based on a denoising diffusion model. The method aims to solve the problems that a traditional single denoising diffusion model is difficult to complete the multi-scale image generation task, and the generated images are locally blurred and lack details by designing a multi-modal image denoising module and a multi-scale denoising strategy. At the same time, infrared images are introduced during training to improve the generation effect of the denoising diffusion model on multi-modal images, effectively solve the problem of images of different scales generated by a single denoising diffusion model, and the problem that the model is limited in the types of generated targets, modalities and environmental conditions, thereby improving the generalization of the model.

[0007] The technical solution of the present invention is as follows:

[0008] A multi-scale multi-modal image conditional generation method based on a denoising diffusion model, comprising:

[0009] Step S1: Send the multimodal image to a denoising module based on image structure fusion complementarity, use the complementary information between the modalities to achieve high-precision denoising and single-modal feature enhancement, and optimize the blurred area of ​​the image;

[0010] Step S2: Send the input conditional text to the CLIP text encoder fine-tuned based on prompt learning to obtain the text embedding features of the conditional control words;

[0011] Step S3: Send the denoised image to a forward denoising module based on multi-scale diffusion for denoising, and gradually add noise at each fuzzy scale to ensure that the structural information of each scale of the image is preserved;

[0012] Step S4: The noisy image and text embedding features are input into the adaptive multi-scale full convolution denoising generation module, and a multi-scale multimodal image is finally generated by denoising at different scales.

[0013] Furthermore, the step S1 comprises:

[0014] Two reconstruction autoencoders are designed to reconstruct input images of different modalities respectively. At the same time, denoised images are obtained through two-stage reconstruction for different reconstruction autoencoders.

[0015] Furthermore, the text encoder in CLIP is fine-tuned based on cue learning to encode text control conditions in multimodal domains.

[0016] Furthermore, the forward noising module can blur the image at different scales and control the blur degree of each scale based on the multi-scale diffusion and blur mechanism, thereby ensuring that the structural information of each scale is retained during the forward noising process.

[0017] Furthermore, the step S4 comprises:

[0018] Through a fully convolutional network, denoising is performed gradually at different scales, and multi-scale multimodal images with high fidelity and good structural consistency are generated under the guidance of text embedding features.

[0019] Furthermore, the step S3 comprises:

[0020] Step S31: blurring different scales of the image by bicubic interpolation;

[0021] Step S32: adding noise at each fuzzy scale;

[0022] Step S32: The images after denoising at each scale are fused into a final noisy image by weighted averaging.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] 1. The present invention can effectively extract and fuse the features of visible light and infrared images, and accurately guide the denoising and restoration of multimodal images. This method can not only handle the denoising of multimodal images, but also guide the denoising of its own image through the deep structure of the image when the modality is single, broadening the scope of application of the module and providing higher resolution input images for subsequent image generation.

[0025] 2. The present invention can learn text information in multimodal fields such as infrared images and visible light images through a small amount of data based on prompt learning, thereby significantly improving the accuracy and reliability of guiding multimodal image generation using text embedding features.

[0026] 3. The present invention can effectively utilize the structural information of images at various scales. This enables the network to generate multi-scale images by denoising at different scales during the generation process, and to generate images from low resolution to high resolution in sequence through multi-step denoising, first generating the global structure, and then gradually adding details, so that the system can generate images with high fidelity and good structural consistency.

[0027] 4. The present invention has a wide range of applications; the ability of multi-modal and multi-scale generation enables the network to complete generation tasks regardless of single-modal input or multi-modal images of visible light and infrared, and to generate multi-modal images of different resolutions, with extremely strong adaptability.

[0028] 5. The present invention has good scalability and adopts a modular design. The denoising model, text encoding model and subsequent generation model can be independently optimized and upgraded, making the system have good scalability and generalization. With the different requirements of input images and conditional texts, the system can be trained based on new data sets to obtain image generation models for different fields, and support future technology upgrades and application expansions. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a schematic diagram of the process of the present invention;

[0030] Figure 2 It is a schematic diagram of the network structure of the present invention;

[0031] Figure 3 It is a flow chart of the training steps of the present invention. DETAILED DESCRIPTION

[0032] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0033] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.

[0034] Embodiment 1

[0035] See also Figure 1-3 , a multi-scale multi-modal image conditional generation method based on a denoising diffusion model, specifically comprising the following steps:

[0036] Step S1: Send the multimodal image to the denoising module based on image structure fusion complementarity, use the complementary information of the deep structure between the modalities to achieve high-precision denoising and enhancement of single-modal features, optimize the blurred area of ​​the image, thereby removing the interference noise in the image background, and obtain an image with more details and texture information, thereby guiding a more refined denoising process; that is, enrich and enhance the single-modal structure with insufficient information, optimize the blurred area of ​​the image, and thereby guide a more refined denoising process;

[0037] Step S2: sending the input conditional text to the CLIP text encoder (Pretrained Text Encoder) fine-tuned based on prompt learning to obtain the text embedding features (Text Embedding) of the conditional control words; that is, sending the input conditional text to the CLIP (Contrastive Language-Image Pretraining) pretrained text encoder (Pretrained Text Encoder) fine-tuned based on prompt learning (Prompt-Learning) to obtain the text embedding features of the conditional control words, which are used to guide the model to generate multimodal images containing specific scenes and targets in the subsequent denoising generation process;

[0038] Step S3: sending the denoised image to a forward denoising module based on multi-scale diffusion for denoising processing, and gradually adding noise at each fuzzy scale to ensure that the structural information of each scale of the image is preserved; that is, sending the denoised image to a forward denoising module based on multi-scale diffusion for denoising processing, accurately controlling the degree of fuzziness at each scale through bicubic interpolation, and gradually adding noise at each fuzzy scale to ensure that the structural information of each scale is preserved; it should be noted that the forward denoising process enables the model to learn the association between multimodal images and ordinary Gaussian white noise, and ensures that the structural information of each scale of the image is preserved by gradually adding noise at each fuzzy scale;

[0039] Step S4: The noisy image and text embedding features are input into the adaptive multi-scale full convolution denoising generation module, and a multi-scale multimodal image is finally generated by denoising at different scales; this module is trained with images of five scales from small to large, so that the model can generate high-quality and high-detail images by learning small-scale input images.

[0040] In this embodiment, specifically, step S1 includes:

[0041] Two reconstruction autoencoders are designed to reconstruct input images of different modalities respectively. At the same time, denoised images are obtained through two-stage reconstruction for different reconstruction autoencoders.

[0042] It should be noted that in step S1, based mainly on the autoencoder representation learning strategy, by designing a deep structure prediction and fusion network, more refined structural information is generated using supplementary information from other modal image structures, thereby more finely guiding the denoising process, which is generally carried out in two stages;

[0043] In the first stage of representation learning, two reconstruction autoencoders need to be trained separately to reconstruct infrared images and visible light images respectively. For the input image, the learning method based on the mask autoencoder first randomly masks a part of the input image, and inputs the remaining masked part into the autoencoder as visible data, and then uses the visible part data to reconstruct the masked part input and output a coarse-precision reconstructed image. During training, the loss function can be simply defined as the reconstruction loss, that is:

[0044]

[0045] Where i is the pixel position and N is the number of pixels;

[0046] In the second stage, the deep structure extraction and fusion module is used to generate structural guidance for the final denoising process. First, after the autoencoder reconstruction, the output of the codecs at different layers needs to be used to extract the deep structure of the predicted image. Due to the different feature representations and structures of visible and infrared images, two deep structure extraction networks need to be trained separately. At this time, the input of the deep structure extraction network is the feature representation of each layer of the first part of the autoencoder. The final prediction loss is as follows:

[0047]

[0048] Where i represents the i-th layer of the autoencoder R, C h i is the number of channels in the deep structure at level i, struct i,c is the predicted depth structure, is the true value of the deep structure, Dist(.,.) represents the cross entropy loss;

[0049] After obtaining the deep structure of infrared and visible light images, it is necessary to further fuse the complementary information of the two to generate fine structure guidance. In order to maximize the use of the deep structure and to be able to perform denoising when the other modality image is missing, the fusion process can use maximum pooling for fusion. In this way, only the other modality data needs to be masked during fusion, and the single modality structure can still be retained as a guide for subsequent denoising.

[0050] In this embodiment, the text encoder in CLIP is fine-tuned based on prompt learning to encode text control conditions in the multimodal field, that is, to achieve feature extraction and interpretation of conditional control words for multiple targets (aircraft, cars, ships, tanks, etc.), multiple modalities (visible light, infrared), and multiple weather conditions (cloudy weather, night, etc.);

[0051] It should be noted that in step S2, the text encoder in CLIP is fine-tuned mainly through prompt learning, so that the encoder can learn the feature extraction and interpretation of conditional control words in multiple scenes (airports, mountains, ports, etc.), multiple modalities (visible light, infrared), and multiple weather conditions (cloudy weather, night, etc.). The specific definition is shown in the following formula:

[0052]

[0053] where f CLIP is the function representation of the CLIP text encoder, θ′ CLIP are the parameters of the CLIP text encoder after fine-tuning. Prompt-Tuning only trains some parameters during training, keeps the parameters of the foundation model unchanged, and trains a small model with a small number of parameters for each specific task. The training method of prompt learning fine-tuning is to add some special conditional control words (Conditioning Token) of a specific length before inputting the original text to increase the probability of generating the expected sequence. By adding the conditional control words (Token) to the input, Prompt-Tuning can ensure that the function itself remains unchanged. Adding the required conditional control content in front of X will change X to X * , thus affecting the probability of X generating the expected Y, so that the model has the ability to extract features under specific input text conditions.

[0054] In this embodiment, specifically, the forward noise addition module can blur the image at different scales and control the blur degree of each scale based on the multi-scale diffusion and blur mechanism, so as to ensure that the structural information of each scale is retained during the forward noise addition process;

[0055] In this embodiment, specifically, step S3 includes:

[0056] Step S31: blurring different scales of the image by bicubic interpolation;

[0057] Step S32: adding noise at each fuzzy scale;

[0058] Step S32: The images after denoising at each scale are fused into a final noisy image by weighted averaging.

[0059] It should be noted that in step S3, during the image denoising process, a forward denoising module based on multi-scale diffusion is used to gradually add noise to the image and accurately control the blur level of each scale. First, bicubic interpolation is used to add noise to each scale I k =D k (I) Fuzzy processing is performed, where D krepresents the diffusion operation at scale k, which preserves the details and structural information of the image. Then, noise is added at each blur scale The variance of the noise is As the scale increases, Ensure that the structural information at the low-resolution scale is well preserved, while the detail information at the high-resolution scale is more affected by noise. The noisy image at each scale is Finally, all scales of the noisy images are fused into the final noisy image by weighted averaging:

[0060]

[0061] where w k is the weighting coefficient of each scale. By gradually adding noise and fusing images of different scales, this process can effectively retain the multi-level feature information of the image, while avoiding excessive interference introduced by noise, ensuring that the structure and details of the image are balanced. It should also be noted that the noise addition process is only performed during model training, which is used to learn how the image is converted into white noise through noise addition, and then reverse the denoising process.

[0062] In this embodiment, specifically, step S4 includes:

[0063] Through a fully convolutional network, denoising is performed gradually at different scales, and multi-scale multimodal images with high fidelity and good structural consistency are generated under the guidance of text embedding features.

[0064] It should be noted that in step S4, the noisy image and text embedding features are input into the adaptive multi-scale full convolution denoising generation module, the goal of which is to generate high-quality multi-scale multimodal images by denoising at different scales. Specifically, the input image I input and the text embedding feature T are combined to form the input I of the module input =I noisy ⊕T, where the symbol ⊕ represents the concatenation of image and text features. This module contains five scales S from small to large 1 ,S 2 ,…,S 5 , represents the image features of different resolutions. After the image at each scale passes through multiple layers of convolution operations, the noise is gradually removed and the details are restored. The output is:

[0065]

[0066] Among them, f k represents the denoising operation performed on scale k, is the denoised image. The model is trained using five scale images from small to large, so that the model can restore higher quality images by learning the noise removal strategy of low-resolution images while gradually increasing the resolution. Finally, the denoising results of the five scales are fused into a multi-scale image I final , the outputs of each scale are fused by weighted average:

[0067]

[0068] Among them, w k is the weighting coefficient of each scale, indicating the contribution of each scale to the final output. Through the multi-scale denoising process, the model can effectively remove noise at different scales while retaining the details of the image, generating the final multi-scale, high-quality multi-modal image.

[0069] The present invention effectively utilizes the spatial and information complementarity of multimodal images through image structure fusion complementary denoising modules, and introduces a fine-tuned CLIP text encoder to solve the problems that the denoising diffusion model cannot effectively generate high-fidelity multimodal images and the types of generated targets and environmental conditions are limited. At the same time, based on the multi-scale denoising strategy, the ability of a single denoising diffusion model to generate images of different scales is realized.

[0070] The above-mentioned embodiments only express the specific implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the protection scope of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the technical solution concept of the present application, and these all belong to the protection scope of the present application.

[0071] This background section is provided to generally present the context of the invention, and the work of the presently named inventors, the work to the extent described in this background section, and aspects of this section that did not constitute prior art at the time of application are neither explicitly nor implicitly admitted to be prior art to the present invention.

Claims

1. A multi-scale multi-modal image conditional generation method based on a denoising diffusion model, characterized in that: include: Step S1: Send the multimodal image to a denoising module based on image structure fusion complementarity, use the complementary information between the modalities to achieve high-precision denoising and single-modal feature enhancement, and optimize the blurred area of ​​the image; Step S2: Send the input conditional text to the CLIP text encoder fine-tuned based on prompt learning to obtain the text embedding features of the conditional control words; Step S3: Send the denoised image to a forward denoising module based on multi-scale diffusion for denoising, and gradually add noise at each fuzzy scale to ensure that the structural information of each scale of the image is preserved; Step S4: The noisy image and text embedding features are input into the adaptive multi-scale full convolution denoising generation module, and a multi-scale multimodal image is finally generated by denoising at different scales.

2. The method for conditional generation of multi-scale multi-modal images based on a denoising diffusion model according to claim 1, characterized in that: The step S1 comprises: Two reconstruction autoencoders are designed to reconstruct input images of different modalities respectively. At the same time, denoised images are obtained through two-stage reconstruction for different reconstruction autoencoders.

3. The multi-scale multi-modal image conditional generation method based on a denoising diffusion model according to claim 2, characterized in that: The text encoder in CLIP is fine-tuned based on cue learning to encode textual control conditions in multimodal domains.

4. The multi-scale multi-modal image conditional generation method based on a denoising diffusion model according to claim 3, characterized in that: The forward denoising module can blur images at different scales and control the blur degree of each scale based on a multi-scale diffusion and blur mechanism, thereby ensuring that structural information at each scale is retained during the forward denoising process.

5. The method for conditional generation of multi-scale multi-modal images based on a denoising diffusion model according to claim 4, characterized in that: The step S4 comprises: Through a fully convolutional network, denoising is performed gradually at different scales, and multi-scale multimodal images with high fidelity and good structural consistency are generated under the guidance of text embedding features.

6. The method for conditional generation of multi-scale multi-modal images based on a denoising diffusion model according to claim 3, characterized in that: The step S3 comprises: Step S31: blurring different scales of the image by bicubic interpolation; Step S32: adding noise at each fuzzy scale; Step S32: The images after denoising at each scale are fused into a final noisy image by weighted averaging.

Citation Information

Cited By

  • Fish image generation and feature enhancement method combining mask learning and diffusion model

    CN121170492A

  • Comparison parameter decoupling multi-modal image synthesis method

    CN121860871A

  • Multi-modal image synthesis method with contrast parameter decoupling

    CN121860871B