Multimodal image fusion method, system and device based on inverse diffusion model and readable storage medium

Through the inverse diffusion model, the visible image features are directed to the infrared image to generate a fusion image with a visible light style, which solves the problem that the fusion image in the prior art is difficult to adapt to the high-level vision pre-trained model, and achieves more efficient downstream task performance.

CN120070202AActive Publication Date: 2025-05-30HARBIN INST OF TECH

Patent Information

Application Number
CN202510124896.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

Existing multimodal image fusion methods are difficult to adapt the fusion image to high-level visual pre-trained models, resulting in the fusion image being poorly performed in downstream tasks.

Method used

The inverse diffusion model is used to reverse the visible light image to the noise potential space, and the reversed visible light image features are used to guide the reversal of the infrared image to generate a fusion image with a visible light style.

Benefits of technology

The fusion images generated by the inverse diffusion model can more effectively adapt to high-level vision pre-trained models, significantly improving the performance of downstream machine perception tasks and reducing training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005259728280000031
    Figure BDA0005259728280000031
  • Figure BDA0005259728280000035
    Figure BDA0005259728280000035
  • Figure BDA0005259728280000038
    Figure BDA0005259728280000038
Patent Text Reader

Abstract

The invention relates to a multi-modal image fusion method of an inverse diffusion model, which comprises the following steps of: 1, reversing a visible light image to a noise potential space by utilizing an inverse diffusion technology, and then guiding an infrared image to be reversed by utilizing the characteristics of the reversed visible light image; 2, guiding through an inverse process in a diffusion model, injecting the appearance attribute of the visible light into an infrared feature, and generating an infrared image with a visible light style by the feature; and step 3, designing a specific fusion rule which is used for fusing the reversed visible light and infrared characteristics of the attention layer in the denoising process, retaining the text interaction capability of the model, and supporting language-driven fusion control. According to the method, the high-quality fusion image can be directly generated without additional training or fine adjustment. The obtained fusion image is highly compatible with the basic model, the problem of difference between data domains is effectively solved, and the performance of a downstream machine perception task is remarkably improved. According to the method, the training cost is remarkably reduced, and an efficient and innovative solution is provided for cross-domain tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a multi-modal image fusion method, system, device and readable storage medium based on an inverse diffusion model. Background Art

[0002] Image fusion technology is widely recognized for its ability to integrate complementary information from multiple source images into a single fused image. Due to the inherent physical limitations of various sensors, each modality has its limitations in capturing all the information in a scene. For example, visible light sensors commonly used in daily applications can effectively capture scene details but are very sensitive to lighting conditions. In poor lighting conditions, such as at night or in overexposed situations, their performance significantly degrades. In contrast, infrared sensors are more robust to different lighting and weather conditions and can capture useful scene information both during the day and at night. However, infrared images lack the detailed structural information provided by visible light sensors. By combining the complementary information of these two modalities, infrared and visible light image fusion technology can generate a fused image that retains as much effective information in the scene as possible.

[0003] With the development of generative learning techniques, numerous fusion methods have been developed for infrared and visible light image fusion tasks. Generally, these methods can be roughly divided into three categories: autoencoders (AEs), generative adversarial networks (GANs), and diffusion models (DMs). AE-based methods usually adopt complex network architectures to improve feature extraction. GAN-based methods utilize an adversarial framework where the generator aims to produce a fused image that can deceive the discriminator, while the discriminator endeavors to distinguish between the generated image and the real image. Recently, DM-based methods have received extensive attention for their ability to produce high-quality results and are generally more stable than GANs. Despite the remarkable achievements of these methods, there is still a key challenge that remains unsolved. This challenge involves the problem of adapting the fused image to downstream tasks. The spectrum captured by infrared sensors is different from that of visible light, resulting in significant differences in the appearance of the images. During the fusion process, most existing methods enforce pixel-level similarity between the fused image and its source images. Consequently, the fused image incorporates the appearance features of both the infrared and visible light modalities. From the perspective of appearance attributes, infrared, visible light, and fused images belong to three different domains. In this context, although existing fusion methods perform well in traditional computer vision tasks, the emergence of foundation models has introduced new challenges. Specifically, the fused image as an independent domain is difficult to seamlessly integrate with pre-trained foundation models. These models are typically trained on large-scale visible light image datasets and to some extent include infrared images. However, few foundation models include fused images in their training data, and there is an inherent domain gap when fused images are directly applied to these models. When the same fused image is directly input into a pre-trained detection model, the results are less satisfactory compared to fused images that are more similar to visible light. Summary of the Invention

[0004] The present invention aims to solve the problem that the fused image is not directly adaptable to high-level vision pre-trained models, and thus proposes a multi-modal image fusion method, system, device, and readable storage medium based on an inverse diffusion model.

[0005] The technical solution adopted by the present invention to solve the above problems is as follows: A multi-modal image fusion method based on an inverse diffusion model, comprising the following steps:

[0006] Step 1: Utilize the inverse diffusion technique to reverse the visible light image to the noise latent space, and then use the features of the reversed visible light image to guide the reversal of the infrared image;

[0007] Step 2: Through guidance in the inverse process of the diffusion model, inject the appearance attributes of visible light into the infrared features, and the features can generate an infrared image with a visible light style;

[0008] Step 3: Design specific fusion rules for fusing and reversing the visible and infrared features in the attention layer during the denoising process, preserving the text interaction ability of the model and supporting language-driven fusion control.

[0009] Further, in Step 1, following the framework of the denoising diffusion probability model, the model includes a diffusion process of T steps and a reverse process of T steps; in the forward process, a clean image x 0 is gradually transformed into white Gaussian noise; it can be expressed as:

[0010]

[0011] where x t represents the image at time t, refers to the hyperparameter related to time t, is a standard normal distribution;

[0012] Its reverse process can be expressed as:

[0013]

[0014] In formula (2), σ t is the hyperparameter regarding α t ,

[0015] predicts x 0 ,

[0016] represents the direction pointing to x t ,

[0017] f t (·) is the network model of the diffusion model, and at step t, the visible features and are generated as follows:

[0018]

[0019] In formula (3), and are generated by the network predicting the noise .

[0020] Further, in Step 2, to integrate the visible features into the infrared image, the update step of the visible-light style infrared is defined as follows:

[0021]

[0022] In formula (4), λ represents the weight factor, That is, a visual cue; it is crucial that the visual cue guides the noisy infrared image towards the visible direction at each time step; according to equations (2), (3), and (4) in It can also be expressed as:

[0023]

[0024] Equation (4) can be rewritten as:

[0025]

[0026] Furthermore, in step three, to maintain the visible appearance, visible features are used as the basic components of the fused image; the iterative generation process of the fused image can be expressed as:

[0027]

[0028] A customized fusion rule is introduced to inject visible-light-like infrared information, specifically: the fusion rule is applied to the self-attention layer Att in the denoising network fuse , expressed as:

[0029]

[0030] T 1 and T 2 represent two hyperparameters that need to satisfy the constraint T 1 > T 2 ;

[0031] To ensure that the fused image mainly retains the visible-light content, the query vector Q fuse is kept unchanged during the iterative process, and the in the self-attention layer is used to retain the visual characteristics of the infrared image and inject infrared information in the early steps of denoising.

[0032] Furthermore, classifier-free guidance is adopted to improve the overall quality of the generated fused image; at each denoising step t, two forward passes are performed through the denoising network:

[0033]

[0034] C text represents the text embedding from the pre-trained text encoder; represents the predicted noise guided by equation (9), while represents the unconditional generation result, and ω is the strength of the guidance vector.

[0035] The present invention also relates to a multi-modal image fusion system for an inverse diffusion model, including a computer module that applies the above-mentioned multi-modal image fusion method for an inverse diffusion model.

[0036] The present invention also relates to a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

[0037] The present invention also relates to a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

[0038] Beneficial effects

[0039] The present invention proposes a fusion model without training, which realizes efficient adaptation to downstream tasks through inversion technology and demonstrates various beneficial effects.

[0040] The present invention integrates the features of visible light images into the inversion process of infrared images to generate fused images with a visible light style. A weight factor is introduced in this process to flexibly control the intensity of visible light appearance attributes and ensure adaptation to different task requirements.

[0041] The present invention combines diffusion-based image editing technology to design customized fusion rules, which not only retain the text interaction ability of the model but also support language-driven fusion control. By making full use of the capabilities of pre-trained diffusion models, the present invention can directly generate high-quality fused images without additional training or fine-tuning. The obtained fused images are highly compatible with the base model, effectively solving the problem of differences between data domains and significantly improving the performance of downstream machine perception tasks.

[0042] The present invention significantly reduces the training cost and provides an efficient and innovative solution for cross-domain tasks. Description of the drawings

[0043] Figure 1 is a schematic diagram of inverse diffusion multimodal fusion in the present invention;

[0044] Figure 2 is a visual comparison diagram of infrared-visible light fusion performed on the MSRS dataset in the present invention. Detailed implementation manners

[0045] The technical solutions of the present invention will be specifically described below.

[0046] As Figure 1 shown, the multimodal image fusion method based on an inverse diffusion model of the present invention specifically includes:

[0047] Using inverse diffusion technology, reverse the visible light image to the noise latent space, and then use the features of the reversed visible light image to guide the reversal of the infrared image.

[0048] Using the framework of the Denoising Diffusion Probability Model (DDPM), which includes a diffusion process of T steps and a reverse process of T steps. In the forward process, a clean image x is gradually 0 converted into white Gaussian noise. The t-th step is as follows:

[0049]

[0050] where x t represents the image at time t, refers to the hyperparameter related to time t, is a standard normal distribution;

[0051] Its reverse process is

[0052]

[0053] In formula (2), σ t is a hyperparameter regarding α t , predicts x 0 , and represents the direction pointing to x t . f t (x t ) represents the predicted output of the network for the input x t . represents the variance, and {z t} represents an independent and identically distributed standard normal vector.

[0054] At the t-th step, the expressions for generating the visible features and are as follows:

[0055]

[0056] In formula (3), and are generated by the network predicting the noise .

[0057] As Figure 1 shown, by guiding through the inverse diffusion model, the appearance attributes of visible light are injected into the infrared features, and the features can generate infrared images with visible light styles.

[0058] To integrate visible light features into infrared images, the present invention defines the update steps for infrared with visible light styles as follows:

[0059]

[0060] In formula (4), λ represents the weight factor, and That is the visual cue. It is crucial for the visual cue to guide the noisy infrared image towards the visible direction at each time step. According to equations (2), (3), and (4), can also be expressed as:

[0061]

[0062] In equation (5), represents the predicted Simplify equation (5) to:

[0063]

[0064] Therefore, equation (4) can be rewritten as:

[0065]

[0066] As Figure 1 shown, design a specific fusion rule to be applied to the attention layer fusion reversal of visible and infrared features during the denoising process, retain the text interaction ability of the model, and support language-driven fusion control. To maintain the visible appearance, use the visible features as the basic components of the fused image. The iterative generation process of the fused image can be expressed as:

[0067]

[0068] It is not enough to solely rely on equation (7) to obtain the fused image because the update process is exactly the same as the reverse process of the visible image, thus excluding the infrared information. To solve this problem, introduce a customized fusion rule to inject infrared information similar to the visible light. Specifically, the designed fusion rule is applied to the self-attention layer Att fuse in the denoising network and can be expressed as:

[0069]

[0070]

[0071] T 1 and T 2 represent two hyperparameters that need to satisfy the constraint condition T 1 > T 2 . In the fusion rule strategy of equation (9), Q fuse does not explicitly specify the information exchange with the infrared or visible image.

[0072] To ensure that the fused image mainly retains the visible light content, keep the query vector Q fuse unchanged during the iterative process.

[0073] Utilize the in the self-attention layer to retain the visual characteristics of the infrared image. During this process, infrared information is injected in the early steps of denoising, which is a crucial stage for the formation of the overall image layout.

[0074] As the iteration progresses, the main structure and content of the fused image have been basically determined, and the subsequent steps are mainly dedicated to the fine-tuning of the image style. After injecting infrared information into the fused image in the initial stage, several additional steps are added to improve the image details by incorporating visible light information. In the final stage, manual modification of the features of the fused image is avoided to generate a more unified and visually consistent fused image.

[0075] Adopt classifier-free guidance to enhance the overall quality of the generated fused image. At each denoising step t, two forward propagations are performed through the denoising network:

[0076]

[0077] C text represents the text embedding from the pre-trained text encoder. represents the predicted noise guided by Equation (9), while represents the unconditional generation result, and ω is the strength of the guidance vector.

[0078] As Figure 1 shown, the strategy for obtaining the fused image is as follows: First, use the image encoder to embed the source infrared image and visible image as and Next, reverse the diffusion process to construct a memory bank that stores the intermediate features and the latent variables , recording the information at each time step t. At this stage, first generate the visible features and then use these features to guide the generation of the corresponding infrared features . To generate the fused image, perform the forward diffusion process conditioned on the memory bank and the text prompt. During this process, a customized fusion rule called appearance feature injection is introduced to obtain the fused feature After that, the fused feature in the forward process is transformed into a fused image with a visible style.

[0079] The present invention adopts three representative datasets: RoadScene, MSRS, and FMB. Among them, the RoadScene dataset contains 221 pairs of images, the MSRS test dataset contains 361 pairs of images, and the FMB test dataset consists of 280 pairs of images.

[0080] Effect verification

[0081] The method proposed by the present invention uses SD v1.5. The hyperparameter λ in formula (4) is set to 0.08, while the ω value in formula (10) is set to 2. In the experiment, T in formula (9) 1 and T 2 are set to 0 and 70 respectively. All experiments are carried out on an NVIDIA GeForce RTX 3090 GPU with 24GB video memory.

[0082] Effect evaluation of the fused dataset

[0083] Four evaluation metrics are selected, including cross-entropy (CE), entropy (EN), standard deviation (SD), and the perceptual score CLIPIQA of CLIP. The higher the value of the above metrics, the better the quality of the fused image. In particular, CLIPIQA simultaneously evaluates the quality perception and abstract perception of the image, aiming to achieve a higher correlation with human visual perception.

[0084] Table 1 shows the comparison of the performance of the fused images, summarizing the comparison of the fusion performance on the RoadScene test dataset, MSRS dataset, and FMB dataset. The best and sub-optimal results in each metric are marked in bold and underlined respectively for easy identification of the model performance.

[0085] Table 1

[0086]

[0087] In Table 1, the quantitative comparison results of the present invention with the current mainstream methods (SOTA) on the RoadScene, MSRS, and FMB datasets are shown in detail.

[0088] The present invention performs excellently in all metrics and shows competitiveness. The highest score is obtained in the standard deviation (SD) metric, indicating that the generated fused image has superior contrast, which is mainly due to the converted infrared image, which provides more detailed information than the original infrared image. In addition, the present invention also performs outstandingly in the CLIPIQA metric, reflecting its ability to effectively transfer the visible appearance to the infrared image, thus generating a fused image with a visible style. In contrast, the comparison methods usually retain the information of the original infrared image, resulting in the generated fused image presenting a unique appearance. The fused image of the present invention has obtained an excellent score in terms of consistency with human visual perception, further highlighting its superiority.

[0089] In Figure 2In this invention, two embodiments in MSRS are presented, and the effects of seven different methods are compared with those of the method of this invention. Scene 1 is a portrayal of a vehicle in the evening, and Scene 2 is a traffic intersection with a building background. Both of these scenes are affected by light degradation, resulting in a low brightness of the visible image. In the fusion results, the fused images generated by U2Fusion tend to be darker, and details are difficult to identify. Other methods, such as MetaFusion, TIMFusion, CDDFuse, and DDFM, can maintain the brightness level of the visible image. More recent diffusion-based methods, such as Text-IF and Diff-IF, although slightly superior to the original image in terms of visual effects, the influence of light degradation is still obvious. In contrast, the fused images generated by the method of this invention are more visually appealing, making the edges sharper and producing clear and visually pleasing fused images. These results highlight the use of a pre-trained stable diffusion model, which has a stronger ability to generate more realistic and high-quality images. The finally generated fused images are characterized by a high and natural contrast, showing vivid colors.

[0090] As described above, it is only a preferred embodiment of the present invention, and there is no any form of limitation to the present invention. Although the present invention has been disclosed as above with the preferred embodiment, it is not used to limit the present invention. Any person skilled in the art can make some changes or modifications to be equivalent embodiments of equivalent changes within the scope of the technical solution of the present invention when using the above-disclosed technical content. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent replacement, and improvement made to the above embodiments according to the technical essence of the present invention within the spirit and principle of the present invention still belong to the protection scope of the technical solution of the present invention.

Claims

1. A multimodal image fusion method based on an inverse diffusion model, characterized in that: The following steps are involved: Step 1: Use the inverse diffusion technique to reverse the visible light image to the noise latent space, and then use the features of the reversed visible light image to guide the reverse of the infrared image; Step 2: Guided by the inverse process in the diffusion model, the appearance attributes of visible light are injected into the infrared features, and the features can generate infrared images with visible light style; Step 3: Design specific fusion rules for the attention layer in the denoising process to fuse the inverted visible light and infrared features, retain the model's text interaction capabilities, and support language-driven fusion control.

2. The multimodal image fusion method of the inverse diffusion model according to claim 1, characterized in that: In step 1, the framework of the denoising diffusion probability model is followed, and the model includes a T-step diffusion process and a T-step reverse process; the forward process gradually transforms a clean image x0 into white Gaussian noise; it can be expressed as: Among them, x t represents the image at time t, refers to the hyperparameters related to time t, is a standard normal distribution; The reverse process is expressed as: In formula (2), σ t It's about α t The hyperparameters of Predict x0, Indicates pointing to x t Directions f t (·) is the network model of the diffusion model, generating visible features at step t and The expression is as follows: In formula (3), and is the network prediction noise Generated.

3. The multimodal image fusion method of the inverse diffusion model according to claim 1, characterized in that: In step 2, in order to integrate the visible light features into the infrared image, the update steps of defining the visible light style infrared are as follows: In formula (4), λ represents the weight factor, That is, the visual cue; the visual cue is crucial to guide the noisy infrared image towards the visible direction at each time step; according to formula (2) and formula (3), formula (4) It can also be expressed as: Formula (4) can be rewritten as:

4. The multimodal image fusion method of the inverse diffusion model according to claim 1, characterized in that: In step 3, in order to maintain the visible appearance, the visible features are used as the basic components of the fused image; the fused image The iterative generation process can be expressed as: A customized fusion rule is introduced to inject infrared information similar to visible light. Specifically, the fusion rule is applied to the self-attention layer Att in the denoising network. fuse , expressed as: T1 and T2 represent two hyperparameters, and the constraint condition T1>T2 must be satisfied; To ensure that the fused image mainly retains the visible light content, the query vector Q is kept during the iteration process. fuse Unchanged, using the self-attention layer Preserve the visual characteristics of infrared images and inject infrared information in the early steps of denoising.

5. The multimodal image fusion method of the inverse diffusion model according to claim 4, characterized in that: The overall quality of the generated fused image is improved by using no classifier guidance; at each denoising step t, two forward propagations are performed through the denoising network: C text represents the text embedding from the pre-trained text encoder; represents the noise predicted according to formula (9), and It represents the unconditional generation result, and ω is the strength of the guidance vector.

6. A multimodal image fusion system based on an inverse diffusion model, characterized in that: The invention comprises a computer module, wherein the module applies a multimodal image fusion method of an inverse diffusion model as described in claims 1 to 5.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Multi-modal visual fusion method and system based on monitoring scene

    CN118096553A

  • Face image generation and recognition method and system based on near-infrared imaging

    CN118097363A

  • Multiband image fusion method based on diffusion model

    CN118154435A

  • Monitoring method, device, equipment, storage medium and computer program product

    CN118865237A

  • Object-level infrared-and-visible-light image fusion method based on fully convolutional neural network

    WO2024174488A1

Cited By

  • Interactive integrated image restoration fusion method and device

    CN121353127A