Method and device for deeply expanding multi-modal image fusion network in combination with diffusion model
By combining the deep expansion of the multimodal image fusion network with the diffusion model and using the parameter generator and diffusion module for adaptive image fusion, the problem of opaque fusion process in the existing technology is solved and efficient multimodal image fusion effect is achieved.
Patent Information
- Application Number
- CN202510952524.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-17
AI Technical Summary
Existing learning-based multimodal image fusion methods have the following problems: the internal working mechanism of the autoencoder is difficult to explain, the fusion process is opaque, and the heuristically designed architecture cannot provide accurate diffusion priors, resulting in unsatisfactory fusion results.
A deep unfolding multimodal image fusion network method combined with a diffusion model is adopted. By constructing a deep fusion and diffusion prior unfolding network, a parameter generator is introduced to generate adaptive parameters, and a denoising diffusion module and a data consistency module are used for image fusion to achieve adaptive and context-aware fusion.
It provides transparent and effective multimodal image fusion, improves fusion performance and adaptability, and the generated fused image can fully reflect the details and overall information of images of different modalities.
Smart Images

Figure CN120807312A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image fusion technology, and in particular to a method and device for a deep-expanded multimodal image fusion network combined with a diffusion model. Background Art
[0002] With the development of deep learning and multimodal technologies, single-modality images are no longer suitable for real-world environments, and multimodal image fusion has attracted widespread attention. The goal of multimodal image fusion is to generate a composite image that combines the strengths of different modal images, improving the completeness and accuracy of the composite image information, thereby providing superior interpretability and practicality for various applications such as object recognition, tracking, and detection.
[0003] Model-based multimodal image fusion methods formulate the image fusion problem as a controllable optimization problem with artificially designed priors, and then solve the optimization problem through an effective iterative algorithm. Although these methods have a transparent fusion process, they cannot handle the complex relationship between different modalities. Learning-based multimodal image fusion methods all benefit from the learning ability of deep networks and can utilize a large amount of training data sets for effective model learning. However, learning-based multimodal image fusion methods still have problems. For example, the internal working mechanism of the autoencoder is difficult to explain and the fusion process is difficult to control; the heuristically designed architecture has the problem of opaque workflow, which greatly hinders the diffusion model from providing accurate diffusion priors, resulting in unsatisfactory fusion results. Summary of the Invention
[0004] In response to the defects of the existing technology, the present invention provides a method and device for a deep expansion multimodal image fusion network combined with a diffusion model.
[0005] In order to achieve the above technical objectives, the specific technical solutions adopted by the present invention are as follows: In one aspect, the present invention provides a method for deep unfolding multimodal image fusion network combined with a diffusion model, comprising the following steps: Acquiring a source image and a pure noise image, wherein the source image includes a first modality image and a second modality image; Constructing and training a deep fusion and diffusion prior expansion network, wherein the deep fusion and diffusion prior expansion network includes a denoising diffusion module and a data consistency module; the data consistency module has a parameter generator, and the parameter generator generates adaptive parameters; The second modality image is filtered and then input into the denoising diffusion module, the residual filter tensor of the pure noise image is calculated and input into the denoising diffusion module; At the time, based on the residual filter tensor of the input filtered second modality image and the pure noise image, the denoising diffusion module uses a diffusion model noise estimation network to calculate the rectified prediction value of the filtered residual image, and takes the rectified prediction value of the filtered residual image as the input of the data consistency module; The rectified prediction value of the filtered residual image is combined with the first modality image by the data consistency module, and after adaptive parameter processing, a prediction fusion image at the current time step is generated; and the residual filter tensor of the generated prediction fusion image is calculated as the next input of the denoising diffusion module. Iterate to the time step length 0, and output the corresponding prediction fusion image as the fusion image.
[0006] Further, the parameter generator generates the adaptive parameter according to the following formula:
[0007] Wherein, is a selection vector obtained by a feature network ; is an adaptive parameter; is a base dictionary; is a first modality image; is a second modality image; is a time step length; is a prediction fusion image at the time step length .
[0008] Further, the rectified prediction value of the filtered residual image is calculated according to the following formula:
[0009] Wherein, is a rectified prediction value of a filtered residual image; is a residual filter tensor at the time step length ; is a noise schedule at the time step length ; represents a predicted noise generated by a noise prediction function .
[0010] Further, the residual filter tensor is obtained according to the following formula:
[0011] Wherein, is a dynamic filter at the time step length ; is a fusion image.
[0012] Further, the prediction fusion image at the time step is obtained according to the following formula:
[0013] wherein, is a prediction fused image with a time step of ; is a convolution operator; is a dynamic filter transpose with a time step of ; is a dynamic filter with a time step of ; is a penalty parameter;
[0014] wherein, is an identity matrix.
[0015] Further, the training of the deep fusion and diffusion prior unfolding network comprises the following steps: obtaining a fusion knowledge prior, generating Gaussian noise; based on the fusion knowledge prior, obtaining a pseudo-target fused image by using a target search method; replacing the fused image with the pseudo-target fused image, and calculating a corresponding residual filter tensor; obtaining a rectified prediction value of a filtered residual image corresponding to the residual filter tensor; generating a prediction fused image according to the rectified prediction value of the filtered residual image, and calculating a loss function; repeating the above steps until convergence.
[0016] Further, the loss function is calculated according to the following steps:
[0017] wherein, is a reconstruction loss; is a noise constraint of a denoising process, , is a sample of a standard Gaussian distribution; is a pixel constraint of the fusion knowledge prior, , is a pseudo-target fused image.
[0018] Further, the reconstruction loss is calculated according to the following formula:
[0019] wherein, is an intensity loss, ; is a gradient loss, ; is a structure similarity loss, ; is an information preservation loss, ; represents a gradient operator; SSIM(·) represents a structure similarity operation; , , , is a penalty parameter.
[0020] Further, the denoising diffusion module calculates a residual filter tensor of the predicted fusion image according to the following formula:
[0021] wherein, is a residual filter tensor at a time step of ; is always set to 0.
[0022] In another aspect, the present application provides a deep unfolding multi-modal image fusion network device combined with a diffusion model, comprising: A first module for collecting source images and pure noise images, the source images including first modal images and second modal images; A second module for constructing a deep fusion and diffusion prior unfolding network, the deep fusion and diffusion prior unfolding network including a denoising diffusion module and a data consistency module; the data consistency module has a parameter generator, and the parameter generator generates adaptive parameters; A third module for inputting the filtered second modal images into the denoising diffusion module, calculating the residual filter tensor of the pure noise images and inputting the residual filter tensor into the denoising diffusion module; at a time step of , based on the input filtered second modal images and the residual filter tensor of the pure noise images, the denoising diffusion module uses a diffusion model noise estimation network to calculate a rectified prediction value of the filtered residual image, and takes the rectified prediction value of the filtered residual image as the input of the data consistency module; A fourth module for generating a predicted fusion image at the current time step by combining the rectified prediction value of the filtered residual image with the first modal image through the adaptive parameter processing of the data consistency module; and calculating the residual filter tensor of the generated predicted fusion image as the next input of the denoising diffusion module; A fifth module for iterative operation to a time step of 0, outputting the corresponding predicted fusion image as a fusion image.
[0023] Compared with the prior art, the present application has the following beneficial technical effects: The deep unfolding multimodal image fusion network method and device combined with the diffusion model provided by the present invention can be used for transparent and effective multimodal image fusion; by introducing adaptive parameters generated by a parameter generator in the data consistency module, adaptive and context-aware image fusion is achieved, providing strong adaptability and fusion capabilities; the deep unfolding technology is integrated with a powerful diffusion prior, thereby improving the fusion performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0025] Figure 1 A schematic diagram of a flow chart of a method for deep unfolding multimodal image fusion network combined with a diffusion model provided in one embodiment; Figure 2 A schematic diagram of a process of a deep unfolding multimodal image fusion network method combined with a diffusion model provided by an embodiment; Figure 3 A schematic diagram of a parameter generator workflow provided by an embodiment; Figure 4 A schematic diagram of a fused image provided by an embodiment, wherein: Figure 4 (a) is a schematic diagram of the fusion image of infrared thermal imaging image and visible light imaging image. Figure 4 (b) is a schematic diagram of the fusion image of computed tomography image and magnetic resonance image. Figure 4 (c) is a schematic diagram of the fusion image of positron emission tomography image and magnetic resonance image. Figure 4 (d) is the fusion image of single photon emission computed tomography image and magnetic resonance image. DETAILED DESCRIPTION
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0027] Reference Figure 1 、 Figure 2 One embodiment provides a method for deep unfolding multimodal image fusion network combined with a diffusion model, comprising the following steps: Acquiring a source image and a pure noise image, wherein the source image includes a first modality image and a second modality image; Constructing and training a deep fusion and diffusion prior expansion network, wherein the deep fusion and diffusion prior expansion network includes a denoising diffusion module and a data consistency module; the data consistency module has a parameter generator, and the parameter generator generates adaptive parameters; The second modality image is filtered and then input into the denoising diffusion module, the residual filter tensor of the pure noise image is calculated and input into the denoising diffusion module; When , based on the residual filter tensor of the input filtered second modality image and the pure noise image, the denoising diffusion module uses the diffusion model noise estimation network to calculate the rectified prediction value of the filtered residual image, and uses the rectified prediction value of the filtered residual image as the input of the data consistency module; The rectified predicted value of the filtered residual image is combined with the first modal image through the data consistency module and processed by adaptive parameters to generate the predicted fused image of the current time step; and the residual filter tensor of the generated predicted fused image is calculated as the next input of the denoising diffusion module; The iterative operation is performed until the time step is 0, and the corresponding predicted fusion image is output as the fusion image.
[0028] The rectified prediction value of the filtered residual image is calculated according to the following formula:
[0029] in, is the rectified prediction value of the residual image after filtering; The time step is The residual filter tensor when ; The time step is Noise scheduling when Represented by the noise prediction function The prediction noise generated.
[0030] The residual filter tensor is obtained according to the following formula:
[0031] in, The time step is Dynamic filter when is the fused image.
[0032] Specifically, we first construct an optimization problem:
[0033] in, is the sparsity promotion function; is the regularization parameter.
[0034] The semi-quadratic splitting algorithm is used to decouple the optimization problem into a priori subproblem and a data consistency subproblem:
[0035] in, is an auxiliary variable; solve it in sequence and subproblems, namely solving the prior subproblem and the data consistency subproblem.
[0036] A denoising diffusion module and a data consistency module are constructed to solve the two sub-problems respectively.
[0037] In the denoising diffusion module, any time step Auxiliary variables Expressed as:
[0038] in, is the noise prediction function; The time step is Noise scheduling at time . It can be regarded as equivalent to Filtered . Therefore, the two sub-problems can be reformulated as:
[0039] in, represents the predicted fused image; is the rectified prediction of the filtered residual image. This rectified prediction is obtained from the input first modality image and fed into the data consistency module. In the denoising diffusion module, the relationship between the prior subproblem and the diffusion mechanism is modeled, and the diffusion model is used to learn and estimate the prior data distribution between the fused image and the input first modality image, thereby providing an efficient diffusion prior for the data consistency module.
[0040] The rectified prediction value of the filtered residual image is processed by the data consistency module in combination with the adaptive parameters to generate the predicted fused image of the current time step. Specifically, the predicted fused image of the time step is obtained according to the following formula:
[0041] in, The time step is The predicted fusion image at the time is the convolution operator; The time step is Dynamic filter transpose when ; The time step is Dynamic filter when is the penalty parameter;
[0042] in, is an identical matrix.
[0043] The denoising diffusion module calculates the residual filter tensor of the predicted fused image according to the following formula:
[0044] in, The time step is The residual filter tensor when ; Always set to 0.
[0045] For the data consistency subproblem, there is a closed-form solution:
[0046] in, .
[0047] Reference Figure 3 , shows the working mechanism of the parameter generator, which generates adaptive parameters according to the following formula:
[0048] in, To select the vector, the feature network Get, feature network The characteristic network architecture is as follows Figure 2 As shown, After channel splicing, at the time step t When , the residual connection network is input to extract feature information and then input into the fully connected network, and the selection vector is obtained after being processed by the normalized exponential function; is the adaptive parameter; As the basic dictionary; is the first modality image; is the second modality image; is the time step; is the time step The predicted fusion image at this time.
[0049] Adaptive parameters generated by the parameter generator The dictionary variable The relaxation is learnable and time step adaptive, and a hypernetwork is incorporated into the data consistency module. The combination of the data consistency module and the hypernetwork provides strong adaptability and fusion capability in a transparent fusion mechanism. The formula for obtaining the prediction fusion image is obtained by correcting the closed-form solution with the generated adaptive parameters. It can be known that the method can realize adaptive and context-aware image fusion.
[0050] The training of the deep fusion and diffusion priori unfolding network comprises the following steps: A fusion knowledge priori is obtained, and Gaussian noise is generated; Based on the fusion knowledge priori, a target search method is used to obtain a pseudo-target fusion image; The pseudo-target fusion image is used to replace the fusion image, and a corresponding residual filter tensor is calculated; The rectified prediction value of the filtered residual image corresponding to the residual filter tensor is obtained; The prediction fusion image is generated according to the rectified prediction value of the filtered residual image, and a loss function is calculated; The above steps are repeated until convergence.
[0051] In the training process, a time step is sampled in [1, T ] in the training process. Then, noise is added to the target labeled image, and the noise variance is determined according to The target fusion image is obtained by using the target search method for training. The marginal distribution is obtained by using the diffusion mechanism:
[0052] Wherein, is obtained by the following formula
[0053]
[0054] At the last time step , it is changed into pure Gaussian noise, and the corresponding pure noise image .
[0055] The loss function is calculated according to the following steps:
[0056] Wherein, is the reconstruction loss; is the noise constraint of the denoising process, , is the sample of the standard Gaussian distribution; is the pixel constraint of the fusion knowledge priori, , is a pseudo-target fusion image.
[0057] the reconstruction loss is calculated according to the following formula:
[0058] wherein, is an intensity loss, ; is a gradient loss, ; is a structural similarity loss, ; is an information preservation loss, ;▽ represents a gradient operator; SSIM(·) represents a structural similarity operation; , , , is a penalty parameter.
[0059] By introducing the information preservation loss in the reconstruction loss to help the module to learn more sufficient information from the diffusion prior.
[0060] Referring to Figure 4 , in an embodiment, four kinds of fusion image schematic diagrams are provided, as shown in the drawings, Figure 4 (a) is a fusion image schematic diagram of an infrared thermal imaging image and a visible light imaging image, Figure 4 (b) is a fusion image schematic diagram of a computed tomography image (CT) and a magnetic resonance image (MRI), Figure 4 (c) is a fusion image schematic diagram of a positron emission computed tomography image (PET) and a magnetic resonance image (MRI), Figure 4 (d) is a fusion image of a single photon emission computed tomography image (SPECT) and a magnetic resonance image (MRI). From the four kinds of fusion images, it can be seen that the method described in the present application can effectively fuse the first modality image and the second modality image together, and the fused image can fully reflect the details and overall information of the first modality image and the second modality image, and the fusion effect is good.
[0061] In another embodiment, a deep unfolding multi-modal image fusion network device combined with a diffusion model is provided, comprising: a first module for acquiring source images and pure noise images, the source images comprising a first modality image and a second modality image; The second module is used to construct a deep fusion and diffusion prior expansion network, which includes a denoising diffusion module and a data consistency module; the data consistency module has a parameter generator, which generates adaptive parameters; The third module is used to filter the second modality image and input it into the denoising diffusion module, calculate the residual filter tensor of the pure noise image and input it into the denoising diffusion module; When , based on the residual filter tensor of the input filtered second modality image and the pure noise image, the denoising diffusion module uses the diffusion model noise estimation network to calculate the rectified prediction value of the filtered residual image, and uses the rectified prediction value of the filtered residual image as the input of the data consistency module; The fourth module is used to combine the rectified prediction value of the filtered residual image with the first modal image through the data consistency module and generate the predicted fusion image of the current time step after adaptive parameter processing; and calculate the residual filter tensor of the generated predicted fusion image as the next input of the denoising diffusion module; The fifth module is used to iterate until the time step is 0, and output the corresponding predicted fused image as the fused image.
[0062] Matters not covered by the present invention are known technologies.
[0063] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0064] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present application, and such modifications and improvements are all within the scope of protection of the present application.
[0065] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A deep multimodal image fusion network method combined with a diffusion model, characterized by: The following steps are involved: Acquiring a source image and a pure noise image, wherein the source image includes a first modality image and a second modality image; Constructing and training a deep fusion and diffusion prior expansion network, wherein the deep fusion and diffusion prior expansion network includes a denoising diffusion module and a data consistency module; The data consistency module has a parameter generator, which generates adaptive parameters; The second modality image is filtered and then input into the denoising diffusion module, the residual filter tensor of the pure noise image is calculated and input into the denoising diffusion module; When , based on the residual filter tensor of the input filtered second modality image and the pure noise image, the denoising diffusion module uses the diffusion model noise estimation network to calculate the rectified prediction value of the filtered residual image, and uses the rectified prediction value of the filtered residual image as the input of the data consistency module; The rectified predicted value of the filtered residual image is combined with the first modal image through the data consistency module and processed by adaptive parameters to generate the predicted fused image of the current time step; and the residual filter tensor of the generated predicted fused image is calculated as the next input of the denoising diffusion module; The iterative operation is performed until the time step is 0, and the corresponding predicted fusion image is output as the fusion image.
2. The method for deep expansion multimodal image fusion network combined with diffusion model according to claim 1, characterized in that: The parameter generator generates adaptive parameters according to the following formula: in, To select the vector, the feature network get; is an adaptive parameter; As the basic dictionary; is the first modality image; is the second modality image; is the time step; is the time step The predicted fusion image at this time.
3. The method for deep expansion multimodal image fusion network combined with diffusion model according to claim 1, characterized in that: The rectified prediction value of the filtered residual image is calculated according to the following formula: in, is the rectified prediction value of the residual image after filtering; The time step is The residual filter tensor when ; The time step is Noise scheduling when Represented by the noise prediction function The prediction noise generated.
4. The method for deep expansion multimodal image fusion network combined with diffusion model according to claim 3, characterized in that: The residual filter tensor is obtained according to the following formula: in, The time step is Dynamic filter when is the fused image.
5. The method for deep expansion multimodal image fusion network combined with diffusion model according to claim 1, characterized in that: The predicted fusion image at the time step is obtained according to the following formula: in, The time step is The predicted fusion image at the time is the convolution operator; The time step is Dynamic filter transpose when ; The time step is Dynamic filter when is the penalty parameter; in, is an identity matrix.
6. The method for deep expansion multimodal image fusion network combined with diffusion model according to claim 1, characterized in that: Training the Deep Fusion and Diffusion Prior Unfolding Network involves the following steps: Obtain fusion knowledge prior and generate Gaussian noise; Based on the fusion knowledge prior, the target search method is used to obtain the pseudo target fusion image; Replace the fused image with the pseudo target fused image and calculate the corresponding residual filter tensor; The residual filter tensor is used to obtain the corresponding rectified prediction value of the filtered residual image; Generate a predicted fusion image based on the rectified prediction value of the filtered residual image and calculate the loss function; Repeat the above steps until convergence.
7. The method for deep expansion multimodal image fusion network combined with diffusion model according to claim 6, characterized in that: The loss function Calculate according to the following steps: in, To rebuild losses; is the noise constraint of the denoising process, , is a sample of standard Gaussian distribution; To integrate pixel constraints with prior knowledge, , is the pseudo target fusion image.
8. The method for deep expansion multimodal image fusion network combined with diffusion model according to claim 7, characterized in that: The reconstruction loss Calculated according to the following formula: in, is the strength loss, ; is the gradient loss, ; is the structural similarity loss, ; To save the information, ; ▽ represents the gradient operator; SSIM(·) represents the structural similarity operation; , , , is the penalty parameter.
9. The method for deep expansion multimodal image fusion network combined with diffusion model according to claim 1, characterized in that: The denoising diffusion module calculates the residual filter tensor of the predicted fused image according to the following formula: in, The time step is The residual filter tensor when ; Always set to 0.
10. A deep-expansion multimodal image fusion network device combined with a diffusion model, characterized in that: include: A first module is configured to acquire a source image and a pure noise image, wherein the source image includes a first modality image and a second modality image; The second module is used to construct a deep fusion and diffusion prior expansion network, which includes a denoising diffusion module and a data consistency module; The data consistency module has a parameter generator, which generates adaptive parameters; The third module is used to filter the second modality image and input it into the denoising diffusion module, calculate the residual filter tensor of the pure noise image and input it into the denoising diffusion module; When , based on the residual filter tensor of the input filtered second modality image and the pure noise image, the denoising diffusion module uses the diffusion model noise estimation network to calculate the rectified prediction value of the filtered residual image, and uses the rectified prediction value of the filtered residual image as the input of the data consistency module; The fourth module is used to combine the rectified prediction value of the filtered residual image with the first modal image through the data consistency module and generate the predicted fusion image of the current time step after adaptive parameter processing; and calculate the residual filter tensor of the generated predicted fusion image as the next input of the denoising diffusion module; The fifth module is used to iterate until the time step is 0, and output the corresponding predicted fused image as the fused image.