Medical image fusion method based on diffusion model and Laplace technology

By combining diffusion model and Laplace technology, the problem of insufficient information integration in multimodal medical image fusion is solved, high-quality multimodal information fusion is achieved, and the visual effect and quantitative evaluation index of the image are improved.

CN120259099APending Publication Date: 2025-07-04WUXI NO 2 PEOPLES HOSPITAL +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510419184.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing multimodal medical image fusion method has limitations when fusing multimodal information, traditional methods lack significant feature attention, and deep learning methods have problems such as training instability and insufficient information integration.

Method used

The fusion method based on diffusion model and Laplace technology is adopted to pre-fusion images through the Laplace fusion module, and combined with the U-net network and the multi-objective loss function optimization model to achieve high-quality fusion of multimodal images.

Benefits of technology

The fusion quality and information integration capabilities of multimodal medical images have been improved, and the visual effect and quantitative evaluation indicators of the fusion images have been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259099A_ABST
    Figure CN120259099A_ABST
Patent Text Reader

Abstract

The invention provides a medical image fusion method based on diffusion and Laplacian technologies, which comprises the following steps of: firstly, converting a to-be-fused two-modal image, namely an enhanced CT image and a corresponding plain-scan CT image, into a gray level image with a specific size, pre-fusing by a Laplacian fusion module LP-F to obtain a pre-fused image with clear texture and clear edge, and then fusing the pre-fused image with the enhanced CT image and the corresponding plain-scan CT image into an image with a specific size; comprising the steps of constructing a Gaussian pyramid, generating a Laplacian pyramid, fusing detail information and the like. And inputting a source image pair and a pre-fused image into a diffusion model, adding noise to pure noise in a forward process, splicing a pure noise image channel, inputting the pure noise image channel into a U-net network, performing reverse denoising iteration to obtain a fused image, and defining a mixed loss function for optimization training. According to the method, the quality of the fused image and the multi-modal information integration capability are improved, and the method has remarkable superiority in key indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and particularly to a medical image fusion method based on a diffusion model and Laplacian technology. Background Art

[0002] In the field of medical imaging, multi-modal medical image fusion technology is of great significance. Due to the theoretical and technical limitations of optical imaging hardware, images obtained under single-sensor or fixed shooting conditions often lack comprehensive information. By fusing images obtained under different sensors or shooting conditions, the information richness of the images can be significantly enhanced. Multi-modal medical image fusion (MIF) can integrate information from multiple medical imaging modalities (such as MRI, CT, PET, etc.), which can not only improve the overall quality of the images, but also detect and locate lesion areas more comprehensively and accurately. It is one of the key directions in modern medical imaging research.

[0003] Existing image fusion methods are divided into traditional methods and deep learning-based methods. Traditional methods such as multi-scale transformation, spatial domain processing, sparse representation, and Laplacian Pyramid (LP), etc., usually use a unified feature representation, lack attention to the significant features of each modality image, and rely on manually designed fusion strategies, resulting in limited fusion effects and adaptability.

[0004] In recent years, the application of deep learning in the field of image fusion has developed rapidly, but different types of deep learning methods also have their own problems. For example, the method based on the autoencoder (AE) requires a time-consuming and energy-consuming two-stage training process; the loss function of the method based on the convolutional neural network (CNN) often needs to be manually designed and lacks robustness; the method based on the generative adversarial network (GAN) has problems of unstable generation and mode collapse; the method based on Transformer has high computational costs and insufficient information integration ability.

[0005] As a new generation of generative technology, the Diffusion Model (DM) originates from the Denoising Diffusion Probability Model (DDPM). Compared with autoencoders and variational autoencoders, DM demonstrates stronger ability to gradually approximate the target data distribution and generate high-fidelity images. Although DM is a generative model, its training process does not rely on adversarial mechanisms, thus effectively avoiding the mode collapse problem that may occur in the training of generative adversarial networks (GANs). In addition, the rigorous mathematical derivation of DM makes its loss function more reasonable than that used by convolutional neural networks (CNNs). In recent years, DM has been widely used in various computer vision tasks due to its powerful generative ability. Dif-fusion is an example of successfully applying the diffusion model to the image and video fusion (IVF) task, which uses DDPM to reconstruct the cascaded source images and designs an additional convolutional network to fuse the diffusion features of multiple channels. Subsequently, DDFM uses the unconditional DDPM as the backbone structure and introduces a likelihood correction module to assist the reverse fusion process. However, these DM-based image fusion methods still face some unresolved challenges. In Dif-fusion, the denoising process is limited to the cascaded source images, which essentially retains the characteristics of a single attribute; and DDFM also only processes a single attribute in the forward process, which limits the ability of the correction module to integrate other modal information. These limitations have, to a certain extent, hindered the full realization of the model's potential in multimodal fusion tasks. Summary of the invention

[0006] The present invention aims to overcome the shortcomings of the prior art in multimodal medical image fusion and provide a medical image fusion method based on diffusion and Laplace technology, which can effectively retain and fuse the texture information and features of multimodal source images and improve the visual effect and quantitative evaluation indicators of the fused image.

[0007] The technical solution adopted by the present invention to solve the technical problem is as follows.

[0008] A medical image fusion method based on diffusion and Laplace technology includes the following steps:

[0009] Step 1: Prepare a data set, convert each image pair to be fused into a grayscale image of a specified pixel size. Each image pair includes an enhanced CT image and a plain scan CT image. The converted image pairs to be fused are recorded as and

[0010] Step 2: Pair the images to be fused and Import the Laplace fusion module to complete the pre-fusion operation and obtain the pre-fused image, which is denoted as

[0011] Step 3: Input the image pair and the pre-fused image into the forward process of the diffusion model, and add Gaussian noise to the images respectively until they become pure noise images where t is the diffusion step size;

[0012] Step 4: Concatenate the pure noise images obtained in Step 3 on the channel dimension to obtain the noise-added image Input it into the U-net network of the diffusion model, and iteratively remove the noise in the reverse process. That is, in the reverse process, denoise the noise-added image in the forward process to generate the denoised fusion image In this way, the final fusion image is obtained through continuous iteration and sampling

[0013] Specifically, the method of the pre-fusion operation in Step 2 includes:

[0014] Step 2.1: Construct Gaussian pyramids G and for the i-th layer of κ,i respectively through Gaussian filtering and downsampling operations, where κ = 1, 2;

[0015] Step 2.2: Generate Laplacian pyramid images L κ,i containing details and edge information at different scales through the differential operation of the Gaussian pyramid;

[0016] Step 2.3: Fuse the generated L κ,i by dynamically adjusting the weight coefficient w κ to fuse the detailed information of different modality images, retain the key features, and generate the fused Laplacian pyramid image of the i-th layer;

[0017] Step 2.4: Reconstruct the image generated in Step 2.3 by upsampling and stacking layer by layer to restore the image to the original size image, that is, the pre-fused image

[0018] Specifically, the operation on the pure noise image in Step 4 includes:

[0019] Step 4.1: Concatenate the pure noise images of the three modalities obtained in the forward process of the diffusion model on the channel dimension to obtain the final noise-added image

[0020] Step 4.2: Input the noise-added image Import it into the U-net network as the input of the U-net network;

[0021] Step 4.3: In the reverse process, the U-net network denoises the noisy image in the forward process to generate a denoised fusion image

[0022] Step 4.4: By continuously iterating and sampling, obtain successively: Finally obtain the fusion image

[0023] Specifically, the loss function in the U-net network used in Step 4 is L = λ1L1 + λ2L2; λ1 and λ2 are hyperparameters used to balance the importance of different loss terms, L1 is the traditional loss function term, and L2 is the information entropy loss function term.

[0024] L1 = L fide + L contr

[0025] L2 = L gene + L local + L step

[0026] where L fide is the fidelity, L contr is the contrast loss, L gene is the global information entropy, L local is the local information entropy, and L step is the time-step information entropy.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] 1. The present invention proposes a multi-modal diffusion model framework, avoiding the training instability of the GAN network and the limitations of the diffusion model in dealing with a single denoising object;

[0029] 2. The present invention designs a Laplacian fusion module before forward noise addition in the traditional diffusion model to pre-fuse medical image pairs and obtain a pre-fused image with clear texture and edges;

[0030] 3. The present invention introduces an information entropy loss function, combines it with the traditional loss function, and constructs a multi-objective loss function to optimize the model training process.

[0031] In summary, the present invention addresses the deficiencies of existing image fusion methods and improves the quality of the fusion image and the multi-modal information integration ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is the overall flowchart of the present invention.

[0033] Figure 2 It is a flowchart of the pre-fusion operation of the present invention.

[0034] Figure 3 It is a schematic structural diagram of the denoising network U-net of the present invention.

[0035] Figure 4 It is a flowchart of the denoising iteration process of the present invention.

[0036] Figure 5 It is the image fusion result of the present invention. Specific implementation manners

[0037] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0038] The medical image fusion method based on diffusion and Laplace techniques according to the present invention, as Figure 1 shown, includes the following steps:

[0039] Step 1, data preparation.

[0040] Make a data set. Each pair of images (source image pair) includes one enhanced CT and one plain CT. Convert the source image pair (the total number of image pairs in the embodiment is 800) into grayscale images with a size of 256*256 pixels, and obtain the image pairs to be fused and and respectively represent the enhanced CT image and the plain CT image.

[0041] Step 2, import the image pairs to be fused and into the Laplace fusion module (LP-F), and use the Laplace fusion module to pre-fuse the source image pair to obtain a pre-fused image, denoted as

[0042] The method of the pre-fusion operation in Step 2 includes:

[0043] Step 2.1, as Figure 2 shown, construct Gaussian pyramids for the i-th layer of the source image pair and respectively through Gaussian filtering and downsampling, where i = 1, 2, 3,..., n; n represents the number of layers of the given Gaussian pyramid, and n = 5 in this embodiment; G κ,i represents the Gaussian pyramid constructed for the i-th layer, and κ represents the image category, κ = 1, 2.

[0044] Step 2.2, through the Gaussian pyramids G κ,iPerform a differential operation to generate the Laplacian pyramid image L of details and edge information at different scales κ,i .

[0045] Equations (1) and (2) describe the generation process of the Laplacian pyramid:

[0046]

[0047] L κ,i = G κ,i - Resize(Upsample(G κ,i+1 )) (2)

[0048] Among them, Equation (1) represents the construction of the Gaussian pyramid in Step 2.1, and Equation (2) represents the differential operation described in Step 2.2. Upsample represents the upsampling operation; Resize represents the resampling operation, which is used to scale the size of the upsampled image so that the size of the upsampled image is the same as that of the Gaussian pyramid image of the current layer.

[0049] Step 2.3. Fusion is performed on the Laplacian pyramid image L generated in Step 2.2 κ,i by dynamically adjusting the weight coefficient w κ to fuse the detail information of different modality images, retain the key features, and obtain the fused Laplacian pyramid image L at the i-th layer fusion,i .

[0050] As shown in Equation (3), a weighted fusion operation is performed on the separately generated L 1,i and L 2,i in Step 2.2.

[0051]

[0052] Among them, w κ is the weight coefficient of category κ.

[0053] Step 2.4. Then, the generated L fusion,i is reconstructed. By upsampling and stacking layer by layer, it is restored to an image of the original size, that is, when the images of all layers are reconstructed, it is the final pre-fusion image.

[0054] Equation (4) describes the reconstruction process of the Laplacian pyramid:

[0055]

[0056] Among them, L n is the Laplacian pyramid image of the highest layer, and L fusion,n-i is the fused Laplacian pyramid image of the (n - 1)-th layer.

[0057] Step 3: Co - input the source image pair and the pre - fused image into the forward process of the diffusion model. Continuously add Gaussian noise to the image until the image becomes a pure noise image where \(t\) is the diffusion step; finally, when \(t = T\) (\(T\) is the set diffusion step, \(T = 1000\) in the embodiment), through the forward process, there are changes: which is the output of the forward process.

[0058] Equations (5) - (8) describe the process of obtaining the pure noise image and the definition of the noisy image:

[0059]

[0060] where \(\epsilon\) is Gaussian noise and \(\epsilon\sim N(0, I)\); \(\beta\) t \(\in(0, 1)\) is the parameter of the Gaussian distribution; \(\alpha\) t \(=(1 - \beta\) t );

[0061]

[0062] Step 4: As shown Figure 4 in the figure, splice the pure noise image obtained in Step 3 on the channels to obtain the noisy image and input it into the U - net network of the diffusion model; in the reverse process, iteratively remove the noise to finally obtain the fused image. The specific steps include:

[0063] Step 4.1: Splice the pure noise images of the three modalities obtained in the forward process of the diffusion model on the channel dimension to obtain the final noisy image

[0064] Step 4.2: Take the noisy image as the input of the U - net network and import it into the denoising network;

[0065] Step 4.3: In the reverse process, the U - net network denoises the noisy image in the forward process to generate the denoised fused image

[0066] Step 4.4: Through continuous iteration and sampling, successively obtain: Finally, obtain the fused image

[0067] the noisy image \(x\) at time step \(t\) tThe noise estimation function obtained through model training is denoted as Remove the estimated noise from the noisy image to obtain the denoised image estimate Use the denoised image and newly sampled noise to generate the image x at the previous time step t-1 ; This process starts from time step T and iterates repeatedly until t = 0, updating the image at each step

[0068] The image update process is as shown in Equations (9) and (10):

[0069]

[0070] Figure 3 This is the structure of the denoising network U-net adopted by the present invention. The U-Net network has an encoder and a decoder, and can effectively capture the context information of the image and generate high-precision results

[0071] The following is a detailed explanation of the important parts in the U-Net network structure

[0072] 1) Encoder: The encoder part is responsible for extracting the features of the image, gradually reducing the spatial dimension of the image (i.e., downsampling) through a series of convolutional layers and pooling layers, while increasing the number of channels of the feature map. The encoder usually consists of multiple blocks, and each block contains the following operations

[0073] Double convolutional layer: Use a 3x3 convolutional kernel, followed by a ReLU activation function, to extract local features

[0074] Pooling layer: Use 2x2 max pooling to halve the spatial size of the image while retaining the most important features

[0075] 2) Decoder: The decoder part is responsible for gradually restoring the feature map extracted by the encoder to the resolution of the original image and generating the result. The decoder also consists of multiple blocks, and each block contains the following operations

[0076] Upsampling layer: Use transposed convolution to enlarge the spatial size of the feature map

[0077] Skip connection: Concatenate the feature map of the corresponding layer in the encoder with the feature map in the decoder to retain more spatial detail information

[0078] Convolutional layer: Similar to the encoder, use a 3x3 convolutional kernel and a ReLU activation function to further process the feature map

[0079] 3) Loss function: The present invention defines a hybrid loss function L = λ1L1 + λ2L2. λ i (i = 1, 2) is a hyperparameter used to balance the importance of different loss terms, and L1 is a traditional loss function term, consisting of the fidelity Lfide and the contrast loss L contr to define; L2 is the information entropy loss function term, which is composed of the global information entropy L gene , the local information entropy L local and the time-step information entropy L step defined. Through the constraint of the new loss function, the reverse process of the diffusion model will learn how to use pure noise to approximate the feature information of all images to the greatest extent, so as to obtain high-quality denoised fusion images.

[0080] Formulas (11)-(17) describe the construction part and process of the loss function:

[0081] L1 = L fide + L contr (11)

[0082] L2 = L gene + L local + L step (12)

[0083]

[0084] where ε is Gaussian noise, is the noise estimation function; f contr represents the contrast function, U represents the U-net network, is the fusion image, is the noisy image; p(x i ) is the probability distribution of the pixel value x i , p(x i , region) is the probability distribution of x i in the local region, and p(x t ) is the probability distribution at the time step t.

[0085] So far, the fusion of two-modal CT images is completed.

[0086] Figure 5 is the fusion result of the image, Figure 5 provides the fusion results of 3 groups of different images. According to Figure 5 the image fusion results shown, it can be seen that the method proposed by the present invention has achieved good fusion effects in terms of extracting texture details, color fidelity, and image contrast.

[0087] To verify the feasibility of the fusion scheme, the researchers compared the method proposed by the present invention (abbreviated as DMF-LP) with 5 other medical image fusion methods published in recent years, including: CNN[1], RCGAN[2], U2Fusion[3], Dif-fusion[4], DDFM[5].

[0088] From relevant literature:

[0089] [1] Liu Y, Chen X, Cheng J, et al. A medical image fusion method based on convolutional neural networks[C] / / 2017 20th international conference on information fusion (Fusion). IEEE, 2017: 1 - 7.

[0090] [2] Li Q, Lu L, Li Z, et al. Coupled GAN with relativistic discriminators for infrared and visible images fusion[J]. IEEE Sensors Journal, 2019, 21(6): 7458 - 7467.

[0091] [3] Xu H, Ma J, Jiang J, et al. U2Fusion: A unified unsupervised image fusion network[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 44(1): 502 - 518.

[0092] [4] Yue J, Fang L, Xia S, et al. Dif - fusion: Towards high color fidelity in infrared and visible image fusion with diffusion models[J]. IEEE Transactions on Image Processing, 2023.

[0093] [5] Zhao Z, Bai H, Zhu Y, et al. DDFM: denoising diffusion model for multi - modality image fusion[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2023: 8082 - 8093.

[0094] To further evaluate the fused images, objective evaluation indicators are used below to provide a more convincing objective evaluation. The evaluation indicators include structural similarity index (SSIM), peak signal-to-noise ratio (PSNR), correlation coefficient (CC), and mutual information (MI). Table 1 shows the quantitative evaluation results of the fused images, using different evaluation indicators to evaluate the fused images obtained by all fusion methods, and the maximum value is bolded.

[0095] Table 1 Quantitative evaluation results of fused images

[0096] Methods SSIM PSNR CC MI CNN

[24] 0.869±0.077 32.668±1.755 0.941±0.012 1.154±0.019 RCGAN

[15] 0.844±0.036 32.879±1.927 0.945±0.043 1.301±0.024 U2Fusion

[12] 0.836±0.106 32.881±1.912 0.916±0.002 1.278±0.023 Dif-fusion

[22] 0.893±0.053 33.883±0.551 0.976±0.003 1.273±0.020 DDFM

[23] 0.897±0.063 34.166±1.093 0.976±0.006 1.159±0.021 DMF-LP 0.912±0.011 34.416±0.568 0.980±0.001 1.274±0.013

[0097] As can be seen from Table 1, the mean and standard deviation of the DMF-LP model in the SSIM index are 0.912 and 0.011, respectively, which is significantly better than other models. This shows that DMF-LP performs well in preserving image structural information and is more robust to input changes. In terms of the PSNR index, the mean and standard deviation of the DMF-LP model are 34.416 and 0.568, respectively, which are also higher than other comparison models. This shows that the DMF-LP model has a better balance between suppressing noise and maintaining details, especially in high dynamic range scenes, it can more stably output high-quality fusion results. For the CC index, the mean of DMF-LP is 0.980 and the standard deviation is 0.001, ranking first among all comparison models. This highlights the superior performance of the model in maintaining the correlation between input images. In terms of MI index, the mean value of DMF-LP is 1.274, slightly lower than the RCGAN model (1.301) and U2Fusion model (1.278), but its standard deviation (0.013) is the lowest in the table and is at the same level as mainstream methods such as U2Fusion and Diffusion. This shows that DMF-LP has achieved a better balance between retaining the amount of source image information and avoiding over-fusion, effectively preventing feature redundancy.

[0098] In general, compared with the mainstream methods in recent years, DMF-LP has reached the SOTA level for the first time in the three core indicators of SSIM, PSNR, and CC, and has shown significantly better robustness in the standard deviation dimension. Especially in the context of the significant progress made in the latest diffusion model-based methods such as Diffusion and DDFM, this model can still achieve a PSNR improvement of 0.019-0.55dB, which fully verifies the advanced nature of the proposed architecture in complex feature fusion tasks.

[0099] Table 2 lists the fusion results of the ablation experiments related to the DMF-LP method proposed in the present invention. × indicates not added, and √ indicates added; LP represents the Laplacian fusion module, and L2 represents the L2 loss function; Ⅰ, Ⅱ, and Ⅲ respectively represent the corresponding three groups of ablation experiments, and the bold indicates the highest value obtained.

[0100] Quantitative evaluation results of the ablation experiments in Table 2

[0101]

[0102] As can be seen from Table 2, removing the Laplacian fusion module LP-F or the L2 loss function shows that the results cannot reach the level of the complete framework in terms of quantitative indicators, which also verifies the collaborative contribution of these modules to improving the overall performance from the side.

[0103] The parts not described in detail in the present invention are applicable to the prior art.

[0104] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A medical image fusion method based on diffusion and Laplace techniques, characterized in that, Including the following steps: Step 1: Create a dataset. Convert each pair of images to be fused into grayscale images of a specified pixel size. Each image pair includes an enhanced CT image and a plain CT image. The converted image pairs to be fused are denoted as and Step 2: The image pair to be fused and are imported into the Laplacian fusion module to complete the pre-fusion operation and obtain a pre-fused image, which is denoted as Step 3: Input the image pair and the pre-fused image into the forward process of the diffusion model together, and add Gaussian noise to the images respectively until the images become pure noise images where t is the diffusion step size; Step 4: Take the pure noise image obtained in Step 3 Stitch it on the channels to obtain the noisy image Input it into the U-net network of the diffusion model, and iteratively remove the noise during the reverse process, that is, during the reverse process, for the noisy image in the forward process Perform denoising to generate the denoised fusion image In this way, the final fusion image is obtained through continuous iteration and sampling 2. The medical image fusion method based on diffusion and Laplace technology according to claim 1, characterized in that, The method of the pre-fusion operation in step 2 includes: Step 2.1: Through Gaussian filtering and downsampling operations, for and respectively construct Gaussian pyramids G for the i-th layer of κ,i , κ = 1, 2; Step 2.2: Generate the Laplacian pyramid image L of details and edge information at different scales through the differential operation of the Gaussian pyramid κ,i ; Step 2.3: Perform fusion on the generated L κ,i by dynamically adjusting the weight coefficient w κ to fuse the detailed information of different modality images, retain the key features, and generate the Laplacian pyramid image after fusion at the i-th layer; Step 2.4: Reconstruct the image generated in Step 2.

3. Through layer-by-layer upsampling and superposition, restore the image to the original size image, that is, the pre-fusion image 3. A medical image fusion method based on diffusion and Laplace techniques according to claim 1, characterized in that The operation on the pure noise image in step 4 includes: Step 4.1: Concatenate the pure noise images of the three modalities obtained in the forward process of the diffusion model on the channel dimension to obtain the final noisy image Step 4.2: Import the noisy image as the input of the U-net network into the U-net network; Step 4.3: In the reverse process, the U-net network denoises the noisy image in the forward process to generate a denoised fused image Step 4.4: Through continuous iteration and sampling, obtain successively: Finally obtain the fused image 4. A medical image fusion method based on diffusion and Laplace techniques according to claim 1, characterized in that, The loss function in the U-net network used in step 4 is L = λ1L1 + λ2L2; λ1 and λ2 are hyperparameters used to balance the importance of different loss terms, L1 is the traditional loss function term, and L2 is the information entropy loss function term. L1 = L fide + L contr L2 = L gene + L local + L step Among them, L fide is the fidelity, L contr is the contrast loss, L gene is the global information entropy, L local is the local information entropy, L step The time-step information entropy.