Unsupervised medical image fusion method based on light-weight gemini GAN network
Through the lightweight Gemini GAN network, learning the basic and detailed layers of multimodal medical images is solved, and the problems of high computing resources and unclear image display in the prior art are achieved, efficient and clear medical images fusion is achieved, suitable for resource-constrained environments and improving the reliability of clinical diagnosis.
Patent Information
- Application Number
- CN202510356119.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-22
Smart Images

Figure BDA0005327393930000056 
Figure BDA0005327393930000058 
Figure BDA0005327393930000061
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image fusion, and particularly to an unsupervised medical image fusion method based on a lightweight dual GAN network. Background Art
[0002] The GAN network automatically learns the latent features of data through adversarial training, that is, the adversarial training between the generator and the discriminator enables the generator to gradually generate high-quality synthetic data. This mechanism has been widely applied in fields such as image generation, style transfer, and data augmentation. In addition, the GAN network performs particularly well in multi-modal data fusion, being able to automatically identify and fuse the key information of different modalities, effectively avoiding the loss or confusion of information.
[0003] Currently, the conditional generative adversarial network is used, which introduces conditional information (usually real samples or additional label data) on the basis of the traditional GAN network to control the characteristics of the generated data and make it more in line with specific requirements. The conditional generative adversarial network consists of two main parts: the generator and the discriminator. In tasks such as medical image fusion, the task of the generator is to generate realistic fake samples (fusion images) according to the input conditional information. The discriminator is responsible for discriminating whether the input is a real sample or a fake sample generated by the generator. In short, the ultimate goal of the discriminator is to improve the ability to identify real and fake samples, so that the generator can be continuously optimized during the training process, thereby generating more and more realistic samples and finally achieving high-quality image fusion results.
[0004] The feature complexity of multi-modal images makes it difficult for a single discriminator to accurately identify the feature information of multiple modal samples at the same time. To solve this problem, Zhang et al. proposed an improved generative adversarial network model - GAN-FM. This model contains a generator and two discriminators, and through their collaborative action, it can effectively fuse the features of multi-modal images with different resolutions, thereby improving the performance of the network in processing multi-modal images. However, its complex working mechanism and excessive parameters lead to large consumption of computing resources and many training parameters, and there is an urgent need for lightweighting. Summary of the Invention
[0005] In view of this, the present invention proposes an unsupervised medical image fusion method based on a lightweight dual GAN network, which learns the base layer and the detail layer of multiple modal medical images respectively through the GAN network, generates a fusion image based on the fusion result of the detail layer and the base layer, realizes the efficient fusion of image information at two levels, and has the characteristics of low computational complexity and few training parameters.
[0006] The technical solution adopted by the embodiments of the present invention to solve its technical problems is:
[0007] An unsupervised medical image fusion method based on a lightweight dual GAN network, comprising:
[0008] Step S1, constructing an unsupervised medical image fusion model based on a lightweight dual GAN network, the unsupervised medical image fusion model including a preprocessing unit, a basic layer GAN network fusion unit, a detail layer GAN network fusion unit, a non-linear enhancement unit, and a fused image output unit;
[0009] Step S2, collecting computed tomography (CT) images and magnetic resonance imaging (MRI) images and establishing a multi-modal medical source image database for training the model, and performing data augmentation on each source image by means of image rotation; wherein, the types of CT images include three-channel pseudo-color medical images and single-channel grayscale medical images;
[0010] Step S3, calling the source images in the source image database to train the unsupervised medical image fusion model, the preprocessing unit preprocessing and image decomposing the CT images and the MRI images, inputting the data of two detail layers obtained by the image decomposition into the detail layer GAN network fusion unit for fusion to obtain a detail layer fusion result, inputting the data of two basic layers obtained by the image decomposition into the basic layer GAN network fusion unit for fusion to obtain a basic layer fusion result, performing fused image reconstruction by using the detail layer fusion result and the basic layer fusion result, and obtaining a fused image of the source images after the initial fused result is subjected to conversion processing by the non-linear enhancement unit, and outputting by the fused image output unit; wherein, the weight parameters of the basic layer GAN network fusion unit and the detail layer GAN network fusion unit are optimized by means of a loss function;
[0011] Step S4, performing medical image fusion by using the trained unsupervised medical image fusion model.
[0012] Preferably, CT images belong to single-channel medical images, and SPECT images and PET images both belong to three-channel pseudo-color medical images;
[0013] The preprocessing unit: for images of the single-channel grayscale medical image type, no preprocessing is required; for images of the three-channel pseudo-color medical image type, it is necessary to convert from the RGB space to the YCbCr color space, and then extract the Y luminance component for subsequent Gaussian filtering decomposition;
[0014] Using Gaussian filtering to decompose the preprocessed CT image I1 and MRI image I2, obtaining the basic layer BL1 and detail layer DL1 of I1, and the basic layer BL2 and detail layer DL2 of I2;
[0015] The base layer BL1 and the base layer BL2 are fed into the base layer GAN network fusion unit;
[0016] The detail layer DL1 and the detail layer DL2 are fed into the detail layer GAN network fusion unit.
[0017] Preferably, the base layer GAN network fusion unit is a generative adversarial network composed of a discriminator Discriminator Generator G BL The data processing process is as follows:
[0018] The generator G BL Performs a fusion process on the base layer BL1 and the base layer BL2 to obtain the base layer fusion result F BL ;
[0019] The base layer fusion result F BL And the base layer BL1 are input into the discriminator For discrimination, and the obtained discrimination result adjusts the generator G BL ;
[0020] The base layer fusion result F BL And the base layer BL2 are input into the discriminator For discrimination, and the obtained discrimination result adjusts the generator G BL ;
[0021] Adversarial training is performed in the form of alternating training between the discriminator and the generator.
[0022] Preferably, the detail layer GAN network fusion unit is a generative adversarial network composed of a discriminator Discriminator Generator G DL The data processing process is as follows:
[0023] The generator G DL Performs a fusion process on the detail layer DL1 and the detail layer DL2 to obtain the detail layer fusion result F DL ;
[0024] The detail layer fusion result F DL And the detail layer DL1 are input into the discriminator For discrimination, and the obtained discrimination result adjusts the generator G DL ;
[0025] The detail layer fusion result F DL And the detail layer DL2 are input into the discriminator For discrimination, and the obtained discrimination result adjusts the generator GDL Adjust;
[0026] Adversarial training is performed in the form of alternating training between the discriminator and the generator.
[0027] Preferably, the generator G DL , D BL Adopt the same framework structure, which is successively composed of an input layer, an encoder, a decoder and an output layer, where:
[0028] The input image is first processed by a 3×3 standard convolutional layer, and the obtained feature map is input into the encoder;
[0029] The encoder is successively composed of 4 CSB modules in series. Each CSB module is composed of a 3×3 standard convolutional layer and a Sandglass Block in series. The Sandglass Block is successively composed of a 3×3 depthwise separable convolution Dwise, a 1×1 dimensionality reduction standard convolutional kernel, a 1×1 dimensionality increase standard convolutional kernel, and a 3×3 depthwise separable convolution Dwise in series;
[0030] The decoder is composed of 3 CONV blocks, and each CONV block is a standard convolutional layer using a 3×3 convolutional kernel;
[0031] The feature map is successively processed by CSB1, CSB2, CSB3, and CSB4. CONV3 receives the downsampling features from CSB1, the downsampling features from CSB2, the skip connection features from CSB3, and the upsampling features from CSB4. CONV2 receives the downsampling features from CSB1, the skip connection features from CSB2, the upsampling features from CONV3, and the upsampling features from CSB4. CONV1 receives the skip connection features from CSB1, the upsampling features from CONV2, the upsampling features from CONV3, and the upsampling features from CSB4;
[0032] The output of CONV1 is sent to the output layer, and after being successively processed by two 3×3 standard convolutional layers, the output result of the generator is obtained.
[0033] Preferably, the discriminator Adopt the same framework structure, which is composed of 5 convolutional blocks connected in sequence, where: the first convolutional kernel is 3×3, the stride is 2, a ReLU activation layer is configured, and the output feature map size is 128×128×16; the second, third, and fourth layers are all convolutional kernels of 3×3, the stride is 2, a BN layer and a ReLU activation layer are configured, and the output feature map sizes are 64×64×32, 32×32×64, and 16×16×128 in sequence; the fifth convolutional kernel is 3×3, the stride is 1, a batch BN layer and a Tanh activation layer, and the output feature map size is 16×16×1; the output result represents the probability that the input image is a real image.
[0034] Preferably, the total loss function Loss of the detailed layer GAN network fusion unit total_DL is expressed as:
[0035]
[0036] where is the loss function of the discriminator , and is the loss function of the discriminator , and its expression is:
[0037]
[0038] Loss(G DL ) is the loss function of the generator G DL , and its expression is:
[0039] Loss(G DL ) = Loss adv_DL + λLoss con_DL
[0040]
[0041] Loss con_DL = α DL Loss grad_DL + β DL Loss ssim_DL + γ DL Loss in_DL
[0042] where λ, α DL , β DL , γ DL are weight parameters for controlling the detailed layer GAN network fusion unit; E represents the mathematical expectation;
[0043] Loss in_DL is the intensity loss function of the detailed layer, and its expression is:
[0044]
[0045]
[0046] where represents the norm, L is the gray value of 256, p l is the gray level probability, is the threshold of saliency, Sign(·) represents the sign function, and represent the significance measurements of DL1 and DL2 respectively, and represent the abundance measurements of DL1 and DL2 respectively, and represent the surface level measurements of DL1 and DL2 respectively;
[0047] Loss ssim_DL is the structural loss function for the detail layer, and the expression is:
[0048] Loss ssim_DL = ε DL (1 - SSIM(F DL , DL1))+(1 - ε DL )(1 - SSIM(F DL , DL2))
[0049] where ε DL is the parameter for controlling the weight; SSIM(·) is the structural similarity formula;
[0050] Loss grad_DL is the enhanced gradient loss function, and the expression is:
[0051]
[0052] where |·|1 represents L1 regularization, represents the Laplacian gradient operator, and a DL is the parameter for controlling the weight; is the loss function term for selecting the maximum gradient.
[0053] Preferably, the expression of the total loss function Loss total_BL of the basic layer GAN network fusion unit is:
[0054]
[0055] where is the loss function of the discriminator , is the loss function of the discriminator , and the expression is:
[0056]
[0057] Loss(G BL ) is the loss function of the basic layer generator, and the expression is:
[0058] Loss(G BL ) = Loss adv_BL + λLoss con_BL
[0059]
[0060] Loss con_BL = α BL Loss grad_BL + β BL Loss ssim_BL + γ BL Loss in_BL
[0061] where λ, α BL , β BL , γ BL are weight parameters for controlling the fusion unit of the basic layer GAN network; E represents the mathematical expectation; Loss in_BL is the intensity loss function of the basic layer, and its expression is:
[0062]
[0063] where represents the norm, L is the gray value of 256, p l is the gray level probability, is the threshold of saliency, Sign(·) represents the sign function, and represent the saliency measurements of BL1 and BL2 respectively, and represent the abundance measurements of BL1 and BL2 respectively, and represent the surface level measurements of BL1 and BL2 respectively;
[0064] Loss ssim_BL is the structural loss function of the basic layer, and its expression is:
[0065] Loss ssim_BL = ε BL (1 - SSIM(F BL , BL1))+(1 - ε BL )(1 - SSIM(F BL , BL2))
[0066] where ε BL is the parameter for controlling the weight; SSIM(·) is the structural similarity formula;
[0067] Loss grad_BL is the enhanced gradient loss function, and its expression is:
[0068]
[0069] Among them, |·|1 represents L1 regularization, represents the Laplacian gradient operator, and a BL is a parameter for controlling the weight; is the loss function term for selecting the maximum gradient.
[0070] Preferably, the model parameter adjustment process is as follows: In the unsupervised medical image fusion model, λ, α BL , β BL , γ BL , ε BL , a BL , α DL , β DL , γ DL , ε DL , a DL are used as hyperparameters. Among them, λ is used to control the balance between the content loss function and the adversarial loss function, λ = 100, and α BL , β BL , γ BL , α DL , β DL , γ DL are used to control the strength of the display of texture detail information in the fused image, and ε BL , ε DL are used to adjust the proportion of the image content in the two modalities, ε BL ∈[0, 1], ε DL ∈[0, 1], and a BL , a DL are coefficients for controlling the enhancement degree of the enhancement loss function, and a BL ∈[1, 5], and a DL ∈[1, 5];
[0071] During the hyperparameter tuning process, the initial learning rate is set to 2×10 -4 , the decay rate is set to 0.9, and Adam and RMSProp are used as optimizers to train the discriminator and the generator respectively;
[0072] When the network inputs the detail layer, λ = 100, and α DL = 1.5, β DL = 3, γ DL = 3, ε DL = 0.3, and a DL = 1.2;
[0073] When the network inputs the base layer, λ = 100, and α BL = 1.2, β BL = 8, γ BL = 5, ε BL = 0.5, and a BL= 1.
[0074] Preferably, the fusion image reconstruction using the detail layer fusion result and the base layer fusion result is specifically as follows:
[0075] Add the fused detail layer and base layer images to obtain the initial fusion result;
[0076] The initial fusion result is processed by the non-linear enhancement unit to normalize the pixel values of the initial fusion result to the range of [0, 255], obtaining the fusion image of the source image; then output, specifically: if it is the fusion of a three-channel pseudo-color medical image and an MRI image, the fusion image of the source image is obtained after restoring to the RGB space; if it is the fusion of a single-channel grayscale medical image and an MRI image, it is directly used as the fusion image of the source image.
[0077] As can be seen from the above technical solutions, the unsupervised medical image fusion method based on the lightweight dual GAN network provided by the embodiments of the present invention proposes an unsupervised medical image fusion framework combining two GAN networks, and uses computed tomography-like images and MRI images for training. The present invention uses the GAN network model to separately learn the detail layer and the base layer of multiple modal samples, improves the generation ability of the generator at different layers, and then obtains a fusion image based on the detail layer fusion result and the base layer fusion result, realizing the effective fusion of information at two levels, which can improve the clarity and quality of the fusion image, and has the characteristics of low computational complexity and few training parameters. In addition, in the results of many existing methods, edge diffusion often occurs in the edges of high-contrast regions, resulting in the contamination of background textures. The present invention adopts an enhanced gradient loss function to strengthen the detail textures and edges of the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 It is the overall framework diagram of the unsupervised medical image fusion model based on the lightweight GAN network.
[0079] Figure 2 It is the schematic diagram of the generator structure in the GAN network fusion unit.
[0080] Figure 3 It is the specific example diagram of the generator structure in the GAN network fusion unit.
[0081] Figure 4 It is the schematic diagram of the discriminator structure in the GAN network fusion unit.
[0082] Figure 5 It is the schematic diagram for comparing the fusion results of PET / MRI images.
[0083] Figure 6 It is the schematic diagram for comparing the fusion results of CT / MRI images.
[0084] Figure 7 It is a schematic diagram for comparing the SPECT / MRI image fusion results. Specific implementation manners
[0085] The technical solutions and technical effects of the present invention will be further elaborated in detail below in conjunction with the accompanying drawings of the present invention.
[0086] The present invention provides an unsupervised medical image fusion method based on a lightweight GAN network, which can be used in tasks of multi-modal medical image fusion such as CT / MRI, PET / MRI, SPECT / MRI, etc.
[0087] Regarding the data set for training the model, 80% of it can be selected from the common brain disease data sets for training, and the remaining 20% is used for testing. The spatial resolution of the experimental source images is all 256×256, and the data set is enhanced by means of image rotation.
[0088] The specific implementation process of the technical solution of the present invention is as follows, and the overall framework diagram is as Figure 1 shown:
[0089] Step S1, construct an unsupervised medical image fusion model based on a lightweight dual GAN network. The unsupervised medical image fusion model includes a preprocessing unit, a basic layer GAN network fusion unit, a detail layer GAN network fusion unit, a non-linear enhancement unit, and a fused image output unit, wherein the two GAN network fusion units optimize the model by using a loss function;
[0090] Step S2, collect computer tomography images and MRI images and establish a multi-modal medical source image database for training the model, and enhance the data of each source image by means of image rotation; wherein, the types of computer tomography images include three-channel pseudo-color medical images and single-channel grayscale medical images;
[0091] Step S3: Call the source images in the source image database to train an unsupervised medical image fusion model. The preprocessing unit preprocesses and decomposes computer tomography (CT) images and magnetic resonance imaging (MRI) images. The data of the two detail layers obtained by image decomposition are input into the detail layer GAN network fusion unit for fusion to obtain the detail layer fusion result. The data of the two base layers obtained by image decomposition are input into the base layer GAN network fusion unit for fusion to obtain the base layer fusion result. The initial fusion result is reconstructed using the detail layer fusion result and the base layer fusion result. The obtained initial fusion result is processed by the nonlinear enhancement unit to obtain the fusion image of the source images, which is output by the fusion image output unit. Among them, the weight parameters of the base layer GAN network fusion unit and the detail layer GAN network fusion unit are optimized through a loss function, enabling the GAN network to achieve dynamic balance between the generator and the discriminator, generate high-quality and diverse samples, and make the losses of the generator and the discriminator tend to be stable during the training process, and achieve ideal effects in both objective and subjective evaluations.
[0092] Step S4: Use the trained unsupervised medical image fusion model for medical image fusion.
[0093] CT images belong to single-channel medical images, while SPECT images and PET images both belong to three-channel pseudo-color medical images.
[0094] For the preprocessing unit of the unsupervised medical image fusion model, single-channel grayscale medical image type images do not require preprocessing (color space conversion); three-channel pseudo-color medical image type images need to be converted from the RGB space to the YCbCr color space, and then the Y luminance component is extracted for subsequent Gaussian filter decomposition. The preprocessed CT images I1 and MRI images I2 (both registered images) are decomposed using Gaussian filters to obtain the base layer BL1 and detail layer DL1 of I1, and the base layer BL2 and detail layer DL2 of I2 (here the size of the Gaussian filter is 3×3). The decomposition process is shown in Equations (1) to (3):
[0095]
[0096] BL1 = I1 * G, BL2 = I2 * G (2)
[0097] After obtaining the base layers BL1 and BL2 through the Gaussian filter, the input images are subtracted from the corresponding base layers to obtain the detail layers DL1 and DL2:
[0098] DL1 = I1 - BL1, DL2 = I2 - BL2 (3)
[0099] The base layer BL1 and the base layer BL2 are fed into the base layer GAN network fusion unit;
[0100] The detail layer DL1 and the detail layer DL2 are fed into the detail layer GAN network fusion unit.
[0101] The basic layer GAN network fusion unit of the unsupervised medical image fusion model is a generative adversarial network composed of a discriminator Discriminator Generator G BL The data processing process is as follows: Generator G BL Performs a fusion process on the basic layer BL1 and the basic layer BL2 to obtain the basic layer fusion result F BL ; The basic layer fusion result F BL And the basic layer BL1 are input into the discriminator For discrimination, and the obtained discrimination result adjusts the generator G BL ; The basic layer fusion result F BL And the basic layer BL2 are input into the discriminator For discrimination, and the obtained discrimination result adjusts the generator G BL ; Adversarial training is carried out in the form of alternating training between the discriminator and the generator.
[0102] The detail layer GAN network fusion unit of the unsupervised medical image fusion model is a generative adversarial network composed of a discriminator Discriminator Generator G DL The data processing process is as follows: Generator G DL Performs a fusion process on the detail layer DL1 and the detail layer DL2 to obtain the detail layer fusion result F DL ; The detail layer fusion result F DL And the detail layer DL1 are input into the discriminator For discrimination, and the obtained discrimination result adjusts the generator G DL ; The detail layer fusion result F DL And the detail layer DL2 are input into the discriminator For discrimination, and the obtained discrimination result adjusts the generator G DL ; Adversarial training is carried out in the form of alternating training between the discriminator and the generator.
[0103] Generator G DL 、D BL Adopt the same lightweight generator framework structure. As shown in Figure 2 , It consists of an input layer, an encoder, a decoder, and an output layer in sequence, where:
[0104] The input image is first processed by a 3×3 standard convolutional layer, and the resulting feature map is input into the encoder. The encoder is composed of 4 CSB (Convolution-Sandglass Block) modules connected in series in sequence. Each CSB module is composed of a 3×3 standard convolutional layer and a Sandglass Block connected in series. The Sandglass Block is composed of a 3×3 depthwise separable convolution (Dwise), a 1×1 dimensionality reduction standard convolutional kernel, a 1×1 dimensionality increase standard convolutional kernel, and a 3×3 depthwise separable convolution (Dwise) connected in series. Among them, the reduction of the number of channels through 1×1 convolution can effectively reduce the computational complexity. A Bottleneck bottleneck layer is set between the 1×1 dimensionality reduction standard convolutional kernel and the 1×1 dimensionality increase standard convolutional kernel to further compress the dimension of the feature map. The number of channels is increased through 1×1 convolution to restore the dimension of the feature map. The 1×1 dimensionality increase standard convolutional kernel increases the number of channels through 1×1 convolution to restore the dimension of the feature map. Then, a 3×3 depthwise separable convolution operation is performed again to further extract features. The parameters of the Sandglass Block can refer to Table 1:
[0105]
[0106] Table 1
[0107] The decoder is composed of 3 CONV blocks, and each CONV block is a standard convolutional layer using a 3×3 convolutional kernel. The feature map is processed by CSB1, CSB2, CSB3, and CSB4 in sequence. CONV3 receives the downsampled features from CSB1, the downsampled features from CSB2, the skip connection features from CSB3, and the upsampled features from CSB4. CONV2 receives the downsampled features from CSB1, the skip connection features from CSB2, the upsampled features from CONV3, and the upsampled features from CSB4. CONV1 receives the skip connection features from CSB1, the upsampled features from CONV2, the upsampled features from CONV3, and the upsampled features from CSB4. CONV1 outputs to the output layer, and after being processed by two 3×3 standard convolutional layers in sequence, the fusion result is obtained. The output layer is used for the final image reconstruction.
[0108] Discriminator Adopt the same framework structure, which consists of 5 convolutional blocks connected in sequence. Among them: the first convolutional kernel is 3×3, the stride is 2, and the ReLU activation layer is configured, and the output feature map size is 128×128×16; the second, third, and fourth layers are all convolutional kernels of 3×3, the stride is 2, and the BN layer and ReLU activation layer are configured, and the output feature map sizes are 64×64×32, 32×32×64, and 16×16×128 in sequence; the fifth convolutional kernel is 3×3, the stride is 1, the batch BN layer and Tanh activation layer are configured, and the output feature map size is 16×16×1; specifically shown in Table 2; the output result of the discriminator is a scalar, indicating the probability that the input image is a real image.
[0109]
[0110] In the present invention, the total loss function Loss of the detailed layer GAN network fusion unit total_DL is expressed as:
[0111]
[0112] Among them, is the loss function of the discriminator , is the loss function of the discriminator , and the expression is:
[0113]
[0114] Loss(G DL ) is the loss function of the generator G DL , and the expression is:
[0115] Loss(G DL ) = Loss adv_DL + λLoss con_DL (7)
[0116]
[0117] Loss con_DL = α DL Loss grad_DL + β DL Loss ssim_DL + γ DL Loss in_DL (9)
[0118] Among them, λ, α DL , β DL , γ DL are the weight parameters for controlling the detailed layer GAN network fusion unit; F DLDenote the obtained fused image of the detail layer; E represents the mathematical expectation, which enables the discriminator to consider the feature information of the detail layer images of image I1 (three-channel pseudo-color image / single-channel grayscale image) and image I2 (MRI image);
[0119] L in_DL is the intensity loss function of the detail layer, representing the medical image function information of the detail layer. L in_DL The expression is:
[0120]
[0121] Among them, represents the norm, L is the gray value 256, p l is the gray level probability, is the threshold of saliency, Sign(·) represents the sign function, and respectively represent the saliency measurements of DL1 and DL2, and respectively represent the abundance measurements of DL1 and DL2, and respectively represent the surface level measurements of DL1 and DL2;
[0122] Loss ssim_DL is the structural loss function of the detail layer, which is used to smooth the pixel granulation caused by overly strong gradient information while ensuring the visual quality of the fused image. The expression is:
[0123] Loss ssim_DL = ε DL (1 - SSIM(F DL , DL1)) + (1 - ε DL )(1 - SSIM(F DL , DL2)) (14)
[0124] Among them, ε DL is the parameter for controlling the weight; SSIM(·) is the structural similarity formula;
[0125] In medical images, texture information can usually reflect the fine structural features of tissues and is an important basis for disease diagnosis. In order to fully capture the texture information in the source images, the present invention proposes an enhanced gradient loss function.
[0126] Loss grad_DL is the enhanced gradient loss function, representing the gradient information of the texture details of the detail layer. The expression is:
[0127]
[0128] Among them, |·|1 represents L1 regularization, represents the Laplacian gradient operator, a DL is a parameter for controlling the weight; is the loss function term for selecting the maximum gradient.
[0129] The total loss function Loss total_BL of the basic layer GAN network fusion unit is expressed as:
[0130]
[0131] Among them, is the loss function of the discriminator , is the loss function of the discriminator , and the expression is:
[0132]
[0133] Loss(G BL ) is the loss function of the basic layer generator, and the expression is:
[0134] Loss(G BL ) = Loss adv_BL + λLoss con_BL (19)
[0135]
[0136] Loss con_BL = α BL Loss grad_BL + β BL Loss ssim_BL + γ BL Loss in_BL (21)
[0137] Among them, λ, α BL , β BL , γ BL are weight parameters for controlling the basic layer GAN network fusion unit; F BL represents the obtained basic layer fusion image; E represents the mathematical expectation, which enables the discriminator to consider the feature information of the basic layer images of image I1 (three-channel pseudo-color image / single-channel grayscale image) and image I2 (MRI image); L in_BL is the intensity loss function of the basic layer, representing the medical image function information of the basic layer, L in_BL The expression is:
[0138]
[0139]
[0140] Among them, represents the norm, L is the gray value of 256, p l is the gray level probability, is the threshold of saliency, Sign(·) represents the sign function, and represent the saliency measurements of BL1 and BL2 respectively, and represent the abundance measurements of BL1 and BL2 respectively, and represent the surface level measurements of BL1 and BL2 respectively;
[0141] Loss ssim_BL is the basic layer structural loss function, which is used to smooth the pixel granulation caused by too strong gradient information while ensuring the visual quality of the fused image. The expression is:
[0142] Loss ssim_BL = ε BL (1 - SSIM(F BL , BL1)) + (1 - ε BL )(1 - SSIM(F BL , BL2)) (26)
[0143] Among them, ε BL is the parameter for controlling the weight; SSIM(·) is the structural similarity formula;
[0144] Loss grad_BL is the enhanced gradient loss function, which represents the gradient information of the texture details of the basic layer. The expression is:
[0145]
[0146] Among them, |·|1 represents the L1 regularization, represents the Laplacian gradient operator, a BL is the parameter for controlling the weight; is the loss function term for selecting the maximum gradient.
[0147] Selection of tuning parameters in the GAN network model:
[0148] The process of adjusting the model parameters is as follows: In the unsupervised medical image fusion model, select λ, α BL , β BL , γ BL , ε BL , a BL , α DL , βDL and γ DL and ε DL and a DL are used as hyperparameters, where λ is used to control the balance between the content loss function and the adversarial loss function, λ = 100, and α BL and β BL and γ BL and α DL and β DL and γ DL are used to control the display strength of the texture detail information in the fused image, and ε BL and ε DL are used to adjust the proportion of the image content in the two modalities, and ε BL ∈[0, 1], and ε DL ∈[0, 1], and a BL and a DL is the coefficient used to control the enhancement degree of the enhancement loss function, and a BL ∈[1, 5], and a DL ∈[1, 5];
[0149] During the hyperparameter tuning process, the initial learning rate is set to 2×10 -4 , the decay rate is set to 0.9, and Adam and RMSProp are used as optimizers to train the discriminator and the generator respectively;
[0150] When the network inputs the detail layer, λ = 100, and α DL = 1.5, β DL = 3, γ DL = 3, ε DL = 0.3, a DL = 1.2;
[0151] When the network inputs the base layer, λ = 100, and α BL = 1.2, β BL = 8, γ BL = 5, ε BL = 0.5, a BL = 1. (Note: The selection of parameters needs to be adjusted accordingly according to the image display quality)
[0152] The specific process of reconstructing the fused image using the fused result of the detail layer and the fused result of the base layer is as follows:
[0153] Add the fused detail layer and the base layer images to obtain the initial fused result;
[0154] The initial fusion result is processed by a non-linear enhancement unit to normalize the pixel values of the initial fusion result to the range of [0, 255], obtaining the fused image of the source image; then the output is performed. Specifically: if it is the fusion of a three-channel pseudo-color medical image and an MRI image, the fused image of the source image is obtained after restoring to the RGB space; if it is the fusion of a single-channel grayscale medical image and an MRI image, it is directly used as the fused image of the source image.
[0155] Based on the above model settings, PET / MRI image fusion, CT / MRI image fusion, and SPECT / MRI image fusion experiments are carried out respectively below.
[0156] Table 3 presents the experimental results of the PET image and MRI image datasets. Figure 5 It is a comparison chart of the fusion results corresponding to Table 3.
[0157]
[0158] Table 3 Comparison of PET / MRI Image Fusion Results
[0159] Regarding the PET / MRI image fusion data results in Table 3, the present invention (the row of Ours in the table represents the fusion result of the present method) has certain advantages in the EN, SF, and AG indicators, indicating that the present invention effectively solves the problem that the fusion algorithm cannot fully retain texture details. Combining the Q AB / F and PSNR values can draw a conclusion. That is, compared with other methods, the algorithm of the present invention can better fuse the key information in the source image and generate more expressive and clearer images. Figure 5In it, two groups of comparison charts of PET / MR fusion results are given. The fusion results of DDcGAN fail to fully preserve the color information of the source PET image. Although the SF and AG index values of DDcGAN are relatively high, the texture details in its fusion result chart are not clearly shown and the contrast of the brain bone part is relatively low; the color information in the fusion result of the EMFusion algorithm covers its texture information, resulting in blurred display of texture details; some texture detail information in the fusion result of GeSeNet is discarded by the network. Although its index values rank among the top, the texture information in the source MRI image is not fully preserved in its fusion image; the color information in the fusion result of the NPAP algorithm (for the specific algorithm, refer to the paper "Medical image fusion with parameter-adaptive pulse coupled neural network in nonsubsampled shearlet transform domain") is changed and can hardly be used as a basis for doctors to diagnose brain diseases; the texture detail information in the fusion result of the LR algorithm (for the specific algorithm, refer to "Laplacian ReDecomposition for Multimodal Medical Image Fusion") is not fully preserved, resulting in blurred display of its image, and the contrast of the brain bone part is also relatively low. In contrast, the fusion method proposed in this paper preserves the detailed texture and color features of the source image to a great extent, significantly enhances the fiber details of the soft tissue structure in the fusion image, and obtains a clearer and more accurate fusion effect.
[0160] Table 4 presents the experimental results of the CT image and MRI image datasets. Figure 6 It is a comparison chart of fusion results corresponding to Table 4.
[0161]
[0162] Comparison of CT / MRI Image Fusion Results in Table 4
[0163] For the CT / MRI image fusion data results in Table 4, our method ranks first in terms of the EN, SF, and AG index results, indicating that the fusion image of this algorithm is rich in texture detail information and retains more feature information in the source image. Figure 6Two groups of comparison diagrams of CT / MRI fusion results are given. Among the fusion metrics of the DDcGAN algorithm, the metric values of EN and AG rank second. However, due to the relatively high contrast of the fused images, the texture display is not clear. In the fusion results of the EMFusion algorithm, color information that did not originally exist appears. In the fusion results of GeSeNet, NPAP, and U2Fusion, the texture details of the brain bone part are blurred, and the overall contrast of the U2Fusion fused image is low, making the image details unclear. In the fusion result diagram of LR, the edge details show a blurring effect. In contrast, the fused image of this algorithm is clearly displayed as a whole, with moderate contrast and good visual effects.
[0164] Table 5 presents the experimental results of the CT image and MRI image datasets. Figure 7 It is a comparison diagram of the fusion results corresponding to Table 5.
[0165]
[0166] Table 5 Comparison of SPECT / MRI Image Fusion Results
[0167] For the SPECT / MRI image fusion data results in Table 5, the metrics of EN, SF, and AG all rank second, indicating that the image fusion method in this paper has certain advantages in retaining details, enhancing edge sharpness, and improving image quality. For Figure 7 , two groups of comparison diagrams of SPECT / MRI fusion results are given. In the DDcGAN fused image, the color information is changed and the algorithm has the effect of enhancing noise, resulting in serious image distortion. In the fused image of EMFusion, the color information in the source SPECT image is not fully retained. In the fused images of GeSeNet and LR algorithms, clear texture details are not displayed. In the fusion results of the NPAP algorithm, the color information is changed. The texture of the U2Fusion fused image is blurred, resulting in reduced soft tissue resolution. In contrast, in the fusion results of this algorithm, the edge texture is clearly displayed and the color information is normally displayed, and it can better perform the medical image fusion task.
[0168] In the present invention, in order to extract deeper image feature information, we design the CSB module in the form of a series connection of a standard convolutional block and The Sandglass Block. Now, the standard convolutional block is denoted as C, and The Sandglass Block is denoted as S, where S has two layers of depthwise separable convolutional layers. In order to reflect the lightweight characteristics of the network framework we designed, we evaluate the computational complexity of the model by measuring the number of training parameters and calculating FLOPs.
[0169]
[0170] Table 6 Model Evaluation
[0171] According to the data in Table 6, the training parameters of the method in this paper are reduced by about 30.6%, and the FLOPs are reduced by about 6.5%.
[0172] The unsupervised medical image fusion method based on the lightweight GAN network provided by the present invention aims to overcome three problems. First, it is unable to fully capture and learn important features in the image. Second, there are too many training parameters. Third, the fused image is not clearly displayed. The specific implementation steps are as follows: 1. Preprocess the input medical images, that is, directly fuse if they are all single-channel images; if there are pseudo-color images, the pseudo-color images need to be converted to the YCbCr space, extract the luminance component, and then fuse with the MRI images. 2. Use the Gaussian filtering technology to decompose the images to be fused into a base layer and a detail layer. 3. Use the designed GAN network framework to fully learn the features of the base layer and the detail layer. At the same time, use the designed loss function to improve the display clarity of the fused image.
[0173] The method of the present invention has the following advantages:
[0174] The present invention significantly reduces the computational amount and the number of parameters of the model through the lightweight GAN network, making the network more suitable for running in resource-constrained environments;
[0175] The present invention uses a Gaussian filter to decompose the image into a base layer and a detail layer, simplifies the image information, makes it easier for the generator to capture image features, and thus improves the clarity and quality of the fused image;
[0176] The present invention significantly improves the visual quality of the fused image through the designed enhanced loss function, thus providing more reliable support for clinical diagnosis.
[0177] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. An unsupervised medical image fusion method based on a lightweight dual GAN network, characterized in that Including: Step S1: Construct an unsupervised medical image fusion model based on a lightweight dual GAN network. The unsupervised medical image fusion model includes a preprocessing unit, a basic layer GAN network fusion unit, a detail layer GAN network fusion unit, a non-linear enhancement unit, and a fused image output unit; Step S2: Collect computer tomography (CT) images and magnetic resonance imaging (MRI) images and establish a multi-modal medical source image database for training the model. Data augmentation is performed on each source image by image rotation. Among them, the types of CT images include three-channel pseudo-color medical images and single-channel grayscale medical images; Step S3: Call the source images in the source image database to train the unsupervised medical image fusion model. The preprocessing unit preprocesses and decomposes the CT images and the MRI images. The data of the two detail layers obtained by the image decomposition are input into the detail layer GAN network fusion unit for fusion to obtain a detail layer fusion result. The data of the two basic layers obtained by the image decomposition are input into the basic layer GAN network fusion unit for fusion to obtain a basic layer fusion result. The fused image is reconstructed using the detail layer fusion result and the basic layer fusion result. The obtained initial fusion result is processed by the non-linear enhancement unit and then the fused image of the source images is obtained, which is output by the fused image output unit. Among them, the weight parameters of the basic layer GAN network fusion unit and the detail layer GAN network fusion unit are optimized through a loss function; Step S4: Use the trained unsupervised medical image fusion model for medical image fusion.
2. The unsupervised medical image fusion method based on the lightweight twin GAN network according to claim 1, wherein CT images belong to single-channel medical images, and SPECT images and PET images both belong to three-channel pseudo-color medical images; The preprocessing unit: For single-channel grayscale medical image type images, no preprocessing is required; For three-channel pseudo-color medical image type images, it is necessary to convert from the RGB space to the YCbCr color space, and then extract the Y luminance component for subsequent Gaussian filter decomposition; Use Gaussian filter decomposition to preprocess the CT image I1 and the MRI image I2 to obtain the basic layer BL1 and the detail layer DL1 of I1, and the basic layer BL2 and the detail layer DL2 of I2; The basic layer BL1 and the basic layer BL2 are sent into the basic layer GAN network fusion unit; The detail layer DL1 and the detail layer DL2 are sent into the detail layer GAN network fusion unit.
3. The unsupervised medical image fusion method based on the lightweight dual GAN network according to claim 2, characterized in that, The basic layer GAN network fusion unit is a generative adversarial network composed of discriminator D BL1 , discriminator D BL2 , generator G BL . The data processing process is as follows: The generator G BL performs a fusion process on the base layer BL1 and the base layer BL2 to obtain the base layer fusion result F BL ; The fusion result F of the base layer BL and the input of the base layer BL1 to the discriminator D BL1 are discriminated, and the obtained discrimination result is used to adjust the generator G BL for adjustment; The basic layer fusion result F BL and the basic layer BL2 are input into the discriminator D BL2 for discrimination, and the obtained discrimination result is used to BL adjust the generator G; Adversarial training is carried out in the form of alternating training between the discriminator and the generator.
4. The unsupervised medical image fusion method based on the lightweight twin GAN network according to claim 3, wherein The detailed layer GAN network fusion unit is a generative adversarial network composed of discriminator D DL1 , discriminator D DL2 , generator G DL . The data processing process is as follows: The generator G DL performs a fusion process on the detail layer DL1 and the detail layer DL2 to obtain the detail layer fusion result F DL ; The fusion result F of the detail layer DL and the input of the detail layer DL1 to the discriminator D DL1 are discriminated, and the obtained discrimination result is used to adjust the generator G DL for adjustment; The detailed layer fusion result F DL and the detailed layer DL2 are input into the discriminator D DL2 for discrimination, and the obtained discrimination result is used to adjust the generator G DL for adjustment; Adversarial training is carried out in the form of alternating training between the discriminator and the generator.
5. The unsupervised medical image fusion method based on the lightweight twin GAN network according to claim 4, wherein, Generator G DL and G BL adopt the same framework structure and are successively composed of an input layer, an encoder, a decoder, and an output layer, where: The input image is first processed by a 3×3 standard convolutional layer, and the resulting feature map is input into the encoder; the encoder is composed of 4 CSB modules connected in series in sequence, and each CSB module is composed of a 3×3 standard convolutional layer and a SandglassBlock connected in series. The Sandglass Block is composed of a 3×3 depthwise separable convolution Dwise, a 1×1 dimensionality reduction standard convolutional kernel, a 1×1 dimensionality increase standard convolutional kernel, and a 3×3 depthwise separable convolution Dwise connected in series in sequence; The decoder is composed of 3 CONV blocks, and each CONV block is a standard convolutional layer using a 3×3 convolutional kernel; the feature map is processed by CSB1, CSB2, CSB3, and CSB4 in sequence. CONV3 receives the downsampled features from CSB1, the downsampled features from CSB2, the skip connection features from CSB3, and the upsampled features from CSB4. CONV2 receives the downsampled features from CSB1, the skip connection features from CSB2, the upsampled features from CONV3, and the upsampled features from CSB4. CONV1 receives the skip connection features from CSB1, the upsampled features from CONV2, the upsampled features from CONV3, and the upsampled features from CSB4; The output of CONV1 is sent to the output layer and is processed by two 3×3 standard convolutional layers in sequence to obtain the output result of the generator.
6. The unsupervised medical image fusion method based on the lightweight twin GAN network according to claim 5, characterized in that, Discriminator It adopts the same framework structure and consists of 5 convolutional blocks connected in sequence. Among them: the first layer has a convolutional kernel of 3×3, a stride of 2, is configured with a ReLU activation layer, and the output feature map size is 128×128×16; the second, third, and fourth layers all have a convolutional kernel of 3×3, a stride of 2, are configured with a BN layer and a ReLU activation layer, and the output feature map sizes are 64×64×32, 32×32×64, and 16×16×128 in sequence; the fifth layer has a convolutional kernel of 3×3, a stride of 1, a batch BN layer and a Tanh activation layer, and the output feature map size is 16×16×1; the output result represents the probability that the input image is a real image.
7. The unsupervised medical image fusion method based on the lightweight dual GAN network according to claim 6, characterized in that, The total loss function Loss of the detailed layer GAN network fusion unit total_DL is expressed as: Among them, is the loss function of the discriminator , is the loss function of the discriminator , and the expression is: Loss(G DL ) is the loss function of the generator G DL , and the expression is as follows: Loss(G DL ) = Loss adv_DL + λLoss con_DL Loss con_DL = α DL Loss grad_DL + β DL Loss ssim_DL + γ DL Loss in_DL Among them, λ, α DL , β DL , γ DL are the weight parameters for controlling the fusion unit of the detailed layer GAN network; E represents the mathematical expectation; Loss in_DL It is the strength loss function of the detail layer, and its expression is: Among them, represents the norm, L is the gray value of 256, p l is the gray level probability, is the threshold of saliency, Sign(·) represents the sign function, and respectively represent the saliency measurements of DL1 and DL2, and respectively represent the abundance measurements of DL1 and DL2, and respectively represent the surface level measurements of DL1 and DL2; Loss ssim_DL It is the structural loss function for the detail layer, and its expression is: Loss ssim_DL = ε DL (1 - SSIM(F DL , DL1))+(1 - ε DL )(1 - SSIM(F DL , DL2)) where ε DL is a parameter for controlling the weight; SSIM(·) is the structural similarity formula; Loss grad_DL To enhance the gradient loss function, the expression is: Among them, |·|1 represents L1 regularization, and ▽ 2 represents the Laplacian gradient operator, and a DL is a parameter for controlling the weight; max(|▽ 2 DL1|, |▽ 2 DL2|) is a loss function term for selecting the maximum gradient.
8. The unsupervised medical image fusion method based on the lightweight twin GAN network according to claim 7, wherein, The total loss function Loss of the basic layer GAN network fusion unit total_BL has the following expression: Among them, is the loss function of the discriminator , is the loss function of the discriminator , and the expression is: Loss(G BL ) generates the loss function for the base layer generator, and the expression is: Loss(G BL ) = Loss adv_BL + λLoss con_BL Loss con_BL = α BL Loss grad_BL + β BL Loss ssim_BL + γ BL Loss in_BL Among them, λ, α BL , β BL , γ BL are weight parameters for controlling the fusion unit of the basic layer GAN network; E represents a number Semester expectation; Loss in_BL It is the strength loss function of the base layer, and the expression is: Among them, represents the norm, L is the gray value of 256, p l is the gray level probability, is the threshold of saliency, Sign(·) represents the sign function, and respectively represent the saliency measurements of BL1 and BL2, and respectively represent the abundance measurements of BL1 and BL2, and respectively represent the surface level measurements of BL1 and BL2; Loss ssim_BL is the basic layer structural loss function, and its expression is: Loss ssim_BL = ε BL (1 - SSM(F BL , BL1)) + (1 - ε BL )(1 - SSIM(F BL , BL2)) where ε BL is a parameter for controlling the weight; SSIM(·) is the structural similarity formula; Loss grad_BL To enhance the gradient loss function, the expression is: Among them, |·|1 represents L1 regularization, and ▽ 2 represents the Laplacian gradient operator, and a BL is a parameter for controlling the weight; max(|▽ 2 BL1|, |▽ 2 BL2|) is the loss function term for selecting the maximum gradient.
9. The unsupervised medical image fusion method based on the lightweight twin GAN network according to claim 8, wherein The process of adjusting model parameters is as follows: Select λ, α in the unsupervised medical image fusion model BL , β BL , γ BL , ε BL , a BL , α DL , β DL , γ DL , ε DL , a DL as hyperparameters. Among them, λ is used to control the balance between the content loss function and the adversarial loss function, λ = 100, α BL , β BL , γ BL , α DL , β DL , γ DL are used to control the strength of the display of texture detail information in the fused image, ε BL , ε DL are used to adjust the proportion of the image content in the two modalities, ε BL ∈[0, 1], ε DL ∈[0, 1], a BL , a DL are the coefficients used to control the enhancement degree of the enhancement loss function, a BL ∈[1, 5], a DL ∈[1, 5]; During the hyperparameter tuning process, the initial learning rate was set to 2×10 -4 , the decay rate was set to 0.9, and Adam and RMSProp were used as optimizers to train the discriminator and generator respectively; When the network inputs the detail layer, λ = 100, α DL = 1.5, β DL = 3, γ DL = 3, ε DL = 0.3, a DL = 1.2; When the network inputs the basic layer, λ = 100, α BL = 1.2, β BL = 8, γ BL = 5, ε BL = 0.5, a BL = 1.
10. The unsupervised medical image fusion method based on the lightweight dual GAN network according to claim 1, wherein, The specific process of reconstructing the fused image by using the fused result of the detail layer and the fused result of the base layer is as follows: The fused detail layer and the base layer images are added together to obtain the initial fused result; The initial fused result is processed by the non-linear enhancement unit to normalize the pixel values of the initial fused result to the range of [0, 255] to obtain the fused image of the source image; then the output is performed. Specifically, if it is the fusion of a three-channel pseudo-color medical image and an MRI image, the fused image of the source image is obtained after restoring to the RGB space; if it is the fusion of a single-channel grayscale medical image and an MRI image, it is directly used as the fused image of the source image.