A medical image fusion method based on artificial intelligence and superpixel segmentation
Through the method based on artificial intelligence and superpixel segmentation, the pre-trained convolutional neural network is used to extract medical image features, which solves the problem of time-consuming and insufficient generalization performance of multimodal medical image fusion in the prior art, and achieves efficient and high-quality image fusion effect.
Patent Information
- Application Number
- CN202211585980.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-12-10
AI Technical Summary
The existing multimodal medical image fusion method requires a large number of specialized data sets for training, which takes a long time and lacks generalization performance, especially when fusion of complex medical images.
Using artificial intelligence and superpixel segmentation methods, pre-trained convolutional neural networks are used to extract image features, and high-quality fusion images are generated through superpixel segmentation dimensionality reduction and fusion weights, avoiding the training process of image modality.
Save training time, improve the efficiency and quality of image fusion, especially the fusion effect of complex images is significantly improved, and the impact of blur and noise is reduced.
Smart Images

Figure CN115731444B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image fusion technology, and more specifically, to a medical image fusion method based on artificial intelligence and superpixel segmentation, which is applicable to the fusion of medical images of different modalities. Background Art
[0002] Due to differences in different imaging technologies, medical devices can capture different organ and tissue characteristics, but can only obtain single-modal medical images, and single-modal medical images cannot provide comprehensive and sufficient information. For example, Computed Tomography (CT) can capture dense structures and bones, while Magnetic Resonance Imaging (MRI) can detect soft tissues [Xu Chen et al., 2021, 274:285.]. Even though CT and MRI are both anatomical images, and Positron Emission Tomography (PET) and Single Photon Emission Computed Tomography (SPECT) are both functional images, the information they focus on is different. CT is suitable for some bone tissue, lung, and coronary artery related examinations, while MRI is suitable for examinations of tissues with higher water content, and MRI is more detailed than CT examinations and can detect small changes in tissues earlier. PET mainly reflects the functional and metabolic information of masses, while SPECT mainly provides the blood flow information of organs and tissues. Multi-modal image fusion allows complementary data to be visualized in a single image. However, sequential analysis of multi-modal images is inconvenient, and due to the limitations of single-modal medical images, multi-modal medical image fusion [Han Xu et al., 2021, 177:186.] is a field that is very necessary to study.
[0003] In the past decade, a variety of medical image fusion models have been proposed, such as the Nonsubsampled Contourlet Transform (NSCT), guided filter fusion (GFF), etc. Most of these methods consist of three steps, namely decomposition, fusion, and reconstruction. Recently, deep networks have been used for multi-fusion problems. The neural network itself can be regarded as a feature extractor, and among them, the intermediate mapping representation can be used to reconstruct the significant features of the fused image. However, although deep learning methods often achieve better performance than traditional methods, they still have deficiencies. First, these methods require a large number of specialized datasets for training and often take a long time [Tianshu Yang et al., 2021, 235.]. Second, the generalization performance of existing fusion methods is insufficient. For relatively complex medical images, the performance of the fusion network is often poor. Summary of the Invention
[0004] To overcome the deficiencies of the prior art, the present invention proposes a medical image fusion method based on artificial intelligence and superpixel segmentation.
[0005] The above object of the present invention is achieved by the following technical solutions:
[0006] A medical image fusion method based on artificial intelligence and superpixel segmentation, comprising the following steps:
[0007] Step 1: Obtain M original medical images of different modalities for the same object, and extract the set of preprocessed block images corresponding to the original medical images of all modalities;
[0008] Step 2: Perform superpixel segmentation on the preprocessed block images obtained in Step 1 to obtain a pseudo-CIELAB image of superpixel segmentation;
[0009] Step 3: Construct an image feature extraction module, obtain the pytorch tensor image corresponding to each modality according to the pseudo-CIELAB image of superpixel segmentation generated in Step 2, and use a pre-trained convolutional neural network to extract the weight map corresponding to the pytorch tensor image of each modality;
[0010] Step 4: Construct an image feature fusion module, fuse the pytorch tensor images corresponding to each modality extracted in Step 3 under the guidance of the weight map, and screen to obtain the final fused image;
[0011] Step 5: Convert the final fused image into other human-readable image formats, then crop it to an appropriate size, and process the overflow area.
[0012] Step 2 described above includes the following steps:
[0013] Step 2.1: Denoise the preprocessed block image and then increase the contrast between the target area and the background area to obtain a set of denoised and enhanced block images;
[0014] Step 2.2: Map each denoised and enhanced block image to the pseudo-color space and then calculate the CIELAB color space to obtain the pseudo-CIELAB image and the five-dimensional vector of each pixel in each pseudo-CIELAB image;
[0015] Step 2.3: Perform superpixel segmentation on the pseudo-CIELAB image to obtain the overall superpixel images of each modality;
[0016] Step 2.4: Adjust the size of the overall superpixel image of each modality to obtain the pseudo-CIELAB image d of superpixel segmentation m , where m ∈ {1, 2,..., M}, representing the m-th modality.
[0017] Step 3 described above includes the following steps:
[0018] Step 3.1: Invert the pseudo-CIELAB image of superpixel segmentation to the pseudo-RGB image format, then continue to convert it to the YCbCr format, and then convert it to the pytorch tensor data format to obtain the corresponding pytorch tensor image d of each modality m , m ∈ {1, 2,..., M};
[0019] Step 3.2: Input the corresponding pytorch tensor image of each modality into the pre-trained convolutional neural network, and extract the corresponding feature maps of the corresponding pytorch tensor image of each modality in each convolutional layer, so as to obtain the set of corresponding feature maps of the pytorch tensor image of each modality in each convolutional layer;
[0020] Step 3.3: Calculate the corresponding norm value for each set of feature maps output by the pre-trained convolutional neural network;
[0021] Step 3.4: Calculate the weight map corresponding to the set of feature maps of the pytorch tensor image of each modality using the norm value obtained in Step 3.3;
[0022] Step 3.5: Perform Gaussian smoothing on each weight map to obtain the smoothed weight map.
[0023] The set of feature maps of the corresponding pytorch tensor image of each modality in each convolutional layer in Step 3.2 above is:
[0024]
[0025] Among them, represents the set of feature maps extracted from the l-th convolutional layer of the pre-trained convolutional neural network for the pytorch tensor image corresponding to the m-th modality, F l (·) represents the convolutional operation of the first l layers of the pre-trained convolutional neural network for the input pytorch tensor image d
[0026] ′
[0027] input pytorch tensor image d m for the first l layers, and max() refers to performing the Relu activation operation, where 0 indicates that the number of zero-padding for convolution is 0.
[0028] As described above, the calculated norm value in step 3.3 is the L1 norm
[0029]
[0030] Among them, represents the L1 norm value corresponding to the m-th modality in the l-th convolutional layer, C l represents the number of output channels of the l-th convolutional layer in the pre-trained convolutional neural network, l ∈ {1, 2,..., L}, where L is the number of convolutional layers of the pre-trained convolutional neural network, represents the c-th feature map extracted from the l-th convolutional layer of the pre-trained convolutional neural network for the m-th modality, c ∈ {1, 2,..., C l}
[0031] As described above, the weight map corresponding to the feature map of the pytorch tensor image of each modality in step 3.4 is
[0032]
[0033] represents the weight map corresponding to the m-th modality in the l-th convolutional layer, represents the L1 norm value corresponding to the j-th modality in the l-th convolutional layer, j ∈ {1, 2,..., M}.
[0034] As described above, in step 4, the pytorch tensor images corresponding to each modality are fused under the guidance of the smoothed weight map, and the fused image obtained before getting the final fused image is
[0035]
[0036] Among them, represents the smoothed weight map corresponding to the l-th convolutional layer The fused image generated by the guidance is smoothed to obtain a smoothed weight map.
[0037] In step 4 as described above, the final fused image is D F :
[0038]
[0039] where torchmax[·] represents the operation of taking the maximum image of the tensor.
[0040] Compared with the prior art, the present invention has the following advantages:
[0041] The present invention proposes a medical image fusion method using a pre-trained convolutional neural network, which does not require training of image modalities, saving a large amount of time; uses a pre-trained convolutional neural network with high performance to detect feature significant regions in the image and extract feature maps of these regions; generates fusion weights by comparing these feature maps to combine the source images.
[0042] In addition, before image fusion, the present invention adds a unique step of superpixel segmentation to achieve image dimensionality reduction and eliminate some abnormal pixel points, accelerating the process of image fusion and reducing the influence of blurred abnormal points in the image fusion process, greatly improving the performance of the fusion model, and being able to fuse high-quality images for complex images. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a schematic diagram of the method flow of the present invention;
[0044] Figure 2 is a schematic diagram of the structure of the vgg19 pre-trained convolutional neural network;
[0045] Figure 3 Table for evaluating and comparing the image fusion method of the present invention with other methods. In the table: GFF is a multi-scale guided filtering fusion method; NSCT is a method based on multi-scale geometric analysis (MGA); NSCT PCNN is a method that uses an improved pulse-coupled neural network for fusion based on NSCT; LP SR is a method that uses sparse coding to fuse the MST low-pass band; LIU CNN is a fusion method based on a siamese neural network; NSST PAPCNN is a fusion method based on non-sub-sampled Shearlet transform. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] For the convenience of those of ordinary skill in the art to understand and implement the present invention, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0047] Embodiment 1:
[0048] Figure 1 It is a flowchart of a medical image fusion method based on artificial intelligence and superpixel segmentation provided by the present invention, which specifically includes the following steps:
[0049] Step 1: Obtain M original medical images of different modalities for the same object, and extract the preprocessing block image sets corresponding to the original medical images of all modalities. In this embodiment, M = 2, and brain MRI images and corresponding brain CT images are used. Extract the Y channels of the original medical images of all modalities to obtain the Y-channel images {I1,..., I m ...,I M}, where I m represents the Y-channel image of the m-th modality, m ∈ {1, 2,..., M}, and in this embodiment, m ∈ {1, 2}; perform image clipping on the Y-channel images {I1,..., I m ...,I M} to remove the areas outside the study, so as to obtain the preprocessing block image sets {J1,..., J m ...,J M}, where J m represents the preprocessing block image of the m-th modality.
[0050] Step 2: Perform superpixel segmentation on the preprocessing block images obtained in Step 1 to obtain a pseudo CIELAB image of superpixel segmentation. It includes the following steps:
[0051] Step 2.1: Obtain a set of denoised and enhanced block images. Use a three-dimensional histogram denoising model to process the preprocessing block images of each modality from Step 1 to reduce the interference of noise in the images, and then use a gamma enhancement model to increase the contrast between the target area and the background area of the preprocessing block images after denoising of each modality, and reduce the probability of incorrect division of the boundary pixel points of the target area, so as to obtain a set of denoised and enhanced block images {S1,..., S m ...,S M}. The formula for gamma transformation is: s = c g *r γ , where r is the pixel gray value of the preprocessing block image after denoising of each modality, s is the pixel gray value of the denoised and enhanced block image, and c gis the grayscale scaling factor, which is used to stretch the image grayscale globally. The value is 1. To increase the contrast between the target region and the background region, the gamma factor γ = 1.5 is taken.
[0052] Step 2.2: Obtain the pseudo-CIELAB image and the five-dimensional vector of each pixel in each pseudo-CIELAB image for the denoised and enhanced block image obtained in Step 2.1. For each denoised and enhanced block image S m , define an expandable mapping to map S m to the pseudo-color space to obtain the pseudo-color image S m '. Then calculate the CIELAB color space of the pseudo-color image S m ' to obtain the pseudo-CIELAB image S m ". For a pixel p = (x, y) in the pseudo-CIELAB image S m ", its value is determined by the brightness li and the color contrasts α and β. These three components are combined with the X-Y coordinates to obtain the five-dimensional vector (x, y, li, α, β) of each pixel, where x and y are the row and column coordinates respectively.
[0053] Step 2.3: Perform superpixel segmentation on the pseudo-CIELAB image S m " to obtain the superpixel global images of each modality. In this embodiment, the publicly available Simple Linear Iterative Clustering (SLIC) algorithm is used to perform superpixel segmentation on the pseudo-CIELAB images S m " from 2 different modalities. The number of segments is set to K sp = 100, and the number of iterations is set to 10 times to obtain the superpixel global image P m . The segmented image is the superpixel global image P m . It reduces the dimension without sacrificing too much accuracy, and the features of each superpixel are more concentrated, which is beneficial for subsequent calculations.
[0054] Step 2.4: Use bilinear interpolation to resize the superpixel global images P m of each modality generated after superpixel generation to a size of 224×224 pixels to adapt to the specification requirements of the pre-trained convolutional neural network in the subsequent steps. The superpixel-segmented pseudo-CIELAB images obtained after processing are denoted as d m , m ∈ {1, 2,..., M}.
[0055] Step 3: Construct an image feature extraction module for extracting the deep features of the image.
[0056] Step 3.1: Input the superpixel-segmented pseudo-CIELAB image d mInverse transform to the pseudo RGB image format, then continue to convert it to the YCbCr format, and then convert it to the pytorch tensor data format. After conversion, the pytorch tensor images corresponding to each modality are d m ’, m ∈ {1, 2, ..., M}.
[0057] Step 3.2: Input the pytorch tensor images d m ’ corresponding to each modality after conversion into a pre-trained convolutional neural network with L convolutional layers, and extract the feature maps of the pytorch tensor images of each modality in each convolutional layer. In this embodiment, the vgg19 pre-trained convolutional neural network is used, and the vgg19 model trained on the ImageNet database in pytorch is called to obtain image features. Among them, the l-th convolutional layer has C l output channels, l ∈ {1, 2, ..., L}, and the corresponding feature map is output is denoted as the c-th feature map extracted by the m-th modality in the l-th convolutional layer of the pre-trained convolutional neural network, c ∈ {1, 2, ..., C l}), where the number of input channels of the (l + 1)-th convolutional layer is the same as the number of output channels of the l-th convolutional layer, and the input of the (l + 1)-th convolutional layer is the output result of the l-th convolutional layer after convolution operation.
[0058] Calculate the set of feature maps extracted by the pytorch tensor image corresponding to each modality in the l-th convolutional layer
[0059]
[0060] represents the set of feature maps extracted by the pytorch tensor image corresponding to the m-th modality in the l-th convolutional layer of the pre-trained convolutional neural network, F l (·) represents the convolutional operation of the first l layers of the pre-trained convolutional neural network on the input pytorch tensor image d m ’, max() refers to performing the Relu activation operation, where 0 means the convolution padding number is 0,
[0061] Step 3.3: For each set of feature maps output by the pre-trained convolutional neural network Calculate the L1 norm
[0062]
[0063] Represents the L1 norm value corresponding to the m-th modality in the l-th convolutional layer.
[0064] Step 3.4: For the l-th convolutional layer of the pre-trained convolutional neural network, the feature maps of the pytorch tensor images of M modalities can generate M weight maps. Denotes the weight map corresponding to the m-th modality in the l-th convolutional layer (m = 1, 2,....M), indicating the contribution amount of the pytorch tensor image of each modality at a specific pixel. For the feature map of the m-th modality in the l-th convolutional layer, use the softmax operator to generate the corresponding weight map:
[0065]
[0066] Represents the L1 norm value corresponding to the j-th modality in the l-th convolutional layer, j ∈ {1, 2,..., M}.
[0067] Step 3.5: For each weight map Perform Gaussian smoothing to obtain the smoothed weight map
[0068] Step 4: Construct an image feature fusion module to fuse the pytorch tensor images dm' corresponding to each modality obtained after transformation in Step 3.1 under the guidance of the smoothed weight map, and screen to obtain the final fused image D F 。
[0069] Step 4.1: According to the smoothed weight map of the l-th layer Guide the image fusion of the pytorch tensor images dm' corresponding to each modality obtained in Step 3.1, and calculate the fused image generated under the guidance of the smoothed weight map corresponding to the l-th convolutional layer Guide
[0070]
[0071] Step 4.2: To reconstruct multi-layer fusion, use the torch.max function of pytorch to compare each tensor-format fused image element by element. After comparison, the image with the largest tensor is the final fused image D F :
[0072]
[0073] Among them, torchmax[·] represents the operation of taking the image with the largest tensor.
[0074] Step 5: According to the need, the final fused image D FConvert it into other human-eye-readable image formats such as jpg or png, and perform cropping to an appropriate size, and finally process the overflow area.
[0075] Figure 3 The performance of the method of the present invention is shown. In the fusion experiments of the original medical images of two modalities, MRI and CT images, the present invention is optimal in terms of EN (information entropy), SSIM (structural similarity index), and Nabf (fusion performance based on noise evaluation); applying the method of the present invention to the fusion experiments of the original medical images of two modalities, MRI T1 and MRI T2 images, shows that it is optimal in terms of EN, Qabf, SSIM, and Nabf. The strategy of the present invention enhances EN by about 0.5%-1%, SSIM by 1.2%-9.9%, MI (mutual information) by 0-22.5%, and improves Qabf (fusion quality) by up to about 34.2% at most, and reduces Nabf by up to about 13 times at most, greatly reducing the noise and artifact levels of the fused images.
[0076] A medical image fusion device based on artificial intelligence and superpixel segmentation, including a first module, a second module, a third module, a fourth module, and a fifth module. The above steps 1 to 5 are respectively implemented by the first to fifth modules.
[0077] It should be noted that the specific embodiments described in the present invention are only examples of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A medical image fusion method based on artificial intelligence and superpixel segmentation, characterized in that, It includes the following steps: Step 1: Obtain M original medical images of different modalities for the same object, and extract the set of preprocessed block images corresponding to the original medical images of all modalities; Step 2: Perform superpixel segmentation on the preprocessed block images obtained in Step 1 to obtain a pseudo-CIELAB image of superpixel segmentation; Step 3: Construct an image feature extraction module, obtain the pytorch tensor images corresponding to each modality based on the pseudo-CIELAB image of superpixel segmentation generated in Step 2, and use a pre-trained convolutional neural network to extract the weight maps corresponding to the pytorch tensor images of each modality; Step 4: Construct an image feature fusion module, fuse the pytorch tensor images corresponding to each modality extracted in Step 3 under the guidance of the weight maps, and screen to obtain the final fused image; Step 5: Convert the final fused image into other human-eye-readable image formats, then crop it to an appropriate size, and process the overflow area; The said Step 2 includes the following steps: Step 2.1: Denoise the preprocessed block images and then increase the contrast between the target area and the background area to obtain a set of denoised and enhanced block images; Step 2.2: Map each denoised and enhanced block image to a pseudo-color space and then calculate the CIELAB color space to obtain a pseudo-CIELAB image and the five-dimensional vectors of each pixel in each pseudo-CIELAB image; Step 2.3: Perform superpixel segmentation on the pseudo-CIELAB image to obtain the overall superpixel images of each modality; Step 2.4: Adjust the size of the superpixel global image of each modality to obtain the pseudo-CIELAB image d of superpixel segmentation m , where m ∈ {1, 2,..., M}, representing the m-th modality; The said Step 3 includes the following steps: Step 3.1: Inverse transform the pseudo-CIELAB image obtained by superpixel segmentation into the pseudo-RGB image format, then continue to convert it into the YCbCr format, and then into the pytorch tensor data format to obtain the pytorch tensor images d corresponding to each modality m ’, m ∈ {1, 2, ..., M}; Step 3.2: Input the pytorch tensor images corresponding to each modality into the pre-trained convolutional neural network, and extract the feature maps corresponding to the pytorch tensor images of each modality in each convolutional layer, so as to obtain the set of feature maps corresponding to the pytorch tensor images of each modality in each convolutional layer; Step 3.3: For each set of feature maps output by the pre-trained convolutional neural network, calculate the corresponding norm value; Step 3.4: Use the norm values obtained in Step 3.3 to calculate the weight maps corresponding to the set of feature maps of the pytorch tensor images of each modality; Step 3.5: Perform Gaussian smoothing on each weight map to obtain the smoothed weight map.
2. The medical image fusion method based on artificial intelligence and superpixel segmentation according to claim 1, characterized in that The set of feature maps of the pytorch tensor images of each modality in each convolutional layer in the said Step 3.2 is: Among them, represents the set of feature maps extracted from the l-th convolutional layer in the pre-trained convolutional neural network for the pytorch tensor image corresponding to the m-th modality, F l (·) represents the convolutional operation of the first l layers of the pre-trained convolutional neural network on the input pytorch tensor image d m ′, max() refers to the Relu activation operation, where 0 indicates that the number of convolutional zero-padding is 0.
3. The medical image fusion method based on artificial intelligence and superpixel segmentation according to claim 2, wherein In step 3.3, the calculated norm value is the L1 norm Among them, represents the L1 norm value corresponding to the m-th modality in the l-th convolutional layer, C l represents the number of output channels of the l-th convolutional layer in the pre-trained convolutional neural network, l ∈ {1, 2,..., L}, where L is the number of convolutional layers of the pre-trained convolutional neural network, represents the c-th feature map extracted by the m-th modality in the l-th convolutional layer of the pre-trained convolutional neural network, c ∈ {1, 2,..., C l}.
4. The medical image fusion method based on artificial intelligence and superpixel segmentation according to claim 3, characterized in that, The weight map corresponding to the feature map of each modal pytorch tensor image in step 3.4 is Denote the weight map corresponding to the \(m\)-th modality in the \(l\)-th convolutional layer, represent the \(L_1\) norm value corresponding to the \(j\)-th modality in the \(l\)-th convolutional layer, where \(j\in\{1, 2, \ldots, M\}\).
5. The medical image fusion method based on artificial intelligence and superpixel segmentation according to claim 4, wherein In step 4, the pytorch tensor images corresponding to each modality are fused under the guidance of the smooth weight map, and the fused image obtained before getting the final fused image is Among them, represents the fused image generated under the guidance of the smoothing weight map corresponding to the l-th convolutional layer , and is the smoothed weight map obtained by performing smoothing processing.
6. The medical image fusion method based on artificial intelligence and superpixel segmentation according to claim 5, characterized in that, In the said step 4, the final fused image is D F : Among them, torchmax[·] represents the operation of taking the maximum image of the tensor.
Citation Information
Patent Citations
A three-dimensional medical image segmentation method based on feature enhancement
CN109685819A
SP-FCN-based MRI brain tumor image segmentation system and method
CN111932549A