A multi-modal medical image segmentation method

Through the multimodal medical image segmentation method, combined with 2D and 3D segmentation networks, high-precision multi-dimensional visualization is achieved, solving the problems of blurred segmentation profiles and high equipment requirements in the prior art, and providing a more accurate medical diagnosis basis.

CN116030043BActive Publication Date: 2025-07-11ANHUI UNIV OF SCI & TECH +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310163065.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-07-11
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

The existing medical image segmentation methods have blurred segmentation profiles, high equipment requirements, are difficult to visualize from multiple aspects, and are difficult to achieve high-precision multi-dimensional display.

Method used

Multimodal medical image segmentation method is used to extract feature information through 2D and 3D segmentation networks, and feature fusion technology is used to fuse features of different dimensions to provide multi-dimensional visual segmentation results.

Benefits of technology

It improves segmentation accuracy, provides more accurate multi-dimensional image basis, and supports medical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030043B_ABST
    Figure CN116030043B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal medical image segmentation method, belonging to the technical field of medical image segmentation, which includes the following steps: S1: Image preprocessing; S2: Constructing 2D and 3D segmentation networks; S3: Image fusion; S4: Slice splicing. By using multi-modal information, the medical images can be more fully utilized in the present invention; through the combination of 2D and 3D segmentation networks and the channel attention mechanism, the information of each modality can achieve intercommunication and complementarity; based on the fusion of 2D and 3D segmentation results, the segmentation result has higher accuracy and clearer edges; the segmentation results of different dimensions and different models are provided, providing a more accurate and multi-dimensional image basis for medical diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and in particular to a multi-modal medical image segmentation method. Background Art

[0002] With the development of medicine, computer technology, and biomedical engineering, medical imaging, as an important research object in visual processing, provides a large number of multi-modal medical images for clinical practice. How to use these multi-modal medical images for medical image segmentation is of great significance for the treatment of patients.

[0003] Currently, most medical image segmentation models aim to obtain biological tissue information such as the brain, lungs, liver, and heart blood contained in medical images. Traditional automatic medical image segmentation methods mainly include segmentation methods based on graph theory, morphology, and deformation models. Deformation models include parametric active contour models and geometric deformation models, etc. The representative method of geometric models is the level set method. Almost all commonly used image segmentation algorithms are based on deterministic methods, but there is uncertainty in the process of image information processing, so to a certain extent, it also has a certain impact on the segmentation accuracy and the generalization of the model.

[0004] In recent years, deep learning methods have developed rapidly, and image segmentation algorithms based on deep learning have achieved remarkable results in the field of medical image segmentation such as the brain, liver, and kidneys. Convolutional neural networks, as commonly used deep learning methods at present, have been widely applied to the image segmentation of various organs or tissues. These methods are mainly divided into patch, semantic, and cascaded architectures. Some scholars have proposed a fully convolutional network structure, which can perform pixel-level classification on images, thereby enabling semantic-level image segmentation. In 2015, some scholars proposed the U-Net network structure, which is a semantic segmentation network based on the fully convolutional network structure. In 2019, some scholars proposed a multi-sequence MRI automatic cardiac segmentation framework based on U-Net to solve the cardiac segmentation problem. In 2020, some scholars proposed a medical image segmentation method based on a multi-receptive field convolutional neural network. In addition, the recurrent neural network contains at least one feedback connection in the constructed structure. The long short-term memory network LSTM, as a special type of RNN, solves the problem of its vanishing gradient. Some scholars have also used 3D LSTM-RNN to segment MRI images, greatly improving the network training efficiency.

[0005] The above-mentioned segmentation methods have problems such as fuzzy segmentation contours, extremely high equipment requirements, and difficulty in visualizing the segmentation results from multiple aspects. Therefore, a method with high segmentation accuracy and multi-dimensional visualization of the segmentation target is an urgent problem to be solved. For this reason, a multi-modal medical image segmentation method is proposed. Summary of the Invention

[0006] The technical problem to be solved by the present invention is as follows: how to solve the problems existing in the existing segmentation methods, such as fuzzy segmentation contours, extremely high equipment requirements, and difficulty in visualizing the segmentation results from multiple aspects. A multi-modal medical image segmentation method is provided. Using multi-modal medical images as input, on the premise of considering the boundaries and accuracy of the segmented objects, the feature information of multi-modal medical images is extracted through two-dimensional and three-dimensional segmentation methods. Then, a feature fusion method is used to fuse the features of different dimensions, so that the final segmentation result can not only obtain context information but also display the segmented objects in multiple dimensions, providing a more accurate and multi-dimensional image basis for medical diagnosis.

[0007] The present invention solves the above technical problems through the following technical solutions. The present invention includes the following steps:

[0008] S1: Image preprocessing

[0009] Read the original image and preprocess the original image;

[0010] S2: Construct 2D and 3D segmentation networks

[0011] Construct 2D and 3D segmentation networks, and use the constructed 2D and 3D segmentation networks to segment the preprocessed image to obtain 2D and 3D segmentation results;

[0012] S3: Image fusion

[0013] Perform slicing processing on the 3D segmentation result for image fusion;

[0014] S4: Slicing splicing

[0015] Perform slicing splicing on the fusion result to obtain the final segmentation result.

[0016] Furthermore, in the step S1, the following process is included:

[0017] S11: After reading the original image, regularize the image through the Z-Score method to obtain image M1:

[0018] M1 = Z_score(M)

[0019] where M is the input image and Z_score is the standard score method;

[0020] S12: Obtain the centered image M2 through the following operations:

[0021]

[0022] where M 1_depth 、M 1_width 、M1_height They are the depth, width, and height of image M1 respectively;

[0023] S13: For the 2D segmentation network, slice image M2 according to the label position. For the 3D segmentation network, chop image M2 according to the depth to obtain M 2D and M 3D :

[0024] M2D = section depth (M2)

[0025] M 3D = Chopping size (M2)

[0026] where section depth means slicing using the number of depths, and Chopping size means chopping according to the size of size;

[0027] S14: For multi-modal data, by combining the slices or merged slices of each modality into multi-channels, and finally saving them in the form of an array and passing them into the corresponding segmentation network. The array form is specifically as follows:

[0028] Numpy 2D = (Width, Height, Modality)

[0029] Numpy 3D = (Width, Height, Size, Modality)

[0030] where Width, Height, Size, and Modeality correspond to the length, height, size, and number of modalities input to the network respectively.

[0031] Furthermore, in the step S2, the 2D segmentation network includes a first encoder and a first decoder; the first encoder includes a first convolutional block and a first downsampling layer, and the tensor is downsampled after passing through the first convolutional block to obtain feature maps of different sizes; the first decoder includes a second convolutional block and a first transposed convolutional layer, and the size is restored through the first transposed convolutional layer and then enters the second convolutional block; between the first encoder and the first decoder, information between different channels is obtained through the CE channel attention mechanism, and features at different levels are obtained through skip connections between corresponding layers.

[0032] Furthermore, the first convolutional block is the same as the second convolutional block, and each convolutional block includes two 2D convolutional layers. After each 2D convolutional layer, activation processing is performed through batch normalization and the ReLU function, as follows:

[0033]

[0034] where x 2Dinput and are the input and output of the 2D convolutional layer, ReLU is the ReLU activation function, BatchNormalization 2D is the batch normalization operation, Conv2D represents the 2D convolution operation, kernel is the convolution kernel size, and padding is the convolution padding length.

[0035] Furthermore, in the step S2, the 3D segmentation network includes a second encoder and a second decoder; the second encoder includes a third convolutional block and a second downsampling layer, and the tensor is downsampled after passing through the third convolutional block to obtain feature maps of different sizes; the second decoder includes a fourth convolutional block and a second transposed convolutional layer, and the size is restored through the second transposed convolutional layer and then enters the fourth convolutional block. The CBAM spatial attention mechanism is used to obtain more spatial information of different feature maps between the second encoder and the second decoder, and skip connections are used between corresponding layers to obtain features of different levels.

[0036] Furthermore, the third convolutional block and the fourth convolutional block are the same. Each convolutional block includes two 3D convolutional layers, and each 3D convolutional layer is activated by batch normalization and the ReLU function as follows:

[0037]

[0038] where x 3Dinput , are the input and output of the 3D convolutional layer, ReLU is the ReLU activation function, BatchNormalization 3D is the batch normalization operation, Conv3D is the 3D convolution operation, kernel is the convolution kernel size, and padding is the convolution padding length.

[0039] Furthermore, in the step S3, the following process is included:

[0040] S31: Slice the 3D segmentation result according to the depth to obtain 2D slices;

[0041] S32: Perform feature fusion with the 2D segmentation result corresponding to the 2D slice to obtain a new 2D fusion result.

[0042] Furthermore, in the step S32, the 2D segmentation result is decomposed by shearlets to obtain the frequency subbands of the image. For the low-frequency coefficients, they are processed using the fusion rule based on the absolute value of the regional coefficients and weights. For the high-frequency coefficients, their support vector values are calculated to determine the high-frequency fusion coefficients. Finally, the image is reconstructed according to the inverse translation-invariant shearlet transform.

[0043] The shearlet formula is as follows:

[0044]

[0045] where det is the determinant of the matrix, A is the scaling matrix, and B is the shear matrix.

[0046] Furthermore, in the step S4, the specific process of stitching is as follows: Create a blank matrix Z1 of the original image, perform central filling according to the serial number of the fused image slices, and fill the remaining blank areas with null values to obtain the stitched image Z2.

[0047] The present invention has the following advantages compared with the prior art: This multi-modal medical image segmentation method can make more full use of medical images by using multi-modal information; through the combination of 2D and 3D segmentation networks and the channel attention mechanism, the information of each modality can be interconnected and complementary; based on the fusion of 2D and 3D segmentation results, the segmentation result has higher accuracy and clearer edges; the segmentation results of different dimensions and different models are provided, providing a more accurate multi-dimensional image basis for medical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a schematic flowchart of the multi-modal medical image segmentation method in Embodiment 1 of the present invention;

[0049] Figure 2 is a schematic structural diagram of the 2D segmentation network in Embodiment 1 of the present invention;

[0050] Figure 3 is a schematic structural diagram of the 3D segmentation network in Embodiment 1 of the present invention;

[0051] Figure 4 is a schematic flowchart of feature fusion in Embodiment 1 of the present invention;

[0052] Figure 5 is a schematic flowchart of the multi-modal medical image segmentation method in Embodiment 2 of the present invention;

[0053] Figure 6 is the output result diagram in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following is a detailed description of the embodiments of the present invention. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.

[0055] Embodiment 1

[0056] As Figure 1 shown, this embodiment provides a technical solution: a multi-modal medical image segmentation method, including the following steps:

[0057] I. Preprocessing of the original image

[0058] 1) After reading the original data, the data is regularized by the Z-Score method to obtain the image M1:

[0059] M1 = Z_score(M)

[0060] where M is the input image and Z_score is the standard score method.

[0061] 2) The background of the medical image accounts for a large proportion in the whole image, and the background is of no help to segmentation. Since the area to be segmented is in the middle of the image, by centering the image M1, the black background that has no impact on segmentation is removed to balance the data. The centered image M2 is obtained through the following operations:

[0062]

[0063] where M 1_depth 、M 1_width 、M 1_height are the depth, width, and height of the image M1 respectively.

[0064] 3) For the 2D segmentation network (the 2D network, that is, the 2D segmentation model in Figure 1 ), the image M2 is sliced according to the label position. For the 3D segmentation network (the 3D network, that is, the 3D segmentation model in Figure 1 ), the image M2 is chopped according to the depth to obtain M 2D and M 3D :

[0065] M 2D = section depth (M2)

[0066] M 3D = Chopping size (M2)

[0067] where section depthIndicates slicing using the depth number, Chopping size Indicates cutting according to the size of size.

[0068] 4) For multi-modal data, by combining the slices or slice combinations of each modality into multi-channels, and finally saving them in the form of an array and passing them into the corresponding segmentation network. The specific form of the array is as follows:

[0069] Numpy 2D =(Width,Height,Modality)

[0070] Numpy 3D =(Width,Height,Size,Modality)

[0071] Among them, Width, Height, Size, and Modality respectively correspond to the length, height, size, and number of modalities of the input network.

[0072] II. Construct 2D and 3D segmentation networks

[0073] 1) Construct a 2D segmentation network. Its segmentation network uses an encoder and decoder structure composed of 2D convolutional blocks; the constructed network uses 4 downsamplings and 4 upsamplings to maintain a U-shaped symmetric structure, as Figure 2 shown.

[0074] When the image data is passed into the 2D segmentation network, convolutional blocks of different scales are used (the convolutional kernel sizes are 1 and 3 respectively), and the activation function is ReLU; in the last layer of the encoder, a channel attention mechanism is introduced to obtain the importance of each channel of the feature map, and then this importance is used to assign a weight value to each feature, so that the neural network focuses on certain feature channels. Improve the channels of the feature map useful for the current task and suppress the feature channels that are not very useful for the current task; finally, make full use of multi-modal data.

[0075] The decoder part uses the method of skip connection to obtain the features of low-level features, and fuses them with the existing features so that the semantics of high-level segmentation features and the details of low-level features are mutually fused, making the segmentation structure have both semantics and more details.

[0076] 2) Construct a 3D segmentation network. Its segmentation network uses an encoder and decoder structure composed of 3D convolutional blocks; the constructed network uses 4 downsamplings and 4 upsamplings to maintain a U-shaped symmetric structure, as Figure 3 shown.

[0077] After the data is input into the 3D segmentation network, convolutional blocks with different scales are used. The convolutional kernel sizes are 1 and 3, and the activation function is ReLU. The CBAM attention mechanism is used in the last layer of the encoder to combine the feature channels and the feature space dimensions, supplement the feature information of different-channel modalities by combining the context features in the spatial channels, and make up for the details of the spatial domain features of the 3D convolution.

[0078] III. Feature Fusion

[0079] Since 3D segmentation is performed based on blocks and there is no communication of context information between blocks, resulting in the lack of context information. Therefore, by fusing the features of the 3D segmentation result and the 2D segmentation result, the final segmentation result can obtain context information while also obtaining more global information, as Figure 4 shown. The specific process is as follows:

[0080] 1) Slice the 3D segmentation result according to the depth to obtain 2D slices;

[0081] 2) Decompose the 2D segmentation result through shearlets to obtain the frequency sub-bands of the image. For the low-frequency coefficients, they are processed using the fusion rule based on the sum of the absolute values of the regional coefficients and weights. For the high-frequency coefficients, their support vector values are calculated to determine the high-frequency fusion coefficients. Finally, the image is reconstructed according to the inverse translation-invariant shearlet transform, where the shearlet formula is:

[0082]

[0083] where det is the determinant of the matrix, A is the scale transformation matrix, and B is the shear matrix;

[0084] 3) Finally, splice the fused image slices, create the blank matrix Z1 of the original image, fill it centrally according to the serial numbers of the fused image slices, and fill the remaining blank areas with null values to obtain the spliced image Z2 and restore the segmentation result.

[0085] Example Two

[0086] The following mainly further illustrates the present invention in combination with the accompanying drawings and specific embodiments;

[0087] In this embodiment, multi-modal brain MRI images are selected for analysis, and four modal medical images, namely Flair, T1, T1ce, and T2, are selected to illustrate the corresponding results after the implementation of the present invention as Figure 6 , and the specific implementation steps are as follows (as Figure 5 shown):

[0088] A. The computer reads the original multi-modal MRI images and first preprocesses the data: First, the data of each modality is normalized using the Z-Score method; then, the background of the medical images is cropped by central cropping to reduce the network computation volume;

[0089] B. The 2D images are sliced and the 3D images are diced. After processing, the data of each model is merged into multi-channel data;

[0090] C. The 2D and 3D preprocessing results are respectively fed into 2D and 3D segmentation networks for segmentation prediction to obtain the 2D and 3D segmentation network prediction results;

[0091] D. The 3D segmentation network segmentation result is sliced and feature fusion is performed with the corresponding 2D segmentation result to obtain a new 2D fusion result;

[0092] E. The newly obtained 2D fusion result is sliced and restored to obtain the fused 3D segmentation result.

[0093] After implementing the above experiments, the final output results obtained are the 2D network prediction result, the 2D fusion result, the 3D network prediction result, and the 3D fusion result. The final output results are as Figure 6 shown.

[0094] It can be seen that the present invention aims at the defects of general medical image algorithms that only optimize single-modal data and do not consider the complementarity of 2D and 3D segmentation networks; general medical image algorithms have single input and single output and do not consider the problem of missing context information in obtaining image features.

[0095] In summary, the multi-modal medical image segmentation method of the above embodiments not only makes full use of multi-modal medical images, but also enables the information between channels to be complementary by adding a channel self-attention mechanism to the 2D and 3D segmentation networks; then, through a feature fusion method, the problem of possible information loss in a single network is optimized, and more effective medical image segmentation is achieved according to multi-modal medical images.

[0096] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A multi-modal medical image segmentation method, characterized in that, It includes the following steps: S1: Image preprocessing Read the original image and preprocess the original image; S2: Construct 2D and 3D segmentation networks Construct 2D and 3D segmentation networks, and use the constructed 2D and 3D segmentation networks to segment the preprocessed image to obtain 2D and 3D segmentation results; S3: Image fusion Slice the 3D segmentation result and perform image fusion; In step S3, it includes the following process: S31: Slice the 3D segmentation result according to the depth to obtain 2D slices; S32: Perform feature fusion with the 2D segmentation result corresponding to the 2D slice to obtain a new 2D fusion result; In step S32, decompose the 2D segmentation result through shearlet to obtain the frequency subbands of the image. For the low-frequency coefficients, process them using the regional coefficient absolute value and weight fusion rule. For the high-frequency coefficients, calculate their support vector values to determine the high-frequency fusion coefficients, and finally reconstruct the image according to the inverse translation-invariant shearlet transform; S4: Slice stitching Perform slice stitching on the fusion result to obtain the final segmentation result.

2. The multimodal medical image segmentation method according to claim 1, wherein: In step S1, it includes the following process: S11: After reading the original image, regularize the image through the Z-Score method to obtain image M1: M1 = Z_score(M) where M is the input image and Z_score is the standard score method; S12: Obtain the centered image M2 through the following operations: Among them, M 1_depth , M 1_width , M 1_height are respectively the depth, width, and height of the image M1; S13: For the 2D segmentation network, slice the image M2 according to the label position, and for the 3D segmentation network, cut the image M2 according to the depth to obtain M 2D and M 3D : M 2D = section depth (M2) M 3D = Chopping size (M2) Among them, section depth means slicing using the depth number, Chopping size means cutting according to the size of size; S14: For multi-modal data, combine the slices or slice combinations of each modality into multi-channels, and finally save them in the form of an array and pass them into the corresponding segmentation network. The specific form of the array is as follows: Numpy 2D =(Width, Height, Modality) Numpy 3D =(Width, Height, Size, Modality) where Width, Height, Size, and Modality correspond to the length, height, size, and number of modalities of the input network respectively.

3. A multimodal medical image segmentation method according to claim 2, characterized in that: In step S2, the 2D segmentation network includes a first encoder and a first decoder; the first encoder includes a first convolutional block and a first downsampling layer, and the tensor is downsampled after passing through the first convolutional block to obtain feature maps of different sizes; the first decoder includes a second convolutional block and a first transposed convolutional layer, and the size is restored through the first transposed convolutional layer and then enters the second convolutional block; between the first encoder and the first decoder, obtain the information between different channels through the CE channel attention mechanism, and obtain the features of different levels through skip connections between the corresponding layers.

4. A multimodal medical image segmentation method according to claim 3, characterized in that: The first convolutional block is the same as the second convolutional block. Each convolutional block includes two 2D convolutional layers, and after each 2D convolutional layer, activation processing is performed through batch normalization and the ReLU function, as follows: Among them, x 2Dinput and are the input and output of the 2D convolutional layer, ReLU is the ReLU activation function, BatchNormalization 2D is the batch normalization operation, Conv2D represents the 2D convolution operation, kernel is the convolution kernel size, and padding is the convolution padding length.

5. A multimodal medical image segmentation method according to claim 3, characterized in that: In step S2, the 3D segmentation network includes a second encoder and a second decoder; the second encoder includes a third convolutional block and a second downsampling layer, and the tensor is downsampled after passing through the third convolutional block to obtain feature maps of different sizes; the second decoder includes a fourth convolutional block and a second transposed convolutional layer, and the size is restored through the second transposed convolutional layer and then enters the fourth convolutional block. More spatial information of different feature maps is obtained through the CBAM spatial attention mechanism between the second encoder and the second decoder, and different levels of features are obtained through skip connections between corresponding layers.

6. A multimodal medical image segmentation method according to claim 5, characterized in that: The third convolutional block and the fourth convolutional block are the same. Each convolutional block includes two 3D convolutional layers, and each 3D convolutional layer is activated through batch normalization and the ReLU function, as follows: where x 3Dinput , is the input and output of the 3D convolutional layer, ReLU is the ReLU activation function, BatchNormalization 3D is the batch normalization operation, Conv3D is the 3D convolution operation, kernel is the convolution kernel size, and padding is the convolution padding length.

7. A multimodal medical image segmentation method according to claim 1, characterized in that: The shearlet formula is as follows: where det is the determinant of the matrix, A is the scaling matrix, and B is the shear matrix.

8. A multimodal medical image segmentation method according to claim 7, characterized in that: In step S4, the specific process of stitching is as follows: create a blank matrix Z1 of the original image, perform central filling according to the serial number of the fused image slice, and fill the remaining blank areas with null values to obtain the stitched image Z2.

Citation Information

Patent Citations

  • 3D medical image segmentation method and device based on hierarchical perception fusion and storage medium

    CN112465754A