Multimodal MRI Brain Tumor Image Segmentation Method and System Robust to Missing Modalities
By adopting the front-to-back interactive learning method in multimodal MRI brain tumor image segmentation, the encoder-decoder 3D-UNet network and attention mechanism are used for characterization fusion, solving the problem of degradation of segmentation effect in the absence of modality, and achieving efficient and robust multimodal segmentation.
Patent Information
- Application Number
- CN202311244917.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-09-25
AI Technical Summary
The existing multimodal MRI brain tumor image segmentation algorithm performs poorly in the absence of modality, resulting in a decrease in segmentation effect.
Using a method based on front-to-back interactive learning, single-modal unique characterization and multimodal fusion characterization are extracted through the encoder-decoder 3D-UNet network, and adaptive dynamic fusion is used to generate robust post-interaction fusion characterization.
In the absence of modality, the method can effectively retain single-modal information and fully integrate multi-modal information, improve segmentation accuracy and robustness, and is better than the performance of the single-modal segmentation model.
Smart Images

Figure CN117237381B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical image technology and is applied to medical image segmentation. Specifically, it relates to a multi-modal MRI brain tumor image segmentation method and system that is robust to missing modalities. More specifically, it relates to a multi-modal MRI brain tumor image segmentation method and system that is robust to missing modalities based on forward and backward interactive learning. Background Art
[0002] Brain tumor segmentation is a key technology for quantitatively analyzing brain tumor regions through medical imaging data, and it is crucial for the diagnosis and prognosis of glioma patients. Clinically, doctors often need to observe, analyze, and diagnose the condition by means of medical images, and precise measurements of the lesion area are also required before surgery. Although manually segmenting the lesion area can also achieve high accuracy, it is often time-consuming and laborious and there is also instability caused by the experience level. Brain tumor segmentation algorithms can use computer assistance to complete operations such as edge detection, clustering, and threshold division, effectively improving the efficiency of lesion area segmentation while ensuring a certain accuracy.
[0003] Multi-modal brain tumor segmentation algorithms can make full use of the complementarity of different modalities of MRI images to improve the segmentation accuracy, because using only a single modality cannot fully express the information of the lesion. Recently, more and more researchers have used multi-modal brain tumor segmentation algorithms to improve the segmentation accuracy of the model, and techniques such as pre-fusion, feature fusion, and output distribution fusion of different modalities of MRI images are used to extract complementary information between different modalities.
[0004] Although using multi-modal data for image segmentation has good effects, in real-world conditions, the problem of missing some modal data often occurs. First, the types of medical images that patients can take are limited by the types of medical devices in the hospital and the corresponding professional medical staff; second, some patients will also consider factors such as contrast time and medical expenses and give up taking some modal medical images; in addition, reasons such as allergies to contrast agents and metal implants in the body will also cause the inability to take contrast-enhanced CT or magnetic resonance imaging, resulting in the missing of some modalities. And the missing of some modal data will lead to insufficient information obtained by existing multi-modal segmentation algorithms, resulting in a decline in the segmentation effect.
[0005] Patent document CN114782350A (application number: 202210393464.8) discloses a method for segmenting MRI brain tumor images based on multi-modal feature fusion with an attention mechanism. In this invention, the backbone network is encoded, then the global and local information is perceived through a hybrid context awareness module, and finally the multi-modal features are fused by an attention fusion module and an image is output. However, the structure designed in this patent cannot be used in the case of missing modalities, and it is prone to be restricted when applied in real medical scenarios, ignoring the robustness of the model to various missing scenarios. Especially in scenarios where only a single modality can be used, the performance of such multi-modal segmentation models is even inferior to that of single-modal segmentation models. Summary of the Invention
[0006] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method and system for segmenting multi-modal MRI brain tumor images that are robust to missing modalities.
[0007] A method for segmenting multi-modal MRI brain tumor images that are robust to missing modalities according to the present invention includes:
[0008] Step S1: Collect multi-modal MRI image data, and preprocess the multi-modal MRI image data to obtain preprocessed multi-modal MRI image data;
[0009] Step S2: Input the single-modal MRI image data in the preprocessed multi-modal MRI image data into multiple different encoder-decoder 3D-UNet networks respectively to obtain discriminative representations of each modality;
[0010] Step S3: Connect the single-modal MRI image data in the preprocessed multi-modal MRI image data in the channel dimension, and input it into the pre-interaction encoder-decoder 3D-UNet network to obtain a pre-interaction fusion representation;
[0011] Step S4: Based on the attention mechanism, adaptively and dynamically fuse the pre-interaction fusion representation and the discriminative representations of each modality to obtain a post-interaction fusion representation;
[0012] Step S5: Input the discriminative representations of each modality, the pre-interaction fusion representation, and the post-interaction fusion representation into different single-layer deep convolutional neural networks respectively to obtain the tumor segmentation maps of each.
[0013] Preferably, in step S1, the following operations are adopted: perform data augmentation processing on the multi-modal MRI image data, including normalization processing, as well as random cropping, mirror flipping, and rotation;
[0014] The multi-modal MRI images include: T1-weighted imaging, T2-weighted imaging, T1ce-weighted imaging, and Flair
[0015] Imaging
[0016] Preferably, the encoder-decoder 3D-UNet network includes: an encoder in four stages, a bottleneck layer, and a decoder in four stages;
[0017] The encoder gradually reduces the spatial resolution and increases the number of representation channels stage by stage to construct a pyramid of representations, obtaining encoded representations Input into the bottleneck layer, and use the bottleneck layer to extract high-dimensional semantic representations. The output of the bottleneck layer is input into the decoder, and the decoder maps the abstract representations extracted by the bottleneck layer back to the resolution of the original input image, obtaining representations containing position and structural information through gradual upsampling and feature fusion
[0018] Preferably, step S3 adopts: concatenating x Flair , x T1 , x T1ce , x T2 in the channel dimension to obtain x Pre ;
[0019] x Pre = CAT([δ m ·x m ) m∈Ω )
[0020] where CAT represents the operation of concatenation in the channel dimension; δ m indicates whether modality m exists; x m represents modality m; Ω = {Flair, T1, T1ce, T2}; x Flair , x T1 , x T1ce , x T2 respectively represent the preprocessed Flair imaging data, preprocessed T1-weighted imaging data, preprocessed T1ce-weighted imaging data, and preprocessed T2-weighted imaging data;
[0021] Input x Pre into the pre-interaction encoder-decoder 3D-UNet network with 4 input channels, obtaining encoded representations in four stages as well as decoded representations in four stages
[0022] Preferably, supervise the representations output by different stages of the encoder-decoder 3D-UNet network;
[0023] The encoder representation of the s-th stage and the decoder representation of the (s + 1)-th stage After upsampling by a factor of two, fusion is performed in the channel dimension, and the result is input into a convolutional neural network with a 1x1x1 convolutional kernel and a softmax function, and the predicted map for brain tumor segmentation at the s-th stage is output.
[0024]
[0025] where ψ m,s A convolutional neural network with a 1x1x1 convolutional kernel and a softmax function; are its learnable parameters, and U represents the operation of upsampling the representation by a factor of two; is supervised by the reference brain tumor segmentation map y;
[0026] Using the standard segmentation loss function for constraint, including the Dice loss and the weighted cross-entropy loss
[0027]
[0028] where ∩ represents the intersection operation between the predicted segmentation map and the reference segmentation map, ||·||1 represents the norm 1, K = {NCR / NET, ED, ET, Background} is a set of different regions, which contains four different regions, namely necrotic and non-enhancing tumor cores, peritumoral edema, and enhancing tumor cores and background; K num is the number of the set K;
[0029]
[0030] is a value used to adjust the weight of each region;
[0031] For supervision is expressed as the formula:
[0032]
[0033] The multi-class segmentation loss function is also defined as the sum of each stage :
[0034]
[0035] where U s is the upsampling operation, which is used to match the size of the prediction at the s-th stage and the actual reference segmentation map.
[0036] Preferably, in step S4, the following is adopted: based on the attention mechanism, the outputs of the last layers of different encoder-decoder 3D-UNet networks are fused to obtain and a representation with a dimension of BxCxHxWxD; using average pooling operation, the average value of HxWxD for each channel is taken to collapse the representation dimension to BxC; the collapsed dimensions are concatenated in the channel dimension to obtain a feature of size Bx5C, which is input into a multi-layer perceptron. The output channel number of the multi-layer perceptron is 5, and then through the Sigmoid function, the weights {w m} m∈Ω and w Pre for the representation are obtained. Based on the weights, a weighted sum of and is performed to obtain the post-interaction fusion representation f Post :
[0037]
[0038]
[0039]
[0040] where f i Pool is the collapsed representation and w i are the weights of the respective corresponding representations.
[0041] Preferably, in step S5, the following is adopted: a convolutional layer with a kernel size of 1x1x1, the input channel number is a preset value, the output channel number is 4, and the output channels respectively correspond to four regions of K={NCR / NET, ED, ET, Background}; the output of the convolutional layer also passes through the softmax function, and the softmax acts on the output channel dimension; and after passing through this convolutional layer and softmax, the respective segmentation maps and are obtained. The segmentation maps are binarized to obtain the binarized segmentation maps of the four regions.
[0042] Preferably, the standard segmentation loss supervises all the output segmentation maps to align them with the reference segmentation map; the total loss function is as follows:
[0043]
[0044] where and represent the standard segmentation loss and the multi-connection segmentation loss applied to each modality m.
[0045] A multi-modal MRI brain tumor image segmentation system robust to missing modalities provided by the present invention includes:
[0046] Module M1: Collect multi-modal MRI image data and preprocess the multi-modal MRI image data to obtain preprocessed multi-modal MRI image data;
[0047] Module M2: Input the single-modal MRI image data in the preprocessed multi-modal MRI image data into multiple different encoder-decoder 3D-UNet networks respectively to obtain discriminative representations of their respective modalities;
[0048] Module M3: Concatenate the single-modal MRI image data in the preprocessed multi-modal MRI image data in the channel dimension and input it into the pre-interaction encoder-decoder 3D-UNet network to obtain a pre-interaction fusion representation;
[0049] Module M4: Adaptively and dynamically fuse the pre-interaction fusion representation and the discriminative representations of their respective modalities based on the attention mechanism to obtain a post-interaction fusion representation;
[0050] Module M5: Input the discriminative representations of their respective modalities, the pre-interaction fusion representation, and the post-interaction fusion representation into different one-layer deep convolutional neural networks respectively to obtain their respective tumor segmentation maps.
[0051] Preferably, the module M1 adopts: performing data augmentation processing on the multi-modal MRI image data including normalization processing and random cropping, mirror flipping, and rotation;
[0052] The multi-modal MRI images include: T1-weighted imaging, T2-weighted imaging, T1ce-weighted imaging, and Flair imaging;
[0053] The encoder-decoder 3D-UNet network includes: an encoder in four stages, a bottleneck layer, and a decoder in four stages;
[0054] The encoder gradually reduces the spatial resolution and increases the number of representation channels stage by stage to construct a pyramid of representations, obtaining encoded representations Input into the bottleneck layer, using the bottleneck layer to extract high-dimensional semantic representations, the output of the bottleneck layer is input into the decoder, and the decoder maps the abstract representations extracted by the bottleneck layer back to the resolution of the original input image, obtaining representations containing position and structural information through gradual upsampling and feature fusion
[0055] The module M3 adopts: taking x Flair , x T1 , x T1ce , x T2 Connect on the channel dimension to obtain x Pre ;
[0056] x Pre = CAT([δ m ·x m ) m∈Ω )
[0057] where CAT represents the operation of concatenation on the channel dimension; δ m indicates whether modality m exists; x m represents modality m; Ω = {Flair, T1, T1ce, T2}; x Flair , x T1 , x T1ce , x T2 respectively represent the preprocessed Flair imaging data, preprocessed T1-weighted imaging data, preprocessed T1ce-weighted imaging data, and preprocessed T2-weighted imaging data;
[0058] Input x Pre into the pre-interaction encoder-decoder 3D-UNet network with 4 input channels to obtain the encoded representations of four stages and the decoded representations of four stages
[0059] Supervise the representations output by the encoder-decoder 3D-UNet network at different stages;
[0060] Upsample the encoder representation of the s-th stage and the decoder representation of the (s + 1)-th stage by a factor of two and fuse them in the channel dimension, then input them into a convolutional neural network with a 1x1x1 convolutional kernel and a softmax function to output the prediction map for brain tumor segmentation at the s-th stage
[0061]
[0062] where ψ m,s is a convolutional neural network with a 1x1x1 convolutional kernel and a softmax function; are its learnable parameters, and U represents the operation of upsampling the representation by a factor of two; It is supervised by the reference brain tumor segmentation map y;
[0063] Use the standard segmentation loss function for constraint, including the Dice loss and the weighted cross-entropy loss
[0064]
[0065] Among them, ∩ represents the intersection operation of the predicted segmentation map and the reference segmentation map, ||·||1 represents the norm 1, and K = {NCR / NET, ED, KT, Backgeound} is a set of different regions, which altogether contains four different regions, namely necrotic and non-enhancing tumor nuclei, peritumoral edema, and enhancing tumor nuclei and background; K num is the number of the set K;
[0066]
[0067] is a value used to adjust the weight of each region;
[0068] For supervision is expressed as the formula:
[0069]
[0070] Multi-class segmentation loss function is also defined as the sum of each stage :
[0071]
[0072] Among them, U s is the upsampling operation, which is used to match the size of the prediction and the actual reference segmentation map in the s-th stage;
[0073] The module M4 adopts: Based on the attention mechanism, the outputs of the last layers of different encoder-decoder 3D-UNet networks are fused to obtain and The characterization dimension size is BxCxHxWxD; Using the average pooling operation, the average value of HxWxD of each channel is taken to collapse the characterization dimension to BxC; The collapsed dimension is concatenated in the channel dimension to obtain a feature size of Bx5C, which is input into a multi-layer perceptron. The output channel number of the multi-layer perceptron is 5, and then through the Sigmoid function, the weights {w m} m∈Ω and w Pre for the characterization are obtained. Based on the weights, and are weighted and summed to obtain the post-interaction fusion characterization f Post :
[0074]
[0075]
[0076]
[0077] Among them, is the collapsed representation, and w i is the weight of each corresponding representation;
[0078] The module M5 adopts: a convolutional layer with a convolutional kernel size of 1x1x1, the number of input channels is a preset value, the number of output channels is 4, and the output channels respectively correspond to four regions of K = {NCR / NET, ED, ET, Background}; the output of the convolutional layer also passes through the softmax function, and the softmax acts on the output channel dimension; and f Post After passing through this convolutional layer and softmax, the respective segmentation maps are obtained and The segmentation maps are binarized to obtain binarized segmentation maps of the four regions;
[0079] The standard segmentation loss will supervise all the output segmentation maps to align them with the reference segmentation map; the total loss function is as follows:
[0080]
[0081] Among them, and represent the standard segmentation loss and the multi-stage segmentation loss applied to each modality m.
[0082] Compared with the prior art, the present invention has the following beneficial effects:
[0083] 1. The present invention not only well preserves the information of a single modality, but also has a good multi-modal fusion effect. Such a characteristic satisfies the robustness of the model under various missing situations and improves the performance of the model in the case of modality loss.
[0084] 2. The present invention fully excavates and preserves the unique information of each modality effective for the current segmentation through the single-modal specific representation learning step; through the pre-interaction representation learning step, the multi-modal information is fully fused so that the multi-modal information can interact sufficiently; finally, through the post-interaction fusion representation, all the representations are fused together according to the missing situation. The post-interaction representation completely preserves the single-modal information and at the same time contains the correlation information between modalities.
[0085] 3. In order to promote the single-modal specific representation learning step and the pre-interaction representation learning step to fully learn the position and structure information of the brain tumor, the present invention designs a standard segmentation loss and a multi-class segmentation loss to promote the model's learning of the position and structure information;
[0086] 4. To ensure the effectiveness of post-interaction representation learning, the present invention uses a standard segmentation loss to constrain the learning of post-interaction representations; by organically combining all loss functions and jointly applying them to the training of the model, each step can perform its own functions. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Other features, objectives, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0088] Figure 1 It is a flowchart of the method in an embodiment of the present invention;
[0089] Figure 2 It is a module structure diagram of post-interaction representation learning in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0090] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0091] Embodiment 1
[0092] As Figure 1-2 shown, a multi-modal MRI brain tumor image segmentation method based on pre-and-post interaction learning and robust to missing modalities provided by the present invention includes:
[0093] MRI image data preprocessing step: The multi-modal MRI brain tumor image contains four nuclear magnetic resonance modalities, namely T1-weighted imaging, T2-weighted imaging, T1ce-weighted imaging, and Flair imaging. Normalize all modal data and perform data augmentation such as random cropping, mirror flipping, and rotation.
[0094] Single-modal specific representation learning step: Each preprocessed MRI modality is input into an independent encoder-decoder 3D-UNet network, and each network learns discriminative representations specific to its own modality.
[0095] Pre-interaction representation learning step: The available preprocessed multi-modal MRI images are concatenated in the channel dimension and input into a pre-interaction encoder-decoder 3D-UNe network. This independent network learns representations of inter-modal correlations, called pre-interaction fusion representations.
[0096] Representation hierarchical supervision step: Supervise the representations output by different levels of the 3D-UNet so that the representations can fully express the structural information and location information of the brain tumor.
[0097] Post - interaction representation learning step: According to the missing conditions of different modalities, adaptively and dynamically fuse the representations of inter - modality correlations and discriminative representations unique to single modalities obtained in the pre - interaction learning step based on the attention mechanism to obtain the post - interaction fusion representation.
[0098] Representation segmentation step: Input the discriminative representations unique to single modalities, the pre - interaction fusion representation, and the post - interaction fusion representation into different single - layer deep convolutional neural networks respectively to obtain their respective brain tumor segmentation maps.
[0099] The present invention fully mines and retains the information of different MRI modalities, captures the correlations between multi - modalities at the same time, and makes the multi - modality fusion more sufficient. According to the modality missing conditions, the model dynamically fuses the representations of different modalities to achieve more robust multi - modality MRI brain tumor segmentation.
[0100] Specifically, the MRI image data pre - processing step includes: pre - processing the input 3D MRI image data. Randomly crop the MRI image with a size of HxWxD (both H, W, and D are guaranteed to be greater than 80) to 80x80x80, and at the same time perform data augmentation operations such as randomly flipping and rotating the image along different axes. Then, truncate the values exceeding 4000 and perform Z - Score normalization on non - zero values. Process the original MRI image into inputs x Flair ,x T1 ,x T1ce ,x T2 。
[0101] Specifically, the single - modality specific representation learning step includes: using different 3D - UNet networks with an input channel number of 1, respectively perform the learning of modality - specific discriminative representations on different modality inputs x Flair ,x T1 ,x T1ce ,x T2 for the learning of modality - specific discriminative representations. The 3D - UNet structure used includes four - stage encoders, a bottleneck layer, and four - stage decoders. The encoder gradually reduces the spatial resolution and increases the number of representation channels stage by stage to construct a pyramid of representations, and the obtained encoded representation is where m refers to the representation of which input modality, Ω = {Flair, T1, T1ce, T2}, s represents the representation at which stage in the encoder, s ∈ {1, 2, 3, 4}. And Input into the bottleneck layer, which further extracts high-dimensional semantic representations. The output of the bottleneck layer will be input into the decoder. The decoder maps the abstract representations extracted by the bottleneck layer back to the resolution of the original input image, obtaining representations containing more location and structural information through gradual upsampling and feature fusion. The obtained representation is
[0102] Specifically, the pre-interaction representation learning step includes: concatenating x Flair , x T1 , x T1ce , x T2 in the channel dimension to obtain x Pre . Meanwhile, considering the case of missing modalities, δ m is introduced to indicate whether modality m exists. δ m is equal to 0 or 1. Therefore, the obtained x Pre can be represented by the following formula:
[0103] x Pre = CAT([δ m ·x m m∈Ω )
[0104] CAT represents the operation of concatenation in the channel dimension, and · represents the operation of element-wise multiplication. The input of the missing modality will be all set to zero, while the non-missing modality remains the original input. Input x Pre into a 3D-UNet model with 4 input channels to obtain the encoded representations of four stages and the decoded representations of four stages
[0105] Specifically, the representation hierarchical supervision step includes: supervising the outputs of the four stages of the 3D-UNet encoder and decoder, aiming to make the learned representations fully express the location and structural information of the brain tumor. Specifically, the encoded representation of the s-th stage and the decoded representation of the (s + 1)-th stage are upsampled by a factor of two and then fused in the channel dimension, and input into a convolutional neural network with a 1x1x1 convolutional kernel and a softmax function to output the prediction map for brain tumor segmentation at the s-th stage The process can be represented by the following formula:
[0106]
[0107] where ψ m,s is a convolutional neural network with a 1x1x1 convolutional kernel and a softmax function, are its learnable parameters, and U represents the operation of upsampling the representation by a factor of two. The is supervised by the reference brain tumor segmentation map y, enabling it to have a better perception of the location and structure of the brain tumor. Using the standard segmentation loss function to perform the constraint, including the Dice loss and the weighted cross-entropy loss Its form is shown in the following formula:
[0108]
[0109] To simplify the writing, the subscript of is omitted, and is used to represent the prediction results of all models for the segmentation map. Among them, ∩ represents the intersection operation of the predicted segmentation map and the reference segmentation map, ||·||1 represents the norm 1, and K = {NCR / NET, ED, ET, Background} is a set of different regions, which includes a total of four different regions, namely necrotic and non-enhancing tumor core (NCR / NET), peritumoral edema (ED), enhancing tumor core (ET), and background. K num is the number of the set K, where K num = 4. Its form is shown in the following formula:
[0110]
[0111] The weighted cross-entropy loss is to improve the segmentation accuracy of only the smaller region area considering the large difference in the area of different segmentation regions. For example, in brain tumor segmentation, the area of the enhancing tumor core is much smaller than the area of the necrotic and non-enhancing tumor core (NCR / NET). is a value used to adjust the weight of each region. Finally, the used for supervision is expressed as the following formula:
[0112]
[0113] The multi-class segmentation loss (Hierarchical-Stage Segmentation Loss) function is also defined as the sum of each stage as shown in the following formula:
[0114]
[0115] where U s is the upsampling operation used to match the size of the prediction at the s-th stage and the actual reference segmentation map.
[0116] Specifically, the post-interaction representation learning step includes: fusing the outputs of the last layers of the decoders of different 3D-UNets based on the attention mechanism, including and The representation dimension size is BxCxHxWxD. Using the average pooling operation, the representation dimension is collapsed to BxC, that is, the average value of HxWxD for each channel is taken. The features obtained by concatenating the collapsed dimensions in the channel dimension have a size of Bx5C and are input into a multi-layer perceptron (MLP). The output channel number of the MLP is 5, and then through the Sigmoid function, the weights {w m} m∈Ω and w Pre for the representation are obtained. Based on these weights, a weighted sum is performed on and to obtain the post-interaction fusion representation f Post , and the specific process is as follows:
[0117]
[0118]
[0119]
[0120] where f i Pool is the collapsed representation, and w i are the weights of the respective corresponding representations.
[0121] Specifically, the representation segmentation step includes: a convolutional layer with a kernel size of 1x1x1, an input channel number of 30, and an output channel number of 4. The output channels correspond to four regions of K={NCR / NET, ED, ET, Background} respectively. The output of the convolutional layer also passes through the softmax function, and the softmax acts on the output channel dimension. and f Post pass through this convolutional layer and softmax to obtain their respective segmentation maps and
[0122] The standard segmentation loss will supervise all the output segmentation maps to align them with the reference segmentation map. Finally, the entire network is trained in an end-to-end manner, and its total loss function is as follows:
[0123]
[0124] and Denote the standard segmentation loss and the multi-connection segmentation loss applied to each modality \(m\). It should be noted that post-processing is required for the output finally, that is, binarization of the segmentation map is performed. At each position, the channel with the largest value is taken as 1, and the remaining channels are 0. Thus, the binarized segmentation maps of the four regions are obtained.
[0125] A multi-modal MRI brain tumor image segmentation system based on forward and backward interaction learning and robust to missing modalities provided by the present invention includes:
[0126] MRI image data preprocessing module: The multi-modal MRI brain tumor image contains four nuclear magnetic resonance modalities, namely T1-weighted imaging, T2-weighted imaging, T1ce-weighted imaging, and Flair imaging. Normalize all modality data and perform data augmentation such as random cropping, mirror flipping, and rotation.
[0127] Single-modal specific feature learning module: Each preprocessed MRI modality is input into an independent encoder-decoder 3D-UNet network, and each network learns discriminative features specific to its respective modality.
[0128] Forward interaction feature learning module: Connect the available preprocessed multi-modal MRI images in the channel dimension and input them into a forward interaction encoder-decoder 3D-UNe network. This independent network learns features representing the inter-modal correlation, called forward interaction fusion features.
[0129] Feature hierarchical supervision module: Supervise the features output by different levels of the 3D-UNet so that the features fully express the structural information and location information of the brain tumor.
[0130] Backward interaction feature learning module: According to the missing situation of different modalities, adaptively and dynamically fuse the features representing the inter-modal correlation obtained by the forward interaction learning module and the discriminative features specific to the single modality based on the attention mechanism to obtain the backward interaction fusion features.
[0131] Feature segmentation module: Input the discriminative features specific to the single modality, the forward interaction fusion features, and the backward interaction fusion features into different layers of a deep convolutional neural network respectively to obtain their respective brain tumor segmentation maps.
[0132] Specifically, the MRI image data preprocessing module includes: Preprocess the input 3D MRI image data. Randomly crop the MRI image with a size of \(H\times W\times D\) (\(H\), \(W\), and \(D\) are all guaranteed to be greater than 80) to \(80\times80\times80\), and at the same time perform data augmentation operations such as randomly flipping and rotating the image along different axes. Then, truncate the values exceeding 4000 and perform Z-Score normalization on non-zero values. Process the original MRI image into the input \(x\) of different modalities.Flair , x T1 , x T1ce , x T2 。
[0133] Specifically, the single-modal specific representation learning module includes: using different 3D-UNet networks with an input channel number of 1 to separately learn the modality-specific discriminative representations for different modality inputs x Flair , x T1 , x T1ce , x T2 The 3D-UNet structure used contains an encoder with four stages, a Bottleneck Layer, and a decoder with four stages. The encoder gradually reduces the spatial resolution and increases the number of representation channels stage by stage to build a pyramid of representations, and the resulting encoded representation is where m refers to the representation of which input modality, Ω = {Flair, T1, T1ce, T2}, s represents the representation of which stage in the encoder, and s ∈ {1, 2, 3, 4}. The is input into the bottleneck layer, and the bottleneck layer further extracts high-dimensional semantic representations. The output of the bottleneck layer will be input into the decoder. The decoder maps the abstract representation extracted by the bottleneck layer back to the resolution of the original input image and obtains a representation containing more position and structural information through gradual upsampling and feature fusion. The resulting representation is
[0134] Specifically, the pre-interaction representation learning module includes: concatenating x Flair , x T1 , x T1ce , x T2 in the channel dimension to obtain x Pre , and considering the case of missing modalities, introducing δ m to indicate whether modality m exists. δ m is equal to 0 or 1. Therefore, the obtained x Pre can be represented by the following formula:
[0135] x Pre = CAT([δ m ·x m ) m∈Ω )
[0136] CAT represents the operation of concatenation in the channel dimension, and · represents the operation of element-wise multiplication. The input of the missing modality will be all set to zero, while the non-missing modality remains the original input. Input x Pre into a 3D-UNet model with an input channel of 4 to obtain the encoded representations of four stages and the decoded representations of four stages
[0137] Specifically, the described hierarchical supervision module includes: Supervising the outputs of the four stages of the 3D-UNet encoder and decoder, aiming to make the learned representations fully express the brain tumor location and structural information. Specifically, the encoder representation of the s-th stage and the decoder representation of the (s + 1)-th stage are upsampled by a factor of two and then fused in the channel dimension, and input into a convolutional neural network with a 1x1x1 convolutional kernel and a softmax function to output the prediction map for brain tumor segmentation at the s-th stage The process can be expressed by the following formula:
[0138]
[0139] where ψ m,s a convolutional neural network with a 1x1x1 convolutional kernel and a softmax function, is its learnable parameter, and U represents the operation of upsampling the representation by a factor of two. This is supervised by the reference brain tumor segmentation map y, enabling it to have a better perception of the brain tumor location and structure. Using the standard segmentation loss function to perform the constraint, including the Dice loss and the weighted cross-entropy loss Its form is shown in the following formula:
[0140]
[0141] To simplify the writing, the subscript of is omitted, and is used to represent the prediction results of all models for the segmentation map. Where ∩ represents the intersection operation of the predicted segmentation map and the reference segmentation map, ||·||1 represents the norm 1, K = {NCR / NET, ED, ET, Background} is a set of different regions, which contains four different regions in total, namely necrotic and non-enhancing tumor core (NCR / NET), peritumoral edema (ED), and enhancing tumor core (ET) and background. K num is the number of the set K, here K num = 4. Its form is shown in the following formula:
[0142]
[0143] The weighted cross - entropy loss is used to improve the segmentation accuracy of smaller - area regions considering the large difference in the area of different segmentation regions. For example, in brain tumor segmentation, the area of the enhanced tumor core is much smaller than the area of the necrotic and non - enhanced tumor core (NCR / NET). It is a value used to adjust the weight of each region. Finally, the is expressed as the following formula:
[0144]
[0145] The multi - class segmentation loss (Hierarchical - Stage Segmentation Loss) function is also defined as the sum of each stage as shown in the following formula:
[0146]
[0147] where U s is the up - sampling operation, used to match the size of the prediction in the s - th stage and the actual reference segmentation map.
[0148] Specifically, the post - interaction representation learning module includes: fusing the outputs of the last layers of the decoders of different 3D - UNets based on the attention mechanism, including and The representation dimension size is BxCxHxWxD. Using the average pooling operation, the representation dimension is collapsed to BxC, that is, taking the average value of HxWxD for each channel. The features obtained by concatenating the collapsed dimensions in the channel dimension have a size of Bx5C and are input into a multi - layer perceptron MLP. The output channel number of the multi - layer perceptron is 5, and then through the Sigmoid function to obtain the weights {w m} m∈Ω and w Pre , based on these weights, perform weighted sums on and to obtain the post - interaction fusion representation f Post , and the specific process is as follows:
[0149]
[0150]
[0151]
[0152] where f i Pool is the collapsed representation, and w i is the weight of each corresponding representation.
[0153] Specifically, the described characterization segmentation module includes: a convolutional layer with a kernel size of 1x1x1, an input channel number of 30, and an output channel number of 4. The output channels respectively correspond to four regions of K = {NCR / NET, ED, ET, Background}. The output of the convolutional layer also passes through the softmax function, and the softmax acts on the output channel dimension. and f Post After passing through this convolutional layer and softmax, respective segmentation maps are obtained. and
[0154] Standard segmentation loss will supervise all the output segmentation maps to align them with the reference segmentation map. Finally, the entire network is trained in an end-to-end manner, and its total loss function is as follows:
[0155]
[0156] and represents the standard segmentation loss and the multi-class segmentation loss applied to each modality m. It should be noted that finally, post-processing is performed on the output, that is, binarization of the segmentation map. The channel with the largest value at each position is 1, and the other channels are 0. Thus, binarized segmentation maps of the four regions are obtained.
[0157] In summary, the present invention uses the method of forward and backward interactive learning to retain the unique characterizations of each single modality and the multi-modal fusion characterization, so as to achieve the robustness of the model in various real modality missing scenarios. The complete single-modal unique characterization learning step ensures that the unique information of each modality effective for the current segmentation is fully mined and retained. The forward interactive characterization learning step fully fuses the multi-modal information, enabling the multi-modal information to fully interact. Finally, the backward interactive fusion characterization fuses all the characterizations together according to the missing situation. The backward interactive characterization completely retains the single-modal information and at the same time contains the inter-modal correlation information. In addition, the present invention maximally learns each single-modal data by simultaneously optimizing the standard segmentation loss and the multi-class segmentation loss acting on each modality, so as to achieve lossless capture of redundant information on each single-modal feature; moreover, the present invention also assigns different weights to each modality branch for each different missing scenario through the attention mechanism, enabling the network to dynamically fuse the information of each branch, thereby improving the performance of the network and the robustness to the missing modality.
[0158] Those skilled in the art know that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a kind of hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structures within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as either software modules for implementing the method or structures within the hardware component.
[0159] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A multi-modal MRI brain tumor image segmentation method robust to missing modalities, characterized in that, Including: Step S1: Collect multi-modal MRI image data and preprocess the multi-modal MRI image data to obtain preprocessed multi-modal MRI image data; Step S2: Input the single-modal MRI image data in the preprocessed multi-modal MRI image data into multiple different encoder-decoder 3D-UNet networks respectively to obtain discriminative representations of their respective modalities; Step S3: Concatenate the single-modal MRI image data in the preprocessed multi-modal MRI image data along the channel dimension and input it into the pre-interaction encoder-decoder 3D-UNet network to obtain a pre-interaction fusion representation; Step S4: Perform adaptive dynamic fusion on the pre-interaction fusion representation and the discriminative representations of their respective modalities based on the attention mechanism to obtain a post-interaction fusion representation; Step S5: Input the discriminative representations of their respective modalities, the pre-interaction fusion representation, and the post-interaction fusion representation into different single-layer deep convolutional neural networks respectively to obtain their respective tumor segmentation maps.
2. The method for multimodal MRI brain tumor image segmentation robust to missing modalities according to claim 1, wherein, The step S1 adopts: performing data augmentation processing on the multi-modal MRI image data including normalization processing and random cropping, mirror flipping, and rotation; The multi-modal MRI images include: T1-weighted imaging, T2-weighted imaging, T1ce-weighted imaging, and Flair imaging.
3. The method for multimodal MRI brain tumor image segmentation robust to missing modalities according to claim 1, wherein, The encoder-decoder 3D-UNet network includes: an encoder with four stages, a bottleneck layer, and a decoder with four stages; The encoder constructs a pyramid of representations by gradually reducing the spatial resolution and increasing the number of representation channels stage by stage to obtain the encoded representation Input into the bottleneck layer, and use the bottleneck layer to extract high-dimensional semantic representations. The output of the bottleneck layer is input into the decoder, and the decoder maps the abstract representations extracted by the bottleneck layer back to the resolution of the original input image, and obtains the representation containing position and structure information through gradual upsampling and feature fusion 4. The method for multimodal MRI brain tumor image segmentation robust to missing modalities according to claim 1, characterized in that, The step S3 adopts: taking x Flair , x T1 , x T1ce , x T2 connecting them in the channel dimension to obtain x Pre ; Among them, CAT represents the operation of concatenation in the channel dimension; δ m indicates whether modality m exists; x m represents modality m; Ω = {Flair, T1, T1ce, T2}; x Flair , x T1 , x T1ce , x T2 respectively represent the preprocessed Flair imaging data, the preprocessed T1-weighted imaging data, the preprocessed T1ce-weighted imaging data, and the preprocessed T2-weighted imaging data; Input x Pre into the pre-interactive encoder-decoder 3D-UNet network with an input channel of 4 to obtain the encoded representations of four stages and the decoded representations of four stages 5. The method for multimodal MRI brain tumor image segmentation that is robust to missing modalities according to claim 1, wherein Supervise the representations output by different stages of the encoder-decoder 3D-UNet network; The encoder representation of the s-th stage and the decoder representation of the (s + 1)-th stage are upsampled by a factor of two and fused in the channel dimension, and then input into a convolutional neural network with a 1x1x1 convolutional kernel and a softmax function to output the prediction map for brain tumor segmentation at the s-th stage Among them, ψ m,s A convolutional neural network with a 1x1x1 convolutional kernel in one layer and a softmax function; are its learnable parameters, and U represents an operation of upsampling the representation by a factor of two; Is supervised by the reference brain tumor segmentation map y; Using a standard segmentation loss function for constraints, including Dice loss and weighted cross-entropy loss Wherein, ∩ represents the intersection operation of the predicted segmentation map and the reference segmentation map, ||·||1 represents the norm 1, K = {NCR / NET, ED, ET, Background} is a set of different regions, which contains a total of four different regions, namely necrotic and non-enhancing tumor nuclei, peritumoral edema, and enhancing tumor nuclei and background; K num is the number of the set K; is a value used to adjust the weight of each region; For supervision Expressed as a formula: Multi-class segmentation loss function is also defined as the sum of each stage as follows: Among them, U s is an upsampling operation used to match the size of the prediction in the s-th stage and the actual reference segmentation map.
6. The method for multimodal MRI brain tumor image segmentation that is robust to missing modalities according to claim 1, wherein The step S4 adopts the following method: based on the attention mechanism, the outputs of the last layers of different encoder-decoder 3D-UNet networks are fused to obtain and The representation dimension size is BxCxHxWxD; using the average pooling operation, the average value of HxWxD for each channel is taken to collapse the representation dimension to BxC; the collapsed dimensions are concatenated in the channel dimension to obtain a feature size of Bx5C, which is input into a multi-layer perceptron. The output channel number of the multi-layer perceptron is 5, and then through the Sigmoid function, the weights {w m} m∈Ω and w Pre for the representation are obtained. Based on the weights, a weighted sum of and is performed to obtain the post-interaction fusion representation f Posy : Among them, is the collapsed representation, and w i is the weight of each corresponding representation.
7. The method for multimodal MRI brain tumor image segmentation that is robust to missing modalities according to claim 1, wherein The step S5 adopts: a convolutional layer with a convolution kernel size of 1x1x1, the number of input channels is a preset value, the number of output channels is 4, and the output channels respectively correspond to four regions of K = {NCR / NET, ED, ET, Background}; the output of the convolutional layer also passes through the softmax function, and the softmax acts on the output channel dimension; and f Post After passing through this convolutional layer and softmax, respective segmentation maps are obtained and The segmentation maps are binarized to obtain binarized segmentation maps of four regions.
8. The method for multimodal MRI brain tumor image segmentation robust to missing modalities according to claim 1, characterized in that, Standard segmentation loss Supervise all output segmentation maps to align them with the reference segmentation map; the total loss function is shown as follows: Among them, and represent the standard segmentation loss and the multi-connection segmentation loss applied to each modality m.
9. A multimodal MRI brain tumor image segmentation system that is robust to missing modalities, characterized in that, Including: Module M1: Collect multi-modal MRI image data and preprocess the multi-modal MRI image data to obtain preprocessed multi-modal MRI image data; Module M2: Input the single-modal MRI image data in the preprocessed multi-modal MRI image data into multiple different encoder-decoder 3D-UNet networks respectively to obtain discriminative representations of their respective modalities; Module M3: Concatenate the single-modal MRI image data in the preprocessed multi-modal MRI image data along the channel dimension and input it into the pre-interaction encoder-decoder 3D-UNet network to obtain a pre-interaction fusion representation; Module M4: Perform adaptive dynamic fusion on the pre-interaction fusion representation and the discriminative representations of their respective modalities based on the attention mechanism to obtain a post-interaction fusion representation; Module M5: Input the discriminative representations of their respective modalities, the pre-interaction fusion representation, and the post-interaction fusion representation into different single-layer deep convolutional neural networks respectively to obtain their respective tumor segmentation maps.
10. The multi-modal MRI brain tumor image segmentation system that is robust to missing modalities according to claim 9, wherein The module M1 adopts: performing data augmentation processing on the multi-modal MRI image data including normalization processing and random cropping, mirror flipping, and rotation; The multi-modal MRI images include: T1-weighted imaging, T2-weighted imaging, T1ce-weighted imaging, and Flair imaging; The encoder-decoder 3D-UNet network includes: an encoder with four stages, a bottleneck layer, and a decoder with four stages; The encoder constructs a pyramid of representations by gradually reducing the spatial resolution and increasing the number of representation channels stage by stage, obtaining an encoded representation Input into the bottleneck layer, and use the bottleneck layer to extract high-dimensional semantic representations. The output of the bottleneck layer is input into the decoder, and the decoder maps the abstract representations extracted by the bottleneck layer back to the resolution of the original input image, obtaining a representation containing position and structural information through gradual upsampling and feature fusion The module M3 adopts: taking x Flair , x T1 , x T1ce , x T2 and connecting them in the channel dimension to obtain x Pre ; x Pre = CAT([δ m ·x m m∈Ω ) Among them, CAT represents the operation of concatenation in the channel dimension; δ m indicates whether modality m exists; x m represents modality m; Ω = {Flair, T1, T1ce, T2}; x Flair , x T1 , x T1ce , x T2 respectively represent the preprocessed Flair imaging data, the preprocessed T1-weighted imaging data, the preprocessed T1ce-weighted imaging data, and the preprocessed T2-weighted imaging data; Input x Pre into the pre-interactive encoder-decoder 3D-UNet network with an input channel of 4 to obtain encoded representations in four stages and decoded representations in four stages Supervise the representations output by different stages of the encoder-decoder 3D-UNet network; The encoder representation of the s-th stage and the decoder representation of the (s + 1)-th stage are upsampled by a factor of two and fused in the channel dimension, and then input into a convolutional neural network with a 1x1x1 convolutional kernel and a softmax function to output the prediction map for brain tumor segmentation at the s-th stage where ψ m,s a convolutional neural network with a 1x1x1 convolutional kernel and a softmax function; are its learnable parameters, and U represents an operation of upsampling the representation by a factor of two; is supervised by the reference brain tumor segmentation map y; Use a standard segmentation loss function for constraints, including Dice loss and weighted cross-entropy loss Where, ∩ represents the intersection operation of the predicted segmentation map and the reference segmentation map, ||·||1 represents the norm 1, K = {NCR / NET, ED, ET, Background} is a set of different regions, which contains a total of four different regions, namely necrotic and non-enhancing tumor cores, peritumoral edema, enhancing tumor cores, and background; K num is the number of the set K; is a value used to adjust the weight of each region; For supervision Expressed as a formula: Multi-class segmentation loss function is also defined as the sum of each stage as follows: Among them, U s is an upsampling operation used to match the sizes of the predictions in the s-th stage and the actual reference segmentation map; The module M4 adopts the following method: based on the attention mechanism, the outputs of the last layers of different encoder-decoder 3D-UNet networks are fused to obtain and a representation with a dimension size of BxCxHxWxD; using average pooling operation, the average value of HxWxD for each channel is taken to collapse the representation dimension to BxC; the collapsed dimensions are concatenated in the channel dimension to obtain a feature size of Bx5C, which is input into a multi-layer perceptron. The output channel number of the multi-layer perceptron is 5, and then through the Sigmoid function, the weights {w m} m∈Ω and w pre for the representation are obtained. Based on the weights, a weighted sum is performed on and to obtain the post-interaction fusion representation f Post : Among them, is the collapsed representation, and w i is the weight of each corresponding representation; The module M5 employs: a convolutional layer with a kernel size of 1x1x1, an input channel number being a preset value, an output channel number being 4, and the output channels corresponding to four regions of K = {NCR / NET, ED, ET, Background} respectively; the output of the convolutional layer also passes through the softmax function, and the softmax acts on the output channel dimension; and f Post Through this convolutional layer and softmax, respective segmentation maps are obtained and The segmentation maps are binarized to obtain binarized segmentation maps of the four regions; Standard segmentation loss Supervise all output segmentation maps to align them with the reference segmentation map; the total loss function is as follows: Among them, and represent the standard segmentation loss and the multi-connected segmentation loss applied to each modality m.
Citation Information
Patent Citations
Multi-modal feature fusion MRI brain tumor image segmentation method based on attention mechanism
CN114782350A
Brain tumor segmentation algorithm based on UNet++ optimization and weight budget
CN112598656A
Vehicle re-identification method based on double sub-networks
CN114067143A