Multimodal Brain Tumor Image Segmentation System Based on Sparse Attention Mechanism
Through the fusion of sparse attention mechanism and multimodal features, the problem of insensitive information and lack of modality in the brain tumor in the shared feature representation is solved, and the accuracy and accuracy of brain tumor segmentation are improved.
Patent Information
- Application Number
- CN202411432181.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-10-14
AI Technical Summary
The existing brain tumor image segmentation method based on shared feature representation is insensitive to brain tumor information, and when feature fusion is missing, information redundancy and important information are problems.
The sparse attention mechanism is adopted to extract multimodal features through MobileNetV3 reverse residual network, and the modal sparse mask fusion Transformer model and multi-scale hollow attention mechanism are combined to perform feature fusion and segmentation.
It improves the accuracy and generalization ability of brain tumor segmentation, accurately captures brain tumor location and boundaries, reduces information redundancy, and enhances attention to important features.
Smart Images

Figure CN119323583B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular, to a multi-modal brain tumor image segmentation system based on a sparse attention mechanism. Background Art
[0002] Brain tumors are one of the top ten malignant tumors. Due to the complex internal structure of the brain, accurate segmentation of brain tumors is of great significance. Magnetic resonance imaging (MRI) modalities include T1-weighted (T1), contrast-enhanced T1-weighted (T1c), T2-weighted (T2), and fluid-attenuated inversion recovery (FLAIR) modalities. The multi-modal information of brain tumors complements each other. The FLAIR modality highlights the key information of the entire tumor, and the T1c modality highlights the core region of the brain tumor. In recent years, great progress has been made in brain tumor segmentation using complete multi-modal magnetic resonance images (MRI). However, in the actual clinical environment, low image quality, patient movement during scanning, acquisition time limitations, etc. can all cause the absence of brain tumor magnetic resonance image modalities.
[0003] For the segmentation of missing brain tumors, researchers have proposed three types of segmentation methods. The first is the segmentation method based on modal image synthesis, which synthesizes the images of the missing modality from the existing modalities, such as generative adversarial networks (GANs), but this method requires cumbersome calculations and cannot synthesize the key features of the missing modality; the second method is the segmentation method based on knowledge distillation. First, a teacher model is trained in the complete brain tumor MRI modality. When the teacher model is fully converged, the student model of the missing modality is trained together with the teacher model. Although this method can improve the feature learning performance in the missing modality student network, it is still difficult to learn the important features of the missing modality from all modalities; the third method is the segmentation method based on shared feature representation. This method fuses the features of the available modalities into a potential common space and projects the features extracted from the available modalities into the segmentation space. This segmentation method does not require cumbersome calculations and a large amount of computing resources.
[0004] However, the following problems exist in this segmentation method: (1) Tumor lesions usually exhibit irregular shapes, and some areas have a wide range of lesions. Therefore, more features need to be extracted for detailed examination. The ordinary feature U-NET model cannot fully capture the context information of brain tumors, is insensitive to brain tumor information, and the perception ability of the model is low; (2) The lack of a certain modality of brain tumors will cause the brain tumors to rely too much on other modalities during feature fusion. Although learnable fusion tokens are introduced for feature fusion of brain tumors, it will also cause the model to prefer the information of a certain modality during fusion and ignore other important information; and during the fusion process, if the query vector and the key vector calculate the weights of all attention, a low attention weight proves a low correlation between the query-key pair. Then, when using the attention weight to perform weighted summation on the value vector, less brain tumor information will be captured, resulting in redundant information exchange, interfering with the process of brain tumor information exchange, and ignoring important brain tumor information. Summary of the Invention
[0005] To solve the problems that the existing segmentation method based on shared feature representation is insensitive to brain tumor information and there is information redundancy and important brain tumor information is ignored during brain tumor feature fusion in the case of missing modalities, the present invention provides a multi-modal brain tumor image segmentation system based on a sparse attention mechanism. The system includes:
[0006] Data acquisition unit: used to acquire multi-modal brain tumor data and preprocess the multi-modal brain tumor data to obtain multi-modal data;
[0007] Feature extraction unit: used to extract the first specific feature F of each modality in the multi-modal data based on the first encoder m ;
[0008] Feature fusion unit: used to obtain the first feature f based on the first specific feature F m and the first fusion token F f ; perform interaction on the first feature f 1 based on the modality sparse mask fusion Transformer model to obtain the second feature Z, extract the weight IM of each modality based on the second feature Z 1 , and obtain the third feature based on the weight IM m ; reshape the third feature m based on the spatial dimension to obtain the fourth feature ; based on the spatial dimension to obtain the fourth feature
[0009] Prediction unit: used to input the fourth feature into the brain tumor segmentation model to obtain the first loss L1, and obtain the predicted segmentation result based on the first loss L1.
[0010] Principle of the present invention: Select the multi-modal features of brain tumors in multi-modal MRI images, focusing on the local detail features of available modalities and the brain tumor modality features of all available modalities; introduce a mask to produce a simulation dataset from common multi-modal brain tumor data, simulating 15 combinations of missing modalities of missing brain tumors, which can improve the brain tumor segmentation effect.
[0011] Use the first encoder to extract features for each modality. The first encoder adds a network structure combined with MobileNetV3 inverted residuals on the basis of the convolutional network, and adopts a hierarchical feature extraction method to gradually extract and enhance features. And the MobileNetV3 inverted residual module includes the SE convolutional attention mechanism (Squeeze-and-Excitation). The SE convolutional attention mechanism adaptively adjusts the weights of channels, enabling it to learn the global information within the modality and emphasize the useful brain tumor information features, dynamically emphasizing the useful brain tumor feature information, focusing on the local detail features of brain tumors, improving the sensitivity to brain tumor information, more accurately capturing the location of brain tumors, suppressing the information of healthy brain tissues, and enhancing the ability to extract the brain tumor features of each modality, thereby improving the brain tumor segmentation accuracy.
[0012] Introduce the intra-modal self-attention mechanism for all available modalities to perform intra-modal feature fusion of brain tumors, introduce learnable fusion tokens for all available modalities to perform cross-modal feature fusion between available modality brain tumors, and introduce a sparse mask matrix to reduce feature information redundancy and increase the attention to important brain tumor features, achieving the improvement of the model segmentation accuracy and generalization ability.
[0013] For the feature fusion of missing modalities, introduce a sparse attention mechanism for the feature fusion of brain tumors in missing modalities. Transformer can efficiently capture global features. The sparse self-attention mechanism sets the weights of invalid positions to 0, and sets the non-attended positions of each attention head to 0 according to the number of categories, reducing the interference of healthy brain tissue information on the fusion of brain tumors in missing modalities, suppressing the information of healthy brain tissues, highlighting valuable brain tumor information, increasing the attention to important brain tumor features, enabling the feature fusion to aggregate brain tumor-related marker information, filtering out irrelevant healthy brain tissue information, and selectively extracting useful feature information for modality fusion, thereby improving the model segmentation accuracy.
[0014] In the decoder, a multi-scale dilated attention mechanism is introduced into the missing modality brain tumor segmentation task to reduce the loss of brain tumor information during the skip connection process, obtain more spatial information of the brain tumor, effectively integrate the feature information from the encoder, increase the receptive field and context information, increase the attention to the brain tumor features, help the model extract the brain tumor information more precisely, better capture the boundary information of the brain tumor, and more accurately determine the boundary of the brain tumor, thereby improving the segmentation performance of the brain tumor.
[0015] The multi-modal images of brain tumors refer to the process of using multiple MRI sequences to obtain information about the characteristics of tumor tissues. It includes T1-weighted (T1), contrast-enhanced T1-weighted (T1c), T2-weighted (T2), and fluid-attenuated inversion recovery (FLAIR) modalities.
[0016] The Sparse Attention mechanism is a technique used in the Transformer model to reduce computational complexity. The core idea is to introduce sparsity in the self-attention calculation, that is, instead of having each position in the sequence perform attention calculations with all other positions, only select some positions for calculation.
[0017] Learnable Fusion Tokens: Used for feature fusion between different modalities. They do not directly represent the data of any specific modality, but as part of the model, they learn how to integrate information from different modalities through optimization during the training process. These fusion tokens can dynamically adjust their parameters to adapt to different data representations and interactions between modalities, thereby obtaining cross-modal brain tumor information.
[0018] The latent space is a low-dimensional space used to represent data in machine learning. It refers to the compressed representation of all useful information contained in the data.
[0019] The MobileNetV3 inverted residual module: includes expansion convolution, depthwise convolution, projection convolution, and SE (Squeeze and Excitation blocks) module.
[0020] Modal Sparse Mask Fusion Transformer: A Transformer architecture for more efficient cross-modal feature fusion and solving the problem of missing modalities. It uses a sparse mask mechanism to handle the missing brain tumor modalities and reduce feature redundancy, and combines the sparse mask mechanism with the self-attention mechanism of traditional Transformers to achieve more effective feature fusion.
[0021] Multi-Head Modal Sparse Mask Attention Mechanism: It is a modal sparse mask attention mechanism formed by combining the sparse mask and the self-attention mechanism. The multi-head modal sparse mask attention mechanism allows the model to learn information in parallel in multiple subspaces, which can enhance the learning ability and expressiveness of the model.
[0022] Modal Sparse Mask Attention Mechanism: It is a modal sparse mask attention mechanism formed by combining the sparse mask and the self-attention mechanism.
[0023] Multi-Scale Dilated Spatial Attention (MSDA) is a multi-scale dilated spatial attention mechanism that can effectively fuse feature information of different scales and use dilated convolution to capture dependencies at a farther distance, thereby enhancing the feature extraction ability of the model.
[0024] Channel Fusion Transformer Layer: A layer that uses a mask mechanism and Transformer combination to reduce channel dimension redundancy.
[0025] Sparse Coding: A technique used to indicate which neurons or parameters are activated or retained during sparse training.
[0026] Sparse Mask Matrix: A sparse mask formed by combining the mask mechanism with the sparse attention mechanism.
[0027] Reshaping: Changing the "shape" of an array (usually multi-dimensional) without changing its data.
[0028] Further, the first specific feature F m includes the first modality T1, the second modality T2, the third modality T1c, and the fourth modality Flair. The first modality T1 is used to extract information about non-enhanced tumor regions, the second modality T2 is used to extract tumor region information about edema regions and necrosis regions, the third modality T1c is used to extract tumor region information about enhanced tumor regions and non-enhanced tumor regions, and the fourth modality Flair is used to extract information about complete tumor regions. An encoder is provided for each modality, and specific features are extracted in each modality. Each modality has a different focus, which improves the feature extraction ability of the encoder.
[0029] Further, the feature extraction unit is further configured to: input the first specific feature F m into a second encoder for segmentation and projection into a latent space, to obtain a second loss L2 of the projected latent space. The decoder is used to decode the modality-specific features and calculate a regularization loss to learn more generalized features and improve the generalization of the model.
[0030] Further, the first encoder includes five encoder layers. The first encoder layer includes three convolutional layers and a residual connection. Each of the second to fifth encoder layers includes three MobileNetV3 inverted residual modules and a residual connection. A hierarchical feature extraction method is adopted to gradually extract and enhance features. The MobileNetV3 inverted residual module includes an SE convolutional attention mechanism. The SE convolutional attention mechanism adaptively adjusts the weights of channels, enabling it to learn the global information within the modality and emphasize the useful brain tumor information features, dynamically emphasizing the useful brain tumor feature information, focusing on the local detail features of the brain tumor, improving the sensitivity to brain tumor information, more accurately capturing the location of the brain tumor, suppressing the information of healthy brain tissue, and enhancing the ability to extract brain tumor features of each modality, thereby improving the accuracy of brain tumor segmentation.
[0031] Further, the second feature Z includes a second fusion token F containing cross-modal information f1 and a second specific feature F containing global information within the modality m1 .
[0032] Further, the modality sparse mask fusion Transformer model includes several layers of modality sparse mask fusion Transformers. Each layer of the modality sparse mask fusion Transformer includes a multi-head modality sparse mask attention mechanism, a feed-forward neural network layer, and a normalization layer.
[0033] Further, the feature fusion unit specifically includes:
[0034] A first fusion unit: configured to reshape the first specific feature F m , and cascade the reshaped first specific feature F m with the position encodings of the first fusion token F f and the first fusion token F f to obtain the first feature f 1 ;
[0035] A second fusion unit: configured to project the first feature f 1 into a preset space based on the modality sparse mask attention mechanism to obtain a fifth feature Z m; Based on the modal sparse mask fusion Transformer model, the fifth feature Z m and the first feature f 1 are fused to obtain the second feature Z, and the preset space includes a query space, a key space, and a value space;
[0036] The third fusion unit: is used to extract the attention information W f1 of each modality in the attention tensor of the second fusion token F j , slice the attention information W j to obtain the spatial weight W m of each modality; reshape the spatial weight W m to obtain the weight IM m ; based on the weight IM m , weight the second specific feature F m1 to obtain the third feature
[0037] The fourth fusion unit: is used to reshape the third feature along the spatial dimension to obtain the sixth feature cascade the sixth feature to obtain the seventh feature reshape the second fusion token F f1 to obtain the eighth feature Based on the modal mask Transformer layer, fuse the seventh feature and the eighth feature along the channel dimension to obtain the fourth feature
[0038] Furthermore, the brain tumor segmentation model includes a decoder, and the decoder includes a multi-scale dilated attention mechanism and a channel fusion Transformer layer. Introducing the multi-scale dilated attention mechanism into the missing modality brain tumor segmentation task reduces the problem of brain tumor information loss during the skip connection process of the brain tumor, obtains more spatial information of the brain tumor, effectively integrates the feature information from the encoder, increases the receptive field and context information, increases the attention to the brain tumor features, helps the model to extract brain tumor information more precisely, better captures the boundary information of the brain tumor, and more accurately determines the boundary of the brain tumor, thereby improving the segmentation performance of the brain tumor.
[0039] Furthermore, the prediction unit specifically includes:
[0040] The first prediction unit: is used to input the fourth feature into the decoder to obtain the third specific feature F m2, based on the channel fusion Transformer layer, fuse the third specific feature F m2 and the fourth feature to obtain the ninth feature D1;
[0041] Second prediction unit: used to cascade the ninth feature D1 to obtain the final feature D, obtain the first loss L1 based on the final feature D, and obtain the predicted segmentation result based on the first loss L1.
[0042] Furthermore, the calculation method for obtaining the second loss L2 is:
[0043] L rgl = ∑ m∈M L Dice (D rgl (E m (x m )), Y g ) + L WCE (D rgl (E m (x m )), Y g );
[0044] Among them, L rgl represents the second loss L2, L Dice represents the Dice loss function, D rgl represents the second encoder, E m represents the first encoder, L WCE represents the weighted cross-entropy loss function, Y g represents the true annotation of the brain tumor, x m represents any 3D input modality path, m represents the modality, and M represents the first specific feature F m , including the first modality T1, the second modality T2, the third modality T1c, and the fourth modality Flair.
[0045] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0046] 1. Feature extraction is performed on each modality using a first encoder. The first encoder adds a network structure combining MobileNetV3 inverse residuals on the basis of a convolutional network and adopts a hierarchical feature extraction method to gradually extract and enhance features. The MobileNetV3 inverse residual module includes a SE convolutional attention mechanism (Squeeze-and-Excitation). The SE convolutional attention mechanism adaptively adjusts the weights of channels, enabling it to learn the global information within the modality and emphasize the useful brain tumor information features, dynamically emphasizing the useful brain tumor feature information, focusing on the local detail features of the brain tumor, improving the sensitivity to brain tumor information, more accurately capturing the location of the brain tumor, suppressing the information of healthy brain tissue, enhancing the ability to extract brain tumor features of each modality, and thus improving the brain tumor segmentation accuracy.
[0047] 2. The intra-modal self-attention mechanism is introduced for all available modalities to perform intra-modal feature fusion of brain tumors. Learnable fusion tokens are introduced for all available modalities to perform cross-modal feature fusion between available modality brain tumors, and a sparse mask matrix is introduced to reduce feature information redundancy, increase the attention to important brain tumor features, and achieve improving the accuracy and generalization ability of the model segmentation.
[0048] 3. For the feature fusion of missing modalities, a sparse attention mechanism is introduced for feature fusion of brain tumors in missing modalities. Transformer can efficiently capture global features. The sparse self-attention mechanism sets the weights of invalid positions to 0 and sets the non-attended positions of each attention head to 0 according to the number of categories, reducing the interference of healthy brain tissue information on the fusion of brain tumors in missing modalities, suppressing the information of healthy brain tissue, highlighting valuable brain tumor information, increasing the attention to important brain tumor features, enabling the feature fusion to aggregate brain tumor-related marker information, filtering out irrelevant healthy brain tissue information, and selectively extracting useful feature information for modality fusion, thereby improving the accuracy of the model segmentation.
[0049] 4. In the decoder, the multi-scale dilated attention mechanism is introduced into the brain tumor segmentation task of missing modalities to reduce the problem of brain tumor information loss during the skip connection process, obtain more spatial information of the brain tumor, effectively integrate the feature information from the encoder, increase the receptive field and context information, increase the attention to brain tumor features, help the model extract brain tumor information more precisely, better capture the boundary information of the brain tumor, and more accurately determine the boundary of the brain tumor, thereby improving the segmentation performance of the brain tumor. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of the present invention, and do not limit the embodiments of the present invention;
[0051] Figure 1 It is a schematic structural diagram of a multi-modal brain tumor image segmentation system based on a sparse attention mechanism in the present invention;
[0052] Figure 2 It is a schematic flowchart of a multi-modal brain tumor image segmentation system based on a sparse attention mechanism in the present invention. Detailed implementation manners
[0053] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0054] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0055] Embodiment 1
[0056] Referring to Figure 1 - Figure 2 , this embodiment provides a multi-modal brain tumor image segmentation system based on a sparse attention mechanism. The system includes:
[0057] A data acquisition unit: used to acquire multi-modal brain tumor data based on a magnetic resonance imaging (MRI) scanner, and preprocess the multi-modal brain tumor data based on a processor to obtain multi-modal data; in this embodiment, the preprocessing operations can include cropping, normalization processing, random rotation, and random mirror flipping, etc. The processor can be a controller, a chip, a computer, and dedicated image processing software including 3D Slicer and FSL, etc. There are computer programs, computational model programs, or related computational functions in these devices; the acquisition device is data-connected to the processor. The processor can be built into the acquisition device or not installed on the acquisition device. The acquisition device transmits the acquired data to the processor for processing. The processor is built with an encoder, a pre-trained modal sparse mask fusion Transformer model, and a pre-trained brain tumor segmentation model. In this embodiment, a memory, a voice acquisition device, and a voice processing chip can also be included. The memory is connected to both the magnetic resonance imaging scanner and the processor, and is used to store multi-modal brain tumor data and multi-modal data; the voice acquisition device and the voice processing chip are connected, and the voice processing chip and the processor are connected. The voice acquisition device is used to acquire voice data, and the voice processing chip is used to process the voice data and convert it into text or computer instructions, etc.
[0058] Feature extraction unit: used to extract the first specific feature F of each modality in the multi-modal data based on the first encoder m ; for example, input the modality information of each modality into the corresponding first encoder to extract the specific feature of each modality;
[0059] Among them, the first specific feature F m includes the first modality T1, the second modality T2, the third modality T1c, and the fourth modality Flair. The first modality T1 is used to extract non-enhanced tumor region information, the second modality T2 is used to extract tumor region information of the edema region and the necrosis region, the third modality T1c is used to extract tumor region information of the enhanced tumor region and the non-enhanced tumor region, and the fourth modality Flair is used to extract complete tumor region information.
[0060] Among them, the first encoder includes five encoder layers. The first encoder layer includes three convolutional layers and a residual connection. Each of the second to fifth encoder layers includes three MobileNetV3 inverted residual modules and a residual connection. The first encoder is composed of ordinary convolution and MobileNetV3 inverted residual. The encoder is divided into five layers, and each layer has three sub-layers. The first layer is composed of three ordinary convolutional layers, which are used to process the input image and extract basic features. The second to fifth layers are composed of MobileNetV3 inverted residual. Each layer is composed of three MobileNetV3 inverted residual modules, which extract features from the results of the previous layer in turn, and retain the information of the input features through the residual connection. This encoder stacks convolutional layers and MobileNetV3 inverted residual modules to perform layer-by-layer feature extraction on the input image, so as to obtain the local detailed features of each modality, that is, the first specific feature F m ;
[0061] Feature fusion unit: used to obtain the first feature f m based on the first specific feature F f and the first fusion token F 1 ; based on the modality sparse mask fusion Transformer model, interact with the first feature f 1 to obtain the second feature Z. Based on the second feature Z, extract the weight IM of each modality m ; based on the weight IM m to obtain the third feature Based on the spatial dimension, reshape the third feature to obtain the fourth feature In this embodiment, the fusion token can be a learnable fusion token.
[0062] Among them, the second feature Z includes the second fusion token F containing cross-modal informationf1 and a second specific feature F that includes global information within the modality m1 .
[0063] Among them, the modality sparse mask fusion Transformer model includes several layers of modality sparse mask fusion Transformers, and each layer of the modality sparse mask fusion Transformer includes a multi-head modality sparse mask attention mechanism, a feed-forward neural network layer, and a normalization layer.
[0064] Among them, the feature fusion unit specifically includes:
[0065] The first fusion unit: used to reshape the first specific feature F m After reshaping the first specific feature F m Concatenate with the first fusion token F f And the first fusion token F f To obtain the first feature f 1 , and its calculation method can be:
[0066] f 1 = Concat(f flair , f t1 , f t1c , f t2 , F f ) + PE;
[0067] Among them, f 1 Represents the first feature f 1 , Concat represents concatenation, f t1 , f t2 , f t1c And f flair Respectively represent the reshaped first modality T1, second modality T2, third modality T1c, and fourth modality Flair, F f Represents the first fusion token F f , PE represents the position of the first fusion token F f ;
[0068] The second fusion unit: used to project the first feature f 5N×5N To a preset space to obtain the fifth feature Z 1 Based on the modality sparse mask attention mechanism and the binary attention mask B ∈ {0, 1} m , and its calculation method can be:
[0069]
[0070] Z m=(Softmax(θ(Q, K, i, j)) * s_mask())V;
[0071] Where E represents converting the scaled dot product of the query vector and the key vector to a positive number, H represents the number of heads of the multi-head modal sparse masked attention mechanism, c represents the number of feature channels, represents the feature dimension of each head, which is used to scale the dot product to make the training more stable; Q, K, and V respectively represent the query vector, key vector, and value vector of the first specific feature F m where i represents the index of the query vector, j represents the index of the key vector, j' represents the index of the accessible modality, T represents the transpose, represents the attention weights after applying the mask, Z m represents the fifth feature Z m , B i,j represents whether the attention score can be calculated between i and j. If not, the modality is missing and B i,j is 0. If so, the modality exists and B i,j is 1; E' represents the attention score of B i,j where B i,j′=1 represents the sum of the attention scores of the key vectors when the index = 1, s_mask represents the sparse mask matrix, which realizes the sparsity of the attention, and * represents element-wise multiplication;
[0072] Based on the modal sparse masked fusion Transformer model, the fifth feature Z m and the first feature f 1 are fused to obtain the second feature Z, and its calculation method can be:
[0073] C2 ← MSDA(Z m ) + f 1 ;
[0074] Z ← (FFN(LN(C2)) + C2);
[0075] where f 1 represents the first feature f 1 Z m represents the fifth feature Z m , MSDA represents the multi-head masked attention mechanism, FFN represents the feed-forward neural network, LN represents the normalization, Z represents the second feature Z, and C2 represents the relevant parameters in the fusion process.
[0076] The preset space includes a query space, a key space, and a value space;
[0077] The third fusion unit: used to extract the attention information W of each modality in the attention tensor of the second fusion token F f1 j Slice the attention information W j to obtain the spatial weight W of each modality m ; The calculation method can be:
[0078]
[0079] W m = Split(W) ∈ R 1×N ;
[0080] where W j represents the attention information W j , H represents the number of heads of the multi-head modality sparse masked attention mechanism, N = h × w × d, h represents the height of the input feature map, w represents the width of the feature map, d represents the depth of the feature map, represents the attention weight matrix obtained in the first layer of modality sparse masked Transformer calculation, W m represents the spatial weight W m , Split represents the slicing operation, and W represents extracting different types of attention information from the attention tensor.
[0081] Reshape the spatial weight W m to obtain the weight IM m ; Based on the weight IM m , weight the second specific feature F m1 to obtain the third feature The calculation method can be:
[0082] IM m = reshape(W m ), IM m ∈ R 1×h×w×d ;
[0083]
[0084] where W m represents the spatial weight W m , IM m represents the weight IM m , reshape represents the reshaping operation, h represents the height of the input feature map, w represents the width of the feature map, d represents the depth of the feature map, represents the third feature F m1 represents the second specific feature F m1 .
[0085] Fourth fusion unit: used to reshape the third feature along the spatial dimension to obtain the sixth feature Cascade the sixth feature to obtain the seventh feature Its calculation method can be:
[0086]
[0087] where represents the sixth feature reshape represents the reshaping operation, represents the third feature represents the seventh feature Concat represents concatenation, N = h × w × d, h represents the height of the input feature map, w represents the width of the feature map, d represents the depth of the feature map, and c represents the number of feature channels.
[0088] Reshape the second fusion token F f1 to obtain the eighth feature Its calculation method can be:
[0089]
[0090] where represents the eighth feature reshape represents the reshaping operation, F f1 represents the second fusion token F f1 , N = h × w × d, h represents the height of the input feature map, w represents the width of the feature map, d represents the depth of the feature map, and c represents the number of feature channels.
[0091] Based on the modality mask Transformer layer, fuse the seventh feature and the eighth feature along the channel dimension to obtain the fourth feature Its calculation method is:
[0092]
[0093] where λ α and ω α both represent pointwise convolutions with a kernel size of 1, represents a depthwise convolution with a kernel size of 3, Ω α∈{q,k,v} (·) represents the mapping function. represents the eighth feature represents the seventh feature Concat represents concatenation, q represents the learnable fusion token, the query vector obtained by mapping; k represents the fourth feature of each modality The key vectors obtained by projection, and the key vectors obtained by concatenating the key vectors of the four modalities in the channel dimension; v represents the fourth feature of each modality The value vectors obtained by projection, and the value vectors obtained by concatenating the value vectors of the four modalities in the channel dimension, m represents the modality, and M represents the first specific feature F m , including the first modality T1, the second modality T2, the third modality T1c, and the fourth modality Flair. Att(q,k) represents the attention weight obtained by the self-attention mechanism, T represents transpose, =h×w×d, h represents the height of the input feature map, w represents the width of the feature map, d represents the depth of the feature map, CFT represents the modality mask attention mechanism along the channel direction, φ i,j respectively correspond to the relationship between the query vector q and the key vector k. If k corresponds to a missing modality, then φ i,j is equal to 0, otherwise φ i,j is equal to 1, j′ represents the index of the accessible modality, i represents the index of the query vector, and j represents the index of the key vector. represents the eighth feature represents a feed-forward neural network that adds a group convolution of size 3 between two fully connected layers in the channel dimension. L2 represents the modality mask Transformer of the L2 layer, and L2 = 2.
[0094] Prediction unit: used to process the fourth feature input into the brain tumor segmentation model to obtain the first loss L1, obtain the predicted segmentation result based on the first loss L1, and display the predicted segmentation result on a display device. The display device is connected to the magnetic resonance imaging scanner, the processor, and the memory. In this embodiment, the display device can be a display screen.
[0095] Among them, the brain tumor segmentation model includes a decoder, and the decoder includes a multi-scale dilated attention mechanism and a channel fusion Transformer layer.
[0096] Among them, the prediction unit specifically includes:
[0097] The first prediction unit: used to process the fourth feature input into the decoder to obtain the third specific feature F m2 , based on the channel fusion Transformer layer, process the third specific feature F m2 and the fourth feature for fusion to obtain the ninth feature D1;
[0098] Second prediction unit: used to cascade the ninth feature D1 to obtain the final feature D, obtain the first loss L1 based on the final feature D, and obtain the predicted segmentation result based on the first loss L1.
[0099] For example, if the fourth feature is input into the first three layers of the decoder, information loss can be reduced, it is easier to capture the boundary information of the brain tumor, the quality of brain tumor segmentation can be improved, and then combined with the fourth feature is input into the fourth layer for weighting and upsampling to the final output layer for brain tumor segmentation. The first loss L1 calculates the error between the predicted segmentation result and the true label, and its calculation method can be:
[0100] L1 = ∑ m∈M L Dice (D f (Concat(E m (x m ))), Y g ) + L WCE (D f (E m (x m ))), Y g );
[0101] Among them, L1 l represents the first loss L1, L Dice represents the Dice loss function, D f represents the decoder, E m represents the first encoder, Concat represents concatenation, L WCE represents the weighted cross-entropy loss function, Y g represents the true annotation of the brain tumor, x m represents any 3D input modality path, m represents the modality, M represents the first specific feature F m , including the first modality T1, the second modality T2, the third modality T1c, and the fourth modality Flair.
[0102] Embodiment 2
[0103] The feature extraction unit is further used to: input the first specific feature F m into the second encoder for segmentation and projection into the latent space, to obtain the second loss L2 of the latent space projection. In this embodiment, the second encoder is a weight-sharing encoder. For example, if the first specific feature F of each modality m is input into the shared second encoder for separate segmentation and projection into the shared latent space, to obtain the second loss L2 of the latent space projection, and the calculation method of its loss can be:
[0104] L rgl = ∑m∈M L Dice (D rgl (E m (x m )),Y g )+L WCE (D rgl (E m (x m )),Y g );
[0105] Among them, L rgl represents the second loss L2, that is, the regularization loss, L Dice represents the Dice loss function, D rgl represents the second encoder, E m represents the first encoder, L WCE represents the weighted cross-entropy loss function, Y g represents the true annotation of the brain tumor, x m represents any 3D input modality path, m represents the modality, and M represents the first specific feature F m , including the first modality T1, the second modality T2, the third modality T1c, and the fourth modality Flair.
[0106] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0107] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A multi-modal brain tumor image segmentation system based on a sparse attention mechanism, characterized in that, The system includes: A data acquisition unit: used to acquire multi-modal brain tumor data, and preprocess the multi-modal brain tumor data to obtain multi-modal data; Feature extraction unit: configured to extract a first specific feature F of each modality in the multimodal data based on the first encoder m ; Feature fusion unit: used to obtain the first feature f based on the first specific feature F m and the first fusion token F f ; based on the modality sparse mask fusion Transformer model, interact with the first feature f 1 to obtain the second feature Z, and based on the second feature Z, extract the weight IM of each modality 1 ; based on the weight IM m to obtain the third feature m ; based on the spatial dimension, reshape the third feature to obtain the fourth feature Prediction unit: used to input the fourth feature into the brain tumor segmentation model to obtain a first loss L1, and obtain a predicted segmentation result based on the first loss L1; The first loss L1 represents the error value between the predicted segmentation result and the true label; The second feature Z and the fourth feature both include a second fusion token F containing cross-modal information f1 and a second specific feature F containing global information within the modality m1 ; The first specific feature F m and the second specific feature F m1 and the first feature f 1 and the third feature all include a first modality T1, a second modality T2, a third modality T1c, and a fourth modality Flair. The first modality T1 is used to extract information on non-enhanced tumor regions, the second modality T2 is used to extract tumor region information in edema regions and necrotic regions, the third modality T1c is used to extract tumor region information in enhanced tumor regions and non-enhanced tumor regions, and the fourth modality Flair is used to extract information on complete tumor regions; The modality sparse mask fusion Transformer model includes several layers of modality sparse mask fusion Transformers, and each layer of the modality sparse mask fusion Transformer includes a multi-head modality sparse mask attention mechanism, a feed-forward neural network layer, and a normalization layer; The feature fusion unit specifically includes: The first fusion unit: for shaping the first specific feature F m After shaping, the first specific feature F after shaping m Is cascaded with the first fusion token F f And the position encoding of the first fusion token F f To obtain the first feature f 1 ; Second fusion unit: for projecting the first feature f 1 into a preset space to obtain a fifth feature Z m ; based on the modality sparse mask fusion Transformer model, fusing the fifth feature Z m and the first feature f 1 to obtain the second feature Z, where the preset space includes a query space, a key space, and a value space; The calculation formula for the second feature Z is: C2 ← MSDA(Z m ) + f 1 ; Z←(FFN(LN(C2))+C2); Among them, f 1 represents the first feature f 1 , Z m represents the fifth feature Z m , MSDA represents the multi-head masked attention mechanism, FFN represents the feed-forward neural network, LN represents normalization, Z represents the second feature Z, and C2 represents the relevant parameter in the fusion process; Third fusion unit: used to extract the attention information W of each modality in the attention tensor of the second fusion token F f1 ; slice the attention information W j to obtain the spatial weight W of each modality j ; reshape the spatial weight W m to obtain the weight IM m ; based on the weight IM m , weight the second specific feature F m to obtain the third feature m1 Fourth fusion unit: for shaping the third feature along the spatial dimension to obtain a sixth feature Cascading the sixth feature to obtain a seventh feature Shaping the second fusion token F f1 to obtain an eighth feature Based on the modality mask Transformer layer, fusing the seventh feature and the eighth feature along the channel dimension to obtain the fourth feature 2. The multi-modal brain tumor image segmentation system based on the sparse attention mechanism according to claim 1, wherein The feature extraction unit is further used for: Input the first specific feature F m into a second encoder for segmentation and project it into the latent space, obtaining a second loss L2 of the latent space projection.
3. The multi-modal brain tumor image segmentation system based on the sparse attention mechanism according to claim 1, wherein The first encoder includes five encoder layers. The first encoder layer includes three convolutional layers and a residual connection. Each of the second to fifth encoder layers includes three MobileNetV3 inverted residual modules and a residual connection.
4. The multimodal brain tumor image segmentation system based on the sparse attention mechanism according to claim 1, characterized in that, The brain tumor segmentation model includes a decoder, and the decoder includes a multi-scale dilated attention mechanism and a channel fusion Transformer layer.
5. The multimodal brain tumor image segmentation system based on the sparse attention mechanism according to claim 4, wherein The prediction unit specifically includes: The first prediction unit: is configured to input the fourth feature into the decoder to obtain a third specific feature F m2 , and based on the channel fusion Transformer layer, fuse the third specific feature F m2 and the fourth feature to obtain a ninth feature D1; A second prediction unit: used to concatenate the ninth feature D1 to obtain the final feature D, obtain the first loss L1 based on the final feature D, and obtain the predicted segmentation result based on the first loss L1.
6. The multimodal brain tumor image segmentation system based on the sparse attention mechanism according to claim 2, wherein The calculation method for obtaining the second loss L2 is: L rgl = ∑ m∈M L Dice (D rgl (E m (x m ))), Y g ) + L WCE (D rgl (E m (x m ))), Y g ); Among them, L rgl represents the second loss L2, L Dice represents the Dice loss function, D rgl represents the second encoder, E m represents the first encoder, L WCE represents the weighted cross-entropy loss function, Y g represents the true annotation of brain tumors, x m represents any 3D input modality path, m represents the modality, and M represents the first specific feature F m , including the first modality T1, the second modality T2, the third modality T1c, and the fourth modality Flair.
Citation Information
Patent Citations
Multi-modal MRI brain tumor image segmentation method based on modal cross attention
CN117635563A
MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on multi-scale feature fusion of improved U-Net
CN117876399A