MRI brain tumor image segmentation method and system based on adaptive feature fusion
By adopting an adaptive feature fusion method in the MRI image segmentation technology of brain tumors, using 3D U-Net and multi-head attention mechanism modules, the high cost and low accuracy problems in the absence of modality are solved, efficient and accurate segmentation effect is achieved, and resource demand is reduced.
Patent Information
- Application Number
- CN202510087163.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
When faced with modal deficiencies, the existing brain tumor MRI image segmentation technology has high training and deployment costs, limited modal synthesis quality, and complex model structure, resulting in low segmentation efficiency and accuracy, making it difficult to effectively deploy in practical application scenarios with limited resources.
The MRI brain tumor image segmentation method is adopted with adaptive feature fusion, and the basic network based on 3D U-Net, multi-head attention mechanism module and adaptive feature fusion module are used to realize image segmentation through feature extraction, attention weighting, adaptive weight allocation and fusion.
It improves segmentation efficiency and accuracy, reduces training and deployment costs, simplifies model structure, reduces resource consumption and training difficulty, and realizes efficient deployment and use in resource-limited scenarios.
Smart Images

Figure CN120014269A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of brain tumor image segmentation, and in particular to an adaptive feature fusion MRI brain tumor image segmentation method and system. Background Art
[0002] Magnetic resonance imaging (MRI) plays a vital role in the diagnosis of brain tumors. MRI can provide multimodal images, such as T1-weighted (T1), contrast-enhanced T1-weighted (T1c), T2-weighted (T2), and T2 fluid-attenuated inversion recovery (FLAIR) images. These different modal images can provide complementary information for brain tumor segmentation and help to more accurately analyze the conditions of different areas of brain tumors. For example, T1 and T1c modalities tend to highlight the core of the tumor, while T2 and FLAIR modalities show the peritumoral edema more clearly.
[0003] At present, brain tumor segmentation is mainly achieved by deep learning methods, and its core goal is to accurately segment the brain tumor area using MRI images of different modalities. For example, some methods will first input images of different modalities into a specific network structure for feature extraction, and then generate the final segmentation result by fusing these features; other methods will first pre-process the image, and then use the codec structure to gradually restore the segmented image. However, in practical applications, due to factors such as image damage, artifacts, acquisition protocols, allergies to contrast agents, or costs, one or more modalities are often missing, which poses a great challenge to brain tumor segmentation. In response to these technical problems, the prior art has proposed many improved methods for modality missing situations to ensure the segmentation effect as much as possible, such as segmentation methods based on specific models and general single model segmentation methods.
[0004] In the field of brain tumor segmentation, some researchers are committed to designing segmentation methods specifically for different modality loss situations. For example, some methods adopt a collaborative training strategy (Blum and Mitchell, 1998) to transfer knowledge between the full modality and the missing modality network. For example, the related work of Hu et al. (2020) and Chen et al. (2021) transfers knowledge from the multimodal teacher network to the unimodal student network at the level of overall image semantics and pixel network output; there is also the adversarial collaborative training network (ACN) proposed by Wang et al. (2021b) to enhance the knowledge distillation from the full modality to the missing modality through entropy and knowledge adversarial learning to align the potential representation; the style matching U-Net (SMU-Net) proposed by Azad et al. (2022) decomposes the common latent space of the full modality and missing modality data into content and style representations, and uses the content and style matching mechanism to transfer useful features in the full modality network to the missing modality network. These methods have achieved certain results in dealing with the situation of missing modalities. However, they have obvious shortcomings, namely, high training and deployment costs, because for N modalities, 2 n -1 model, which will consume a lot of time and space resources in practical applications.
[0005] In addition, there is a special method that uses generative adversarial networks (GANs; Goodfellow et al., 2014) to synthesize images of missing modalities and then achieve full modality segmentation, such as the related research work of Lee et al. (2020) and Yu et al. (2019). However, the generative adversarial network itself has some limitations. It is difficult to train, especially for three-dimensional image generation, and it will generate additional overhead both in the training stage and the deployment stage, which affects its application effect in actual brain tumor segmentation.
[0006] In addition to the above-mentioned methods based on specific models, there are also some methods that try to use a single general model to handle all modal missing situations. The usual practice is to use specific encoders to embed each modality into a shared latent space, and then perform feature fusion and subsequent segmentation operations on this basis. For example, Havaei et al. (2016) adopted this idea; Dorent et al. (2019) proposed a heteromodal variational encoder-decoder (HVED) that combines multimodal variational autoencoders to remodel modalities from common latent variables to form a truly shared latent representation; Shen and Gao (2019) proposed adversarial training to adapt the feature map of the missing modality to the feature map of the full modality; Zhou et al. (2021b) modeled the correlation between modalities through latent correlation representation learning to estimate the representation of the missing modality in the latent space; Zhou et al. (2021a) explicitly generated feature-enhanced images to provide the necessary feature representation for the missing modality; Ding et al. (2021) proposed a region-aware fusion network (RFNet) that relies on a region-aware fusion module to adaptively fuse features of available image modalities according to different regions. Although these methods can use one model to deal with multiple modal loss combinations, they often have more complex designs, with multiple encoders and even decoders, and the interactions between the various parts are also relatively complex, which is not conducive to model understanding, training and optimization in practical applications.
[0007] However, in the context of existing brain tumor segmentation technology, the above methods have some obvious defects.
[0008] 1. High training and deployment costs
[0009] First, some methods adopt the strategy of designing a separate model for each modality combination, such as some methods based on collaborative training (Hu et al., 2020; Chen et al., 2021; Wang et al., 2021b; Azad et al., 2022). They need to train corresponding models for different modality missing situations. When there are N modalities, it is necessary to train 2 n -1 model. This method consumes a lot of computing resources during the training phase, including processor operation time and memory usage, resulting in long training time and low efficiency. At the same time, when these models are actually deployed, due to the large number of models, a large amount of storage space will be occupied, and the requirements for hardware devices are high. It is difficult to promote and use them efficiently in some practical application scenarios with limited resources. For example, in some primary medical units or medical equipment with limited computing resources, it is difficult to carry so many models to realize brain tumor segmentation function.
[0010] In addition, the method of using generative adversarial networks (GANs) to synthesize missing modality images and achieve full modality segmentation (Lee et al., 2020; Yu et al., 2019) has the problem of high training difficulty. Especially when dealing with the task of generating three-dimensional brain tumor MRI images, it is difficult to achieve ideal training results, resulting in uneven quality of synthesized missing modality images, which in turn affects the final accuracy of brain tumor segmentation. Moreover, during the training and deployment process, the additional introduction of GANs modules will increase the overall computational overhead and resource usage, which is not conducive to the efficient operation of the entire segmentation process.
[0011] Furthermore, those methods that attempt to use a single general model to handle all modality loss situations (Havaei et al., 2016; Dorent et al., 2019; Shen and Gao, 2019; Zhou et al., 2021a; Zhou et al., 2021b; Ding et al., 2021), although avoiding the problem of training multiple models, often have complex network structure designs, with multiple encoders and even decoders, and the interaction between the various parts is also relatively complex. This complexity makes the model prone to problems such as overfitting during training, making it difficult to train to the optimal state; on the other hand, it also increases the difficulty of understanding, debugging and optimizing the model. For R&D personnel, they need to spend more energy to adjust model parameters and improve the structure. In practical applications, complex structures may also lead to slower computing speeds and affect segmentation efficiency.
[0012] 2. Limited quality of modal synthesis
[0013] Among the related methods that use image synthesis to deal with the problem of missing modalities in brain tumor segmentation, some use generative adversarial networks (GANs) to synthesize images of missing modalities and then achieve full modality segmentation, such as the related research work carried out by Lee et al. (2020) and Yu et al. (2019).
[0014] Among these methods, GANs need to try to generate missing modal images based on existing modal information and preset generation rules. However, due to the complexity of brain tumor MRI images themselves, their internal tissue structure, texture features and other information vary among different individuals and at different stages of the disease. It is difficult for GANs to accurately grasp these characteristics to generate high-quality missing modal images. The generated images may deviate from the real missing modal images in terms of boundary clarity and internal grayscale features of the tumor area, which means that the quality of the generated images is difficult to guarantee and may not be able to accurately restore the real modal information.
[0015] The subsequent brain tumor segmentation work is based on these synthesized modal images and other existing modal images to perform feature extraction, fusion and other operations to generate segmentation results. If the quality of the synthesized missing modality images is poor, the information obtained at the feature level will be erroneous, which will affect the accuracy of the final brain tumor segmentation and fail to achieve the ideal segmentation effect. It is difficult to play a reliable role in actual clinical diagnosis and other application scenarios.
[0016] 3. Complex model structure
[0017] In the existing brain tumor segmentation methods, some general single models adopt a more complex design, with multiple encoders, decoders and complex interactions. For example, some models use multiple encoders to extract features from images of different modalities, and then use complex interaction mechanisms to fuse and process these features, and finally generate corresponding segmentation results by multiple decoders. On the one hand, this complex structure increases the difficulty of model training, because multiple encoders, decoders and their interactions require careful adjustment of parameters, which makes it easy to encounter problems such as gradient disappearance, gradient explosion or difficulty in converging to the optimal state during training. On the other hand, in actual application scenarios, complex model structures are not conducive to understanding and maintenance. When the segmentation effect is poor, it is difficult for R&D personnel to quickly locate the problem and make targeted improvements. Moreover, complex structures often mean more computing resource consumption, the operation speed may be affected when performing segmentation tasks, and problems such as overfitting are prone to occur, resulting in poor generalization ability when facing new brain tumor MRI image data, and the inability to output segmentation results stably and accurately. Summary of the invention
[0018] In view of the shortcomings of the prior art, the present invention overcomes the defects of the prior art in segmenting brain tumor MRI images, improves segmentation efficiency and accuracy, avoids the high cost problem faced by traditional multi-model methods in brain tumor MRI image segmentation, adopts a simpler and more efficient structural design, reduces the resource consumption and training difficulty caused by multiple encoders, decoders and complex interactions, and reduces training and deployment costs.
[0019] The present invention provides an MRI brain tumor image segmentation method with adaptive feature fusion. The method uses a network model to implement MRI brain tumor image segmentation. The network model includes: a basic network based on 3D U-Net, a multi-head attention mechanism module and an adaptive feature fusion module. The basic network based on 3D U-Net includes an encoder and a decoder. The method includes:
[0020] S1, the encoder extracts features from the MRI brain tumor images of different modalities to be segmented, and inputs the first feature maps of different tumor areas in the extracted MRI brain tumor images of different modalities into the multi-head attention mechanism module;
[0021] S2, the multi-head attention mechanism module performs attention weighted combination on the first feature map, and inputs the second feature map after the attention weighted combination into the adaptive feature fusion module;
[0022] S3, the adaptive feature fusion module adaptively weights and fuses the second feature map according to the contribution of the second feature map to brain tumor segmentation, and inputs the adaptively fused third feature map into the decoder;
[0023] S4. The decoder restores the original image resolution of the third feature map to obtain segmentation results of different tumor areas.
[0024] Preferably, before using the network model, the network model needs to be trained, and the training process specifically includes:
[0025] A1. Collecting data sets of MRI brain tumor images of different modalities for preprocessing to obtain preprocessed data sets;
[0026] A2. Based on the preprocessed data set, the network model is pre-trained by self-supervised learning to obtain the pre-trained network model, wherein the pre-training process specifically includes: randomly selecting MRI brain tumor images of some modalities from the pre-processed data set for masking to simulate the missing conditions of different modalities, and also masking some MRI brain tumor images of the remaining modalities, reconstructing the masked MRI brain tumor image through the encoder, the decoder, the multi-head attention mechanism module and the adaptive feature fusion module, calculating the difference between the reconstructed MRI brain tumor image and the original MRI brain tumor image using the mean square error as the loss function, and updating the parameters of the network model through back propagation to obtain the pre-trained network model;
[0027] A3. Based on the preprocessed data set, the pre-trained network model is further trained in a supervised manner to obtain the trained network model. The further training process specifically includes: each time a sample is selected from the preprocessed data set, different modality missing conditions are randomly generated as network inputs to obtain segmentation results under different modality missing conditions. Based on the segmentation results under different modality missing conditions, the segmentation loss of the network model is calculated by a combined loss function. Based on the segmentation loss and the consistency loss between different modality missing conditions, the parameters of the network model are further fine-tuned through back propagation. When the network model can adapt to the segmentation tasks under different modality missing conditions, the training is stopped, and the final parameters of the network model are saved to obtain the trained network model.
[0028] Preferably, the combined loss function is a weighted sum of a cross entropy loss and a Dice loss; the cross entropy loss is used to measure the probability difference between the predicted segmentation result and the actual segmentation label at each pixel, and the calculation formula is:
[0029]
[0030] Among them, L CE represents the cross entropy loss, N represents the total number of pixels, C represents the number of segmentation categories, i represents the i-th pixel, j represents the j-th segmentation category, i and j are positive integers, and y ij Indicates the true value of the i-th pixel belonging to the j-th category in the true segmentation label, p ij Indicates the probability value that the i-th pixel in the predicted segmentation result belongs to the j-th category;
[0031] The Dice loss is used to measure the overlap between the predicted segmentation result and the actual segmentation label, and the calculation formula is:
[0032]
[0033] Among them, L Dice represents the Dice loss, p i Indicates the probability value of the i-th pixel in the predicted segmentation result belonging to the target category, y i Represents the probability value that the i-th pixel in the real segmentation label belongs to the target category, ∈ is a positive number to avoid the situation where the denominator is 0; the combined loss function L total is the cross entropy loss L CE With the Dice loss L Dice The weighted sum of total =αL CE +(1-α)L Dice , where α represents the trade-off between the cross entropy loss LCE With the Dice loss L Dice Importance weight parameter.
[0034] Preferably, the preprocessing includes normalization operations and data enhancement operations; the different modality MRI brain tumor images include images of T1, T1c, T2, and FLAIR modalities; during the training process of the network model, the preprocessed data set is divided into a training set, a validation set, and a test set, and a batch of images are randomly selected from the training set each time and input into the network model for forward propagation to calculate the loss, and the gradient is calculated by back propagation and the parameters of the network model are updated using the Adam optimizer. After each training round, the performance of the network model is evaluated on the validation set, and it is determined whether the network model is overfitted or whether it needs to continue training based on the indicators of the validation set. When the indicators of the validation set reach the preset optimal indicators or the training reaches the preset maximum rounds, the training is stopped and the final parameters of the network model are saved;
[0035] During the pre-training process of the network model, a full-modality replacement image is optimized to replace the MRI brain tumor image with the missing modality when performing segmentation reasoning on the new MRI brain tumor image. Specifically, when performing segmentation reasoning on the new MRI brain tumor image, if there is a modality missing, the MRI brain tumor image with the missing modality is directly replaced with the full-modality replacement image obtained in the pre-training process, and the complete or filled multi-modality MRI brain tumor image is input into the trained network model, and passes through the encoder, the multi-head attention mechanism module, the adaptive feature fusion module, the decoder and the segmentation head in sequence, and finally outputs the segmentation results of different tumor areas.
[0036] Preferably, the multi-head attention mechanism module is configured with four heads, and step S2 includes:
[0037] S201, annotating the first feature map with feature dimensions to obtain an annotated feature map, wherein the feature dimensions are [batch_size, num_channels, height, width, depth], where batch_size represents the data batch size of each input feature map, num_channels represents the number of channels of the feature map, and height, width, and depth represent the height, width, and depth of the feature map in three-dimensional space, respectively;
[0038] S202, evenly dividing the annotated feature map into four subspaces along the channel dimension, each subspace corresponds to a head, and each head corresponds to a linear transformation layer;
[0039] S203, perform linear transformation and dot product operation on the query vector, key vector, and value vector of each head through the linear transformation layer corresponding to each head to obtain the attention weight of each head. The calculation formula is:
[0040]
[0041] Among them, Attention(Q,K,V) represents the attention weight of each head, Q, K, and V represent the query vector, key vector, and value vector of each head after linear transformation, respectively. k is the dimension of the key vector;
[0042] S204, performing attention weighting on the feature map input to the head according to the attention weight of each head, to obtain the feature map after attention weighting of each head;
[0043] S205. The feature maps after the four heads' attention weighting are spliced along the channel dimension, and the spliced feature maps are linearly transformed to restore the spliced feature maps to the same number of channels as the first feature map, so as to obtain a second feature map after the attention weighted combination.
[0044] Preferably, assuming that the second feature map comes from n modes, the feature dimension of the feature map of each mode in the second feature map is annotated as [batch size ,num_channels i ,height i ,width i ,depth i ], where i represents the i-th mode and batch size Indicates the data batch size of each input feature map, num_channels i Indicates the number of channels of the feature map of the i-th mode, height i 、width i 、depth i Respectively represent the height, width and depth of the feature graph of the i-th mode in the three-dimensional space, n is a positive integer, and i is a positive integer less than or equal to n; step S3 includes:
[0045] S301, fuse the feature map of each modality with the average feature map of all modalities along the channel dimension to obtain the fused feature map of each modality, wherein the feature dimension of the fused feature map of each modality is marked as [batch size ,num_channels i +num_channels avg ,height i ,width i ,depthi ], num_channels avg represents the number of channels of the average feature map, where the average feature map is obtained by averaging the feature maps of all modalities in the channel dimension;
[0046] S302: The feature map after fusion of each modality is passed through the convolution layer corresponding to the modality to generate the initial attention weight map of each modality. The calculation formula is:
[0047]
[0048] in, represents the average feature map, W i represents the initial attention weight map of the i-th modality, F i represents the convolutional layer corresponding to the i-th modality, θ i represents the parameters of the convolutional layer corresponding to the i-th mode, σ represents the Sigmoid function, which is used to map the generated weight value to the (0,1) interval;
[0049] S303, normalize the initial attention weight map of each modality through the Softmax function to obtain the normalized attention weight map of each modality, and the calculation formula is:
[0050]
[0051] in, represents the normalized attention weight map of the i-th modality, N avail Indicates the number of currently available modalities, and the sum of the normalized attention weight maps of all modalities is 1;
[0052] S304, multiply the feature map of each modality by the normalized attention weight map of the modality point by point in the voxel dimension, and then sum the feature maps after weighted multiplication of all modalities to obtain the third feature map after adaptive fusion. The calculation formula is:
[0053]
[0054] in, represents the third feature map after adaptive fusion, Represents a point-wise multiplication operation in the voxel dimension.
[0055] Preferably, the encoder includes a plurality of convolutional layers and downsampling layers; the decoder adopts a structure of alternating a plurality of upsampling layers and a plurality of 3×3×3 convolutional layers, each upsampling layer is connected to a 3×3×3 convolutional layer, and the last layer of the decoder is a 1×1×1 convolutional layer. Step S4 includes:
[0056] S401, using a skip connection mechanism, concatenate and fuse the feature maps of the third feature map after being amplified by each upsampling layer with the feature maps of the same resolution stage after being processed by the downsampling layer along the channel dimension, input the concatenated and fused feature maps into the 3×3×3 convolutional layer corresponding to the current upsampling layer for processing, restore the third feature map to a size close to the original image resolution, and obtain a feature map of a size close to the original image resolution;
[0057] S402, converting the number of channels of the feature map having a size close to the original image resolution into the number of segmentation categories 3 through the 1×1×1 convolution layer, and then outputting the segmentation results of different tumor regions through an activation function.
[0058] The present invention also provides an MRI brain tumor image segmentation system with adaptive feature fusion, which is used to implement the MRI brain tumor image segmentation method with adaptive feature fusion. The system uses a network model to implement MRI brain tumor image segmentation. The network model includes: a basic network based on 3D U-Net, a multi-head attention mechanism module and an adaptive feature fusion module. The basic network based on 3D U-Net includes an encoder and a decoder; the encoder is used to extract features of MRI brain tumor images of different modalities to be segmented, and input the first feature maps of different tumor areas in the extracted MRI brain tumor images of different modalities into the multi-head attention mechanism module; the multi-head attention mechanism module is used to perform attention weighted combination on the first feature map, and input the second feature map after the attention weighted combination into the adaptive feature fusion module; the adaptive feature fusion module is used to adaptively weight and fuse the second feature map according to the contribution of the second feature map to brain tumor segmentation, and input the third feature map after adaptive fusion into the decoder; the decoder is used to restore the original image resolution of the third feature map to obtain segmentation results of different tumor areas.
[0059] Preferably, the system further includes: a training module, used to train the network model before using the network model, and the training process specifically includes: collecting data sets of MRI brain tumor images of different modalities for preprocessing to obtain a preprocessed data set; based on the preprocessed data set, pre-training the network model using a self-supervised learning method to obtain the pre-trained network model, and the pre-training process specifically includes: randomly selecting MRI brain tumor images of some modalities from the preprocessed data set for masking to simulate the missing conditions of different modalities, and at the same time, masking is also performed on some MRI brain tumor images of the remaining modalities, and the masked MRI brain tumor image is reconstructed through the encoder, the decoder, the multi-head attention mechanism module and the adaptive feature fusion module, and the mean square error is used as the loss function to calculate the difference between the reconstructed MRI brain tumor image and the original MRI brain tumor. The difference between the images is obtained by updating the parameters of the network model through back propagation to obtain the pre-trained network model; based on the pre-processed data set, the pre-trained network model is further trained in a supervised manner to obtain the trained network model, and the further training process specifically includes: each time a sample is selected from the pre-processed data set, different modality missing conditions are randomly generated as network input to obtain segmentation results under different modality missing conditions, based on the segmentation results under different modality missing conditions, the segmentation loss of the network model is calculated through a combined loss function, based on the segmentation loss and the consistency loss between different modality missing conditions, the parameters of the network model are further fine-tuned through back propagation, and when the network model can adapt to the segmentation tasks under different modality missing conditions, the training is stopped, the final parameters of the network model are saved, and the trained network model is obtained.
[0060] Preferably, the combined loss function is a weighted sum of a cross entropy loss and a Dice loss; the cross entropy loss is used to measure the probability difference between the predicted segmentation result and the actual segmentation label at each pixel, and the calculation formula is:
[0061]
[0062] Among them, L CE represents the cross entropy loss, N represents the total number of pixels, C represents the number of segmentation categories, i represents the i-th pixel, j represents the j-th segmentation category, i and j are positive integers, and y ij Indicates the true value of the i-th pixel belonging to the j-th category in the true segmentation label, p ij Indicates the probability value that the i-th pixel in the predicted segmentation result belongs to the j-th category;
[0063] The Dice loss is used to measure the overlap between the predicted segmentation result and the actual segmentation label, and the calculation formula is:
[0064]
[0065] Among them, L Dice represents the Dice loss, p i Indicates the probability value of the i-th pixel in the predicted segmentation result belonging to the target category, y i Represents the probability value that the i-th pixel in the real segmentation label belongs to the target category, ∈ is a positive number to avoid the situation where the denominator is 0; the combined loss function L total is the cross entropy loss L CE With the Dice loss L Dice The weighted sum of total =αL CE +(1-α)L Dice , where α represents the trade-off of the cross entropy loss L CE With the Dice loss L Dice Importance weight parameter.
[0066] Compared with the prior art, the present invention has the following beneficial effects:
[0067] 1. Improve segmentation efficiency and accuracy. The present invention overcomes the defects of the prior art in segmenting brain tumor MRI images, and improves segmentation efficiency and accuracy. Improve segmentation efficiency: The present invention improves the network structure to make it more concise and efficient, reduce unnecessary computing resource consumption and training time, improve segmentation efficiency, and facilitate more convenient deployment and use in actual application scenarios. Improve segmentation accuracy: The present invention optimizes the network structure to more effectively utilize the information contained in each modality, making the feature fusion between different modalities more reasonable and accurate, thereby improving the accuracy of brain tumor segmentation and providing more reliable segmentation results for application scenarios such as clinical diagnosis.
[0068] 2. Reduce training and deployment costs. The present invention avoids the high cost problem faced by traditional multi-model methods in brain tumor MRI image segmentation. The present invention adopts a simpler and more efficient structural design to reduce the resource consumption and training difficulty caused by multiple encoders, decoders and complex interactions, and avoids problems such as slow computing speed and overfitting due to complex structure. It can achieve efficient segmentation of brain tumor MRI images with less resource investment, and is convenient for deployment and use in various practical application scenarios such as primary medical units and general medical equipment, reducing overall training and deployment costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1A schematic diagram of a flow chart of an MRI brain tumor image segmentation method with adaptive feature fusion provided by the present invention;
[0070] Figure 2 A schematic diagram of the training process flow of an MRI brain tumor image segmentation method with adaptive feature fusion provided by the present invention;
[0071] Figure 3 A schematic flow chart of a multi-head attention mechanism module of an MRI brain tumor image segmentation method with adaptive feature fusion provided by the present invention;
[0072] Figure 4 A schematic diagram of the framework of the multi-head attention mechanism module provided by the present invention;
[0073] Figure 5 A schematic diagram of the flow of an adaptive feature fusion module of another adaptive feature fusion MRI brain tumor image segmentation method provided by the present invention;
[0074] Figure 6 A schematic diagram of the framework of the adaptive feature fusion module provided by the present invention;
[0075] Figure 7 A schematic diagram of a decoder flow of an MRI brain tumor image segmentation method with adaptive feature fusion provided by the present invention;
[0076] Figure 8 A schematic diagram of the structure of an MRI brain tumor image segmentation system with adaptive feature fusion provided by the present invention. DETAILED DESCRIPTION
[0077] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0078] like Figure 1 As shown, an embodiment of the present invention provides an MRI brain tumor image segmentation method with adaptive feature fusion. The method uses a network model to implement MRI brain tumor image segmentation. The network model includes: a basic network based on 3D U-Net, a multi-head attention mechanism module and an adaptive feature fusion module. The basic network based on 3D U-Net includes an encoder and a decoder. The method includes:
[0079] S1. The encoder extracts features from the MRI brain tumor images of different modalities to be segmented, and inputs the first feature maps of different tumor areas in the extracted MRI brain tumor images of different modalities into the multi-head attention mechanism module;
[0080] S2, the multi-head attention mechanism module performs attention weighted combination on the first feature map, and inputs the second feature map after the attention weighted combination into the adaptive feature fusion module;
[0081] S3, the adaptive feature fusion module adaptively weights and fuses the second feature map according to the contribution of the second feature map to brain tumor segmentation, and inputs the adaptively fused third feature map into the decoder;
[0082] S4. The decoder restores the original image resolution of the third feature map to obtain segmentation results of different tumor areas.
[0083] In terms of the overall network architecture design, the present invention selects a 3D U-Net structure as the basic network framework. 3D U-Net itself has good feature extraction capabilities. Its encoder part can gradually extract deep-level feature information from the input MRI brain tumor image through a combination of a series of convolutional layers and downsampling layers. These features cover global semantic information from local to more abstract as the number of network layers increases. The decoder part can gradually restore the deep-level features to a size close to the original image resolution through upsampling operations, and through the jump connection mechanism, the features of the encoder at different stages are fused with the features of the decoder at the corresponding stage, effectively supplementing the spatial information lost in the downsampling process, thereby showing certain advantages in the field of medical image segmentation. This basic network framework provides a stable and reliable feature processing foundation for subsequent improvements, and can well adapt to the characteristics of brain tumor MRI images and the needs of segmentation tasks.
[0084] like Figure 2 As shown, in the embodiment of the present invention, before using the network model, the network model needs to be trained, and the training process specifically includes:
[0085] A1. Collect data sets of MRI brain tumor images of different modalities for preprocessing to obtain preprocessed data sets; first collect multi-modal MRI brain tumor image data sets, including image data of commonly used modalities such as T1, T1c, T2, and FLAIR. Preprocess these image data, including normalization operations, normalizing the pixel values of the image to a specific interval (such as the [0,1] interval), so that the network can converge more stably during training. At the same time, data enhancement operations such as random rotation, flipping, scaling, etc. are also performed to increase the diversity of training data and improve the generalization ability of the network.
[0086] A2. Based on the preprocessed data set, the network model is pre-trained by self-supervised learning to obtain a pre-trained network model. The pre-training process specifically includes: randomly selecting some modal MRI brain tumor images from the preprocessed data set for masking to simulate the missing conditions of different modalities, and also masking some MRI brain tumor images of the remaining modalities, reconstructing the masked MRI brain tumor images through the encoder, decoder, multi-head attention mechanism module and adaptive feature fusion module, using the mean square error as the loss function, calculating the difference between the reconstructed MRI brain tumor image and the original MRI brain tumor image, updating the parameters of the network model through back propagation, and obtaining the pre-trained network model;
[0087] A3. Based on the preprocessed data set, the pre-trained network model is further trained in a supervised manner to obtain a trained network model. The further training process specifically includes: each time a sample is selected from the preprocessed data set, different modal missing conditions are randomly generated as network inputs to obtain segmentation results under different modal missing conditions. Based on the segmentation results under different modal missing conditions, the segmentation loss of the network model is calculated through a combined loss function. Based on the segmentation loss and the consistency loss between different modal missing conditions, the parameters of the network model are further fine-tuned through back propagation. When the network model can adapt to the segmentation tasks under different modal missing conditions, the training is stopped, the final parameters of the network model are saved, and the trained network model is obtained.
[0088] The combined loss function includes cross entropy loss and Dice loss; cross entropy loss is used to measure the probability difference between the predicted segmentation result and the actual segmentation label at each pixel, and the calculation formula is:
[0089]
[0090] Among them, L CE represents the cross entropy loss, N represents the total number of pixels, C represents the number of segmentation categories, i represents the i-th pixel, j represents the j-th segmentation category, i and j are positive integers, and y ij Indicates the true value (0 or 1) of the i-th pixel belonging to the j-th category in the true segmentation label, p ij It represents the probability value that the i-th pixel in the predicted segmentation result belongs to the j-th category; Dice loss is used to measure the overlap between the predicted segmentation result and the actual segmentation label, and the calculation formula is:
[0091]
[0092] Among them, L Dice represents the Dice loss, p iIndicates the probability value of the i-th pixel in the predicted segmentation result belonging to the target category, y i Represents the probability value of the i-th pixel in the true segmentation label belonging to the target category, ∈ represents a very small positive number, which is used to avoid the situation where the denominator is 0; the combined loss function L total is the cross entropy loss L CE With Dice loss L Dice The weighted sum of total =αL CE +(1-α)L Dice , where α represents the trade-off between cross entropy loss L CE With Dice loss L Dice Importance weight parameter.
[0093] In an embodiment of the present invention, preprocessing includes normalization operation and data enhancement operation; different modality MRI brain tumor images include images of T1, T1c, T2, and FLAIR modalities. During the training process, the data set is divided into a training set, a validation set, and a test set (for example, divided in a ratio of 8:1:1), and a batch of images are randomly selected from the training set each time to input the network model for forward propagation to calculate the loss, and the gradient is calculated by back propagation and the parameters of the network model are updated using the Adam optimizer. After each training round, the performance of the network model is evaluated on the validation set, and the network model is judged whether it is overfitting or whether it needs to continue training according to the indicators of the validation set (loss value or accuracy) When the indicators of the validation set reach the preset optimal indicators or the training reaches the preset maximum rounds (such as 200 rounds), stop training and save the parameters of the final network model. In this embodiment, the Adam optimizer is selected to update the parameters of the network, which has the advantage of adaptive learning rate, and can dynamically adjust the learning rate of each parameter according to the gradient of the parameter, accelerate the convergence speed and improve the stability of training. The initial learning rate is set to 0.001, and a learning rate decay strategy is adopted. For example, after a certain number of training rounds (epochs), the learning rate is decayed according to a certain ratio (such as 0.9) to avoid the situation in which the learning rate is too large in the later stage of training and cannot converge to the optimal solution.
[0094] In this embodiment, the network model includes a training process and a reasoning process. The training of the entire network is divided into two main stages. The first stage: pre-training is performed using a self-supervised learning method combined with the improved structure of the present invention. Specifically, when inputting an MRI brain tumor image, some modalities are randomly selected for masking (simulating the lack of modality), and some three-dimensional image blocks of the remaining modalities are also masked, just like in the masked autoencoder (MAE) task of natural images, except that this is for multimodal medical image scenes. The network must reconstruct the masked content through the codec, the middle multi-head attention mechanism, and the final adaptive feature fusion module, using the mean square error (MSE) as the loss function, that is, calculating the difference between the reconstructed image and the original image, and updating the network parameters through back propagation. At the same time, in this process, a full-modal replacement image is optimized (through model inversion, according to certain regularization constraints, a full-modal image representation that can minimize the reconstruction error is found. This image can be used to replace the missing modal information in the subsequent reasoning stage). Phase II: Fine-tune the network parameters obtained from the pre-training in Phase I, and use a supervised approach to train with labeled brain tumor segmentation image data. Each time a sample is selected from the data set, different modality loss situations are randomly generated as network inputs (by randomly discarding some modalities), and the features processed by the multi-head attention mechanism and the adaptive feature fusion module are input to the segmentation head (composed of convolutional layers, etc., for the final generation of segmentation results). At the same time, the segmentation loss (such as the combination of Dice loss and cross entropy loss to measure the difference between the segmentation result and the true annotation) and the consistency loss based on different modality loss situations (by encouraging the consistency between the features extracted in different missing modal situations, using the idea of knowledge distillation to improve the network's generalization ability for different modality combinations) are calculated. These losses are combined to further fine-tune the network parameters through back propagation, so that the network can better adapt to the brain tumor segmentation task in different modality loss situations.
[0095] For the reasoning process of the network model, when performing segmentation reasoning on a new MRI brain tumor image, if there is a missing modality, the missing modality information is directly replaced with the full-modality replacement image obtained in the pre-training stage, and the complete (or filled) multimodal image is input into the trained network, and then passes through the encoder, multi-head attention mechanism, adaptive feature fusion module, decoder and segmentation head in sequence, and finally outputs the segmentation result of the brain tumor to obtain a predicted map of the tumor area, such as distinguishing different areas such as the tumor core, peritumoral edema and enhanced tumor.
[0096] like Figure 3-4 As shown, in this embodiment of the present invention, the multi-head attention mechanism module sets four heads, and step S2 includes:
[0097] S201, annotating the feature dimension of the first feature map to obtain an annotated feature map;
[0098] The feature dimension is [batch_size, num_channels, height, width, depth], where batch_size represents the data batch size of each input feature map, num_channels represents the number of channels of the feature map, height, width, and depth represent the height, width, and depth of the feature map in three-dimensional space respectively;
[0099] S202, evenly divide the annotated feature map into four subspaces along the channel dimension, each subspace corresponds to a head, and each head corresponds to a linear transformation layer;
[0100] S203, perform linear transformation and dot product operation on the query vector, key vector, and value vector of each head through the linear transformation layer corresponding to each head to obtain the attention weight of each head; the calculation formula is:
[0101]
[0102] Among them, Attention(Q,K,V) represents the attention weight of each head, Q, K, and V represent the query vector, key vector, and value vector of each head after linear transformation, respectively. k is the dimension of the key vector;
[0103] S204, performing attention weighting on the feature map input to the head according to the attention weight of each head, to obtain the feature map after attention weighting of each head;
[0104] S205. Concatenate the feature maps after the four heads' attention weighting along the channel dimension, perform a linear transformation on the concatenated feature maps, restore the concatenated feature maps to the same number of channels as the first feature maps, and obtain a second feature map after the attention weighted combination.
[0105] In the present invention, the effective application of the multi-head attention mechanism is a key technical point. It is crucial to reasonably set the relevant parameters of the multi-head attention mechanism, which determines whether it can accurately capture the association between the features of the codecs without introducing too much computation and overfitting risks. Specifically, parameters such as the number of heads and the dimension of each head need to be adapted to the specific brain tumor segmentation task. For example, if the number of heads is set too small, it may not be possible to fully explore the complex relationship between features from multiple angles, and it is not possible to fully capture the valuable information for brain tumor segmentation in different modality images; and if the number of heads is too large, it will greatly increase the amount of calculation, resulting in a significant increase in training and reasoning time, and may even cause overfitting due to the complexity of the model, making the segmentation effect of new brain tumor MRI images worse in practical applications. After a large number of experimental verifications, we set the number of heads to 4. Under this setting, the model can not only better learn the association between different modalities and different position features of the same modality, but also ensure that the overall computing resource consumption is within an acceptable range, ensuring that brain tumor MRI images can be quickly segmented in practical application scenarios. At the same time, the dimension corresponding to each head has also been fine-tuned to match the dimension of the input features and the overall network structure, so that the features output by the multi-head attention mechanism can be accurately passed to the decoder part, providing more discriminative feature information for subsequent segmentation operations, helping to improve the accuracy of brain tumor segmentation.
[0106] like Figure 5-6 As shown, in the embodiment of the present invention, assuming that the second feature map comes from n modes, the feature dimension of the feature map of each mode in the second feature map is marked as [batch size ,num_channels i ,height i ,width i ,depth i ], where i represents the i-th mode and batch size Indicates the data batch size of each input feature map, num_channels i Indicates the number of channels of the feature map of the i-th mode, height i 、width i 、depth i Respectively represent the height, width and depth of the feature graph of the i-th mode in the three-dimensional space, n is a positive integer, and i is a positive integer less than or equal to n. Step S3 includes:
[0107] S301, fusing the feature map of each modality with the average feature map of all modalities along the channel dimension to obtain a fused feature map of each modality;
[0108] Among them, the feature dimension of each modality fusion feature map is marked as [batch size ,num_channels i +num_channels avg ,height i ,width i ,depth i ], num_channels avg Indicates the number of channels of the average feature map. The average feature map is obtained by averaging the feature maps of all modalities in the channel dimension;
[0109] S302: Generate an initial attention weight map of each modality through the convolution layer corresponding to the modality by passing the fused feature map of each modality; the calculation formula is:
[0110]
[0111] in, represents the average feature map, W i represents the initial attention weight map of the i-th modality, F i represents the convolutional layer corresponding to the i-th modality, θ i represents the parameters of the convolutional layer corresponding to the i-th mode, σ represents the Sigmoid function, which is used to map the generated weight value to the (0,1) interval;
[0112] S303, normalizing the initial attention weight map of each modality through the Softmax function to obtain the normalized attention weight map of each modality; the calculation formula is:
[0113]
[0114] in, represents the normalized attention weight map of the i-th modality, N avail Indicates the number of currently available modalities, and the sum of the normalized attention weight maps of all modalities is 1;
[0115] S304, multiply the feature map of each modality by the normalized attention weight map of the modality point by point in the voxel dimension, and then sum the feature maps after weighted multiplication of all modalities to obtain a third feature map after adaptive fusion; the calculation formula is:
[0116]
[0117] in, represents the third feature map after adaptive fusion, Represents a point-wise multiplication operation in the voxel dimension.
[0118] Specifically, suppose the input is the feature maps from different modalities T1, T1c, T2, and FLAIR, denoted as F T1 、F T1c 、F T2 、F FLAIR , whose dimensions are [batch size ,num_channels,height,width,depth]. First, the feature map of each modality is fused with the average feature map of all modalities along the channel dimension to obtain the fused feature map of each modality, which are represented as Its dimension is [batch size ,num_channels sum ,height,width,depth], where num_channels sum is the sum of the number of feature channels of each modality. Next, for the combined features of each modality, a convolution operation is performed through a specific convolution layer (the convolution kernel size, step size and other parameters are optimized and selected based on multiple experiments) to obtain a preliminary weight feature map. For example, for the T1 modality, after the convolution layer (the corresponding parameter is θ T1 )get The same logic can be applied to other modes to obtain W T1c , W T2 , W FLAIR , whose dimension is [batch size ,num_modalities,height,width,depth], num_modalities represents the number of modalities. The value of each position in this initial weight map represents the initial importance of the corresponding modality at that position. Then, considering the possibility of missing modalities, different numbers of feature maps will be used for fusion, so these weights need to be normalized, and the Softmax function is used for normalization to obtain the final attention weight map Finally, the feature maps F of each modality are T1 、F T1c 、F T2 、F FLAIR And the corresponding attention weight map The weights of the corresponding modes in are multiplied element by element, and the multiplication results are added in sequence to obtain the feature map after adaptive fusion In this way, through the adaptive feature fusion module, the network can reasonably fuse the features of each modality according to the actual importance of each modality in different input images, and provide more discriminative feature representation for subsequent brain tumor segmentation.
[0119] like Figure 7As shown, in this embodiment of the present invention, the encoder includes multiple convolutional layers and downsampling layers. The decoder adopts a structure in which multiple upsampling layers and multiple 3×3×3 convolutional layers are alternated. Each upsampling layer is followed by a 3×3×3 convolutional layer, and the last layer of the decoder is a 1×1×1 convolutional layer. Step S4 includes:
[0120] S401, using a skip connection mechanism, concatenate and fuse the feature map of the third feature map after being amplified by each upsampling layer with the feature map of the same resolution stage after being processed by the downsampling layer along the channel dimension, input the concatenated and fused feature map into the 3×3×3 convolution layer corresponding to the current upsampling layer for processing, restore the third feature map to a size close to the resolution of the original image, and obtain a feature map of a size close to the resolution of the original image;
[0121] S402, converting the number of channels of the feature map with a size close to the original image resolution into the number of segmentation categories 3 through a 1×1×1 convolution layer, and then outputting the segmentation results of different tumor areas through an activation function.
[0122] In this embodiment, the decoder is mainly responsible for gradually restoring the features processed by the multi-head attention mechanism and the adaptive feature fusion to a size close to the resolution of the original input image to generate the final brain tumor segmentation result. The decoder part adopts a structure of alternating upsampling layers and convolution layers. First, the size of the feature map is gradually enlarged by upsampling methods such as transposed convolution (also called deconvolution) or interpolation, for example, doubling the size in the height, width, and depth directions each time. After each upsampling, a convolution layer is connected. The convolution kernel size is usually set to 3×3×3 (corresponding to a three-dimensional image), the step size is 1, and the filling method selects a suitable mode according to the actual situation (such as same filling to keep the feature map size unchanged). The purpose is to further adjust and refine the features of the enlarged feature map so that it contains more accurate spatial information and semantic information. At the same time, the decoder also uses a jump connection mechanism to splice and fuse the feature map of the corresponding resolution stage in the encoder with the feature map of the current stage of the decoder. The feature map in the encoder, which is downsampled to 1 / 8 of the original image, will be concatenated with the feature map upsampled to the same 1 / 8 size in the decoder along the channel dimension, and then input into the subsequent convolutional layer for processing. This can effectively supplement the detail information lost during the encoder downsampling process, so that the final generated segmentation result can take into account both the deep semantic features and the local detail features in the original image, and more accurately locate and divide the brain tumor area. Finally, in the last layer of the decoder, a 1×1×1 convolution layer is used to convert the number of channels of the feature map to the number of segmentation categories 3, and then an activation function (such as Sigmoid function for binary classification or Softmax function for multi-classification) is passed to output the final brain tumor segmentation probability map. The higher the probability value, the more likely it is that the area belongs to the corresponding brain tumor category.
[0123] The data processing flow of the present invention is as follows: MRI brain tumor image data is input, and these data contain images of different modalities (such as T1, T1c, T2, FLAIR). Before being sent to the network, a series of preprocessing operations will be performed. For example, the image is normalized, and the pixel values are normalized to a specific interval (such as ([0,1]) interval) to ensure the consistency of the numerical range of different image data, which is convenient for subsequent network calculations; at the same time, the image will be cropped according to the set size, for example, cropped to a uniform size of [128,128,128] (the size can be adjusted and optimized according to actual conditions and hardware resources, etc.), and irrelevant background areas in the image are removed, etc., so as to reduce the amount of data while highlighting the brain tumor area of focus. The preprocessed data is sent to the network and firstly extracted by the encoder part. The encoder consists of multiple convolutional layers and downsampling layers. As the number of network layers increases, the resolution of the image gradually decreases, the number of feature channels gradually increases, and abstract features from local to global are gradually extracted. For example, the local texture, edge and other features in the image are initially extracted, and more abstract semantic features such as the entire brain tumor area and its relationship with surrounding tissues are obtained when the deep network is reached. Then, the extracted features are passed to the multi-head attention mechanism module between the encoder and decoder. According to the calculation method of the multi-head attention mechanism mentioned above, the features are processed to enhance the ability to capture the associated information between the features and output more semantically rich features. The features processed by the multi-head attention mechanism enter the adaptive feature fusion module, and adaptive weights are allocated and fused according to the importance of different modal features to obtain the fused features. The fused features are sent to the decoder, which gradually restores the resolution of the features to a size close to the original image through the upsampling layer. At the same time, it uses jump connections to fuse the features of different stages of the encoder, supplement the spatial information, and finally output the segmentation result. The segmentation result is presented in the form of a probability map or a binary map (depending on the specific output settings and post-processing methods), indicating the area where the brain tumor is located in the image.
[0124] like Figure 8As shown, the present invention also provides an MRI brain tumor image segmentation system with adaptive feature fusion, which realizes MRI brain tumor image segmentation by using a network model, wherein the network model includes: a basic network based on 3D U-Net, a multi-head attention mechanism module 200 and an adaptive feature fusion module 300, and the basic network based on 3D U-Net includes an encoder 100 and a decoder 400. The encoder 100 is used to extract features of MRI brain tumor images of different modalities to be segmented, and input the feature maps of different tumor areas in the extracted MRI brain tumor images of different modalities into the multi-head attention mechanism module. The multi-head attention mechanism module 200 is used to perform attention weighted combination on the feature maps input into the multi-head attention mechanism module, and input the feature maps after the attention weighted combination into the adaptive feature fusion module. The adaptive feature fusion module 300 is used to adaptively weight and fuse the feature maps input into the adaptive feature fusion module according to the contribution of the feature maps input into the adaptive feature fusion module to brain tumor segmentation, and input the adaptively fused feature maps into the decoder. The decoder 400 is used to restore the original image resolution of the feature maps input into the decoder to obtain segmentation results of different tumor areas.
[0125] The present invention adopts a multi-head attention mechanism module to realize multi-view feature attention and feature fusion enhancement. Multi-view feature attention: a reasonable number of heads 4 are set so that the network can simultaneously focus on input features from multiple different perspectives. Each head independently calculates the attention weight to capture the key information of the feature in different representation subspaces, which helps to fully mine the diverse features of the tumor area in the multimodal MRI image, including the texture, shape, position and other information of the tumor under different modalities, and provide rich feature basis for subsequent precise segmentation. Feature fusion enhancement: multiple weighted feature representations output by the multi-head attention mechanism are fused through splicing and linear transformation to integrate the key information under different perspectives into a more representative feature vector. This fusion method can fully retain and integrate the feature advantages captured by each head, further enhance the expressive power of the feature, so that the features transmitted to the decoder contain more comprehensive and in-depth semantic information, which helps the decoder to generate segmentation results more accurately.
[0126] The present invention adopts an adaptive feature fusion module to realize dynamic weight calculation, weight normalization processing and adaptive fusion strategy. When calculating the attention weight, appropriate linear transformation and dot product operation methods are adopted to ensure that each head can adaptively allocate weights according to the characteristics of the input features. For features that are outstanding in a certain modality or region, the corresponding head will be given a higher weight, so that the network can focus on the most valuable feature part for segmentation, avoid the influence of irrelevant or interfering information, and effectively improve the accuracy and effectiveness of feature expression. Dynamic weight calculation: According to the contribution of different modal MRI images to brain tumor segmentation, a dynamic weight calculation mechanism is constructed through a specific convolution layer and activation function (such as Sigmoid function). The features of each modality are spliced with the average fused features, and then processed by the convolution layer to obtain the initial attention weight, which can reflect the importance of each modality in the current image scene. For example, for a modality that can clearly display the tumor boundary, its weight will be increased accordingly, so as to play a greater role in the fusion process. Weight normalization processing: Considering the uncertainty of modality missing, the attention weights of different modalities are normalized by using the Softmax function. This ensures that no matter how the number of modalities involved in the fusion changes, the sum of the weights of all modalities is always 1, making the fusion process consistent and stable. The normalized weights can reasonably allocate the contribution ratio of each modality in the final fusion feature, ensuring that the fused features can effectively reflect the characteristic information of the tumor under different modality combinations, avoiding segmentation errors caused by modality loss or weight imbalance. Adaptive fusion strategy: According to the calculated normalized weights, each modality feature is multiplied pixel by pixel (voxel-level multiplication) and summed to achieve adaptive feature fusion. This fusion strategy enables the network to dynamically adjust the proportion of different modalities in the fusion feature according to their actual performance in the image, and make full use of the advantageous information of each modality. For example, in some images, the T2-weighted image contributes greatly to the peritumoral edema feature, so its feature will be given a higher weight during fusion, thereby highlighting the features of the edema area and improving the segmentation accuracy.
[0127] The present invention realizes the coordinated optimization of the overall network structure. Seamless connection between modules: ensure that the multi-head attention mechanism and the adaptive feature fusion module are seamlessly connected in the codec structure, so that data can flow and process smoothly between modules. Coordinated adjustment of parameters: during the training process, the parameters of the overall network structure are coordinated and adjusted, including the number of heads in the multi-head attention mechanism, the dimensions of each head, the convolutional layer parameters in the adaptive feature fusion module, the learning rate of the network, the regularization parameters, etc. Training and inference optimization: design efficient training and inference processes to give full play to the advantages of the improved network structure.
[0128] Example 1 Experimental setting: 200 brain tumor MRI image data from the BraTS2020 dataset were selected for the experiment. First, the images were preprocessed and cropped to a uniform size of 128×128×128 pixels, and the pixel values of the images were normalized to the interval [0,1]. During network training, the initial learning rate was set to 0.001, the Adam optimizer was used, and the number of iterations was set to 500. The number of multi-head attention mechanism heads was set to 4, and the weight initialization in the adaptive feature fusion module adopted a random normal distribution, with the mean set to 0 and the standard deviation set to 0.01. Experimental results: After network training and testing, the segmentation results were obtained. As shown in the data of Table 1, the method AdapterMedSeg adopted by the present invention is compared with the existing methods U-HVED (Hetero-modal variational encoder-decoder), RFNet (Region-aware fusion network), ModGen (Models Genesis: Generic autodidactic models for 3D medical image analysis), and M3AE (Multimodal Masked Autoencoder). The present invention has certain advantages in segmenting DSC (Dice similarity coefficient) and HD96 (Hausdorff distance), especially in the segmentation effect of the tumor core area and the enhanced tumor area. It is more obvious, which reflects that the present invention can more accurately segment brain tumor MRI images in practical applications, and provide more valuable reference information for subsequent clinical diagnosis and other work.
[0129]
[0130]
[0131] Table 1
[0132] Example 2 Experimental settings: In this example, we selected 200 MRI brain tumor image data from the BraTS2020 dataset for the experiment. For data preprocessing, we changed the normalization method and adopted the method of normalizing the image intensity value to a standard normal distribution with a mean of 0 and a variance of 1, replacing the normalization operation in the previous example. At the same time, the hyperparameters of the network are fine-tuned. In view of the changes in the amount of data and the preprocessing method, we appropriately reduce the learning rate to 0.0001 in the hope that the network can converge more stably. In terms of parameter adjustment of the network structure, we change the number of heads of the multi-head attention mechanism to 6, hoping to observe its influence on the degree of attention to different modal features and the final segmentation effect by adjusting this parameter. In addition, for the adaptive feature fusion module, we appropriately adjust the key parameters that control the weight allocation of different modalities, such as adjusting the initial weight value of the weight corresponding to the T1 weighted modal feature fusion from 0.2 to 0.25, in order to explore the role of different initial weight settings on feature fusion and segmentation results under this instance data. Experimental results: After model training and testing under the above experimental settings, we obtained the corresponding brain tumor segmentation results. As shown in Table 2, compared with the segmentation accuracy indicators in the previous examples, it is found that the average Dice coefficient of the overall tumor segmentation has slightly decreased from 0.908 to 0.906, the average Dice coefficient of the core part of the tumor has decreased from 0.867 to 0.863, and the average Dice coefficient of the enhanced tumor area has decreased from 0.816 to 0.812. From the segmentation results, the segmentation accuracy is lower than that of the previous examples, and some boundary blurring has occurred. This may be due to the change in the number of heads of the multi-head attention mechanism and the adjustment of the weight parameters of the adaptive feature fusion module. The network's capture and fusion effects of different modal features have changed, resulting in the failure to fully utilize the advantageous information of each modality to accurately determine the tumor boundary when generating the segmentation results. However, in the segmentation of some peritumoral edema areas, compared with the previous examples, the accuracy of the segmentation results has not changed much, indicating that after adjusting the parameters of the adaptive feature fusion module, the fusion and utilization of the relevant modal features of the edema area can still maintain a certain stability. This example further verifies the stability and effectiveness of the present invention under different parameter settings. It also reflects that the parameter adjustment of each module has an important influence on the final segmentation effect. In practical applications, reasonable parameter optimization is required according to specific data conditions and task requirements.
[0133]
[0134] Table 2
[0135] Example 3 Experimental settings: In Example 3, we reset the appropriate preprocessing process. First, the image is normalized so that its pixel value range is unified to the [0,1] interval to eliminate the intensity differences that may exist when different devices collect images. Then, according to the size distribution of the image in the data set, it is uniformly cropped to a fixed size, such as a cube image of 128×128×128 pixels, to facilitate subsequent network input and processing. In terms of network parameters, we adjusted them according to the scale of the data set and the complexity of the image. For example, the initial learning rate is set to 0.001, the Adam optimizer is used to update the network parameters, and the learning rate attenuation strategy is set. After a certain number of training rounds, the learning rate is attenuated according to a specific ratio, such as attenuating to 0.9 times the original every 50 rounds. At the same time, considering the number of samples in the data set, we set the batch size to 4 and the total training rounds to 500 rounds to make full use of the data for model training and ensure that the model can learn sufficiently effective features for brain tumor segmentation. Experimental results: By running the MRI brain tumor image segmentation method containing adaptive feature fusion on the new data set, we obtained the corresponding segmentation results. From the data in Table 3, the present invention can still accurately segment the brain tumor area from the MRI image, whether it is the tumor core area, the peritumoral edema area or the enhanced tumor area. Compared with some traditional segmentation methods on this data set, the method proposed by the present invention has obvious advantages in accuracy, and can achieve higher Dice similarity coefficient values and excellent HD96, indicating that the segmentation results are more closely aligned with the real annotation. This result fully demonstrates that the present invention has good applicability in different tumor areas. In the face of brain tumor MRI image data collected from different medical institutions and different equipment, it has the potential to play the advantages of accurate segmentation and provide reliable image segmentation basis for the diagnosis and subsequent treatment of brain tumors. After a linear transformation layer, the spliced feature dimension is restored to the dimension consistent with the input feature, and the feature output processed by the multi-head attention mechanism is obtained, which is passed to the decoder to continue the subsequent segmentation operation. Such a multi-head attention mechanism enables the network to capture the relationship between the features of different modal images more comprehensively and meticulously, which helps to improve the accuracy of brain tumor segmentation.
[0136]
[0137] Table 3
[0138] Combined with the experimental comparison of Examples 1, 2 and 3, it can be seen that the method of the present invention improves the segmentation accuracy. In order to verify the effect of the present invention in MRI brain tumor image segmentation, we conducted experiments on the BraTS2020 dataset, a public brain tumor MRI image segmentation dataset, and compared it with the relevant methods in the existing background technology. The experimental results show that the present invention shows significant advantages in segmentation accuracy. For example, using the common Dice similarity coefficient (DSC) and HD96 as evaluation indicators, the present invention is compared with some existing technologies based on specific model training (requires training of multiple models of collaborative training strategy-related methods) and some complex single model structures (there are multiple encoders, decoders and complex interactive methods). In the segmentation of different areas of brain tumors (the entire tumor area, the core area of the tumor, and the enhanced tumor area), the DSC value is significantly improved. This means that the segmentation results generated by the present invention are closer to the actual tumor area situation, and can more accurately outline the boundaries and ranges of brain tumors in MRI images, providing a more reliable basis for subsequent clinical diagnosis, disease monitoring, etc.
[0139] Compared with the prior art, the present invention has the following advantages:
[0140] 1. Improve segmentation accuracy. By adding a multi-head attention mechanism to the network structure and adopting an adaptive feature fusion module, the present invention has significantly improved the accuracy of brain tumor segmentation. The multi-head attention mechanism can capture the complex correlation information between images of different modalities and focus on key feature areas. The adaptive feature fusion module can adaptively assign weights to perform feature fusion based on the importance of each modality, making full use of the effective information of each modality. For example, when testing the BraTS2020 dataset, compared with traditional methods, the present invention improved the segmentation accuracy of the entire tumor area by about 2.8%, the core tumor area by about 10.2%, and the enhanced tumor area by about 10.7%, making the segmentation results more in line with the actual tumor boundaries and internal structures, providing a more reliable basis for clinical diagnosis.
[0141] 2. Improve training and deployment efficiency. This invention avoids the need to train a large number of models in traditional methods (such as training 2 models for N modes). n -1 model), only one improved network structure can be used to cope with multiple modal combinations, greatly reducing training time and storage space usage. Compared with some complex single-model methods, the present invention has a simple structure, does not have multiple complex encoders, decoders and interactions, is not prone to overfitting, gradient disappearance and other problems during training, can converge to a better state faster, and has relatively low hardware resource requirements during actual deployment. It can be easily applied to equipment in various medical scenarios, whether it is a high-end imaging diagnostic workstation or an ordinary computer device in a primary medical unit, and can be efficiently deployed and operated.
[0142] 3. Enhance the generalization ability of the model. Since the present invention can more reasonably utilize the feature information of each modality, the generalization ability of the model is enhanced when facing brain tumor MRI images of different individuals, different stages of the disease, and different modality missing situations. By testing on data sets from multiple different sources and simulated different modality missing scenarios, the model of the present invention can maintain relatively stable segmentation performance under different data distributions. Unlike some traditional methods, the segmentation accuracy drops significantly when faced with new data features or modality missing situations. It shows good versatility and adaptability, and can better meet the diverse needs in actual clinical applications.
[0143] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An adaptive feature fusion MRI brain tumor image segmentation method, characterized in that: The method uses a network model to implement MRI brain tumor image segmentation, the network model includes: a basic network based on 3DU-Net, a multi-head attention mechanism module and an adaptive feature fusion module, the basic network based on 3D U-Net includes an encoder and a decoder, and the method includes: S1, the encoder extracts features from the MRI brain tumor images of different modalities to be segmented, and inputs the first feature maps of different tumor areas in the extracted MRI brain tumor images of different modalities into the multi-head attention mechanism module; S2, the multi-head attention mechanism module performs attention weighted combination on the first feature map, and inputs the second feature map after the attention weighted combination into the adaptive feature fusion module; S3, the adaptive feature fusion module adaptively weights and fuses the second feature map according to the contribution of the second feature map to brain tumor segmentation, and inputs the adaptively fused third feature map into the decoder; S4. The decoder restores the original image resolution of the third feature map to obtain segmentation results of different tumor areas.
2. The MRI brain tumor image segmentation method based on adaptive feature fusion according to claim 1, characterized in that: Before using the network model, the network model needs to be trained. The training process specifically includes: A1. Collecting data sets of MRI brain tumor images of different modalities for preprocessing to obtain preprocessed data sets; A2. Based on the preprocessed data set, the network model is pre-trained by self-supervised learning to obtain the pre-trained network model. The pre-training process specifically includes: Randomly select some modal MRI brain tumor images from the preprocessed data set for masking to simulate the missing of different modalities, and also mask some MRI brain tumor images of the remaining modalities, reconstruct the masked MRI brain tumor images through the encoder, the decoder, the multi-head attention mechanism module and the adaptive feature fusion module, calculate the difference between the reconstructed MRI brain tumor image and the original MRI brain tumor image by taking the mean square error as the loss function, and update the parameters of the network model by back propagation to obtain the pre-trained network model; A3. Based on the preprocessed data set, the pre-trained network model is further trained in a supervised manner to obtain the trained network model. The further training process specifically includes: Each time a sample is selected from the preprocessed data set, different modality missing conditions are randomly generated as network inputs to obtain segmentation results under different modality missing conditions. Based on the segmentation results under different modality missing conditions, the segmentation loss of the network model is calculated through a combined loss function. Based on the segmentation loss and the consistency loss between different modality missing conditions, the parameters of the network model are further fine-tuned through back propagation. When the network model can adapt to segmentation tasks under different modality missing conditions, the training is stopped, the final parameters of the network model are saved, and the trained network model is obtained.
3. The MRI brain tumor image segmentation method based on adaptive feature fusion according to claim 2, characterized in that: The combined loss function is a weighted sum of the cross entropy loss and the Dice loss; The cross entropy loss is used to measure the probability difference between the predicted segmentation result and the actual segmentation label at each pixel, and the calculation formula is: Among them, L CE represents the cross entropy loss, N represents the total number of pixels, C represents the number of segmentation categories, i represents the i-th pixel, j represents the j-th segmentation category, i and j are positive integers, and y ij Indicates the true value of the i-th pixel belonging to the j-th category in the true segmentation label, p ij Indicates the probability value that the i-th pixel in the predicted segmentation result belongs to the j-th category; The Dice loss is used to measure the overlap between the predicted segmentation result and the actual segmentation label, and the calculation formula is: Among them, L Dice represents the Dice loss, p i Indicates the probability value of the i-th pixel in the predicted segmentation result belonging to the target category, y i Indicates the probability value that the i-th pixel in the true segmentation label belongs to the target category, ∈ is a positive number to avoid the situation where the denominator is 0; The combined loss function L total is the cross entropy loss L CE With the Dice loss L Dice The weighted sum of total =αL CE +(1-α)L Dice , where α represents the trade-off between the cross entropy loss L CE With the Dice loss L Dice Importance weight parameter.
4. The MRI brain tumor image segmentation method based on adaptive feature fusion according to claim 2, characterized in that: The preprocessing includes normalization operation and data enhancement operation; the different modality MRI brain tumor images include T1, T1c, T2, and FLAIR modality images; During the training process of the network model, the preprocessed data set is divided into a training set, a validation set and a test set. A batch of images are randomly selected from the training set each time and input into the network model for forward propagation to calculate the loss. The gradient is calculated by back propagation and the parameters of the network model are updated using the Adam optimizer. After each training round, the performance of the network model is evaluated on the validation set. It is determined whether the network model is overfitted or whether further training is required based on the indicators of the validation set. When the indicators of the validation set reach the preset optimal indicators or the training reaches the preset maximum rounds, the training is stopped and the final parameters of the network model are saved. During the pre-training process of the network model, a full-modality replacement image is optimized to replace the MRI brain tumor image of the missing modality when performing segmentation reasoning on the new MRI brain tumor image, specifically including: When performing segmentation reasoning on a new MRI brain tumor image, if there is a missing modality, the MRI brain tumor image with the missing modality is directly replaced with the full-modality replacement image obtained in the pre-training process, and the complete or filled multi-modal MRI brain tumor image is input into the trained network model, and passes through the encoder, the multi-head attention mechanism module, the adaptive feature fusion module, the decoder and the segmentation head in sequence, and finally outputs the segmentation results of different tumor areas.
5. The MRI brain tumor image segmentation method based on adaptive feature fusion according to claim 4, characterized in that: The multi-head attention mechanism module sets four heads, and step S2 includes: S201, annotating the first feature map with feature dimensions to obtain an annotated feature map, wherein the feature dimensions are [batch_size, num_channels, height, width, depth], where batch_size represents the data batch size of each input feature map, num_channels represents the number of channels of the feature map, and height, width, and depth represent the height, width, and depth of the feature map in three-dimensional space, respectively; S202, evenly dividing the annotated feature map into four subspaces along the channel dimension, each subspace corresponds to a head, and each head corresponds to a linear transformation layer; S203, perform linear transformation and dot product operation on the query vector, key vector, and value vector of each head through the linear transformation layer corresponding to each head to obtain the attention weight of each head. The calculation formula is: Among them, Attention(Q,K,V) represents the attention weight of each head, Q, K, and V represent the query vector, key vector, and value vector of each head after linear transformation, respectively. k is the dimension of the key vector; S204, performing attention weighting on the feature map input to the head according to the attention weight of each head, to obtain the feature map after attention weighting of each head; S205. The feature maps after the four heads' attention weighting are spliced along the channel dimension, and the spliced feature maps are linearly transformed to restore the spliced feature maps to the same number of channels as the first feature map, so as to obtain a second feature map after the attention weighted combination.
6. The adaptive feature fusion MRI brain tumor image segmentation method according to claim 5, characterized in that: Assuming that the second feature map comes from n modes, the feature dimension of the feature map of each mode in the second feature map is labeled as [batch size ,num_channels i ,height i ,width i ,depth i ], where i represents the i-th mode and batch size Indicates the data batch size of each input feature map, num_channels i Indicates the number of channels of the feature map of the i-th mode, height i 、width i 、depth i They represent the height, width and depth of the feature graph of the i-th mode in three-dimensional space, respectively, n is a positive integer, and i is a positive integer less than or equal to n; Step S3 includes: S301, fuse the feature map of each modality with the average feature map of all modalities along the channel dimension to obtain the fused feature map of each modality, wherein the feature dimension of the fused feature map of each modality is marked as [batch size ,num_channels i +num_channels avg ,height i ,width i ,depth i ], num_channels avg represents the number of channels of the average feature map, where the average feature map is obtained by averaging the feature maps of all modalities in the channel dimension; S302: The feature map after fusion of each modality is passed through the convolution layer corresponding to the modality to generate the initial attention weight map of each modality. The calculation formula is: in, represents the average feature map, W i represents the initial attention weight map of the i-th modality, F i represents the convolutional layer corresponding to the i-th modality, θ i represents the parameters of the convolutional layer corresponding to the i-th mode, σ represents the Sigmoid function, which is used to map the generated weight value to the (0,1) interval; S303, normalize the initial attention weight map of each modality through the Softmax function to obtain the normalized attention weight map of each modality, and the calculation formula is: in, represents the normalized attention weight map of the i-th modality, N avail Indicates the number of currently available modalities, and the sum of the normalized attention weight maps of all modalities is 1; S304, multiply the feature map of each modality by the normalized attention weight map of the modality point by point in the voxel dimension, and then sum the feature maps after weighted multiplication of all modalities to obtain the third feature map after adaptive fusion. The calculation formula is: in, represents the third feature map after adaptive fusion, Represents a point-wise multiplication operation in the voxel dimension.
7. The MRI brain tumor image segmentation method based on adaptive feature fusion according to claim 6, characterized in that: The encoder includes multiple convolutional layers and downsampling layers; the decoder adopts a structure of alternating multiple upsampling layers and multiple 3×3×3 convolutional layers, each upsampling layer is connected to a 3×3×3 convolutional layer, and the last layer of the decoder is a 1×1×1 convolutional layer. Step S4 includes: S401, using a skip connection mechanism, concatenate and fuse the feature maps of the third feature map after being amplified by each upsampling layer with the feature maps of the same resolution stage after being processed by the downsampling layer along the channel dimension, input the concatenated and fused feature maps into the 3×3×3 convolutional layer corresponding to the current upsampling layer for processing, restore the third feature map to a size close to the original image resolution, and obtain a feature map of a size close to the original image resolution; S402, converting the number of channels of the feature map having a size close to the original image resolution into the number of segmentation categories 3 through the 1×1×1 convolution layer, and then outputting the segmentation results of different tumor regions through an activation function.
8. An adaptive feature fusion MRI brain tumor image segmentation system, used to implement the adaptive feature fusion MRI brain tumor image segmentation method according to any one of claims 1 to 7, characterized in that: The system uses a network model to implement MRI brain tumor image segmentation, and the network model includes: a basic network based on 3D U-Net, a multi-head attention mechanism module and an adaptive feature fusion module, and the basic network based on 3D U-Net includes an encoder and a decoder; The encoder is used to extract features from the MRI brain tumor images of different modalities to be segmented, and input the first feature maps of different tumor areas in the extracted MRI brain tumor images of different modalities into the multi-head attention mechanism module; The multi-head attention mechanism module is used to perform attention weighted combination on the first feature map, and input the second feature map after the attention weighted combination into the adaptive feature fusion module; The adaptive feature fusion module is used to adaptively weight and fuse the second feature map according to the contribution of the second feature map to brain tumor segmentation, and input the adaptively fused third feature map into the decoder; The decoder is used to restore the original image resolution of the third feature map to obtain segmentation results of different tumor areas.
9. The adaptive feature fusion MRI brain tumor image segmentation system according to claim 8, characterized in that: The system further includes: a training module, which is used to train the network model before using the network model. The training process specifically includes: Collecting data sets of MRI brain tumor images of different modalities for preprocessing to obtain preprocessed data sets; Based on the preprocessed data set, the network model is pre-trained by self-supervised learning to obtain the pre-trained network model. The pre-training process specifically includes: Randomly select some modal MRI brain tumor images from the preprocessed data set for masking to simulate the missing of different modalities, and also mask some MRI brain tumor images of the remaining modalities, reconstruct the masked MRI brain tumor images through the encoder, the decoder, the multi-head attention mechanism module and the adaptive feature fusion module, calculate the difference between the reconstructed MRI brain tumor image and the original MRI brain tumor image by taking the mean square error as the loss function, and update the parameters of the network model by back propagation to obtain the pre-trained network model; Based on the preprocessed data set, the pre-trained network model is further trained in a supervised manner to obtain the trained network model. The further training process specifically includes: Each time a sample is selected from the preprocessed data set, different modality missing conditions are randomly generated as network inputs to obtain segmentation results under different modality missing conditions. Based on the segmentation results under different modality missing conditions, the segmentation loss of the network model is calculated through a combined loss function. Based on the segmentation loss and the consistency loss between different modality missing conditions, the parameters of the network model are further fine-tuned through back propagation. When the network model can adapt to segmentation tasks under different modality missing conditions, the training is stopped, the final parameters of the network model are saved, and the trained network model is obtained.
10. The adaptive feature fusion MRI brain tumor image segmentation system according to claim 9, characterized in that: The combined loss function is a weighted sum of the cross entropy loss and the Dice loss; The cross entropy loss is used to measure the probability difference between the predicted segmentation result and the actual segmentation label at each pixel, and the calculation formula is: Among them, L CE represents the cross entropy loss, N represents the total number of pixels, C represents the number of segmentation categories, i represents the i-th pixel, j represents the j-th segmentation category, i and j are positive integers, and y ij Indicates the true value of the i-th pixel belonging to the j-th category in the true segmentation label, p ij Indicates the probability value that the i-th pixel in the predicted segmentation result belongs to the j-th category; The Dice loss is used to measure the overlap between the predicted segmentation result and the actual segmentation label, and the calculation formula is: Among them, L Dice represents the Dice loss, p i Indicates the probability value of the i-th pixel in the predicted segmentation result belonging to the target category, y i Indicates the probability value that the i-th pixel in the true segmentation label belongs to the target category, ∈ is a positive number to avoid the situation where the denominator is 0; The combined loss function L total is the cross entropy loss L CE With the Dice loss L Dice The weighted sum of total =αL CE +(1-α)L Dice , where α represents the trade-off between the cross entropy loss L CE With the Dice loss L Dice Importance weight parameter.
Citation Information
Cited By
Brain tumor detection method based on attention mechanism and MRI (Magnetic Resonance Imaging) multi-modal fusion
CN121353272A
A brain tumor detection method based on attention mechanism and MRI multi-modal fusion
CN121353272B