Enhanced Attention-based Encoder-Decoder Network Glioma Fusion Segmentation System and Method
By introducing an enhancing codec network to glioma image segmentation, the problems of large storage space, long training time and poor depth information segmentation of 2D models in the prior art are solved, and more efficient and accurate glioma segmentation is achieved.
Patent Information
- Application Number
- CN202211214257.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-09-30
AI Technical Summary
The existing 3D models of glioma image segmentation have large storage space and long training time, and the 2D model depth information segmentation effect is not good.
A glioma fusion segmentation system based on enhancing attention is proposed. By extracting 2D image data from the image data in the three directions of axis, crown and vector, a corresponding enhancing network sub-model is established, and an enhancing attention module is introduced to compensate for the receptive field limitation of convolutional neural networks.
Effectively utilize global context structure information, improve the efficiency and accuracy of glioma segmentation, reduce network parameters, save computing resources, and improve training and parameter debugging efficiency.
Smart Images

Figure CN115661165B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for fused segmentation of glioma MR images, and particularly to a glioma fused segmentation system and method based on an encoder-decoder network with enhanced attention. Background Art
[0002] Glioma is the most common type of brain tumor, with sub-regions of varying degrees of invasiveness and high heterogeneity, including infiltrative edema tissue (ED) and tumor core regions, including necrotic core (NCR), non-enhancing tumor (NET), and enhancing tumor (ET). Early detection and accurate diagnosis are the keys to the treatment of brain glioma cancer. Magnetic resonance imaging (MRI) and computed tomography (CT) for pathological organs are the most commonly used diagnostic tools. Clinically, doctors and professionals check the tumor condition by visualizing 3D data layer by layer, resulting in a doubling of patient image data. Manually tracking and annotating tumor regions is quite time-consuming. At the same time, due to differences in professional levels and fatigue, the manual segmentation results of the same sample can vary greatly. Therefore, accurate and reproducible automatic detection and segmentation of glioma will help diagnose and segment tumors in a timely manner in a clinical environment, liberate manpower and improve efficiency, and can be applied to large-scale pathological data analysis for statistical analysis and learning of brain glioma, and construct a pathological data feature database.
[0003] However, the inherent histological variations of gliomas are further complicated by the heterogeneity of MRI scan tumor features (such as intensity). The different sub-regions of gliomas vary greatly in appearance and shape according to their biological conditions, which makes glioma segmentation very challenging, and there are often high annotation differences even in the segmentation of different datasets by expert doctors. In the past few decades, image segmentation technology has changed greatly; from traditional image processing techniques (such as watershed segmentation, k-means clustering) to semi-automatic machine learning methods (including extracting features from ROIs to train classifiers), and then to fully automatic data-driven deep learning methods (such as popular CNNs). Recently, using deep learning technology for brain tumor segmentation has shown good results because they can learn complex features from data. Deep learning models based on convolutional neural networks (CNNs) and fully convolutional networks (FCNs), such as SegNet, deep neural networks, U-Net, QuickNAT, DenseNet, and their variants, have shown remarkable performance in segmentation tasks.
[0004] The variant U-Net of the convolutional neural network (CNN) has become a powerful tool for medical image segmentation. It consists of a symmetric encoder-decoder network with skip connections to enhance the retention of details. However, due to the limitations caused by its own mechanism, it has not been effectively solved. For example, due to the characteristics of weight sharing in convolutional operations, it has problems such as limited receptive fields and inability to capture the feature dependencies of distant pixels. 3D U-Net utilizes the volume characteristics of brain MR data, but at the cost of large storage space and longer training time. At the same time, it also faces challenges such as class imbalance and higher computational costs. The imbalance problem can be solved by various loss functions, but due to the high computational costs of these models, model learning and parameter overshooting remain the main challenges. The 2D U-Net has fewer parameters and can be trained quickly, so parameter over-tuning can be performed. The 2D U-Net has poor segmentation effects due to the lack of depth information. However, studies have confirmed that under the premise of efficiently performing hyperparameter tuning, 2D models can achieve similar or more advanced performance to 3D models.
[0005] Although the segmentation method based on convolutional neural network has remarkable effects, in complex tasks such as glioma segmentation, due to problems such as local visual blurring and low contrast between tumor sub-regions, the inherent characteristics of convolutional neural network limit its performance in global structure analysis. Summary of the Invention
[0006] The purpose of the present invention is to solve the technical problems of large storage space and long training time of the existing 3D models for glioma image segmentation, and poor depth information segmentation effects of 2D models, and to propose a glioma fusion segmentation system and method based on an encoder-decoder network with enhanced attention.
[0007] The general idea of the present invention is as follows: 2D image data is extracted from the image data in the axial, coronal, and sagittal directions, and an axial direction encoder-decoder network sub-model, a coronal direction encoder-decoder network sub-model, a sagittal direction encoder-decoder network sub-model, and a binary classification segmentation auxiliary network model are established. In each model, the problem of limited receptive fields and inability to capture the feature dependencies of distant pixels in the original convolutional neural network is compensated by an enhanced attention module, better utilizing the global context structure information, and improving the efficiency and accuracy of glioma segmentation.
[0008] The technical solution provided by the present invention is as follows:
[0009] A glioma fusion segmentation system based on an encoder-decoder network with enhanced attention, characterized in that:
[0010] It includes a four-class encoder-decoder network model and a binary classification segmentation auxiliary network model;
[0011] The four-class encoding and decoding network model is used for feature extraction and region segmentation based on four sub-regions: the core necrosis region, the edema region, the enhanced tumor region, and the normal region; the four-class encoding and decoding network model includes three encoding and decoding network sub-models with the same structure, and the encoding and decoding network sub-model includes an axial direction encoding and decoding network sub-model, a coronal direction encoding and decoding network sub-model, and a sagittal direction encoding and decoding network sub-model;
[0012] The structure of the two-class segmentation auxiliary network model is the same as that of the encoding and decoding network sub-model, and the two-class segmentation auxiliary network model is used for feature extraction and region segmentation based on two sub-regions: the tumor core region and the normal region in the axial direction;
[0013] The encoding and decoding network sub-model includes an encoding module, an enhanced attention module, and a decoding module; the encoding module is used to extract high-level semantic representations of the input image and perform downsampling at the same time to generate the bottommost feature map and transmit it to the enhanced attention module;
[0014] The enhanced attention module includes a feature encoding unit, a position encoding unit, 4 to 8 feature enhancement modules with the same structure connected in series in sequence, and a double-layer convolution output module; the feature encoding unit and the position encoding unit respectively receive the bottommost feature map and are used to add feature encoding and position encoding to each feature vector in the bottommost feature map to form encoded data information; the feature enhancement module includes a multi-channel self-attention unit, a first splicing unit, a fully connected feed-forward network unit, and a second splicing unit;
[0015] The multi-channel self-attention unit receives the encoded data information and is used to add attention scores to each feature vector in the encoded data information to form feature aggregation data information; the first splicing unit is used to receive the encoded data information and the feature aggregation data information and perform splicing to form the first spliced data information;
[0016] The fully connected feed-forward network unit receives the first spliced data information and is used to perform feature transformation on the feature vector to form feature transformation data information; the second splicing unit is used to receive the first spliced data information and the feature transformation data information and perform splicing to form the second spliced data information; the double-layer convolution output module is connected to the last feature enhancement module and is used to receive the information output by the last feature enhancement module, process it to form global feature encoding information, and output the global feature encoding information to the decoding module; the decoding module is used to perform upsampling on the global feature encoding information to restore it to the same spatial resolution as the input data.
[0017] Further, the multi-channel self-attention unit includes an attention normalization layer, three parallel linear layers, and a multi-channel calculation unit;
[0018] The attention normalization layer is used to accelerate the convergence of the model. The three parallel linear layers all receive the data output by the attention normalization layer to form a query matrix, a key matrix, and a value matrix respectively. The parameters in the query matrix, key matrix, and value matrix are split and then passed to the multi-channel calculation unit to calculate the attention scores corresponding to each channel, which are added to the feature vector.
[0019] Further, the parameters of the query matrix, key matrix, and value matrix are split and then passed to the multi-channel calculation unit to calculate the self-attention scores specifically as follows:
[0020] The parameters of the query matrix, key matrix, and value matrix are split 8 - 12 times. The parameters after each split are passed to the multi-channel calculation unit to calculate the attention scores corresponding to each channel, generating 8 - 12 groups of attention scores. The 8 - 12 groups of self-attention scores are fused to obtain the final attention scores, which are added to the feature vector.
[0021] Further, the encoding module sequentially includes an input convolution module, a first downsampling module, a second downsampling module, a third downsampling module, and a fourth downsampling module; the input convolution module includes two input convolution layers, and a batch normalization and an activation function are set after each input convolution layer. The input convolution module is used to receive the input tumor image information, and the output data of the fourth downsampling module is output to the enhanced attention module; the first downsampling module, the second downsampling module, the third downsampling module, and the fourth downsampling module are sequentially connected, and the input channels and output channels of the sampling increase exponentially;
[0022] The decoding module sequentially includes a fourth upsampling module, a third upsampling module, a second upsampling module, a first upsampling module, and an output segmentation module; the fourth upsampling module is used to receive the output data of the enhanced attention module, and the output segmentation module is used to output multi-modal data; the fourth upsampling module, the third upsampling module, the second upsampling module, and the first upsampling module are sequentially connected, and the input channels and output channels of the sampling decrease exponentially in the same way as each downsampling module.
[0023] Further, the input convolution module of the encoding module outputs a first feature map, the first downsampling module outputs a second feature map, the second downsampling module outputs a third feature map, the third downsampling module outputs a fourth feature map, and the fourth downsampling module outputs the bottommost feature map;
[0024] Each upsampling module of the decoding module includes an upsampling layer, a splicing unit and a double-layer convolution unit; the splicing unit of the fourth upsampling module receives the output information of the corresponding upsampling layer and the fourth feature map, splices them, and outputs them to the corresponding double-layer convolution unit; the splicing unit of the third upsampling module receives the output information of the corresponding upsampling layer and the third feature map, splices them, and outputs them to the corresponding double-layer convolution unit; the splicing unit of the second upsampling module receives the output information of the corresponding upsampling layer and the second feature map, splices them, and outputs them to the corresponding double-layer convolution unit; the splicing unit of the first upsampling module receives the output information of the corresponding upsampling layer and the first feature map, splices them, and outputs them to the corresponding double-layer convolution unit.
[0025] The present invention also provides a method for glioma fusion segmentation based on codec network with enhanced attention, which is special in that the method adopts the above-mentioned codec network glioma fusion segmentation system based on enhanced attention, and comprises the following steps:
[0026] S1. Preprocess the standard tumor MR image set, extract the 2D image data of tumor tissue in the axial, coronal and sagittal directions, and remove the 2D slice data of non-tumor tissue, respectively, and construct the axial direction training data set, coronal direction training data set, and sagittal direction training data set, and construct the axial direction verification data set, coronal direction verification data set, and sagittal direction verification data set;
[0027] S2, constructing an initial codec network sub-model: constructing three initial codec network sub-models with the same network structure, wherein the initial codec network sub-models include an axial direction initial codec network sub-model, a crown direction initial codec network sub-model and a sagittal direction initial codec network sub-model;
[0028] Construct an initial binary classification segmentation auxiliary network model;
[0029] S3. Train the initial encoding and decoding network sub-model and the initial binary classification segmentation auxiliary network model and debug the network hyperparameters.
[0030] Input the axial direction training data set in step S1 into the axial direction initial encoding and decoding network sub-model, input the crown direction training data set into the crown direction initial encoding and decoding network sub-model, input the sagittal direction training data set into the sagittal direction initial encoding and decoding network sub-model, respectively train and debug the network hyperparameters to obtain the axial direction encoding and decoding network sub-model, the crown direction encoding and decoding network sub-model and the sagittal direction encoding and decoding network sub-model;
[0031] Input the training data set in the axial direction in step S1 into the initial binary classification segmentation auxiliary network model, perform training and debug the network hyperparameters to obtain the binary classification segmentation auxiliary network model;
[0032] S4. Verify the encoding and decoding network sub-model and the binary classification segmentation auxiliary network model
[0033] Input the axial direction verification dataset, coronal direction verification dataset, and sagittal direction verification dataset in step S1 into the encoding and decoding network sub-models in the corresponding directions to obtain sub-region segmentation probability masks in three directions. Perform "rounded averaging" on the probability masks corresponding to each pixel position to obtain the network segmentation result that fuses the three directions. Input the axial direction verification dataset into the binary classification segmentation auxiliary network model to obtain the binary classification segmentation result of the tumor core;
[0034] Calculate the DICE score α1 between the network segmentation result that fuses the three directions and the ground truth label mask. If the DICE score α1 remains unchanged or the increase is less than 1% of the current DICE score α1, then obtain the final four-class encoding and decoding network model. If the DICE score α1 increases and the increase is greater than or equal to 1% of the current DICE score α1, then return to step S3 to continue training each sub-model of the four-class encoding and decoding network model until the final four-class encoding and decoding network model is obtained;
[0035] Calculate the DICE score α2 between the binary classification segmentation result of the tumor core and the ground truth label mask. If the DICE score α2 remains unchanged or the increase is less than 1% of the current DICE score α2, then obtain the final binary classification segmentation auxiliary network model. If the DICE score α2 increases and the increase is greater than or equal to 1% of the current DICE score α2, then return to step S3 to continue training the binary classification segmentation auxiliary network model until the final binary classification segmentation auxiliary network model is obtained;
[0036] S5. Extract the information of each image in the axial, coronal, and sagittal directions from the MR image set to be measured to form three datasets to be measured. Input the three datasets to be measured into the final fusion segmentation encoding and decoding network model in S4 to obtain the network segmentation result that fuses the three directions of the MR image set to be measured. Input the axial direction dataset to be measured into the final binary classification segmentation auxiliary network model in S4 to obtain the binary classification segmentation result of the MR image set to be measured. Correct the network segmentation result that fuses the three directions of the MR image set to be measured according to the binary classification segmentation result of the MR image set to be measured to obtain the final fusion segmentation result of the MR image set to be measured.
[0037] Furthermore, in step S3, the specific process of training the initial encoding and decoding network sub-model is as follows:
[0038] Randomly shuffle the training dataset in the corresponding direction and input it into the initial encoder-decoder network sub-model to obtain a preliminary predicted segmentation map, which includes the core necrosis area, the edema area, the enhanced tumor area, and the normal area. Calculate the cross-entropy loss and DICE score between the preliminary predicted segmentation map and the ground truth label mask. Combine the cross-entropy loss and DICE score between the preliminary predicted segmentation map and the ground truth label mask as the loss function of the initial encoder-decoder network sub-model for backpropagation, and iteratively update the weights of the initial encoder-decoder network sub-model. After reaching the convergence condition of the initial encoder-decoder network sub-model, save the network parameters, and finally obtain the encoder-decoder network sub-model.
[0039] In step S3, the specific process of training the initial binary classification segmentation auxiliary network model is as follows:
[0040] In the axial training dataset, merge the label masks of the two sub-regions of the enhanced tumor area and the core necrosis area into a tumor core label mask, and merge the label masks of the two sub-regions of the normal area and the edema area into a background label mask. The tumor core label mask and the background label mask are collectively referred to as the binary classification label mask. Randomly shuffle the axial training dataset and input it into the initial binary classification segmentation auxiliary network model to obtain a preliminary tumor core segmentation prediction map. Calculate the cross-entropy loss and DICE score between the preliminary tumor core segmentation prediction map and the binary classification label mask. Combine the cross-entropy loss and DICE score between the preliminary tumor core segmentation prediction map and the binary classification label mask as the loss function of the initial binary classification segmentation auxiliary network for backpropagation, and iteratively update the weights of the initial binary classification segmentation auxiliary network. After reaching the convergence condition of the initial binary classification segmentation auxiliary network, save the network parameters, and finally obtain the binary classification segmentation auxiliary network.
[0041] Furthermore, in step S1, the specific preprocessing of the standard tumor MR image set is as follows:
[0042] S1.1: Take the center of each 3D image in the standard tumor MR image set as the center of the cropped 3D image, and crop the original 3D image and the corresponding label mask with a preset window size to remove the redundant blank boundaries that are not related to the actual imaging data, reduce the data size, and improve the model inference efficiency;
[0043] S1.2: Perform Gaussian normalization on the intensity of the cropped image in step S1.1 to increase the model convergence speed;
[0044] S1.3: Simultaneously extract four-modal 2D slices from the image data processed in step S1.2 in the three directions of the axis, coronal, and sagittal. The four modalities include fluid-attenuated inversion recovery sequence, T1-weighted, post-contrast T1-weighted, and T2-weighted. Use the four-modal data as the four channels of the new 2D image data to generate a 2D tumor MR multi-modal dataset with four channels in three directions. Randomly select 75% of the dataset as the training dataset, and the remaining 25% as the validation dataset.
[0045] Further, in step S3, when training the initial encoder-decoder network sub-model, the axial direction training data set, the coronal direction training data set, and the sagittal direction training data set are randomly shuffled, and the image data set is divided into multiple groups according to a preset quantity, and each group is respectively input into the corresponding initial encoder-decoder network sub-model for training;
[0046] When training the initial binary classification segmentation auxiliary network model, the training data set in the axial direction is randomly shuffled, and the image data set is divided into multiple groups according to a preset quantity, and each group is respectively input into the binary classification segmentation auxiliary network model for training;
[0047] In step S4, when validating the encoder-decoder network sub-model, the axial direction validation data set, the coronal direction validation data set, and the sagittal direction validation data set are randomly shuffled, and each time the data set of a single image is input into the corresponding initial encoder-decoder network sub-model for validation;
[0048] When validating the binary classification segmentation auxiliary network model, the validation data set in the axial direction is randomly shuffled, and each time the data set of a single image is input into the initial binary classification segmentation auxiliary network model for validation.
[0049] Advantages of the present invention:
[0050] 1. The present invention proposes a four-class encoder-decoder network model designed with a U-shaped network architecture, and an enhanced attention mechanism with long-distance global structural information and interaction relationships between features is embedded at the bottom layer of the network, which can make up for the characteristics such as weight sharing of convolution operations in the original U-shaped fully convolutional neural network method, and solves the problems of limited receptive field and inability to capture the feature dependence of distant pixels existing in the existing U-shaped fully convolutional neural network.
[0051] 2. In order to improve the performance of brain glioma MR image segmentation and reduce the complexity of the network structure, the present invention proposes a network model structure combining a four-class encoder-decoder network model and a binary classification segmentation auxiliary network model, which greatly reduces the network parameters and saves computing resources, effectively improving the training and parameter debugging efficiency of the network; this network model uses a more lightweight U-shaped as the basic network architecture, and then introduces multi-scale and enhanced attention mechanisms, which can effectively solve the problems of complex three-dimensional segmentation network models and high computational costs.
[0052] 3. Based on the fact that the segmentation accuracy of the tumor core is generally higher than that of the four sub-regions, the present invention proposes to post-process the results processed by the four-class encoder-decoder network model by combining the binary classification segmentation auxiliary network model, and use the reliable tumor core segmentation results to correct the segmentation results of the four-class encoder-decoder network model, improving the segmentation accuracy of the tumor boundary.
[0053] 4. Reconstruct the original three-dimensional dataset into two-dimensional datasets in the axial, coronal, and sagittal directions, doubling the data for network training. Therefore, when reconstructing the dataset, the present invention preprocesses the data, uses the brain MR image center detection method to crop the redundant boundary data during imaging, reduces the brain image window, and during the network training process, to overcome class imbalance and enable the model to converge quickly, slices of data that do not contain diseased tissues are removed.
[0054] 5. The encoding module of the method provided by the present invention extracts high-level semantic representations by using cascaded convolutional layers, and at the same time downsamples the input image to generate a compact feature map, effectively capturing local context information. The decoding module uses skip connections to reuse the high-resolution feature maps of the encoder to recover the lost spatial information from the high-level semantic representations; the self-attention module uses the global interaction between the semantic features at the end of the encoder to explicitly model the complete context information, flattens the feature map output by the encoding module in vector form, and the attention module linearly projects these vector feature sequences into the feature encoding module, while adding position encoding to characterize position information, adopts a multi-channel attention module to maintain multiple complex relationships between different positions in the sequence, and uses 4 - 8 self-attention modules, each module containing its corresponding learnable weight matrix, splices and projects the output of the multi-head attention module to obtain the final weight matrix, while making the network model lightweight and improving the operation speed on the premise of ensuring accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 Schematic diagram of the structures of the encoding module and the decoding module of the axial direction encoding and decoding network sub-model in the glioma fusion segmentation system of the encoding and decoding network based on enhanced attention according to the embodiment of the present invention;
[0056] Figure 2 Schematic diagram of the enhanced attention module structure of the axial direction encoding and decoding network sub-model in the glioma fusion segmentation system of the encoding and decoding network based on enhanced attention according to the embodiment of the present invention;
[0057] Figure 3 Schematic diagram of the feature enhancement module structure in the embodiment of the present invention;
[0058] Figure 4 Schematic diagram for comparing the segmentation results with the ground truth labels using the segmentation method of the embodiment of the present invention, where (a), (b), and (c) are schematic diagrams of the ground truth label results in the axial, coronal, and sagittal directions respectively; (d), (e), and (f) are schematic diagrams of the segmentation results in the axial, coronal, and sagittal directions respectively using the segmentation method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] In this embodiment, the axial, coronal, and sagittal directions are the direction definitions in human anatomy. Specifically, the coronal direction is the extension direction of the axis that is parallel to the ground in the left-right direction and perpendicular to the sagittal plane; the sagittal direction is the extension direction of the axis that is parallel to the ground in the front-back direction and perpendicular to the coronal plane; the axial direction is the axis that is perpendicular to the ground in the up-down direction.
[0060] See Figure 1 , this embodiment provides a glioma fusion segmentation system based on an enhanced attention encoder-decoder network. The system includes a four-class encoder-decoder network model and a two-class segmentation auxiliary network model.
[0061] The four-class encoder-decoder network model is used to extract features and perform region segmentation based on four sub-regions: the core necrosis region, the edema region, the enhanced tumor region, and the normal region. The four-class encoder-decoder network model includes three encoder-decoder network sub-models with the same structure. The encoder-decoder network sub-model includes an axial direction encoder-decoder network sub-model, a coronal direction encoder-decoder network sub-model, and a sagittal direction encoder-decoder network sub-model. The structure of the two-class segmentation auxiliary network model is the same as that of the encoder-decoder network sub-model. The two-class segmentation auxiliary network model is used to extract features and perform region segmentation based on two sub-regions: the tumor core region and the normal region in the axial direction.
[0062] Each encoder-decoder network sub-model includes an encoding module, an enhanced attention module, and a decoding module.
[0063] The encoding module is used to extract high-level semantic representations of the input image and perform downsampling at the same time to generate the bottommost feature map and transmit it to the enhanced attention module. Specifically, the encoding module sequentially includes an input convolution module, a first downsampling module, a second downsampling module, a third downsampling module, and a fourth downsampling module.
[0064] The first downsampling module starts to perform downsampling of the feature map, including a max-pooling downsampling layer, two two-dimensional convolutions, batch normalization, and an activation function. The second downsampling module, the third downsampling module, and the fourth downsampling module adopt the same network structure as the first downsampling module, but the feature channel parameters increase exponentially. The encoding module can extract deep features with a large receptive field through a series of convolution operations and continuous downsampling layers.
[0065] The input convolution module includes two input convolution layers, and batch normalization and an activation function are set after each input convolution layer. The input convolution module is used to receive the input tumor image information and output the first feature map. The first downsampling module outputs the second feature map, the second downsampling module outputs the third feature map, the third downsampling module outputs the fourth feature map, and the fourth downsampling module outputs the bottommost feature map and outputs it to the enhanced attention module.
[0066] The enhanced attention module includes a feature encoding unit, a position encoding unit, 4 to 8 feature enhancement modules with the same structure connected in series in sequence, and a double-layer convolution output module; the feature encoding unit and the position encoding unit respectively receive the bottom-layer feature map, and are used to add feature encoding and position encoding to each feature vector in the bottom-layer feature map to form encoded data information.
[0067] The feature enhancement module includes a multi-channel self-attention unit, a first splicing unit, a fully-connected feed-forward network unit, and a second splicing unit; the multi-channel self-attention unit includes an attention normalization layer, three parallel linear layers, and a multi-channel calculation unit; the attention normalization layer receives the encoded data information and is used to accelerate the convergence of the model; the three parallel linear layers all receive the data output by the attention normalization layer to form a query matrix, a key matrix, and a value matrix respectively. The parameters in the query matrix, the key matrix, and the value matrix are split and then passed to the multi-channel calculation unit to calculate the attention scores corresponding to each channel, and are added to the feature vector to form feature aggregation data information.
[0068] Among them, the parameters in the query matrix, the key matrix, and the value matrix are split and then passed to the multi-channel calculation unit to calculate the attention scores corresponding to each channel specifically as follows: the parameters of the query matrix, the key matrix, and the value matrix are split 8 - 12 times, and the parameters after each split are passed to the multi-channel calculation unit to calculate the attention scores corresponding to each channel, generating 8 - 12 groups of attention scores. The 8 - 12 groups of self-attention scores are fused to obtain the final attention score, and are added to the feature vector to form feature aggregation data information.
[0069] The first splicing unit is used to receive the encoded data information and the feature aggregation data information and splice them to form the first spliced data information.
[0070] The fully-connected feed-forward network unit includes a fully-connected normalization layer and a fully-connected feed-forward layer; the fully-connected normalization layer is used to receive the first spliced data information, perform normalization processing, and then transmit it to the fully-connected feed-forward layer; the fully-connected feed-forward network unit is used to perform feature transformation on the feature vector to form feature transformation data information.
[0071] The second splicing unit is used to receive the first spliced data information and the feature transformation data information and splice them to form the second spliced data information; the double-layer convolution output module is connected to the last feature enhancement module, and is used to receive the information output by the last feature enhancement module, process it, and form global feature encoding information, and output the global feature encoding information to the decoding module;
[0072] The decoding module is used to upsample the global feature encoding information and restore it to the same spatial resolution as the input data. Specifically, the decoding module includes a fourth upsampling module, a third upsampling module, a second upsampling module, a first upsampling module and an output segmentation module in sequence; the fourth upsampling module is used to receive the output data of the enhanced attention module, and the output segmentation module is used to output multimodal data; the fourth upsampling module, the third upsampling module, the second upsampling module, and the first upsampling module are connected in sequence, and the sampled input channels and output channels are reduced exponentially with each downsampling module; in this embodiment, the decoding module and the encoding module adopt a codec network structure with four-level jump connections; each upsampling module of the decoding module includes an upsampling layer, a splicing unit and a dual convolution layer; the splicing unit of the fourth upsampling module receives the output information of the corresponding upsampling layer and the fourth feature map, splices them and outputs them to the corresponding double convolution layer; the splicing unit of the third upsampling module receives the output information of the corresponding upsampling layer and the third feature map, splices them and outputs them to the corresponding double convolution layer; the splicing unit of the second upsampling module receives the output information of the corresponding upsampling layer and the second feature map, splices them and outputs them to the corresponding double convolution layer; the splicing unit of the first upsampling module receives the output information of the corresponding upsampling layer and the first feature map, splices them and outputs them to the corresponding double convolution layer, and the decoding module reuses the high-resolution feature map of the encoding module by using skip connections to recover the lost spatial information from the high-level semantic representation extracted by the encoding module using cascaded convolution layers.
[0073] This embodiment provides the segmentation method of the above-mentioned codec network glioma fusion segmentation system based on enhanced attention, including the following steps:
[0074] S1. Preprocess the standard tumor MR image set, extract the 2D image data of tumor tissue in the axial, coronal and sagittal directions, and remove the 2D slice data of non-tumor tissue, respectively, and construct the axial direction training data set, coronal direction training data set, and sagittal direction training data set, and construct the axial direction verification data set, coronal direction verification data set, and sagittal direction verification data set;
[0075] In this embodiment, the standard tumor MR image set includes preoperative multimodal MRI scan files of brain tumors, with a total of 369 complete 3D brain scan case data containing manually annotated masks. Each 3D scan data consists of data in four modalities, namely fluid-attenuated inversion recovery sequence (FLAIR), T1-weighted (T1), post-contrast T1-weighted (T1_CE), and T2-weighted (T2), and the corresponding four-class ground truth label masks. The shape of each modality data and the ground truth label mask for a single case is uniformly registered to the size: 240*240*155 (voxels). The label data includes four types of voxel masks: lesion-free background - corresponding label 0, necrotic core of tumor (NCR) - corresponding label 1, infiltrated edematous tissue or peritumoral edema (ED) - corresponding label 2, and enhanced tumor (ET) - corresponding label 4.
[0076] The preprocessing of the standard tumor MR image set is specifically as follows:
[0077] S1.1: Taking the center of each 3D image in the standard tumor MR image set as the center of the cropped 3D image, cropping the original 3D image and the corresponding ground truth label mask with a preset window size to remove redundant blank boundaries unrelated to the actual imaging data, reducing the data size, and improving the model inference efficiency;
[0078] S1.2: Performing Gaussian normalization on the intensity of the cropped images in step S1.1 to increase the model convergence speed;
[0079] S1.3: Simultaneously extracting 2D slices in four modalities from the image data processed in step S1.2 in the three directions of axial, coronal, and sagittal. The four modalities include fluid-attenuated inversion recovery sequence, T1-weighted, post-contrast T1-weighted, and T2-weighted; using the four-modal data as the four channels of the new 2D image data to generate a 2D tumor MR multimodal dataset with four channels in three directions, and randomly extracting 75% of the dataset as the training dataset and the remaining 25% as the validation dataset.
[0080] S2: Constructing an initial encoder-decoder network submodel: Constructing three initial encoder-decoder network submodels with the same network structure. The initial encoder-decoder network submodel includes an axial-direction initial encoder-decoder network submodel, a coronal-direction initial encoder-decoder network submodel, and a sagittal-direction initial encoder-decoder network submodel; constructing an initial binary classification segmentation auxiliary network model.
[0081] See Figure 1 and Figure 2 , for the axial-direction initial encoder-decoder network submodel, adopting an encoder-decoder network structure with four-level skip connections and embedding an enhanced attention mechanism at the bottom layer of the network, Figure 1 The a end and b end in Figure 2 are respectively connected to the a end and b end in
[0082] ①. Encoding module
[0083] The encoding module consists of five modules: The first module is the input convolution module, which is composed of a double-layer convolution. The double-layer convolution consists of two two-dimensional convolution layers with the same structure, namely, a two-dimensional convolution with an input channel of 4, an output channel of 64, a stride of 1, a padding of 1, and a convolution kernel of 3*3, batch normalization, an activation function, and a two-dimensional convolution with an input channel and an output channel of 64, a stride of 1, a padding of 1, and a convolution kernel of 3*3, batch normalization, and an activation function; The second module is the first downsampling module, which is composed of a max pooling module and a double-layer convolution module. The max pooling module is a max pooling layer with a stride of 2. The double-layer convolution module consists of a two-dimensional convolution with an input channel of 64, an output channel of 128, a stride of 1, a padding of 1, and a convolution kernel of 3*3, batch normalization, an activation function, a two-dimensional convolution with an input channel and an output channel of 128, a stride of 1, a padding of 1, and a convolution kernel of 3*3, batch normalization, and an activation function; The third module is the second downsampling module, the fourth module is the third downsampling module, and the fifth module is the fourth downsampling module. These three modules have the same structure as the first downsampling module. The differences are as follows: In the second downsampling module, the input and output channel numbers of the two two-dimensional convolution layers of the double-layer convolution are 128, 256 and 256, 256 respectively; In the third downsampling module, the input and output channel numbers of the two two-dimensional convolution layers of the double-layer convolution are 256, 512 and 512, 512 respectively; In the fourth downsampling module, the input and output channel numbers of the two two-dimensional convolution layers of the double-layer convolution are 512, 1024 and 1024, 1024 respectively.
[0084] ②. Enhanced attention module
[0085] See Figure 2 , the enhanced attention module embedded at the bottom layer of the network includes a feature encoding module, a position encoding module, a feature enhancement module composed of 6 identical layers after fusing the encoding, and a double-layer convolution output module, as shown in Figure 3As shown, each feature enhancement module includes a multi-channel self-attention unit and a fully-connected feed-forward network unit. Each unit adopts skip connection and layer normalization operations. The multi-channel self-attention unit is used to perform feature aggregation, and the fully-connected feed-forward network unit is used to complete feature transformation. The enhanced attention module first takes the depth features extracted by the encoding module as the input feature sequence for embedded feature encoding and adds position encoding. The feature encoding and position encoding generate an encoded representation for each feature vector in the sequence, and then it is fed into three linear layers with the same structure, which are used to generate the query matrix, key matrix, and value matrix respectively to calculate the self-attention scores. When passing through all the feature enhancement modules, the attention scores of each feature vector will be added to it. To represent the complex relationships and subtle differences between feature vectors, the above calculations are repeated in parallel using multiple channels, that is, the parameters in the query matrix, key matrix, and value matrix are split 12 times and passed through their respective channels, and then all the attention scores are fused to obtain the final attention scores. At this time, the output of the multi-channel self-attention unit also needs to pass through the fully-connected feed-forward network unit. Similar to the multi-channel self-attention unit, the first layer is a normalization layer, the second layer is a fully-connected feed-forward module, followed by a concatenation operation. The fully-connected feed-forward module consists of a linear fully-connected layer with input and output channels of 768 and 1024, an activation function, dropout, a linear fully-connected layer with input and output channels of 1024 and 768, and dropout.
[0086] ③. Decoding module
[0087] The decoding layer is composed of five modules, as Figure 1Shown as follows: The bottom layer is the fourth upsampling module, which consists of an upsampling layer, a splicing unit, and a double-layer convolutional unit. The upsampling layer is composed of a transposed convolution with an input channel of 1024, an output channel of 512, a stride of 2, a padding of 0, and a convolution kernel of 2*2. The splicing unit fuses the output of the transposed convolution with the output of the same size from the corresponding encoding module. The structure of the double-layer convolutional unit is, in sequence, a two-dimensional convolution with an input channel of 1024, an output channel of 512, a stride of 1, a padding of 1, and a convolution kernel of 3*3, batch normalization, an activation function, a two-dimensional convolution with an input channel and an output channel of 512, a stride of 1, a padding of 1, and a convolution kernel of 3*3, batch normalization, and an activation function. The second module, the third module, and the fourth module are the third upsampling module, the third upsampling module, and the first upsampling module respectively. These three modules have the same structure as the fourth upsampling module. The differences are as follows: In the third upsampling module, the input and output channel numbers of the two two-dimensional convolutional layers in the double-layer convolution are 512, 256 and 256, 256 respectively; In the third upsampling module, the input and output channel numbers of the two two-dimensional convolutional layers in the double-layer convolution are 256, 128 and 128, 128 respectively; In the first upsampling module, the input and output channel numbers of the two two-dimensional convolutional layers in the double-layer convolution are 128, 64 and 64, 64 respectively. The top layer is the output segmentation module, which is composed of a single-layer two-dimensional convolution with an input channel of 64, an output channel of 4, a stride of 1, a padding of 1, and a convolution kernel of 3*3.
[0088] The initial binary classification segmentation auxiliary network model also adopts an encoder-decoder network with four-level skip connections embedded with an enhanced attention module. The difference is that this auxiliary network is a binary classification network model only for the segmentation of the glioma core area and the background. The output segmentation module at the top layer of the decoding module is composed of a single-layer two-dimensional convolution with an input channel of 64, an output channel of 2, a stride of 1, a padding of 1, and a convolution kernel of 3*3.
[0089] S3. Train the initial encoder-decoder network sub-model and the initial binary classification segmentation auxiliary network model and debug the network hyperparameters.
[0090] Input the axial training data set in step S1 into the axial initial encoder-decoder network sub-model, input the coronal training data set into the coronal initial encoder-decoder network sub-model, and input the sagittal training data set into the sagittal initial encoder-decoder network sub-model. Train them respectively and debug the network hyperparameters to obtain the axial encoder-decoder network sub-model, the coronal encoder-decoder network sub-model, and the sagittal encoder-decoder network sub-model.
[0091] Taking the training and debugging of the network hyperparameters of the axial initial encoder-decoder network sub-model as an example:
[0092] After randomly shuffling the axial direction training data set, the image data set is divided into multiple groups according to a preset quantity. In this embodiment, 24 images are used as a group. Each group is respectively input into the axial direction initial encoding and decoding network sub-model to obtain a preliminary predicted segmentation map. The preliminary predicted segmentation map includes a core necrosis area, an edema area, an enhanced tumor area, and a normal area. Calculate the cross-entropy loss and DICE score between the preliminary predicted segmentation map and the ground truth label mask. Combine the cross-entropy loss and DICE score between the preliminary predicted segmentation map and the ground truth label mask as the loss function of the axial direction initial encoding and decoding network sub-model for backpropagation, and iteratively update the weights of the axial direction initial encoding and decoding network sub-model. After reaching the convergence condition of the axial direction initial encoding and decoding network sub-model, save the network parameters, and finally obtain the axial direction encoding and decoding network sub-model.
[0093] Specifically, the axial training dataset is randomly shuffled, and a portion of the data is read in batches of size 24. The shape of each case of four-modal MR data is adjusted to 4 * 192 * 192, and the shape of the corresponding ground truth label mask is adjusted to 1 * 192 * 192. Subsequently, the input data and labels are randomly rotated by 90 degrees, 180 degrees, and 270 degrees, and then random Gaussian noise and a random intensity value transformation in the range (0.75, 1.25) are added to the input data. The label category 4 is replaced with 3. Then it is fed into the input convolution module of the axial initial encoder-decoder network sub-model with an enhanced attention module, and a feature map with a shape of 64 * 192 * 192 is obtained as the initial input of the encoding layer. At the same time, the first feature map at this time is saved. After passing through the first downsampling module, the size of the feature map is reduced by half compared to the input image, and the feature dimension is doubled, and the shape becomes 128 * 192 * 192. At the same time, the second feature map output by the module is saved. It passes through the second downsampling module, the third downsampling module, and the fourth downsampling module in sequence. The size of the feature map is reduced, and the depth is increased. The third feature map, the fourth feature map, and the bottommost feature map output by the module are saved. The input data undergoes four-level downsampling feature extraction and finally obtains a depth feature map with a shape of 1024 * 12 * 12. The depth feature map (the bottommost feature map) is fed into the enhanced attention module at the bottom layer of the network. The enhanced attention module encodes the interaction information between multiple feature vectors in the sequence by integrating the global information from the complete input sequence, that is, it is achieved by defining three learnable weight matrices. First, a two-layer convolution is used to encode this feature map to obtain a sequence of feature vectors with a size of 768 * 144 and add a learnable position encoding matrix to generate an encoded representation for each feature vector. Then it is fed into three linearly connected layers with the same structure, which are used to generate the query matrix, the key matrix, and the value matrix respectively. Calculate the dot product of the query and all keys, and then use the softmax operator to normalize the dot product to obtain the attention scores. Each vector becomes the weighted sum of all vectors in the sequence, where the weights are given by the attention scores; when passing through 6 enhanced attention modules, each self-attention unit will add its own attention score to each feature vector. To represent the complex relationships and subtle differences between feature vectors, the above calculations are repeated in parallel using multiple channels, that is, the query, key, and value parameters are split 12 times and passed through their respective channels, and then all the attention scores are fused to obtain the final attention scores. The output of the self-attention unit at this time also has to pass through a fully connected feed-forward network module. Similar to the multi-channel self-attention unit, the first layer is a normalization layer, the second layer is a fully connected feed-forward module, followed by a concatenation operation. The fully connected feed-forward module consists of a linear fully connected layer with input and output channels of 768 and 1024, an activation function, random dropout, a linear fully connected layer with input and output channels of 1024 and 768, and random dropout.Then, through a double-layer convolutional unit, the feature encoding injected with global information is restored to a depth feature map with a shape of 1024*12*12, which is fed into the decoding module. It passes through four levels of upsampling modules in sequence, and within each module, a splicing operation with the corresponding output feature map saved in the encoding layer is implemented to restore the resolution of the depth feature map to the same shape as the initial input resolution after downsampling, that is, 64*192*192. The feature map is fed into the output segmentation module, and a single-layer two-dimensional convolution with an input channel of 64, an output channel of 4, a stride of 1, a padding of 1, and a convolution kernel of 3*3 is used to finally perform segmentation on four sub-regions of the brain glioma, namely the non-lesion background, tumor core necrosis (NCR), infiltrated edema tissue or peritumoral edema (ED), and enhanced tumor (ET), obtaining a feature map with a shape of 4*192*192. The softmax operator is used to convert the feature map into a probability map corresponding to each category, and the cross-entropy loss between the output probability map and the corresponding ground-truth label mask and the dice loss of categories 1, 2, and 3 are calculated. The two losses are added as the gradient backpropagation function, and the Adam optimization algorithm is used to update the network weights. The above process is repeated until the entire axial training dataset iteration is completed. The number of traversals is set to 60, that is, the training dataset is fully trained 60 times, and the weight parameters of the axial branch network are saved. The coronal and sagittal training datasets are used to iteratively optimize the coronal and sagittal branch networks respectively, and the optimization process is the same as that of the axial branch network. Finally, the weight parameters of the coronal and sagittal branch networks are saved.
[0094] The axial training dataset in step S1 is input into the initial binary classification segmentation auxiliary network model, and the network hyperparameters are trained and debugged to obtain the binary classification segmentation auxiliary network model; in the axial training dataset, the label masks of the two sub-regions of the enhanced tumor area and the core necrosis area are merged into a tumor core label mask, and the label masks of the two sub-regions of the normal area and the edema area are merged into a background label mask. The tumor core label mask and the background label mask are collectively referred to as the binary classification label mask; after randomly shuffling the axial training dataset, the image dataset is divided into multiple groups according to a preset quantity. In this embodiment, 24 images are used as a group, and each group is input into the initial binary classification segmentation auxiliary network model to obtain a preliminary tumor core segmentation prediction map. The cross-entropy loss and DICE score between the preliminary tumor core segmentation prediction map and the binary classification label mask are calculated, and the cross-entropy loss and DICE score between the preliminary tumor core segmentation prediction map and the binary classification label mask are combined as the loss function of the initial binary classification segmentation auxiliary network for backpropagation, and the weights of the initial binary classification segmentation auxiliary network are iteratively updated. After reaching the convergence condition of the initial binary classification segmentation auxiliary network, the network parameters are saved, and finally the binary classification segmentation auxiliary network is obtained.
[0095] Specifically in this embodiment: Use the axial training dataset to optimize the binary classification segmentation auxiliary network. Randomly shuffle the axial training dataset, read part of the data in batches of 24, adjust the shape of each case of four-modal MR data to 4*192*192, adjust the shape of the corresponding ground truth label mask to 1*192*192. Subsequently, randomly rotate the input data and labels by 90 degrees, 180 degrees, and 270 degrees, and then add random Gaussian noise and a random intensity value conversion with a range of (0.75, 1.25) to the input data. Merge and convert the corresponding label masks 1 and 3 of the enhanced tumor and necrotic tissue sub-regions into the tumor core label mask 1, and convert the corresponding label masks 0 and 2 of the original lesion-free background and peritumoral edema sub-regions into the background label mask 0. Then feed it into the binary classification segmentation auxiliary network. The network optimization process is similar to the axial branch network optimization process. The difference lies in the output segmentation module at the end of the network. Use a single-layer two-dimensional convolution with an input channel of 64, an output channel of 2, a stride of 1, a padding of 1, and a convolution kernel of 3*3 to perform the binary classification process of the glioma tumor core in the brain, and obtain a feature map with a shape of 2*192*192.
[0096] S4. Verify the encoder-decoder network sub-model and the binary classification segmentation auxiliary network model
[0097] After randomly shuffling the axial direction validation dataset, coronal direction validation dataset, and sagittal direction validation dataset in step S1, input each image one by one into the encoder-decoder network sub-model in the corresponding direction each time to obtain the sub-region segmentation probability masks in three directions. Perform "rounding and averaging" on the probability masks corresponding to each pixel position to obtain the network segmentation result that fuses the three directions; calculate the DICE score α1 between the network segmentation result that fuses the three directions and the ground truth label mask. If the DICE score α1 remains unchanged or the increase is less than 1% of the current DICE score α1, then obtain the final four-class encoder-decoder network model; if the DICE score α1 increases and the increase is greater than or equal to 1% of the current DICE score α1, then return to step S3 to continue training each sub-model of the four-class encoder-decoder network model until the final four-class encoder-decoder network model is obtained;
[0098] Specifically in this embodiment: Use three test data sets in the axial, coronal, and sagittal directions. Each time, read an image data, adjust the shape of the four-modal MR data to 4*192*192, and input it into the four-classification encoding and decoding network model in step S3 to obtain three feature maps corresponding to the axial, coronal, and sagittal directions with shapes of 4*192*192, 4*192*144, and 4*192*144. After the softmax transformation, convert the feature maps into probability maps for classifying the corresponding sub-regions. Stack the two-dimensional probability maps with four channels in the axial, coronal, and sagittal directions respectively to restore the three-dimensional size of the data with four channels, that is, 4*144*192*192. Then, round the three-dimensional probability maps in the three directions with a threshold of 0.5, that is, set the values greater than or equal to 0.5 to 1, and the values less than 0.5 to 0. Obtain the network segmentation result that fuses the three directions by taking the mean of the three probability maps.
[0099] Randomly shuffle the axial direction validation data set. Each time, input the data set of a single image into the binary classification segmentation auxiliary network model to obtain the binary classification segmentation result of the tumor core. Calculate the DICE score α2 between the binary classification segmentation result of the tumor core and the ground truth label mask. If the DICE score α2 remains unchanged or the increase is less than 1% of the current DICE score α1, then obtain the final binary classification segmentation auxiliary network model; if the DICE score α2 increases and the increase is greater than or equal to 1% of the current DICE score α1, then return to step S3 to continue training the binary classification segmentation auxiliary network model until the final binary classification segmentation auxiliary network model is obtained. Specifically in this embodiment: Use the axial test data set. Each time, read an image data, adjust the shape of the four-modal MR data to 4*192*192, input it into the binary classification segmentation auxiliary network model in step S3 and perform the softmax transformation on the result to obtain a binary classification probability map with a shape of 2*192*192, and obtain the tumor core segmentation map through the argmax transformation. The part of the enhanced tumor region in the network segmentation result that fuses the three directions with a voxel set less than 200 cubic millimeters is corrected to the tumor core necrosis region. Then, subtract the tumor core mask in the tumor core auxiliary segmentation map from the enhanced tumor mask in the three-plane fusion segmentation result, and the difference between the two is corrected to the tumor core necrosis region. The missing voxels in the adjacent part between the tumor core mask and the peritumoral edema region are filled with peritumoral edema. Finally, remove the independent distributed noise voxels in the peritumoral edema region of the three-plane fusion segmentation map with a voxel set less than 200 cubic millimeters and a distance greater than 75 millimeters from the center of the largest part region.
[0100] S5. Extract the information of each image in the axial, coronal, and sagittal directions from the MR image set to be measured to form three data sets to be measured; input the three data sets to be measured into the final fusion segmentation four-class encoding and decoding network model in S4 to obtain the network segmentation results of the MR image set to be measured fused in three directions; input the axial-direction data set to be measured into the final binary segmentation auxiliary network model in S4 to obtain the binary segmentation results of the MR image set to be measured, and correct the network segmentation results of the MR image set to be measured fused in three directions according to the binary segmentation results of the MR image set to be measured to obtain the final fusion segmentation results of the MR image set to be measured; see Figure 4 , it can be seen that the results obtained by using the method provided in this embodiment are in good agreement with the ground truth labels and the segmentation accuracy is comparable.
Claims
1. A glioma fusion segmentation system based on an enhanced attention encoder-decoder network, characterized in that: it includes a four-class encoder-decoder network model and a two-class segmentation auxiliary network model; the four-class encoder-decoder network model is used for feature extraction and region segmentation based on four sub-regions: the core necrosis region, the edema region, the enhanced tumor region, and the normal region; the four-class encoder-decoder network model includes three encoder-decoder network sub-models with the same structure, and the encoder-decoder network sub-model includes an axial encoder-decoder network sub-model, a coronal encoder-decoder network sub-model, and a sagittal encoder-decoder network sub-model; the structure of the two-class segmentation auxiliary network model is the same as that of the encoder-decoder network sub-model, and the two-class segmentation auxiliary network model is used for feature extraction and region segmentation based on two sub-regions: the tumor core region and the normal region in the axial direction; the encoder-decoder network sub-model includes an encoding module, an enhanced attention module, and a decoding module connected in sequence; the enhanced attention module includes a feature encoding unit, a position encoding unit, 4 to 8 feature enhancement modules with the same structure connected in series in sequence, and a double-layer convolution output module; the feature encoding unit and the position encoding unit respectively receive the output information of the encoding module and form encoded data information; the feature enhancement module includes a multi-channel self-attention unit, a first splicing unit, a fully connected feed-forward network unit, and a second splicing unit; the multi-channel self-attention unit receives the encoded data information and is used to add attention scores to each feature vector in the encoded data information to form feature aggregation data information; the first splicing unit is used to receive the encoded data information and the feature aggregation data information and perform splicing to form first-spliced data information; the fully connected feed-forward network unit receives the first-spliced data information and is used to perform feature transformation on the feature vectors to form feature transformation data information; the second splicing unit is used to receive the first-spliced data information and the feature transformation data information and perform splicing to form second-spliced data information; the double-layer convolution output module is connected to the last feature enhancement module and is used to receive the information output by the last feature enhancement module, process it to form global feature encoding information, and output the global feature encoding information to the decoding module; the decoding module is used to perform upsampling on the global feature encoding information to restore it to the same spatial resolution as the input data.
2. The glioma fusion segmentation system based on an enhanced attention encoder-decoder network according to claim 1, characterized in that: the encoding module sequentially includes an input convolution module, a first downsampling module, a second downsampling module, a third downsampling module, and a fourth downsampling module; the input convolution module includes two input convolution layers, and each input convolution layer is provided with batch normalization and an activation function. The input convolution module is used to receive the input tumor image information, and the output data of the fourth downsampling module is output to the enhanced attention module; the first downsampling module, the second downsampling module, the third downsampling module, and the fourth downsampling module are connected in sequence, and the input channels and output channels of the sampling increase exponentially; The decoding module includes a fourth upsampling module, a third upsampling module, a second upsampling module, a first upsampling module and an output segmentation module in sequence; the fourth upsampling module is used to receive the output data of the enhanced attention module, and the output segmentation module is used to output multimodal data; the fourth upsampling module, the third upsampling module, the second upsampling module, and the first upsampling module are connected in sequence, and the sampled input channels and output channels are reduced at the same exponential rate as each downsampling module.
3. The attention-enhanced codec network glioma fusion segmentation system according to claim 1 or 2, Features: The input convolution module of the encoding module outputs a first feature map, the first downsampling module outputs a second feature map, the second downsampling module outputs a third feature map, the third downsampling module outputs a fourth feature map, and the fourth downsampling module outputs a bottom-level feature map; Each upsampling module of the decoding module includes an upsampling layer, a splicing unit and a double-layer convolution unit; the splicing unit of the fourth upsampling module receives the output information of the corresponding upsampling layer and the fourth feature map, splices them, and outputs them to the corresponding double-layer convolution unit; the splicing unit of the third upsampling module receives the output information of the corresponding upsampling layer and the third feature map, splices them, and outputs them to the corresponding double-layer convolution unit; the splicing unit of the second upsampling module receives the output information of the corresponding upsampling layer and the second feature map, splices them, and outputs them to the corresponding double-layer convolution unit; the splicing unit of the first upsampling module receives the output information of the corresponding upsampling layer and the first feature map, splices them, and outputs them to the corresponding double-layer convolution unit.
4. A glioma fusion segmentation method based on enhanced attention encoder-decoder network, It is characterized in that The glioma fusion segmentation system based on enhanced attention encoding and decoding network described in any one of claims 1 to 3 is adopted, comprising the following steps: S1. Preprocess the standard tumor MR image set, extract the 2D image data of tumor tissue in the axial, coronal and sagittal directions, and remove the 2D slice data of non-tumor tissue, respectively, and construct the axial direction training data set, coronal direction training data set, and sagittal direction training data set, and construct the axial direction verification data set, coronal direction verification data set, and sagittal direction verification data set; S2. Constructing initial codec network sub-models: constructing three initial codec network sub-models with the same network structure, wherein the initial codec network sub-models include an axial initial codec network sub-model, a coronal initial codec network sub-model, and a sagittal initial codec network sub-model; Construct an initial binary classification segmentation auxiliary network model; S3, train the initial encoding and decoding network sub-model and the initial binary classification segmentation auxiliary network model and debug the network hyperparameters; Input the axial direction training data set in step S1 into the axial direction initial encoding and decoding network sub-model, input the crown direction training data set into the crown direction initial encoding and decoding network sub-model, input the sagittal direction training data set into the sagittal direction initial encoding and decoding network sub-model, respectively train and debug the network hyperparameters to obtain the axial direction encoding and decoding network sub-model, the crown direction encoding and decoding network sub-model and the sagittal direction encoding and decoding network sub-model; Input the training dataset in the axial direction in step S1 into the initial binary classification segmentation auxiliary network model, train and debug the network hyperparameters to obtain the binary classification segmentation auxiliary network model; S4. Verify the encoder-decoder network sub-model and the binary classification segmentation auxiliary network model Input the axial direction validation dataset, coronal direction validation dataset, and sagittal direction validation dataset in step S1 into the corresponding direction encoder-decoder network sub-models to obtain sub-region segmentation probability masks in three directions. Perform "rounded averaging" on the probability masks corresponding to each pixel position to obtain a network segmentation result that fuses the three directions; Input the axial direction validation dataset into the binary classification segmentation auxiliary network model to obtain the tumor core binary classification segmentation result; Calculate the DICE score α1 between the network segmentation result that fuses the three directions and the ground truth label mask. If the DICE score α1 remains unchanged or the increase is less than 1% of the current DICE score α1, then obtain the final four-class encoder-decoder network model; If the DICE score α1 increases and the increase is greater than or equal to 1% of the current DICE score α1, then return to step S3 to continue training each sub-model of the four-class encoder-decoder network model until the final four-class encoder-decoder network model is obtained; Calculate the DICE score α2 between the tumor core binary classification segmentation result and the ground truth label mask. If the DICE score α2 remains unchanged or the increase is less than 1% of the current DICE score α2, then obtain the final binary classification segmentation auxiliary network model; If the DICE score α2 increases and the increase is greater than or equal to 1% of the current DICE score α2, then return to step S3 to continue training the binary classification segmentation auxiliary network model until the final binary classification segmentation auxiliary network model is obtained; S5. Extract the information of each image in the axial, coronal, and sagittal directions from the MR image set to be tested to form three datasets to be tested; Input the three datasets to be tested into the final fusion segmentation encoder-decoder network model in S4 to obtain the network segmentation result that fuses the three directions of the MR image set to be tested; Input the axial direction dataset to be tested into the final binary classification segmentation auxiliary network model in S4 to obtain the binary classification segmentation result of the MR image set to be tested, and correct the network segmentation result that fuses the three directions of the MR image set to be tested according to the binary classification segmentation result of the MR image set to be tested to obtain the final fusion segmentation result of the MR image set to be tested.
5. The method for glioma fusion segmentation based on an enhanced attention encoder-decoder network according to claim 4, characterized in that: In step S1, the preprocessing of the standard tumor MR image set is specifically as follows: S1.
1. Take the center of each 3D image in the standard tumor MR image set as the center of the cropped 3D image, and crop the original 3D image and the corresponding label mask with a preset window size to remove redundant blank boundaries that are not relevant to the actual imaging data, reduce the data size, and improve the model inference efficiency; S1.
2. Perform Gaussian normalization on the intensities of the cropped images in step S1.1 to increase the model convergence speed; S1.
3. Simultaneously extract four-modal 2D slices from the image data processed in step S1.2 in the three directions of the axial, coronal, and sagittal planes. The four modalities include fluid-attenuated inversion recovery sequence, T1-weighted, post-contrast T1-weighted, and T2-weighted. Use the four-modal data as the four channels of the new 2D image data to generate a 2D tumor MR multi-modal dataset with four channels in three directions. Randomly select 75% of the dataset as the training dataset and the remaining 25% as the validation dataset.
Citation Information
Patent Citations
Cascaded U-N Net brain tumor segmentation method combined with wavelet transform
CN112634192A
Image semantic segmentation method based on coding and decoding structure
CN113807355A