A Brain Tumor Medical Image Segmentation Method Based on Spatial Information and Feature Channels

Through the improved U-Net network structure, combined with the RepVGG module and the dual attention mechanism, the problem of long segmentation time and poor effect in brain tumor medical image segmentation is solved, and more efficient feature extraction and precise lesion segmentation are achieved.

CN114549538BActive Publication Date: 2025-07-11HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210172181.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-24
Publication Date
2025-07-11
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

The existing brain tumor medical image segmentation method has limitations in the combination of feature extraction and classifiers, resulting in high cost of segmentation time and limited improvement in segmentation effect. The single attention mechanism fails to fully utilize the correlation between feature channels.

Method used

The improved method based on U-Net network structure is adopted, and the feature extraction capability is enhanced using the RepVGG module, and a dual attention mechanism is added to the decoder part. Through the dual attention mechanism composed of extrusion and excitation modules, the feature utilization rate is enhanced and the number of complex model parameters is reduced.

Benefits of technology

It improves the accuracy of brain tumor segmentation, reduces the cost of segmentation time, enhances feature extraction capabilities, and solves the problems of weak segmentation capabilities of small targets and unclear edges of multiple targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114549538B_ABST
    Figure CN114549538B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for segmenting brain tumor medical images based on spatial information and feature channels. The present invention proposes an improved U-Net network using RepVGG and a dual attention mechanism. RepVGG enables good module generalization performance without increasing parameters through a branch structure and parameter reconstruction method. This module can solve the disadvantages of complex fully convolutional neural networks, such as having many parameters and long processing time. The attention mechanism has been widely applied to segmentation tasks and has the characteristic of automatically focusing on the target and suppressing irrelevant regions in the input image. The dual attention mechanism effectively extracts and utilizes the focal information in space and the correlation information between feature channels to achieve precise small target segmentation ability. The present invention solves the limitations of large overhead of complex network models and weak small target segmentation ability, enabling the model to obtain precise segmentation of the lesion area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image segmentation. Specifically, it relates to a convolutional neural network model based on spatial information and feature channels. Background Art

[0002] Glioblastoma is a common primary or secondary brain tumor disease, which is a stubborn disease with high treatment difficulty, high mortality, and high recurrence rate. In the early diagnosis stage, by virtue of precise medical imaging technology and improving the precise segmentation ability of the lesion area, it can help doctors give accurate diagnostic results in a timely manner and carry out subsequent treatments. Computer-aided medical image analysis can provide precise guidance for medical professionals to understand diseases and study clinical challenges, so as to improve the diagnostic quality. Clinically, the segmentation of lesions in medical images is mostly diagnosed manually by experienced doctors, which is very time-consuming for doctors. After a long time and a large amount of diagnostic work, physical and mental fatigue is likely to lead to an increase in the probability of misdiagnosis.

[0003] Currently, traditional image processing techniques mainly include two parts: feature extraction and classifier. The design complexity, application limitations, stability of feature extraction algorithms, and the limitations of combining specific feature extraction algorithms with specific classifiers limit the development of image processing techniques. The structure of a computer neural network mimics the structure of the human brain and is composed of a large number of parallel nodes. Each node can perform some basic calculations. Through the training of learning samples, the connection relationship between nodes and the weights of the connections can be obtained to improve the segmentation accuracy. Therefore, the segmentation technology based on deep learning has better performance in this field than other traditional computer vision methods.

[0004] Domestic and foreign research scholars have done a lot of research work on the problem of brain tumor segmentation based on convolutional networks. The network model branch structure and the attention mechanism are relatively mature improved network technologies among them. In the network branch structure, different-sized receptive fields are obtained through convolutional kernels of different sizes, and then different-scale fusions can be obtained through the splicing process, thereby increasing the feature information volume of this layer of structure. The attention mechanism quickly scans the global image to obtain the target area that needs to be focused on, that is, the attention focus, and then invests more attention resources in this area to obtain more detailed information about the target that needs to be focused on, while suppressing other irrelevant information.

[0005] However, there are certain limitations in using only one of these two technologies alone, which are mainly reflected in: The principle of the network model branch structure is to obtain more different-scale information in the sample during the training process. This method will increase the time cost of model segmentation. Currently, the single attention mechanism is only limited to enhancing the weight of the target area and does not utilize the correlation between feature channels, ultimately resulting in limited improvement in the model segmentation effect. Summary of the Invention

[0006] In view of the deficiencies of the prior art, the present invention proposes a method for segmenting brain tumor medical images based on spatial information and feature channels.

[0007] The present invention is an improved method based on the U-Net network structure. In the encoder part, the RepVGG module is used to reduce the number of complex model parameters and enhance the feature extraction ability by utilizing the advantages of its branched network structure and the characteristics of reconfigurability. At the same time, a dual attention mechanism composed of a squeeze-and-excitation module and an attention module is added to the skip connection part of the decoder of the network to enhance the feature extraction ability, suppress irrelevant regions, and improve the utilization rate of features. The improvement of the entire network aims to reduce the time cost of lesion segmentation and improve the accuracy of brain tumor segmentation.

[0008] The method of the present invention specifically includes the following steps:

[0009] Step 1: Preprocess the three-dimensional MRI data set;

[0010] Step 2: The data set processed through the above preprocessing step is used as the model training data set A;

[0011] Step 2-1: In the encoder part, the image x0 in the preprocessed data set A first passes through 4 RepVGG convolution modules and downsampling processing to increase the number of feature channels and reduce the image size. The 5th time only passes through 1 RepVGG convolution module for processing, and each processing obtains results g with different image sizes and different numbers of feature channels. i ;

[0012] Step 2-2: The decoder part processes the result g5 of the above 5th time through 4 upsamplings and VGG convolution modules with an added dual attention mechanism, and finally performs feature fusion through 1×1 convolution to obtain the probability prediction value of each pixel belonging to the target classification;

[0013] Step 2-3: Calculate the loss value by using the BCEDiceLoss function for the prediction results and the true labels of multiple regions, and update the parameters of the neural network in Step 2-1 and 2-2 through backpropagation; when the value of the BCEDiceLoss function tends to be stable, obtain the final set of model parameters;

[0014] Step 3: Reconstruct the model structure and parameters to reduce the number of parameters;

[0015] In the RepVGG module in the encoder, the 1x1 branched convolution and the identity mapping are merged into the 3x3 convolution stack in a reconstructed manner to reduce the number of model parameters and speed up the model's image segmentation speed.

[0016] Preferably, the three-dimensional MRI data set is preprocessed; specifically, the following steps are included:

[0017] Step 1-1: Perform two-dimensional slicing on the three-dimensional MRI data to obtain a two-dimensional picture sequence with four feature channels of T1, T2, FLAIR, and T1CE respectively, and delete the pictures with pixel values of zero;

[0018] Step 1-2: Normalize the two-dimensional picture sequence with four feature channels;

[0019] Step 1-3: Center crop the two-dimensional picture sequence after normalization in step 1-2, and crop the picture size from 240×240 to 160×160.

[0020] Preferably, the convolutional module of the encoder is improved by the RepVGG module adopted in step 2-1, achieving the effect of increasing the feature information amount while preventing gradient disappearance and gradient explosion; the convolutional formula for each layer is as follows;

[0021] g i = f(x i ) + g(x i ) + x i (1)

[0022] x i represents the input of the i-th layer of convolution, f(x i ) represents a 3x3 convolution operation, g(x i ) represents a 1x1 convolution operation, and g i represents the result after convolution fusion.

[0023] Preferably, the VGG module improved by the dual attention mechanism adopted in step 2-2 enables the model to not only extract target features more accurately in terms of spatial information but also make full use of the correlation information between different channels; specifically:

[0024]

[0025]

[0026] g represents the skip connection of the encoder, x l represents the upsampled image data, f1 and f2 represent two different 1x1 convolutions, f3 represents a combination of a 1x1 convolution, Sigmoid activation function, and ReLU activation function, F sq is the global average pooling convolution, F ex is a 1x1 convolution and ReLU activation function, F scale is the matrix multiplication operation.

[0027] Preferably, the BCEDiceLoss loss function is adopted in step 2-3 because the combination of soft dice loss and cross-entropy loss can achieve stability. At the same time, since the evaluation of segmentation is carried out on three partially overlapping regions, multi-region optimization is selected simultaneously. The formula is as follows:

[0028] L'(x,y) = L' dice (x,y)+0.5*L' bce (x,y) (4)

[0029] L(x,y) = L whole (x,y)+L core (x,y)+L enh (x,y) (5)

[0030] L' dice represents the dice loss function, and L' bce represents the cross loss function. The combination of the two is BCEDiceLoss; L whole is the loss function of the entire tumor region, and L core represents the loss function of the tumor core region, and L enh represents the loss function of the tumor enhancement region; x represents the predicted value, y represents the true label value, L'(x,y) represents the loss value of a single region, and L(x,y) represents the combination of the loss values of the three regions.

[0031] Preferably, in step 3, the RepVGG module in the encoder is merged with the 1x1 branch convolution and the identity mapping into the 3x3 convolution stack by reconstruction; first, the bn layer and the conv layer are converted into a conv with a bias vector, obtaining a 3x3 convolution kernel, two 1x1 convolution kernels, and three bias values; then, the two 1x1 convolutions are equivalently converted into 3x3 convolutions by zero-padding; the three convolution kernels are stacked into a 3x3 convolution using the formula and the three bias values are added; finally, a VGG structure with only one 3x3 convolution is obtained, as follows:

[0032]

[0033] bn(x*W,μ,σ,γ,β)=(x*W')+b'(7)

[0034] Conv(x,w1)+Conv(x,w2)+Conv(x,w3)=Conv(x,w1+w2+w3)(8)

[0035] In Formulas (6) and (7), μ, σ, γ, and β represent the mean, standard deviation, learning scale factor, and deviation respectively, x is the input data, W is the weight of the bn layer, W' is the weight of the convolution kernel after conversion between the bn and conv layers, b' is the converted deviation, and * represents the convolution operation; in Formula (8), x is the input data, w1, w2, and w3 are the weight parameters of three convolutions respectively, and the three convolutions are linearly superimposed to finally obtain the VGG structure with only one 3x3 convolution.

[0036] Advantages of the present invention:

[0037] 1. The present invention enables the network to increase the amount of feature information during the training stage while preventing gradient vanishing and gradient explosion through the proposed reconfigurable branch network, enabling the network to obtain five different-scale features and enhancing the model's learning ability. During the prediction stage, the parameter reconstruction can reduce the number of network model parameters and improve the segmentation efficiency.

[0038] 2. The present invention enables the model to not only extract target features more accurately in terms of spatial information through attention adaptive weighting but also make full use of the correlation information between different channels. Therefore, the dual attention mechanism can effectively solve the problems of weak small target segmentation ability and unclear edge segmentation between multiple targets in convolutional neural networks. Description of the drawings

[0039] Figure 1 It is the structural diagram of the entire encoder-decoder network model.

[0040] Figure 2 It is the structural diagram of the training stage and inference stage of ResNet and RepVGG.

[0041] Figure 3 It is the flowchart of the spatial attention mechanism in the dual attention mechanism.

[0042] Figure 4 It is the flowchart of the channel attention mechanism in the dual attention mechanism.

[0043] Figure 5 It is the flowchart of the RepVGG structure and parameter reconstruction.

[0044] Figure 6 It is the comparison diagram between the present invention and other improved methods for brain tumor medical image segmentation. Detailed implementation manners

[0045] In order to make the technical solutions and advantages of the present invention clearer, the present invention will be described in detail below in combination with the drawings and examples.

[0046] Step 1: Preprocess the three-dimensional MRI dataset.

[0047] Step 1-1: Perform two-dimensional slicing on the three-dimensional MRI data to obtain a sequence of pictures with four feature channels. Delete the pictures in the sample label map where all pixels are 0 to reduce the amount of irrelevant samples.

[0048] Step 1-2: Normalize each picture in the sequence of pictures of the four modalities using the method in formula (2) to reduce the sample imbalance. Here, u is the average value of the pixel points, a i is the value of each pixel point before normalization, and b i is the value of each pixel point after normalization.

[0049]

[0050] b i =(a i -u) / s (2)

[0051] Step 1-3: Since there is a black background unrelated to the segmentation target in the dataset pictures, in order to avoid unnecessary video memory consumption, center-crop the two-dimensional picture sequence normalized in step 1-2, and crop the picture size from 240×240 to 160×160.

[0052] Step 2: Put the preprocessed dataset into the model for training.

[0053] Step 2-1: The encoder part is as shown in the appendix Figure 1 First, pass the image x0 in the preprocessed dataset A through 4 RepVGG convolution modules and downsampling to achieve the purpose of increasing the number of feature channels and reducing the picture size. The 5th time, it is only processed through 1 RepVGG convolution module. Each processing obtains results g with different picture sizes and different numbers of feature channels. i .

[0054] One RepVGG contains two groups of modules with the same structure. The modules have 3x3 convolutions, 1x1 convolutions, and identity mapping modules. The appendix Figure 2 shows that the information flow during training is:

[0055] g i =f(x i )+g(x i )+x i (3)

[0056] where f represents a 3×3 convolution, g represents a 1×1 convolution, and x irepresents the identity mapping. This structure extracts feature information of five different scales through five convolutions, and processes the features of each scale separately. The entire network can process the features of the lower layer while retaining the features of the upper layer through skip connections. In this way, by introducing the information between layers and using the image feature information of five different scales and different numbers of channels, the information content contained in the extracted features is increased. The encoder part realizes the functions of increasing the information content of each layer of features, preventing gradient disappearance and gradient explosion by introducing RepVGG.

[0057] Step 2-2: The decoder part processes the above-mentioned result g5 of the fifth time through four upsamplings and a VGG convolution module with a dual attention mechanism, and finally performs feature fusion through a 1×1 convolution to obtain the probability prediction value of each pixel belonging to the target classification.

[0058] Among them, the process of the dual attention mechanism is as follows, and the structure is as attached Figure 3 Skip connection feature g i And the upsampling result x l Enter the 1x1 convolution block for processing respectively to obtain a feature map with the result of H×W, then splice them to get y, and then pass y through a 1x1 convolution block, ReLU activation function, and Sigmoid activation function to obtain the spatial attention coefficient Multiply the attention coefficient With the upsampling result x l Do matrix multiplication, and its result selectively focuses on the target area in the feature map.

[0059] As attached Figure 4 The result of the attention module will be used as the input x of the squeeze-and-excitation module. The squeeze-and-excitation module selectively emphasizes the interdependent channel mappings by integrating the relevant features in all channel mappings. The input x passes through the squeeze F sq (Global average pooling), and then through F ex (Convolutional network and ReLU activation function) to obtain the channel attention coefficient Finally Multiply with the x matrix to further improve the feature representation in the channel dimension, which helps to obtain a more accurate segmentation result.

[0060] Step 2-3: Calculate the error between the prediction result and the true label through the loss function, and then update the parameters through backpropagation. When the value of the loss function tends to be stable, the final set of model parameters is obtained. Among them, the sum of the soft dice loss and the cross-entropy loss is used as the loss function during training to obtain stability:

[0061] L'(x,y)=L' dice (x,y)+0.5*L' bce (x,y) (4)

[0062] The labels provided for training are "edema", "non-enhanced tumor and necrosis", and "enhanced tumor". However, the evaluation of the segmentation is carried out on three partially overlapping regions: the whole tumor consists of all labels, the tumor core consists of "non-enhanced tumor and necrosis" and "enhanced tumor", and the enhanced tumor is a single label. Therefore, instead of optimizing individual classes, these regions are chosen to improve the performance of the segmentation, and the optimization objective is changed to three tumor sub-regions:

[0063] L(x,y) = L whole (x,y) + L core (x,y) + L enh (x,y) (5)

[0064] The Adam optimizer was used during the training process, with a learning rate of 0.003, a momentum parameter of 0.9, and a weight decay of 0.0001.

[0065] Step 3: Reconstruct the model parameters to reduce the number of parameters. The RepVGG module in the encoder can merge the 1x1 branch convolution and the identity mapping into the 3x3 convolution stack through reconstruction, reducing the number of model parameters and accelerating the speed of the model to segment images.

[0066] The structural reparameterization technique is used to remove redundant branches. Its principle is to use the linear characteristics of convolution for simple algebraic transformations as shown in the appendix Figure 2 . Before stacking, bn is used in each branch. Let represent a convolution kernel of size i×i, c2 represent the number of output channels, c1 represent the number of input channels, and use μ (i) , σ (i) , γ (i) , β (i) represent the mean, standard deviation, learned scale factor, and bias respectively, and use x to represent the input data. From the perspective of parameter reconstruction, first, we fuse the convolutional layer and the bn layer. Although the bn layer plays a positive role during training, however, during network inference, there are more layer operations, which affect the performance of the model and occupy more memory or video memory space. Therefore, it is necessary to merge the parameters of the bn layer into the convolutional layer to reduce the calculation and improve the speed of model inference. First, the expression of the bn layer is:

[0067]

[0068] First, convert each bn and its previous conv layer into a conv with a bias vector.

[0069]

[0070] The final fusion result is as follows:

[0071] bn(x*W,μ,σ,γ,β) = (x*W') + b' (8)

[0072] Since the eigenvalues between the input layer and the output layer remain unchanged due to the identity mapping feature, it can be constructed as a 1x1 unit convolution. After the above transformation, a 3x3 convolution kernel, two 1x1 convolution kernels, and three bias values will be obtained. Then, the two 1x1 convolutions are equivalently transformed into a 3x3 convolution by zero-padding. Using the formula Conv(x,w1)+Conv(x,w2)+Conv(x,w3) = Conv(x,w1+w2+w3) (9), the three convolution kernels are stacked into a 3x3 convolution and the three bias values are added together, where x is the input data, and w1, w2, and w3 are the weight parameters of the three convolutions respectively. Finally, a VGG structure with only one 3x3 convolution is obtained.

[0073] According to the appendix Figure 6 , the DualAttentionRepUnet of the present invention is significantly superior to other models in terms of the dice value. The proposed model can well detect all three regions, and its prediction results coincide highly with the official true mask. At the same time, the reconstructed model effectively reduces the number of parameters and the segmentation time, reducing the number of model parameters from 8.83MB to 8.32MB and the time for segmenting cases from 7.12s to 6.58s.

Claims

1. A brain tumor medical image segmentation method based on spatial information and feature channels, characterized in that It includes the following steps: Step 1: Preprocess the three-dimensional MRI dataset; Step 2: The dataset processed by the above preprocessing steps is used as the model training dataset A; Step 2-1: In the encoder part, the image x0 in the preprocessed dataset A first undergoes 4 RepVGG convolutional modules and downsampling to increase the number of feature channels and reduce the image size. The 5th time, it only undergoes 1 RepVGG convolutional module for processing. Each time of processing obtains results g with different image sizes and different numbers of feature channels i ; Step 2-2: The decoder part processes the result g5 of the above 5th time through 4 times of upsampling and a VGG convolutional module with a dual attention mechanism, and finally performs feature fusion through a 1×1 convolution to obtain the probability prediction value of each pixel belonging to the target classification; Step 2-3: Calculate the loss value by using the BCEDiceLoss function for the prediction results and the true labels of multiple regions, and update the parameters of the neural networks in Steps 2-1 and 2-2 through backpropagation; when the value of the BCEDiceLoss function tends to be stable, obtain the final set of model parameters; Step 3: Reconstruct the model structure and parameters to reduce the number of parameters; In the RepVGG module in the encoder, the 1x1 branch convolution and the identity mapping are merged into the 3x3 convolution stack by reconstruction: First, convert the bn layer and the conv layer into a conv with a bias vector, and obtain a 3x3 convolution kernel, two 1x1 convolution kernels, and three bias values; then equivalently convert the two 1x1 convolutions into 3x3 convolutions by zero-padding; use the formula to stack the three convolution kernels into a 3x3 convolution and add the three bias values; finally, obtain a VGG structure with only one 3x3 convolution, specifically as follows: bn(x*W,μ,σ,γ,β)=(x*W')+b'(7) Conv(x,w1)+Conv(x,w2)+Conv(x,w3)=Conv(x,w1+w2+w3)(8) In Formulas (6) and (7), μ, σ, γ, and β respectively represent the mean value, standard deviation, and learned scale factor and bias. x is the input data, W is the weight of the bn layer, W' is the weight of the convolution kernel after the conversion of the bn and conv layers, b' is the converted bias, and * is the convolution operation; in Formula (8), x is the input data, w1, w2, and w3 are the weight parameters of the three convolutions respectively. The three convolutions are linearly superimposed to finally obtain a VGG structure with only one 3x3 convolution, reducing the number of model parameters and accelerating the speed of the model for segmenting images.

2. The method for segmenting brain tumor medical images based on spatial information and feature channels according to claim 1, wherein: The preprocessing of the three-dimensional MRI dataset mentioned above; specifically includes the following steps: Step 1-1: Perform two-dimensional slicing on the three-dimensional MRI data to obtain a two-dimensional picture sequence with four feature channels of T1, T2, FLAIR, and T1CE respectively, and delete the pictures with pixel values of zero; Step 1-2: Normalize the two-dimensional picture sequence with four feature channels; Step 1-3: Center-crop the two-dimensional picture sequence normalized in the above 1-2, and crop the picture size from 240×240 to 160×160.

3. A method for segmenting brain tumor medical images based on spatial information and feature channels according to claim 1, characterized in that: The RepVGG module adopted in Step 2-1 improves the convolutional module of the encoder, achieving the effect of increasing the feature information amount while preventing gradient disappearance and gradient explosion; the convolutional formula for each layer is as follows; g i = f(x i ) + g(x i ) + x i (1) x i represents the input of the i-th layer of convolution, f(x i ) represents a 3x3 convolution operation, g(x i ) represents a 1x1 convolution operation, g i represents the result after convolution fusion.

4. A method for segmenting brain tumor medical images based on spatial information and feature channels according to claim 1, characterized in that: The VGG module improved by the double attention mechanism adopted in step 2-2 enables the model to not only extract target features more accurately in terms of spatial information, but also make full use of the correlation information between different channels. Specifically: g represents the skip connection of the encoder, x l represents the upsampled image data, f1 and f2 represent two different 1x1 convolutions, f3 represents a combination of a 1x1 convolution, Sigmoid activation function and ReLU activation function, F sq is the global average pooling convolution, F ex is a 1x1 convolution and ReLU activation function, F scale is the matrix multiplication operation.

5. A method for segmenting brain tumor medical images based on spatial information and feature channels according to claim 1, characterized in that: The BCEDiceLoss loss function is adopted in step 2-3 because the combination of the soft dice loss and the cross-entropy loss can achieve stability. At the same time, since the evaluation of the segmentation is carried out on three partially overlapping regions, multi-region optimization is selected simultaneously. The formula is as follows: L'(x,y) = L' dice (x,y) + 0.5 * L' bce (x,y) (4) L(x,y) = L whole (x,y) + L core (x,y) + L enh (x,y) (5) L' dice represents the Dice loss function, L' bce represents the cross-entropy loss function, and the combination of the two is BCEDiceLoss; L whole the loss function for the entire tumor region, L core represents the loss function for the tumor core region, L enh represents the loss function for the tumor enhancement region; x represents the predicted value, y represents the true label value, L'(x,y) represents the loss value of a single region, and L(x,y) represents the combination of the loss values of the three regions.

Citation Information

Patent Citations

  • CT image kidney segmentation algorithm based on residual double-attention deep network

    CN110675406A

  • Brain glioma medical image segmentation method based on U-Net network

    CN112446891A