Brain tumor image segmentation method based on Transform multi-scale fusion network

By adopting a multi-scale fusion network based on Transformer in medical image segmentation, the problem of difficulty in capturing global and long-distance semantic information interactions in the prior art is solved, and higher segmentation fineness and accuracy are achieved.

CN120013969AInactive Publication Date: 2025-05-16XUZHOU MEDICAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510480279.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing medical image segmentation technology is difficult to effectively capture global and long-distance semantic information interactions, resulting in insufficient precision of segmentation results.

Method used

A multi-scale fusion network based on Transformer is adopted to extract multi-scale features through a multi-scale network and input them into the Transformer structure, capturing local information and global information, combining the cascaded upsampling and jump connection of the decoder, restoring image space details to enhance the local details of the segmentation results.

Benefits of technology

It improves the precision of medical image segmentation, can capture local semantic and texture information more effectively, and steadily captures geometric and structural information in medical data, improving segmentation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013969A_ABST
    Figure CN120013969A_ABST
Patent Text Reader

Abstract

The invention discloses a brain tumor image segmentation method based on a Transform multi-scale fusion network, and belongs to the field of pattern recognition, and the method comprises the steps: firstly, selecting image data needing to be segmented, and carrying out the image preprocessing, and obtaining the amplified image data; then, a clear long-range dependency relationship is captured by introducing a Transform structure in an encoder part, so that context information of an image can be better understood; the space details of the image are recovered through cascade up-sampling and jump connection in a decoder part, so that efficient brain tumor image segmentation is realized; and finally, outputting a segmentation result of the brain tumor image through a segmentation head. The method can provide objective and efficient assistance for brain tumor image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for realizing image segmentation by utilizing a Transformer-based multi-scale fusion network model, and belongs to the field of pattern recognition. Background Art

[0002] Brain tumors are a type of malignant tumor that originates in brain tissue or metastasizes to the brain from other parts of the body. Early diagnosis plays a vital role in improving patients' chances of survival. In order to help clinicians make accurate diagnoses, it is necessary to segment some key objects in medical images and extract features from the segmented areas. With the development of artificial intelligence, deep learning has been widely used in the field of brain tumor image segmentation. This is mainly because the convolutional neural network used for feature learning is insensitive to image noise, blur, contrast, etc., and can therefore provide excellent segmentation effects for medical images. U-Net can effectively fuse low-level and high-level image features by adopting a symmetrical encoder-decoder structure and jump connections, which makes it the benchmark architecture for most medical image segmentation tasks and has inspired many meaningful improvements. For example, 3D networks, recurrent neural networks, and jump connections. Although the experimental results show that the above improvement strategy is a perfect solution for the medical image segmentation task, for 3D networks, due to the huge number of parameters, the network model faces problems such as high computational cost and excessive GPU memory usage; the design of using recurrent neural networks to improve the performance of medical image segmentation is not common, mainly because good medical image quality is required to obtain complete and effective temporal information, but the actual medical image data is not satisfactory; although the jump connection can fuse low-resolution and high-resolution information to improve feature representation, it has the problem of a large semantic gap between low-resolution features and high-resolution features, which will cause the feature map to be unclear. Based on this, in order to obtain better performance, researchers have successfully used a variety of techniques such as network architecture design, loss function design, transfer learning, interactive segmentation, graph convolutional neural network, etc. to improve the accuracy of medical image segmentation. However, although the convolutional neural network as the backbone network has a relative advantage in extracting underlying features, it cannot learn global and long-distance semantic information interactions well due to the limitations of the convolution operation. In addition, the convolution operation usually produces lower-resolution features, and these detailed information is difficult to recover by simple upsampling methods, thus affecting the fineness of the segmentation results. Summary of the invention

[0003] Purpose of the invention: In order to improve the precision of medical image segmentation, the present invention proposes a brain tumor image segmentation method based on the Transformer multi-scale fusion network, which can use long-range dependencies for encoding to capture local and global information, thereby providing objective and efficient assistance for solving brain tumor image segmentation.

[0004] Technical solution: A brain tumor image segmentation method based on Transformer multi-scale fusion network, including the following steps:

[0005] Step 1: construct a multi-scale fusion network to segment the target in the image to be processed;

[0006] The multi-scale fusion network is a U-Net structure, including an encoder and a decoder;

[0007] A segmentation header is attached to the last layer of the decoder to output the segmentation results of the target;

[0008] Step 2, training the multi-scale fusion network to obtain a trained multi-scale fusion network;

[0009] Step 3: preprocess the image to be processed, input the preprocessed image into the trained multi-scale fusion network, and segment the target in the image.

[0010] Furthermore, the encoder includes a multi-scale network and a Transformer structure;

[0011] The multi-scale network includes four combination modules, namely a first combination module, a second combination module, a third combination module and a fourth combination module, and the four combination modules are connected in sequence;

[0012] The multi-scale network is used to extract multi-scale features, which are sequentially expanded into a one-dimensional tensor and input into the Transformer structure to capture local and global information, thereby obtaining a reconstructed feature map.

[0013] The reconstructed feature map is input into a decoder for decoding.

[0014] Furthermore, cascade upsampling is performed in the decoder, and the multi-scale network of the encoder and the upsampling jump connection of the decoder are used to fuse multi-scale features and restore the image spatial details to enhance the local details of the segmentation results; the image size is restored through transposed convolution to achieve accurate positioning and segmentation of the target, and the segmentation result is output.

[0015] Furthermore, the first combination module includes a stacked convolution layer and a pooling layer;

[0016] The second combination module includes a stacked first bottleneck module and two second bottleneck modules, the first bottleneck module has different numbers of input and output channels, and the second bottleneck module has the same number of input and output channels;

[0017] The third combined module includes a stacked first bottleneck module and three second bottleneck modules;

[0018] The fourth combined module includes a first bottleneck module and five second bottleneck modules which are stacked.

[0019] Furthermore, the first combination module first performs a convolution operation with a stride of 2 and a convolution kernel size of 7×7 on the image, and uses the result as the input of the MaxPool layer; then performs a pooling operation with a stride of 2 and a convolution kernel size of 2×2; and outputs the data to the second combination module.

[0020] Further, the first bottleneck module includes two branches;

[0021] The first branch of the first bottleneck module includes three convolutional layers connected in sequence, namely a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer;

[0022] The second branch of the first bottleneck module includes a convolutional layer, which is a 1×1 convolutional layer;

[0023] The output results of the two branches of the first bottleneck module are obtained by the ReLU activation function to obtain the output result of the first bottleneck module;

[0024] The second bottleneck module includes three convolutional layers connected in sequence, namely a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer; the output results after the three convolutions are combined with the input data of the second bottleneck module through the ReLU activation function to obtain the output result of the second bottleneck module;

[0025] The first bottleneck module in the second combined module performs a convolution operation with a step size of 1, without downsampling, the input size and the output size are the same, and the number of input channels C is equal to the number of channels C1 of the 1×1 convolution layer in the first branch of the first bottleneck module;

[0026] The first bottleneck module in the third combination module performs a convolution operation with a step size of 2 and performs downsampling. The input size is twice the output size. The number of input channels C is not equal to the number of channels C1 of the 1×1 convolution layer in the first branch of the first bottleneck module, C=2×C1.

[0027] Furthermore, the convolutional layers in the four combined modules are all the input feature maps that have undergone convolution operations, BN layers, and ReLU activation functions in sequence.

[0028] Furthermore, the Transformer structure includes Transformer modules connected in sequence, totaling 12 Transformer modules; each Transformer module has the same structure;

[0029] The Transformer module consists of multiple sublayers, including convolutional layers, layer normalization, attention modules, and feedforward neural networks. Each sublayer is connected through residual connections. The channel attention module is used to replace the traditional self-attention module to reduce the parameter scale and improve computational efficiency. The channel attention module consists of only one global average pooling and convolution layer.

[0030] Furthermore, the cross entropy loss function and the Dice loss function are combined by weighted summation to train the Transformer-based multi-scale fusion network, and the network parameters are analyzed to complete the brain tumor image segmentation.

[0031] Furthermore, the multi-scale fusion network is trained using the LGG Segmentation dataset; the LGG Segmentation dataset is amplified, including: subtracting the mean of the image from each pixel in the image and dividing it by the standard deviation so that the image has zero mean and unit variance; randomly cropping image blocks and randomly horizontally mirroring half of the image for data amplification.

[0032] Beneficial effects: This method combines the advantages of multi-scale representation and Transformer fusion network structure, and has the following advantages: (1) The multi-scale fusion network can effectively improve the ability to capture local semantic and texture information; (2) The channel attention mechanism can robustly capture the geometric and structural information present in medical data. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a flow chart of a fusion network construction method for brain tumor image segmentation of the present invention;

[0034] Figure 2 It is a schematic diagram of the structure of a multi-scale network; Figure 3 It is the specific structure diagram of each combination module in the multi-scale network; Figure 4 This is the specific structure diagram of the two bottleneck modules;

[0035] Figure 5 It is a schematic diagram of the structure of the Transformer module;

[0036] Figure 6 It is a visualization diagram of the brain tumor image segmentation experiment used in the present invention. DETAILED DESCRIPTION

[0037] The present invention will be further explained below in conjunction with the accompanying drawings.

[0038] like Figure 1 As shown, a brain tumor image segmentation method based on Transformer multi-scale fusion network includes the following steps:

[0039] Step 1: Select the image data to be segmented and perform image preprocessing to obtain the amplified image data.

[0040] Take the LGG Segmentation dataset as an example. This dataset consists of 3929 images of brain magnetic resonance imaging (MRI) images and manually annotated FLAIR (Fluid Attenuated Inversion Recovery) abnormal segmentation masks. The preprocessing of MRI images includes the following steps:

[0041] Step 1.1: For each pixel in the MRI image, subtract the image mean and divide by the standard deviation to make it have zero mean and unit variance.

[0042] Step 1.2: Randomly crop an image patch of size 224×224 and randomly mirror half of the image horizontally for data augmentation.

[0043] Step 2: In the encoder part, a multi-scale network structure is constructed by stacking several different combination modules. This implementation is completed based on the practice of LGG Segmentation data. The constructed multi-scale network includes 5 combination modules, such as Figure 2 The figure shows that each combination module is connected in sequence. Figure 3 For each combined module specific structure, Figure 4 (a) and Figure 4 (b) in the figure shows the specific structure diagrams of the first bottleneck module and the second bottleneck module. The construction process includes the following specific steps:

[0044] Step 2.1: Figure 3 After the preprocessed image is input to the first combination module, a convolution operation with a stride of 2 and a convolution kernel size of 7×7 is first performed on the image, and the result is used as the input of the MaxPool layer; then a pooling operation with a stride of 2 and a convolution kernel size of 2×2 is performed; and the data is output to the second combination module;

[0045] For the convolutional layers in the five combined modules, the input feature maps must undergo convolution operations, BN layers, and ReLU activation functions in sequence.

[0046] Step 2.2: Given two different bottleneck modules (the first bottleneck module has different numbers of input and output channels, and the second bottleneck module has the same number of input and output channels), the stacking process of the second combination module, the third combination module, and the fourth combination module includes:

[0047] The second combination module includes a stacked first bottleneck module and two second bottleneck modules. The first bottleneck module in the second combination module performs a convolution operation with a step size of 1 without downsampling. The input size and output size are the same. The number of input channels C is equal to the number of channels C1 of the 1×1 convolution layer in the first branch of the first bottleneck module (C=C1=64). The parameter rules of the second bottleneck modules in the second combination module, the third combination module, and the fourth combination module are the same, except for the input size.

[0048] The first bottleneck module includes two branches, a first branch and a second branch, the first branch includes three convolutional layers connected in sequence, namely a 1×1 convolutional layer, a 3×3 convolutional layer and a 1×1 convolutional layer, and the second branch includes one convolutional layer, which is a 1×1 convolutional layer; the results of the first branch and the second branch are subjected to a ReLU activation function to obtain the output result of the first bottleneck module;

[0049] The third combination module includes a stacked first bottleneck module and three second bottleneck modules. The first bottleneck module in the third combination module performs a convolution operation with a step size of 2 and downsampling. The input size is twice the output size. The number of input channels C is not equal to the number of channels C1 of the 1×1 convolution layer in the first branch of the first bottleneck module (C=2×C1).

[0050] The fourth combination module includes a stacked first bottleneck module and five second bottleneck modules. The parameter rules of the first bottleneck module and the second bottleneck module in the fourth combination module are the same as those in step 2.2.2, except that the input size is different.

[0051] The output of the fourth combination module is 14×14. The output of the fourth combination module is sequentially expanded into a one-dimensional tensor, and then converted into a one-dimensional vector with a length of 196 after expansion, and then positionally embedded to retain the image spatial information to complete the construction of the multi-scale network.

[0052] Step 3: By introducing Transformer into the U-Net structure as an encoder, local and global information is captured.

[0053] In the encoder, the Transformer includes Transformer modules connected in sequence, totaling 12 Transformer modules.

[0054] like Figure 5As shown in the figure, each Transformer module consists of multiple sub-layers, including convolutional layers, layer normalization, attention modules, and feed-forward neural networks, and each sub-layer is connected by residual connections.

[0055] The channel attention module is used to replace the traditional self-attention module in the Transformer module to reduce the parameter scale and improve the computational efficiency. The channel attention module consists of only one global average pooling and convolution layer.

[0056] Step 4: In the decoder part, the spatial details of the image are restored through cascade upsampling and skip connections to achieve accurate positioning and segmentation of brain tumor targets.

[0057] The output after the reconstruction of the Transformer structure is jump-connected with the output of the middle layer of the multi-scale network of different spatial dimensions to enhance the local details of the segmentation result. In this specific implementation, the middle layer of the multi-scale network of different spatial dimensions refers to the second combination module, the third combination module, and the fourth combination module.

[0058] The image size is restored through transposed convolution to achieve accurate positioning and segmentation of the target.

[0059] A segmentation head is attached to the last layer of the decoder to output accurate segmentation results of brain tumor images.

[0060] Step 5: Combine the cross entropy loss function and the Dice loss function by weighted summation to train the Multi-scale Fusion Network Based on Transformer (MFNTrans), analyze the network parameters, and complete the brain tumor image segmentation.

[0061] In order to verify the effectiveness and superiority of the Transformer-based multi-scale fusion network, the Transformer-based multi-scale fusion network MFNTrans of the present invention is compared with the existing Res-UNet, ResNeXt-UNet, VGG-UNet, and Attention U-Net through the LGG Segmentation data experiment. The LGG Segmentation dataset has a total of 3929 224×224 pixel images, of which 3219 images are used for training and 710 images are used for testing.

[0062] The effects of different loss function combinations on the segmentation performance of MFNTrans are analyzed, as shown in Table 1. "DSC" represents the Dice similarity coefficient. As can be seen from Table 1, different ratios of the cross entropy loss function and the Dice loss function have a significant impact on the segmentation performance of the model. When the weight ratio of the cross entropy loss function to the Dice loss function is 0.5:0.5, the DSC of the model reaches a maximum value of 90.77%. Therefore, the weight ratio of 0.5:0.5 will be used as the standard combination in subsequent experiments.

[0063] Table 1. Comparison of MFNTrans performance under different loss function combinations

[0064]

[0065] Analyzing the impact of different numbers of skip connections on the segmentation performance of MFNTrans, we can see from Table 2 that as the number of skip connections increases, the accuracy of model segmentation also increases. It is worth noting that when the number of skip connections increases from 0 to 3, the DSC value of the model increases by 4.95%, which shows that it is effective to improve the segmentation performance by increasing the number of skip connections.

[0066] Table 2 Comparison of MFNFTrans performance under different hop connection numbers

[0067]

[0068] To evaluate the performance of various U-Net networks in brain tumor image segmentation, we can see from Table 3 that: (1) Res-UNet, ResNeXt-UNet and Attention U-Net with shortcut connections show stronger generalization ability than the single deep network VGG-UNet; (2) MFNTrans shows the effectiveness of brain tumor image segmentation and obtains the largest DSC value compared with the existing advanced models. Therefore, using MFNTrans to segment brain tumor images not only improves the model's ability to extract complex image features, but also improves the segmentation accuracy.

[0069] Table 3 Comparison of brain tumor image segmentation performance of various U-Net networks

[0070]

[0071] Combination Figure 6 , showing the visualization of MFNTrans’s results for brain tumor image segmentation. Figure 6 It can be seen that MFNTrans can better capture the boundaries of the tumor area, and the prediction results are highly consistent with the true labels.

[0072] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A Transformer-based image segmentation method, characterized in that: The steps include: Step 1: Build a multi-scale fusion network to segment the target in the image to be processed The multi-scale fusion network is a U-Net structure, including an encoder and a decoder; A segmentation header is attached to the last layer of the decoder to output the segmentation results of the target; Step 2, training the multi-scale fusion network to obtain a trained multi-scale fusion network; Step 3: preprocess the image to be processed, input the preprocessed image into the trained multi-scale fusion network, and segment the target in the image.

2. According to claim 1, a Transformer-based image segmentation method is characterized in that: The encoder includes a multi-scale network and a Transformer structure; The multi-scale network includes four combination modules, namely a first combination module, a second combination module, a third combination module and a fourth combination module, and the four combination modules are connected in sequence; The multi-scale network is used to extract multi-scale features, which are sequentially expanded into a one-dimensional tensor and input into the Transformer structure to capture local and global information, thereby obtaining a reconstructed feature map. The reconstructed feature map is input into a decoder for decoding.

3. According to claim 1, a Transformer-based image segmentation method is characterized in that: Cascade upsampling is performed in the decoder, and the multi-scale network of the encoder and the upsampling jump connection of the decoder are used to fuse multi-scale features, restore image spatial details, and output segmentation results.

4. According to claim 2, a Transformer-based image segmentation method, characterized in that: The first combination module includes a stacked convolution layer and a pooling layer; The second combination module includes a stacked first bottleneck module and two second bottleneck modules, the first bottleneck module has different numbers of input and output channels, and the second bottleneck module has the same number of input and output channels; The third combined module includes a stacked first bottleneck module and three second bottleneck modules; The fourth combined module includes a first bottleneck module and five second bottleneck modules which are stacked.

5. The Transformer-based image segmentation method according to claim 4, characterized in that: The first combination module first performs a convolution operation with a stride of 2 and a convolution kernel size of 7×7 on the image, and uses the result as the input of the MaxPool layer; then performs a pooling operation with a stride of 2 and a convolution kernel size of 2×2; and outputs the data to the second combination module.

6. The Transformer-based image segmentation method according to claim 4, characterized in that: The first bottleneck module includes two branches; The first branch of the first bottleneck module includes three convolutional layers connected in sequence, namely a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer; The second branch of the first bottleneck module includes a convolutional layer, which is a 1×1 convolutional layer; The output results of the two branches of the first bottleneck module are obtained by the ReLU activation function to obtain the output result of the first bottleneck module; The second bottleneck module includes three convolutional layers connected in sequence, namely a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer; the output results after the three convolutions are combined with the input data of the second bottleneck module through the ReLU activation function to obtain the output result of the second bottleneck module; The first bottleneck module in the second combined module performs a convolution operation with a step size of 1, without downsampling, the input size and the output size are the same, and the number of input channels C is equal to the number of channels C1 of the 1×1 convolution layer in the first branch of the first bottleneck module; The first bottleneck module in the third combination module performs a convolution operation with a step size of 2 and performs downsampling. The input size is twice the output size. The number of input channels C is not equal to the number of channels C1 of the 1×1 convolution layer in the first branch of the first bottleneck module, C=2×C1.

7. The Transformer-based image segmentation method according to claim 6, characterized in that: In the convolution layer, the input feature map undergoes convolution operation, BN layer and ReLU activation function successively.

8. The Transformer-based image segmentation method according to claim 6, characterized in that: The Transformer structure includes Transformer modules connected in sequence, totaling 12 Transformer modules; The structure of each Transformer module is the same; The Transformer module consists of multiple sub-layers, including convolutional layers, layer normalization, attention modules, and feed-forward neural networks, and each sub-layer is connected by residual connections.

9. The Transformer-based image segmentation method according to claim 1, characterized in that: The multi-scale fusion network is trained by adopting a weighted combination of a cross entropy loss function and a Dice loss function to optimize network parameters.

10. The method for constructing a multi-scale fusion network based on Transformer according to claim 1, characterized in that: The multi-scale fusion network is trained using the LGG Segmentation dataset; The LGG Segmentation dataset was expanded, including: Subtract the mean of the image from each pixel and divide by the standard deviation to make the image have zero mean and unit variance; Randomly crop image patches and randomly mirror half of the image horizontally for data augmentation.

Citation Information

Patent Citations

  • Cascade Transform-based brain tumor segmentation method

    CN117036380A

  • MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on multi-scale feature fusion of improved U-Net

    CN117876399A

  • MRI brain tumor segmentation method based on attention bottleneck fusion

    CN118314350A