Brain tumor image segmentation method and system
By using the inverted residual bottleneck structure driven by 3D U-Net network and Transformer driven in the brain tumor image segmentation model, and combining the mixed loss function for training, the problems of insufficient perception of global tumor information and high demand for computing resources in the existing technology are solved, and the segmentation effect of high precision and low computational volume is achieved.
Patent Information
- Application Number
- CN202510507749.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to simulate long-distance dependencies in brain tumor image segmentation, resulting in insufficient perception of tumor global information. At the same time, the computing resources based on ViT model are high and a large amount of labeled data is required, resulting in high computing costs and overfitting.
A U-shaped encoder-decoder structure based on 3D U-Net network is adopted, and an inverted residual bottleneck structure driven by Transformer is introduced. The tumor details are preserved through jump connections, memory usage is reduced, and model training is performed through mixed loss functions (cross entropy loss, Dice loss and orthogonal regularization loss).
It improves the segmentation accuracy of the brain tumor image segmentation model, while reducing the amount of model calculation, reducing the demand for resources, and enhancing the generalization ability of the model.
Smart Images

Figure CN120032134A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image segmentation, and in particular to a method and system for brain tumor image segmentation. Background Art
[0002] Gliomas are one of the most common primary brain malignant tumors and account for a high proportion of all central nervous system tumors. For example, studies have shown that gliomas account for about 27% of central nervous system tumors and as high as 81% of malignant tumors. The classification of brain tumors is relatively complex. The World Health Organization (WHO) divides them into two categories according to cell origin and behavior: non-malignant (grade I or II, also called low-grade tumors) and malignant (grade III or IV, also called high-grade tumors). Different types of brain tumors have different growth potential and invasiveness, which increases the difficulty of accurate segmentation. With the continuous advancement of medical imaging technology, medical image segmentation technology is also developing.
[0003] Brain tumor segmentation technology based on deep learning has made significant progress in the past few years, thanks to the powerful capabilities of deep learning in image recognition and segmentation tasks. These advances have not only improved the accuracy of segmentation, but also accelerated the transformation process from imaging data to clinical applications. The fusion of multimodal MRI data (such as T1-weighted, T2-weighted, FLAIR, etc.) provides rich information for brain tumor segmentation. Deep learning models can learn more comprehensive feature representations from multimodal data, thereby improving the accuracy of segmentation. Convolutional Neural Networks (CNNs) can automatically learn rich feature representations from images. The multi-layer convolution structure can capture features of different scales, which helps to identify complex tumor boundaries and achieve high-precision segmentation. Through a large amount of training data, CNNs can learn the characteristics of different types of tumors and background tissues, have strong generalization capabilities, and can easily process multimodal MRI data (such as T1, T2, FLAIR, etc.), and improve the accuracy of segmentation through multimodal fusion. CNNs are divided into two-dimensional and three-dimensional convolutions. Two-dimensional convolutions have less computational complexity when processing a single slice, and are faster in training and reasoning. For mobile devices with limited resources and real-world applications, two-dimensional convolution is more efficient and can complete the task quickly. However, it cannot capture the complete structure and contextual information of the tumor. This may lead to inaccurate segmentation results in some cases, especially when the tumor boundaries are blurred or the shape is complex. Three-dimensional convolution can process the spatial relationship between adjacent slices, provide richer contextual information, and help the model better understand the overall structure of the tumor and its relationship with surrounding tissues. However, the complexity of the three-dimensional convolution model is high, which increases the training time and memory consumption of the model. Although 3D convolution can capture local spatiotemporal information, it may not be as effective as self-attention mechanisms such as Transformer when processing long-distance dependencies.
[0004] Transformer can effectively model long-distance dependencies and extract features at different scales through the self-attention mechanism to capture multi-level information of tumors. This is because it allows the model to establish connections between different locations, which helps to identify the complex structure of the tumor. For multimodal data, Transformer can easily handle and improve the accuracy of segmentation by fusing information from different modalities. In addition, its self-attention mechanism can be calculated in parallel during the training process, which can improve the training and inference efficiency of the model. Vision Transformer (ViT) is a model based on the Transformer architecture, which was originally used for computer vision tasks such as image classification. Unlike traditional convolutional neural networks, ViT divides the image into a series of "patches" and treats these patches as sequence data, capturing global dependencies through the self-attention mechanism. This mechanism makes ViT perform well in handling long-distance dependencies and global context information. Some studies have applied ViT to the task of brain tumor segmentation and achieved remarkable results. For example, the TransBTS model combines ViT and 3D convolutional neural networks to achieve high-precision brain tumor segmentation through multi-scale feature extraction and global context modeling. ViT has high computational resource requirements, especially when processing high-resolution medical imaging data. This may limit its application on resource-limited devices. In addition, due to the lack of inductive bias in CNN, it usually requires more training data to avoid overfitting, and its performance on small data sets is not as good as that of pure CNN segmentation networks. The acquisition of high-quality annotated data is costly and time-consuming, which has brought great obstacles to the development of ViT-based segmentation models.
[0005] As can be seen above, convolutional neural network-based segmentation models cannot simulate long-range dependencies, which hinders their perception of global information about tumors. In addition, the computational resource requirements of ViT-based segmentation models are high, requiring a large number of annotated datasets to obtain optimal segmentation performance, resulting in high computational costs and overfitting on small datasets.
[0006] Therefore, how to improve the segmentation accuracy of the brain tumor image segmentation model while reducing the amount of model calculation is a technical problem that needs to be urgently solved by technical personnel in this field. Summary of the invention
[0007] To solve the above technical problems, the present application provides a brain tumor image segmentation method, which can improve the segmentation accuracy of the brain tumor image segmentation model and reduce the model calculation amount. The present application also provides a brain tumor image segmentation system, which has the same technical effect.
[0008] The first objective of the present application is to provide a brain tumor image segmentation method.
[0009] The above-mentioned application objective 1 of the present application is achieved through the following technical solutions: A brain tumor image segmentation method, comprising: Obtain a pre-trained brain tumor image segmentation model, wherein the brain tumor image segmentation model is constructed based on a U-shaped encoder-decoder structure of a 3DU-Net network, including an encoder consisting of one input convolution layer and four lower convolution modules connected in sequence, and a decoder consisting of four upper convolution modules connected in sequence and one output convolution layer, wherein the output of each lower convolution module in the encoder is jump-connected to the input of the corresponding upper convolution module in the decoder; The lower convolution module is composed of 1 lower convolution layer and 2 convolution layers, and the upper convolution module is composed of 1 upper convolution layer and 2 convolution layers; The lower convolution layer adopts a first inverted residual bottleneck structure, which includes one depthwise convolution layer and two pointwise convolution layers connected in sequence, and one pointwise convolution layer for residual connection; The upper convolution layer adopts a second inverted residual bottleneck structure, which includes one depthwise transposed convolution layer and two pointwise convolution layers connected in sequence, and one pointwise transposed convolution layer for residual connection; The convolution layer adopts a third inverted residual bottleneck structure and a fourth inverted residual bottleneck structure connected in sequence, wherein the third inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence, and one pointwise convolution layer for residual connection, and the fourth inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence; A brain tumor image to be segmented is acquired, and the brain tumor image is segmented using the brain tumor image segmentation model to obtain a segmentation result.
[0010] Preferably, in the brain tumor image segmentation method, in the first inverted residual bottleneck structure, the input feature is input into the depthwise convolution layer and the pointwise convolution layer used for residual connection, the output of the depthwise convolution layer is group normalized and input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise convolution layer used for residual connection to obtain the output feature of the first inverted residual bottleneck structure.
[0011] Preferably, in the brain tumor image segmentation method, in the second inverted residual bottleneck structure, the input features are input into the depthwise transposed convolution layer and the pointwise transposed convolution layer for residual connection, the output of the depthwise transposed convolution layer is group normalized and input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise transposed convolution layer for residual connection to obtain the output feature of the second inverted residual bottleneck structure.
[0012] Preferably, in the brain tumor image segmentation method, in the third inverted residual bottleneck structure, the input feature is input into the depthwise convolution layer and the pointwise convolution layer used for residual connection, the output of the depthwise convolution layer is group normalized and input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise convolution layer used for residual connection to obtain the output feature of the third inverted residual bottleneck structure.
[0013] Preferably, in the brain tumor image segmentation method, in the fourth inverted residual bottleneck structure, the output features of the previously connected third inverted residual bottleneck structure are input into the depthwise convolution layer, the output of the depthwise convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the input of the depthwise convolution layer as the output feature of the fourth inverted residual bottleneck structure.
[0014] Preferably, in the brain tumor image segmentation method, the convolution kernel size of the brain tumor image segmentation model is 7×7.
[0015] Preferably, in the brain tumor image segmentation method, obtaining a pre-trained brain tumor image segmentation model comprises: Obtain sample brain tumor images; Construct a hybrid loss function composed of the cross entropy loss function, the Dice loss function and the orthogonal regularization loss function; The initial brain tumor image segmentation model is trained using the brain tumor image samples and the mixed loss function to obtain the trained brain tumor image segmentation model.
[0016] Preferably, in the brain tumor image segmentation method, the mixed loss function is specifically: ; In the formula, represents the hybrid loss function, represents the cross entropy loss function, represents the Dice loss function, represents the orthogonal regularization loss function, and Indicates adjustment weight.
[0017] The second objective of the present application is to provide a brain tumor image segmentation system.
[0018] The second application objective of the present application is achieved through the following technical solutions: A brain tumor image segmentation system, comprising: An acquisition unit is used to acquire a pre-trained brain tumor image segmentation model, wherein the brain tumor image segmentation model is constructed based on a U-shaped encoder-decoder structure of a 3D U-Net network, including an encoder consisting of one input convolution layer and four lower convolution modules connected in sequence, and a decoder consisting of four upper convolution modules connected in sequence and one output convolution layer, wherein the output of each lower convolution module in the encoder is jump-connected to the input of the corresponding upper convolution module in the decoder; The lower convolution module is composed of 1 lower convolution layer and 2 convolution layers, and the upper convolution module is composed of 1 upper convolution layer and 2 convolution layers; The lower convolution layer adopts a first inverted residual bottleneck structure, which includes one depthwise convolution layer and two pointwise convolution layers connected in sequence, and one pointwise convolution layer for residual connection; The upper convolution layer adopts a second inverted residual bottleneck structure, which includes one depthwise transposed convolution layer and two pointwise convolution layers connected in sequence, and one pointwise transposed convolution layer for residual connection; The convolution layer adopts a third inverted residual bottleneck structure and a fourth inverted residual bottleneck structure connected in sequence, wherein the third inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence, and one pointwise convolution layer for residual connection, and the fourth inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence; The segmentation unit is used to obtain a brain tumor image to be segmented, and perform image segmentation on the brain tumor image using the brain tumor image segmentation model to obtain a segmentation result.
[0019] Preferably, in the brain tumor image segmentation system, in the first inverted residual bottleneck structure, input features are input into a depthwise convolution layer and a pointwise convolution layer for residual connection, the output of the depthwise convolution layer is group normalized and then input into a first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by a GELU activation function and then input into a second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise convolution layer for residual connection to obtain the output features of the first inverted residual bottleneck structure; In the second inverted residual bottleneck structure, the input feature is input into the depthwise transposed convolution layer and the pointwise transposed convolution layer for residual connection, the output of the depthwise transposed convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise transposed convolution layer for residual connection, and the output feature of the second inverted residual bottleneck structure is obtained; In the third inverted residual bottleneck structure, the input feature is input into the depthwise convolution layer and the pointwise convolution layer for residual connection, the output of the depthwise convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise convolution layer for residual connection to obtain the output feature of the third inverted residual bottleneck structure; In the fourth inverted residual bottleneck structure, the output features of the previously connected third inverted residual bottleneck structure are input into the depthwise convolution layer, the output of the depthwise convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the input of the depthwise convolution layer as the output feature of the fourth inverted residual bottleneck structure.
[0020] The above technical solution performs brain tumor image segmentation through a pre-trained brain tumor image segmentation model, wherein the brain tumor image segmentation model is constructed based on a U-shaped encoder-decoder structure of a 3D U-Net network, including an encoder consisting of one input convolution layer and four down-convolution modules connected in sequence, and a decoder consisting of four up-convolution modules connected in sequence and one output convolution layer, and the output of each down-convolution module in the encoder is jump-connected to the input of the corresponding up-convolution module in the decoder; specifically, the high-resolution image from the encoder is integrated with the upsampled low-resolution image of the decoder through the jump connection, thereby retaining more tumor details and improving the segmentation accuracy.
[0021] In addition, the lower convolution module consists of 1 lower convolution layer and 2 convolution layers, and the upper convolution module consists of 1 upper convolution layer and 2 convolution layers; the lower convolution layer adopts the first inverted residual bottleneck structure, which includes 1 depthwise convolution layer and 2 pointwise convolution layers connected in sequence, and 1 pointwise convolution layer for residual connection; the upper convolution layer adopts the second inverted residual bottleneck structure, which includes 1 depthwise transposed convolution layer and 2 pointwise convolution layers connected in sequence, and 1 pointwise convolution layer for residual connection. The convolution layer adopts the third inverted residual bottleneck structure and the fourth inverted residual bottleneck structure connected in sequence. The third inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence, and one pointwise convolution layer for residual connection. The fourth inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence. Specifically, by introducing the inverted residual bottleneck structure driven by Transformer into the 3D U-Net network, a depthwise convolution or depthwise transposed convolution is relocated to the first layer to reduce the number of channels, and then two pointwise convolutions are performed for expansion and projection operations. This structural design aims to perform complex and inefficient convolution operations before efficient convolution operations to improve memory usage efficiency and reduce excessive parameters. At the same time, residual connections of pointwise convolution or pointwise transposed convolution are also adopted, which further promotes the propagation of gradients in the convolution layer and realizes easier gradient flow.
[0022] In summary, in the above brain tumor image segmentation model, both downsampling and upsampling use an improved inverted residual bottleneck structure to reduce memory usage while maintaining the semantic richness of 3D multimodal tumor information. Compared with the existing convolutional neural network-based segmentation model and ViT-based segmentation model, the above technical solution can improve the segmentation accuracy of the brain tumor image segmentation model while reducing the model's computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0024] Figure 1Schematic flowchart of a brain tumor image segmentation method provided in an embodiment of the present application; Figure 2 Schematic network architecture diagram of a brain tumor image segmentation model provided in an embodiment of the present application; Figure 3 Schematic diagram of an inverted residual bottleneck structure provided in an embodiment of the present application, where Fig. (a) is a schematic diagram of the structure of a convolutional block in ResNeXt, Fig. (b) is a schematic diagram of the inverted residual bottleneck structure proposed by MobileNetV2, and Fig. (c) is a schematic diagram of the inverted residual bottleneck structure driven by Transformer; Figure 4 Schematic diagrams of a down-convolution layer, an up-convolution layer, and a convolution layer provided in an embodiment of the present application; Figure 5 Schematic diagram of a brain tumor image segmentation system provided in an embodiment of the present application. Detailed implementation manners
[0025] In order to enable those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0026] In the embodiments provided by the present application, it should be understood that the disclosed methods and systems can be implemented in other ways. The system embodiments described below are only illustrative. For example, the division of units and modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or modules can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces, indirect coupling or communication connection of devices or modules, and can be electrical, mechanical, or other forms.
[0027] In addition, each functional unit in the embodiments of the present application can be all integrated in a processor, or each unit can be separately used as a device, or two or more units can be integrated in a device; each functional unit in the embodiments of the present application can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0028] A person skilled in the art can understand that all or part of the steps of the following method embodiments can be completed by program instructions and related hardware. The aforementioned program instructions can be stored in a computer-readable storage medium. When the program instructions are executed, the steps of the following method embodiments are executed; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a magnetic disk or an optical disk, and other media that can store program codes.
[0029] It should be understood that the use of "system", "device", "unit" and / or "module" in this application is only a method for distinguishing different components, elements, parts, parts or assemblies at different levels. However, if other words can achieve the same purpose, the word can be replaced by other expressions.
[0030] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, "multiple" and "several" mean two or more, unless otherwise clearly and specifically defined.
[0031] If a flow chart is used in the present application, the flow chart is used to illustrate the operations performed by the system according to the embodiment of the present application. It should be understood that the preceding or following operations are not necessarily performed accurately in order. On the contrary, each step can be processed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or a certain step or several steps of operations can be removed from these processes.
[0032] It should also be noted that, in this article, terms such as "comprises", "includes" or any other variations thereof are intended to cover non-exclusive inclusion, so that an article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such articles or devices. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the article or device including the above elements.
[0033] The embodiments of the present application are written in a progressive manner.
[0034] like Figure 1 As shown, the embodiment of the present application provides a brain tumor image segmentation method, comprising: S101. Obtain a pre-trained brain tumor image segmentation model; In S101, specifically, the brain tumor image segmentation model is constructed based on a U-shaped encoder-decoder structure of a 3D U-Net network, including an encoder consisting of one input convolution layer and four down-convolution modules connected in sequence, and a decoder consisting of four up-convolution modules connected in sequence and one output convolution layer, wherein the output of each down-convolution module in the encoder is jump-connected to the input of the corresponding up-convolution module in the decoder; Among them, 3D U-Net is an architecture based on convolutional neural networks, which is specially used for the segmentation task of three-dimensional volume images, especially in the field of medical imaging. 3D U-Net is an extended version of U-Net, inheriting the encoder-decoder structure of U-Net and optimized for three-dimensional data. Its main features are as follows: Encoder (Contracting Path): The encoder part gradually reduces the spatial dimension through convolutional layers and maximum pooling layers, while increasing the number of feature channels, thereby extracting high-level features and contextual information of the image. Decoder (Expansive Path): The decoder part gradually restores the spatial resolution of the image through upsampling (or transposed convolution), and splices the features in the encoder with the features in the decoder through skip connections, thereby retaining spatial detail information. Skip connections: Skip connections connect the corresponding layers of the encoder and decoder, allowing the network to simultaneously utilize high-level features and low-level detail information, thereby achieving more accurate segmentation.
[0035] In this embodiment, combined with Figure 2 As shown in the figure, the brain tumor image segmentation model is built based on the U-shaped encoder-decoder structure of the 3D U-Net network. The encoder and decoder perform downsampling and upsampling operations, respectively, for tumor feature extraction and reconstruction, respectively.
[0036] In the encoder, an input convolution layer (denoted as InConv) is retained to preprocess brain tumor features. The input convolution layer can be Figure 2 The convolutional network structure shown in or other existing convolutional network structures are not specifically limited in this application; 4 down-convolution modules (denoted as DownTDBlock) are used to extract multimodal feature representations from brain tumor MR images; through downsampling, the resolution of the features is gradually reduced, while the number of channels in the feature map is increased.
[0037] In the decoder, four up-convolution modules (denoted as UpTDBlock) are used to gradually restore the features extracted by the encoder to the original resolution of the input image. The decoder also fuses the high-resolution feature map from the encoder with the low-resolution feature map from the decoder through skip connections to retain more detailed tumor information. Finally, the image segmentation result is output through the output convolution layer (denoted as OutConv). The output convolution layer can be used Figure 2 The convolutional network structure shown in or other existing convolutional network structures are not specifically limited in this application. By upsampling, the resolution of features is gradually improved, and the number of channels in the feature map is reduced.
[0038] In addition, in order to further improve the segmentation accuracy of the brain tumor image segmentation model and reduce the model calculation amount, in this embodiment, the applicant introduces the Transformer-driven inverted residual bottleneck structure into the 3D U-Net network.
[0039] Specifically, combined Figure 3 As shown in the figure, Figure (a) is a structural diagram of the convolution block in ResNeXt; Figure (b) is a schematic diagram of the inverted residual bottleneck structure proposed by MobileNetV2; Figure (c) is a schematic diagram of the inverted residual bottleneck structure driven by Transformer; Dw represents the depthwise convolution layer, ReLU (Rectified Linear Unit), ReLU6 (Rectified Linear Unit 6) and GELU (Gaussian Error Linear Unit) are activation functions, GN (Group Normalization) and BN (Batch Normalization) are group normalization and batch normalization respectively, 1×1×1 and 3×3×3 are the convolution kernel sizes, 64→256 represents the number of input feature and output feature channels is 64 and 256, 64→64, 256→256 and 256→256 are not repeated here.
[0040] from Figure 3It can be seen that the bottleneck layer (projection layer, i.e. the dark part) in MobileNetV2 contains all the necessary information, while the expansion layer (i.e. the increase in the number of channels) only provides the implementation of the nonlinear transformation of the tensor. MobileNetV2 uses jump connections directly between bottlenecks to promote gradient propagation. Inspired by MobileNet, the applicant uses the Transformer-driven inverted residual bottleneck to improve memory usage efficiency and reduce excessive parameters. The 3×3 convolution kernel in the conventional convolutional network is the gold standard and provides the best performance on GPU hardware. The non-local self-attention mechanism in the Transformer architecture allows it to better model global dependencies. In order to better model the global information of the tumor and retain the inductive bias present in CNN, the applicant made the following adjustments to the inverted residual bottleneck: reposition a depthwise convolution to the first layer, reduce the number of channels, and then perform two pointwise convolutions for expansion and projection operations. The design aims to perform complex and inefficient convolution operations before efficient convolution operations, which is inspired by the Transformer-based architecture, where the MSA (Multi-head self-attention) block is placed before the MLP (multi-layer perceptron) layer. It also uses residual connections of pointwise convolution or transposed convolution. This further promotes the propagation of gradients in the convolution layer and achieves easier gradient flow.
[0041] Specifically, in this embodiment, the lower convolution module is composed of 1 lower convolution layer and 2 convolution layers, and the upper convolution module is composed of 1 upper convolution layer and 2 convolution layers; the lower convolution layer adopts the first inverted residual bottleneck structure, which includes 1 depthwise convolution layer and 2 pointwise convolution layers connected in sequence, and 1 pointwise convolution layer for residual connection; the upper convolution layer adopts the second inverted residual bottleneck structure, which includes 1 depthwise transposed convolution layer and 2 pointwise convolution layers connected in sequence, and 1 pointwise transposed convolution layer for residual connection; the convolution layer adopts the third inverted residual bottleneck structure and the fourth inverted residual bottleneck structure connected in sequence, which includes 1 depthwise convolution layer and 2 pointwise convolution layers connected in sequence, and 1 pointwise convolution layer for residual connection, and the fourth inverted residual bottleneck structure includes 1 depthwise convolution layer and 2 pointwise convolution layers connected in sequence.
[0042] Specifically, the brain tumor image segmentation model consists of three main convolutional modules: convolutional layer (denoted as TDConv), down convolutional layer (denoted as DownTDConv) and up convolutional layer (denoted as UpTDConv). The cross-stage calculation distribution based on 3D U-Net is adopted to form the stage calculation ratio of the encoder and decoder, that is, the number of inverted residual bottlenecks in each convolutional layer is (2,2,2,2).
[0043] The encoder includes 4 down-convolution modules DownTDBlock, each of which consists of 1 down-convolution layer DownTDConv and 2 convolution layers TDConv. The down-convolution layer DownTDConv is used for feature dimension reduction, and the convolution layer TDConv is used to extract brain tumor features. The decoder includes 4 up-convolution modules UpTDBlock, each of which consists of 1 up-convolution layer UpTDConv and 2 convolution layers TDConv. The up-convolution layer UpTDConv is used for feature recovery, and the resolution is gradually adjusted back to the size of the original brain tumor image. The convolution layer TDConv is applied to the feature fusion of the encoder and decoder networks. The jump connection integrates the high-resolution image from the encoder with the upsampled low-resolution image of the decoder, thereby retaining more tumor details and improving segmentation accuracy. The following table summarizes the detailed network structure of the proposed brain tumor image segmentation model.
[0044] Table 1 Detailed network structure of brain tumor image segmentation model
[0045] In the table, DB 1 Indicates the first down-convolution module DownTDBlock, DB connected in sequence in the encoder 2 Indicates the second down-convolution module DownTDBlock, DB connected in sequence in the encoder 3 Indicates the third down-convolution module DownTDBlock, DB connected in sequence in the encoder 4 Indicates the fourth down-convolution module DownTDBlock connected in sequence in the encoder; UB 1 Indicates the first up-convolution module UpTDBlock, UB connected in sequence in the decoder 2 Indicates the second up-convolution module UpTDBlock, UB connected in sequence in the decoder 3 Indicates the third up-convolution module UpTDBlock, UB connected in sequence in the decoder 4 represents the fourth up-convolution module UpTDBlock connected in sequence in the decoder. Stride represents the step size of the convolution module, R represents the expansion ratio, and C OutIndicates the number of output channels.
[0046] The input-output relationship of the down-convolution module DownTDBlock can be expressed by the following expression: ; In the formula, represents the output feature map of the i-th down-convolution module DownTDBlock (i.e., ) connected in sequence in the encoder, represents the output feature map of the (i + 1)-th down-convolution module DownTDBlock (i.e., ) connected in sequence in the encoder, represents the data processing operation of the (i + 1)-th down-convolution module DownTDBlock connected in sequence in the encoder, represents the data processing operation of the convolutional layer TDConv, represents the data processing operation of the down-convolutional layer DownTDConv; The input-output relationship of the up-convolution module UpTDBlock can be expressed by the following expression: ; In the formula, represents the output feature map of the i-th up-convolution module UpTDBlock (i.e., ) connected in sequence in the decoder, represents the output feature map of the (i + 1)-th up-convolution module UpTDBlock (i.e., ) connected in sequence in the decoder, represents the output feature map of the (4 - i)-th down-convolution module DownTDBlock (i.e., ) connected in sequence in the encoder, represents the data processing operation of the (i + 1)-th up-convolution module UpTDBlock connected in sequence in the decoder, represents the data processing operation of the up-convolutional layer UpTDConv, represents a set of features.
[0047] In this embodiment, for the initially constructed brain tumor image segmentation model, existing brain tumor image samples, such as an existing brain tumor image dataset composed of multi-modal MR images, can be pre-employed, and based on existing model training methods, the brain tumor image segmentation model can be trained to obtain a trained brain tumor image segmentation model. This application does not make specific restrictions on this.
[0048] S102. Obtain the brain tumor image to be segmented, and use the brain tumor image segmentation model to perform image segmentation on the brain tumor image to obtain a segmentation result.
[0049] In S102, specifically, after obtaining the trained brain tumor image segmentation model, the brain tumor image to be segmented can be input into the brain tumor image segmentation model, so as to use the model to perform image segmentation on the brain tumor image to obtain a segmentation result.
[0050] Currently, convolutional neural network-based segmentation models cannot simulate long-range dependencies, which hinders their perception of global information about tumors. In addition, ViT-based segmentation models have high computational resource requirements and require a large number of annotated datasets to obtain optimal segmentation performance, resulting in high computational costs and overfitting on small datasets.
[0051] In the above embodiment, brain tumor image segmentation is performed through a pre-trained brain tumor image segmentation model, wherein the brain tumor image segmentation model is constructed based on a U-shaped encoder-decoder structure of a 3D U-Net network, including an encoder consisting of one input convolution layer and four down-convolution modules connected in sequence, and a decoder consisting of four up-convolution modules connected in sequence and one output convolution layer, and the output of each down-convolution module in the encoder is jump-connected to the input of the corresponding up-convolution module in the decoder; specifically, the high-resolution image from the encoder is integrated with the upsampled low-resolution image of the decoder through the jump connection, thereby retaining more tumor details and improving the segmentation accuracy.
[0052] In addition, the lower convolution module consists of 1 lower convolution layer and 2 convolution layers, and the upper convolution module consists of 1 upper convolution layer and 2 convolution layers; the lower convolution layer adopts the first inverted residual bottleneck structure, which includes 1 depthwise convolution layer and 2 pointwise convolution layers connected in sequence, and 1 pointwise convolution layer for residual connection; the upper convolution layer adopts the second inverted residual bottleneck structure, which includes 1 depthwise transposed convolution layer and 2 pointwise convolution layers connected in sequence, and 1 pointwise convolution layer for residual connection. The convolution layer adopts the third inverted residual bottleneck structure and the fourth inverted residual bottleneck structure connected in sequence. The third inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence, and one pointwise convolution layer for residual connection. The fourth inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence. Specifically, by introducing the inverted residual bottleneck structure driven by Transformer into the 3D U-Net network, a depthwise convolution or depthwise transposed convolution is relocated to the first layer to reduce the number of channels, and then two pointwise convolutions are performed for expansion and projection operations. This structural design aims to perform complex and inefficient convolution operations before efficient convolution operations to improve memory usage efficiency and reduce excessive parameters. At the same time, residual connections of pointwise convolution or pointwise transposed convolution are also adopted, which further promotes the propagation of gradients in the convolution layer and realizes easier gradient flow.
[0053] In summary, in the above brain tumor image segmentation model, both downsampling and upsampling use an improved inverted residual bottleneck structure to reduce memory usage while maintaining the semantic richness of 3D multimodal tumor information. Compared with the existing convolutional neural network-based segmentation model and ViT-based segmentation model, the above embodiment can improve the segmentation accuracy of the brain tumor image segmentation model while reducing the model calculation amount.
[0054] In other embodiments of the present application, in the first inverted residual bottleneck structure, the input features are input into the depthwise convolution layer and the pointwise convolution layer for residual connection, the output of the depthwise convolution layer is group normalized and input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise convolution layer for residual connection to obtain the output feature of the first inverted residual bottleneck structure.
[0055] Specifically, combined Figure 4 As shown in Figure 2, the down convolution layer DownTDConv adopts the first inverted residual bottleneck structure, including a depthwise convolution layer (Dw) with a stride of 2 and two pointwise convolution layers (Cov 1×1×1 , the convolution kernel size is 1×1×1), the number of channels is C In , C In ×R and C In R is the channel expansion ratio of the second pointwise convolution layer. The residual connection in the down convolution layer DownTDConv uses a pointwise convolution layer with a stride of 2. Its input-output relationship can be expressed as follows: ; In the formula, represents the input features of the lower convolutional layer DownTDConv, Represents the output features of the lower convolutional layer DownTDConv, Represents the data processing operation of the pointwise convolution layer in the down convolution layer DownTDConv, Represents the data processing operation of the GELU activation function, represents the data processing operation of group normalization, Represents the data processing operation of the depthwise convolution layer in the down convolution layer DownTDConv, Represents the output features of the pointwise convolution layer used for residual connection in the down convolution layer DownTDConv.
[0056] In other embodiments of the present application, in the second inverted residual bottleneck structure, the input features are input into the depthwise transposed convolution layer and the pointwise transposed convolution layer for residual connection, the output of the depthwise transposed convolution layer is group normalized and input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise transposed convolution layer for residual connection to obtain the output feature of the second inverted residual bottleneck structure.
[0057] Specifically, combined Figure 4 As shown in Figure 2, the upper convolution layer UpTDConv adopts the second inverted residual bottleneck structure, which includes a depthwise transposed convolution layer (DwTransCov) with a stride of 2, followed by two pointwise convolution layers (TransCov 1×1×1 , the convolution kernel size is 1×1×1, a TransCov layer consists of two pointwise convolution layers, and the pointwise convolution layer is a regular convolution layer with a convolution kernel size of 1×1), the number of channels is C In , C In ×R and C In . R is the channel expansion ratio of the second pointwise convolution layer. The residual connection in the upper convolution layer UpTDConv uses a pointwise transposed convolution layer with a stride of 2. Among them, the pointwise transposed convolution layer and the depthwise transposed convolution layer are the pointwise and depthwise versions of the transposed convolution, respectively, corresponding to the transposed convolution with a convolution kernel size of 1×1 and the parameter GROUPS set to the number of channels C. The decoder uses skip connections to connect the output features of each downsampling stage in the encoder with the corresponding upsampled output features in the corresponding upper convolution layer UpTDConv in the channel dimension to facilitate subsequent upsampling. Its input-output relationship can be expressed by the following expression: ; In the formula, Represents the input features of the upper convolution layer UpTDConv, Represents the output features of the upper convolution layer UpTDConv, Represents the data processing operation of the pointwise convolution layer in the upper convolution layer UpTDConv. Represents the data processing operation of the GELU activation function, represents the data processing operation of group normalization, Represents the data processing operation of the depthwise transposed convolution layer in the upper convolution layer UpTDConv, Represents the output features of the pointwise transposed convolution layer used for residual connection in the upper convolution layer UpTDConv.
[0058] In other embodiments of the present application, in the third inverted residual bottleneck structure, the input features are input into the depthwise convolution layer and the pointwise convolution layer for residual connection, the output of the depthwise convolution layer is group normalized and input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise convolution layer for residual connection to obtain the output features of the third inverted residual bottleneck structure.
[0059] In the fourth inverted residual bottleneck structure, the output features of the previously connected third inverted residual bottleneck structure are input into the depthwise convolution layer. The output of the depthwise convolution layer is group normalized and input into the first pointwise convolution layer connected in sequence. The output of the first pointwise convolution layer is processed by the GELU activation function and input into the second pointwise convolution layer connected in sequence. The output of the second pointwise convolution layer is added to the input of the depthwise convolution layer as the output feature of the fourth inverted residual bottleneck structure.
[0060] Specifically, combined Figure 4 As shown in Figure 1, the convolutional layer TDConv uses two inverted residual bottleneck structures connected in sequence: the first inverted residual bottleneck structure in the convolutional layer TDConv, that is, the third inverted residual bottleneck structure, includes a depthwise convolution layer (DwCov) and two pointwise convolution layers (Cov 1×1×1 , the convolution kernel size is 1×1×1), and the number of channels is C In , C In ×R and C Out , stride is 1. R is the channel expansion ratio of the second pointwise convolution layer. For activation function and normalization, the operation is simplified by using group normalization GN after the first depthwise convolution layer and applying GELU activation function after the first pointwise convolution layer. The third inverted residual bottleneck structure uses C channels in the residual connection. OutThe second inverted residual bottleneck structure in the convolution layer TDConv, that is, the fourth inverted residual bottleneck structure, also includes a depthwise convolution layer (DwCov) and two pointwise convolution layers (Cov1×1×1, with a convolution kernel size of 1×1×1). The key difference between it and the third inverted residual bottleneck structure is that the input features (that is, the input features of the depthwise convolution layer in the fourth inverted residual bottleneck structure) are directly added to the output features of the second pointwise convolution layer to obtain the final output features of the convolution layer TDConv.
[0061] The input-output relationship of the convolutional layer TDConv can be expressed as follows: ; ; In the formula, represents the input features of the convolutional layer TDConv, Represents the output features of the convolutional layer TDConv, Represents the data processing operation of the pointwise convolution layer in the convolution layer TDConv. Represents the data processing operation of the GELU activation function, represents the data processing operation of group normalization, Represents the data processing operation of the depthwise convolution layer in the convolution layer TDConv. Represents the output features of the pointwise convolution layer used for residual connection in the convolution layer TDConv, Represents the intermediate features of the convolutional layer TDConv, that is, the output features of the third inverted residual bottleneck structure.
[0062] In the above embodiment, in view of the excellent performance of the GELU activation function in the Transformer, the applicant uses GELU instead of ReLU to provide stronger smoothness and nonlinearity, thereby enhancing the adaptability of the segmentation scheme. In addition, fewer activation functions are used in the brain tumor image segmentation model. Among them, the GELU activation function is approximated as follows: ; In the formula, Represents the input of the GELU activation function, and the output range of the GELU activation function is ,when When positive, the output is close to ;when When negative, the output is close to 0.
[0063] In addition, due to the high information content and memory requirements of 3D MR images, larger batch sizes cannot be used when training brain tumor image segmentation models. In this case, the applicant replaced batch normalization in 3D U-Net with group normalization, which is more suitable for small batch training and reduces the utilization of normalization layers.
[0064] In summary, the above brain tumor image segmentation model uses fewer activation functions and normalization layers in the downsampling and upsampling blocks, and combines GELU activation and group normalization to achieve a more concise and accurate segmentation of brain tumors.
[0065] In other embodiments of the present application, inspired by the Transformer architecture, the applicant uses larger convolution kernels in the brain tumor image segmentation model to capture global dependencies in the tumor image to enhance the long-distance modeling capability of the pure convolutional network while ensuring the extraction of local features. In the segmentation network, an adjustable convolution kernel is used to improve the segmentation accuracy. By experimenting with convolution kernels of different sizes, it is found that under a small training set, the 7×7 convolution kernel reaches performance saturation in tumor segmentation. In other embodiments of the present application, the convolution kernel size of all convolution kernels in the brain tumor image segmentation model is 7×7.
[0066] In other embodiments of the present application, one implementation of the step of obtaining a pre-trained brain tumor image segmentation model specifically includes: S201. Obtain a brain tumor image sample; In S201, specifically, the brain tumor image samples may adopt an existing brain tumor image dataset composed of multi-modal MR images, and the present application does not impose any specific limitation on this.
[0067] S202. construct a hybrid loss function composed of a cross entropy loss function, a Dice loss function and an orthogonal regularization loss function; In S202, specifically, in order to facilitate back propagation, orthogonal regularization is used to constrain the model weights, and the cross-entropy loss function and the Dice loss function are combined to construct a hybrid loss function in a weighted manner. The specific expression is: ; In the formula, represents the mixed loss function, represents the cross entropy loss function, represents the Dice loss function, represents the orthogonal regularization loss function, and Indicates adjustment weight.
[0068] Among them, the calculation formula of the cross entropy loss function is as follows: ; In the formula, Represents the number of categories, is an indicator function (0 or 1), if the sample The true category is equal to , then it takes 1, otherwise it takes 0. Represents the prediction sample Belongs to category probability.
[0069] Among them, the calculation formula of Dice loss function is as follows: ; In the formula, is the tumor area in the MR image location, To segment the tumor results, is the true value of the tumor area, is the smoothing parameter.
[0070] Among them, the calculation formula of the orthogonal regularization loss function is as follows: ; In the formula, represents the sum of all filter banks in the neural network, represents a filter bank, yes The transpose of Represents the identity matrix.
[0071] S203. Using the brain tumor image samples and the mixed loss function, the initial brain tumor image segmentation model is trained to obtain a trained brain tumor image segmentation model.
[0072] In S203, specifically, the brain tumor image samples are divided according to the set division ratio to obtain the training set, the test set and the validation set, and then the model training parameters are set, such as the number of epochs for training all data, the batch size, etc., random initialization is enabled, and the Adam optimization algorithm is used for training. The specific training process is not repeated here. The model in each training process is retained, and the mixed loss function is used to evaluate the model. Finally, the model with the best evaluation index is saved as the trained brain tumor image segmentation model.
[0073] In this embodiment, an orthogonal regularization loss is used to impose constraints on the model weights, which reduces the risk of overfitting, enhances the generalization of tumor segmentation, reduces the performance degradation caused by overfitting, and can achieve better brain tumor segmentation accuracy with fewer model parameters.
[0074] like Figure 5 As shown, in another embodiment of the present application, a brain tumor image segmentation system is provided, comprising: An acquisition unit 10 is used to acquire a pre-trained brain tumor image segmentation model, wherein the brain tumor image segmentation model is constructed based on a U-shaped encoder-decoder structure of a 3D U-Net network, including an encoder consisting of one input convolution layer and four down-convolution modules connected in sequence, and a decoder consisting of four up-convolution modules connected in sequence and one output convolution layer, wherein the output of each down-convolution module in the encoder is jump-connected to the input of the corresponding up-convolution module in the decoder; The down-convolution module consists of 1 down-convolution layer and 2 convolution layers, and the up-convolution module consists of 1 up-convolution layer and 2 convolution layers; The lower convolution layer adopts the first inverted residual bottleneck structure, which includes one depthwise convolution layer and two pointwise convolution layers connected in sequence, and one pointwise convolution layer for residual connection; The upper convolution layer adopts the second inverted residual bottleneck structure, which includes one depthwise transposed convolution layer and two pointwise convolution layers connected in sequence, and one pointwise transposed convolution layer for residual connection; The convolution layer adopts the third inverted residual bottleneck structure and the fourth inverted residual bottleneck structure connected in sequence. The third inverted residual bottleneck structure includes 1 depthwise convolution layer and 2 pointwise convolution layers connected in sequence, and 1 pointwise convolution layer for residual connection. The fourth inverted residual bottleneck structure includes 1 depthwise convolution layer and 2 pointwise convolution layers connected in sequence. The segmentation unit 11 is used to obtain a brain tumor image to be segmented, and perform image segmentation on the brain tumor image using a brain tumor image segmentation model to obtain a segmentation result.
[0075] In other embodiments of the present application, in the above-mentioned brain tumor image segmentation system, in the first inverted residual bottleneck structure, the input feature is input into the depthwise convolution layer and the pointwise convolution layer for residual connection, the output of the depthwise convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise convolution layer for residual connection to obtain the output feature of the first inverted residual bottleneck structure; In other embodiments of the present application, in the above-mentioned brain tumor image segmentation system, in the second inverted residual bottleneck structure, the input feature is input into the depthwise transposed convolution layer and the pointwise transposed convolution layer for residual connection, the output of the depthwise transposed convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise transposed convolution layer for residual connection, and the output feature of the second inverted residual bottleneck structure is obtained; In other embodiments of the present application, in the above-mentioned brain tumor image segmentation system, in the third inverted residual bottleneck structure, the input feature is input into the depthwise convolution layer and the pointwise convolution layer for residual connection, the output of the depthwise convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise convolution layer for residual connection to obtain the output feature of the third inverted residual bottleneck structure; In other embodiments of the present application, in the above-mentioned brain tumor image segmentation system, in the fourth inverted residual bottleneck structure, the output features of the previously connected third inverted residual bottleneck structure are input into the depthwise convolution layer, the output of the depthwise convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the input of the depthwise convolution layer as the output feature of the fourth inverted residual bottleneck structure.
[0076] In other embodiments of the present application, in the above-mentioned brain tumor image segmentation system, the convolution kernel size of the brain tumor image segmentation model is 7×7.
[0077] In other embodiments of the present application, in the above-mentioned brain tumor image segmentation system, when the acquisition unit 10 executes the acquisition of the pre-trained brain tumor image segmentation model, it is specifically used to: Obtain sample brain tumor images; Construct a hybrid loss function composed of the cross entropy loss function, the Dice loss function and the orthogonal regularization loss function; The initial brain tumor image segmentation model is trained using brain tumor image samples and a mixed loss function to obtain a trained brain tumor image segmentation model.
[0078] In other embodiments of the present application, in the above-mentioned brain tumor image segmentation system, the hybrid loss function is specifically: ; In the formula, represents the mixed loss function, represents the cross entropy loss function, represents the Dice loss function, represents the orthogonal regularization loss function, and Indicates adjustment weight.
[0079] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A brain tumor image segmentation method, characterized in that: include: Obtaining a pre-trained brain tumor image segmentation model, wherein the brain tumor image segmentation model is constructed based on a U-shaped encoder-decoder structure of a 3D U-Net network, including an encoder consisting of one input convolution layer and four lower convolution modules connected in sequence, and a decoder consisting of four upper convolution modules connected in sequence and one output convolution layer, wherein the output of each lower convolution module in the encoder is jump-connected to the input of the corresponding upper convolution module in the decoder; The lower convolution module is composed of 1 lower convolution layer and 2 convolution layers, and the upper convolution module is composed of 1 upper convolution layer and 2 convolution layers; The lower convolution layer adopts a first inverted residual bottleneck structure, which includes one depthwise convolution layer and two pointwise convolution layers connected in sequence, and one pointwise convolution layer for residual connection; The upper convolution layer adopts a second inverted residual bottleneck structure, which includes one depthwise transposed convolution layer and two pointwise convolution layers connected in sequence, and one pointwise transposed convolution layer for residual connection; The convolution layer adopts a third inverted residual bottleneck structure and a fourth inverted residual bottleneck structure connected in sequence, wherein the third inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence, and one pointwise convolution layer for residual connection, and the fourth inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence; A brain tumor image to be segmented is acquired, and the brain tumor image is segmented using the brain tumor image segmentation model to obtain a segmentation result.
2. The method according to claim 1, characterized in that In the first inverted residual bottleneck structure, the input feature is input into the depthwise convolution layer and the pointwise convolution layer for residual connection. The output of the depthwise convolution layer is group normalized and input into the first pointwise convolution layer connected in sequence. The output of the first pointwise convolution layer is processed by the GELU activation function and input into the second pointwise convolution layer connected in sequence. The output of the second pointwise convolution layer is added to the output of the pointwise convolution layer for residual connection to obtain the output feature of the first inverted residual bottleneck structure.
3. The method according to claim 1, characterized in that In the second inverted residual bottleneck structure, the input features are input into the depthwise transposed convolution layer and the pointwise transposed convolution layer for residual connection. The output of the depthwise transposed convolution layer is group normalized and input into the first pointwise convolution layer connected in sequence. The output of the first pointwise convolution layer is processed by the GELU activation function and input into the second pointwise convolution layer connected in sequence. The output of the second pointwise convolution layer is added to the output of the pointwise transposed convolution layer for residual connection to obtain the output features of the second inverted residual bottleneck structure.
4. The method according to claim 1, characterized in that In the third inverted residual bottleneck structure, the input features are input into the depthwise convolution layer and the pointwise convolution layer for residual connection. The output of the depthwise convolution layer is group normalized and input into the first pointwise convolution layer connected in sequence. The output of the first pointwise convolution layer is processed by the GELU activation function and input into the second pointwise convolution layer connected in sequence. The output of the second pointwise convolution layer is added to the output of the pointwise convolution layer for residual connection to obtain the output features of the third inverted residual bottleneck structure.
5. The method as claimed in claim 4, characterized in that In the fourth inverted residual bottleneck structure, the output features of the previously connected third inverted residual bottleneck structure are input into the depthwise convolution layer, the output of the depthwise convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the input of the depthwise convolution layer as the output feature of the fourth inverted residual bottleneck structure.
6. The method according to claim 1, characterized in that The convolution kernel size of the brain tumor image segmentation model is 7×7.
7. The method according to claim 1, characterized in that The step of obtaining a pre-trained brain tumor image segmentation model comprises: Obtain brain tumor image samples; Construct a hybrid loss function composed of the cross entropy loss function, the Dice loss function and the orthogonal regularization loss function; The initial brain tumor image segmentation model is trained using the brain tumor image samples and the mixed loss function to obtain the trained brain tumor image segmentation model.
8. The method as claimed in claim 7, characterized in that The hybrid loss function is specifically: ; In the formula, represents the hybrid loss function, represents the cross entropy loss function, represents the Dice loss function, represents the orthogonal regularization loss function, and Indicates adjustment weight.
9. A brain tumor image segmentation system, characterized in that: include: An acquisition unit is used to acquire a pre-trained brain tumor image segmentation model, wherein the brain tumor image segmentation model is constructed based on a U-shaped encoder-decoder structure of a 3D U-Net network, including an encoder consisting of one input convolution layer and four lower convolution modules connected in sequence, and a decoder consisting of four upper convolution modules connected in sequence and one output convolution layer, wherein the output of each lower convolution module in the encoder is jump-connected to the input of the corresponding upper convolution module in the decoder; The lower convolution module is composed of 1 lower convolution layer and 2 convolution layers, and the upper convolution module is composed of 1 upper convolution layer and 2 convolution layers; The lower convolution layer adopts a first inverted residual bottleneck structure, which includes one depthwise convolution layer and two pointwise convolution layers connected in sequence, and one pointwise convolution layer for residual connection; The upper convolution layer adopts a second inverted residual bottleneck structure, which includes one depthwise transposed convolution layer and two pointwise convolution layers connected in sequence, and one pointwise transposed convolution layer for residual connection; The convolution layer adopts a third inverted residual bottleneck structure and a fourth inverted residual bottleneck structure connected in sequence, wherein the third inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence, and one pointwise convolution layer for residual connection, and the fourth inverted residual bottleneck structure includes one depthwise convolution layer and two pointwise convolution layers connected in sequence; The segmentation unit is used to obtain a brain tumor image to be segmented, and perform image segmentation on the brain tumor image using the brain tumor image segmentation model to obtain a segmentation result.
10. The system as claimed in claim 9, characterized in that In the first inverted residual bottleneck structure, input features are input into a depthwise convolution layer and a pointwise convolution layer for residual connection, the output of the depthwise convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by a GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise convolution layer for residual connection, and the output features of the first inverted residual bottleneck structure are obtained; In the second inverted residual bottleneck structure, the input feature is input into the depthwise transposed convolution layer and the pointwise transposed convolution layer for residual connection, the output of the depthwise transposed convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise transposed convolution layer for residual connection, and the output feature of the second inverted residual bottleneck structure is obtained; In the third inverted residual bottleneck structure, the input feature is input into the depthwise convolution layer and the pointwise convolution layer for residual connection, the output of the depthwise convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the output of the pointwise convolution layer for residual connection to obtain the output feature of the third inverted residual bottleneck structure; In the fourth inverted residual bottleneck structure, the output features of the previously connected third inverted residual bottleneck structure are input into the depthwise convolution layer, the output of the depthwise convolution layer is group normalized and then input into the first pointwise convolution layer connected in sequence, the output of the first pointwise convolution layer is processed by the GELU activation function and then input into the second pointwise convolution layer connected in sequence, the output of the second pointwise convolution layer is added to the input of the depthwise convolution layer as the output feature of the fourth inverted residual bottleneck structure.
Citation Information
Patent Citations
Brain tumor image segmentation system based on Unet variant network
CN114372988A
Power load prediction method and system based on multi-scale and depth separable convolution
CN118228872A
Image segmentation method and system based on residual error reverse bottleneck and sparse attention
CN118967722A
Brain tumor segmentation method and device based on improved 3DU-Net, and electronic equipment
CN119832234A