Liver tumor segmentation method and device based on axial deep convolution and Transform
By using axial depth convolution and Transformer methods in liver tumor CT image segmentation, the problem of segmentation inaccuracy caused by blurred boundary and variable scales is solved, and the parameter quantity and calculation complexity of the model are reduced, achieving more efficient liver tumor segmentation performance.
Patent Information
- Application Number
- CN202510267538.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art faces the problem of segmentation incorrect segmentation caused by blurred boundary and variable scales in liver tumor CT image segmentation, as well as the calculation complexity caused by excessive model parameters.
Using a liver tumor segmentation method based on axial depth convolution and Transformer, a segmentation model including encoder, decoder and bottleneck layer is constructed, and axial depth convolution module is used to replace traditional convolution, and the interactive attention Transformer module and marker-aware MLP module are designed to improve feature learning capabilities.
The segmentation performance of liver tumor images is significantly improved, especially when dealing with fuzzy boundaries and multi-scale features, the segmentation performance is better than the existing technology, while significantly reducing the number of parameters and calculation complexity of the model, improving the segmentation efficiency.
Smart Images

Figure CN120219401A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision, and specifically, to a liver tumor segmentation method and device based on axial depth convolution and Transformer. Background Art
[0002] In the field of medical image segmentation, especially for the CT image segmentation of liver tumors, many challenges are often faced. The medical images of liver tumors generally present these characteristics: on the one hand, the characteristics of tumors are extremely similar to those of normal liver tissues, which makes the tumor boundaries blurred; on the other hand, the sizes of tumors are different and their shapes are irregular, thus resulting in variable scales.
[0003] In the prior art, some traditional segmentation methods, such as region growing method, threshold segmentation method, etc. are used to segment liver tumors. However, limited by the expressiveness of features of traditional segmentation methods, the segmentation of these models for liver tumors is difficult to reach an ideal performance level. Many liver tumor segmentation models have been developed based on deep learning technology in the prior art, and these models have achieved remarkable results in improving the segmentation accuracy. However, models with better segmentation performance often rely on more and more complex modules and continuously increase the depth of the network. Although the above-mentioned models have achieved significant improvements in the accuracy of liver tumor segmentation, due to the limitation of computing resources in the clinical environment, they are difficult to be directly applied to actual scenarios.
[0004] Therefore, there are mainly two problems in the prior art: one is the problem of inaccurate segmentation caused by blurred boundaries and variable scales; the other is the problem of too large model parameter quantity. Summary of the Invention
[0005] To overcome at least one deficiency in the prior art, this application provides a liver tumor segmentation method and device based on axial depth convolution and Transformer.
[0006] In a first aspect, a liver tumor segmentation method based on axial depth convolution and Transformer is provided, including:
[0007] Obtain a model training data set; the samples in the training data set are liver tumor images;
[0008] Train a liver tumor segmentation model based on the model training dataset to obtain a trained liver tumor segmentation model; the liver tumor segmentation model includes an encoder, a decoder, and a bottleneck layer; the encoder includes a plurality of sequentially connected encoding layers, a first token-aware MLP module, and an interaction-aware Transformer module, and the outputs of two encoding layers among the plurality of encoding layers and the output of the first token-aware MLP module are all input into the interaction-aware Transformer module for integration; the decoder includes a second token-aware MLP module and a plurality of sequentially connected decoding layers, the output of the interaction-aware Transformer module is input into the bottleneck layer for feature compression, and the output of the bottleneck layer is input into the second token-aware MLP module; the plurality of encoding layers and the plurality of decoding layers correspond one by one and there are skip connections;
[0009] Input the liver tumor image to be segmented into the trained liver tumor segmentation model to obtain a segmented image of the liver tumor image to be segmented.
[0010] In one embodiment, the encoding layer includes an axially depth convolutional layer, a first pointwise convolutional layer, a max pooling layer, and a GeLU activation function layer connected in sequence; the output of the axially depth convolutional layer is skip-connected to the corresponding decoding layer.
[0011] In one embodiment, the decoding layer includes an upsampling layer, a second pointwise convolutional layer, an axially depth convolutional layer, a third pointwise convolutional layer, and a GeLU activation function layer connected in sequence; the output of the upsampling layer is fused with the output of the axially depth convolutional layer in the encoding layer.
[0012] In one embodiment, the bottleneck layer includes a fourth pointwise convolutional layer, a plurality of parallelly distributed axially depth convolutional layers, a fifth pointwise convolutional layer, and a GeLU activation function layer connected in sequence; after the input of the bottleneck layer passes through the fourth pointwise convolutional layer, it is respectively input into the plurality of axially depth convolutional layers, the outputs of each axially depth convolutional layer are fused, and the fused features are input into the fifth pointwise convolutional layer.
[0013] In one embodiment, the interaction-aware Transformer module is used to respectively adopt cross-scale attention for the input multiple feature maps F3, F4, F5, and correspondingly obtain two query vectors Q 34 and Q 35 、 key-value pairs {K4, V4}, key-value pairs {K5, V5}, where K4 and V4 are the key vector and value vector in the key-value pair {K4, V4} respectively, and K5 and V5 are the key vector and value vector in the key-value pair {K5, V5} respectively; the resolution of the feature map F3 is greater than the resolution of the feature map F4, and the feature map F5 is the output of the first token-aware MLP module;
[0014] Transpose the feature map K4 and multiply it with Q 34 to obtain the first multiplication result; multiply the first multiplication result with V4 to obtain the second multiplication result;
[0015] Transpose K5 and multiply it with Q 35 to obtain the third multiplication result; multiply the third multiplication result with V5 to obtain the fourth multiplication result;
[0016] Fuse the second multiplication result and the fourth multiplication result, and input the fused features into a feed-forward network to obtain the output of the interactive perception Transformer module.
[0017] In one embodiment, the first token perception MLP module includes a Shifted MLP and a Norm layer. The Shifted MLP shifts the input in the width and height directions and processes it through an MLP. The output of the Shifted MLP is layer-normalized by the Norm layer to obtain normalized features. The normalized features are added to the input to obtain the output of the first token perception MLP module.
[0018] In a second aspect, a liver tumor segmentation device based on axial depth convolution and Transformer is provided, including:
[0019] A dataset acquisition module for acquiring a model training dataset; the samples in the training dataset are liver tumor images;
[0020] A model training module for training a liver tumor segmentation model based on the model training dataset to obtain a trained liver tumor segmentation model; the liver tumor segmentation model includes an encoder, a decoder, and a bottleneck layer; the encoder includes a plurality of consecutive encoding layers, a first token perception MLP module, and an interactive perception Transformer module. The outputs of two encoding layers among the plurality of encoding layers and the output of the first token perception MLP module are all input into the interactive perception Transformer module for integration; the decoder includes a second token perception MLP module and a plurality of decoding layers. The output of the interactive perception Transformer module is input into the bottleneck layer for feature compression, and the output of the bottleneck layer is input into the second token perception MLP module; there is a one-to-one correspondence and a skip connection between the plurality of encoding layers and the plurality of decoding layers;
[0021] A prediction module for inputting the liver tumor image to be segmented into the trained liver tumor segmentation model to obtain a segmentation image of the liver tumor image to be segmented.
[0022] In one embodiment, the encoding layer includes an axial depth convolutional layer, a first pointwise convolutional layer, a max pooling layer, and a GeLU activation function layer connected in sequence; the output of the axial depth convolutional layer is skip-connected to the corresponding decoding layer.
[0023] In one embodiment, the interaction-aware Transformer module is used to apply cross-scale attention to the input multiple feature maps F3, F4, and F5 respectively, and correspondingly obtain two query vectors Q 34 and Q 35 , key-value pairs {K4, V4}, key-value pairs {K5, V5}, where K4 and V4 are the key vector and value vector in the key-value pair {K4, V4} respectively, and K5 and V5 are the key vector and value vector in the key-value pair {K5, V5} respectively; the resolution of the feature map F3 is greater than that of the feature map F4, and the feature map F5 is the output of the first token-aware MLP module;
[0024] Transpose the feature map K4 and multiply it with Q 34 to obtain a first multiplication result; multiply the first multiplication result with V4 to obtain a second multiplication result;
[0025] Transpose K5 and multiply it with Q 35 to obtain a third multiplication result; multiply the third multiplication result with V5 to obtain a fourth multiplication result;
[0026] Fuse the second multiplication result and the fourth multiplication result, and input the fused feature into the feed-forward network to obtain the output of the interaction-aware Transformer module.
[0027] In one embodiment, the first token-aware MLP module includes a Shifted MLP and a Norm layer. The Shifted MLP shifts the input in the width and height directions and processes it through the MLP; the output of the Shifted MLP is layer-normalized by the Norm layer to obtain a normalized feature; the normalized feature is added to the input to obtain the output of the first token-aware MLP module.
[0028] Compared with the prior art, the present application has the following beneficial effects: The liver tumor segmentation method and device based on axial depth convolution and Transformer of the present application construct a liver tumor segmentation model, replace the original traditional convolution with an axial depth convolution module, reduce the computational complexity, design an interactive attention Transformer module to jointly utilize multiple small-scale feature maps, obtain richer global features with a lower computational complexity, and adopt a label-aware MLP module to process features of different dimensions through a shifted axial operation, enhancing the learning ability of local detailed information. The present application can effectively enhance the feature representation ability of liver tumors, significantly improve the feature learning effect of the model, especially when dealing with fuzzy boundaries and multi-scale features, and the segmentation performance is significantly better than that of the prior art. In addition, by introducing axial depth convolution, the present application significantly reduces the number of parameters and computational complexity of the liver tumor segmentation model, thereby effectively improving the segmentation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The present application can be better understood by referring to the description given below in conjunction with the accompanying drawings. The drawings, together with the following detailed description, are included in this specification and form a part of this specification. In the drawings:
[0030] Figure 1 shows a schematic diagram of the liver tumor segmentation model;
[0031] Figure 2 shows a schematic diagram of the encoding layer;
[0032] Figure 3 shows a schematic diagram of the decoding layer;
[0033] Figure 4 shows a schematic diagram of the bottleneck layer;
[0034] Figure 5 shows a schematic diagram of the interactive attention Transformer module;
[0035] Figure 6 shows a schematic diagram of the first label-aware MLP module;
[0036] Figure 7 shows a schematic diagram of the Shifted MLP. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] Exemplary embodiments of the present application will be described below in conjunction with the accompanying drawings. For clarity and conciseness, not all features of the actual embodiments are described in the specification. However, it should be understood that many specific decisions specific to the embodiments can be made during the development of any such actual embodiment to achieve the specific goals of the developer, and these decisions may vary with different embodiments.
[0038] Here, it should also be noted that in order to avoid obscuring the present application with unnecessary details, only the device structures closely related to the solution according to the present application are shown in the drawings, while other details less relevant to the present application are omitted.
[0039] It should be understood that the present application is not limited to the described embodiments only due to the following description with reference to the drawings. In this document, where feasible, embodiments can be combined with each other, features can be replaced or borrowed between different embodiments, and one or more features can be omitted in one embodiment.
[0040] An embodiment of the present application provides a liver tumor segmentation method based on axial depth convolution and Transformer, which mainly includes the following steps:
[0041] Step S1, obtain a model training data set; the samples in the training data set are liver tumor images.
[0042] First, perform data screening. Based on the gold standard (annotation mask) of the liver, determine the region range where the liver is located, and only retain the three-dimensional image data containing the liver and liver tumor parts. Since there are differences in slice spacing in the data collected by different CT devices, the present application uniformly adjusts the interval of the CT images of all cases along the Z-axis to 1 mm to ensure that the distance between slices is consistent. The original data contains liver labels. In view of the fact that the present application only focuses on liver tumors, the original labels are re-encoded, only the tumor labels are retained, and the liver is merged into the background area. The specific encoding is that 0 represents the background and 1 represents the liver tumor.
[0043] Then, resample the data. In order to make the input CT liver tumor images adapt to the model, they need to be reprocessed to adjust the images to the appropriate size required by the model. Since the convolutional neural network itself cannot understand the voxel spacing, resampling is performed on the medical liver tumor CT image data set. The purpose is to unify the images with different resolutions and voxel spacings to a consistent spatial scale for subsequent analysis and segmentation tasks. The resampled three-dimensional images are sliced. After uniformly adjusting the image size, a total of 25,472 two-dimensional images are obtained, including 18,538 normal images and 6,934 tumor images.
[0044] Then, data normalization. To improve the quality of two-dimensional CT slices and facilitate more efficient network training, this application uses window technology to preprocess the original liver CT images. Specifically, the HU (Hounsfield Unit) value range of the window is set to [-200, 200] to enhance the contrast and details of the images. This range is selected based on the typical HU value distribution of the liver tissue and its lesion areas, which can effectively highlight the detailed features of the liver and its tumors while suppressing the interference of background noise and other irrelevant information. Finally, to further optimize the image data, this application normalizes the pixel values of the images to the range of [0, 1].
[0045] Step S2: Train the liver tumor segmentation model based on the model training dataset to obtain the trained liver tumor segmentation model; Figure 1 The schematic diagram of the liver tumor segmentation model is shown. Refer to Figure 1 , train the liver tumor segmentation model based on the model training dataset to obtain the trained liver tumor segmentation model; the liver tumor segmentation model includes an encoder, a decoder, and a bottleneck layer; the encoder includes a plurality of encoding layers connected in sequence, a first token-aware MLP module, and an interaction-aware Transformer module. The outputs of two encoding layers and the output of the first token-aware MLP module in the plurality of encoding layers are all input into the interaction-aware Transformer module for integration; the decoder includes a second token-aware MLP module and a plurality of decoding layers connected in sequence. The output of the interaction-aware Transformer module is input into the bottleneck layer for feature compression, and the output of the bottleneck layer is input into the second token-aware MLP module; there is a one-to-one correspondence and skip connections between the plurality of encoding layers and the plurality of decoding layers.
[0046] Here, in the encoder and decoder, the last downsampling layer and the first upsampling layer are replaced with token-aware MLP. After mapping the convolutional features to the abstract token space, these tokens are processed by the MLP to extract information crucial for the segmentation task.
[0047] Design an interaction attention Transformer module to jointly utilize multiple small-scale feature maps to obtain richer global features with lower computational complexity. Use the token-aware MLP module to process features of different dimensions through shifted axial operations, which improves the learning ability of local detailed information.
[0048] In the decoder, perform upsampling operations on the downsampled data in sequence, and perform feature fusion through skip connections between the corresponding layers of the encoder and decoder, and finally output the segmentation result of the liver tumor.
[0049] Furthermore, a Maxpooling layer, a BatchNorm layer, and a GeLU activation function layer are provided after each encoding layer; a bilinear interpolation layer (Bilinear Interpolate) is provided before each decoding layer.
[0050] Step S3: Input the liver tumor image to be segmented into the trained liver tumor segmentation model to obtain the segmentation image of the liver tumor image to be segmented.
[0051] In this embodiment, the feature representation ability of liver tumors can be effectively enhanced, and the feature learning effect of the model can be significantly improved. Especially when dealing with fuzzy boundaries and multi-scale features, the segmentation performance is significantly better than that of the prior art. In addition, by introducing axial depth convolution, the present application significantly reduces the number of parameters and computational complexity of the liver tumor segmentation model, thereby effectively improving the segmentation efficiency.
[0052] In one embodiment, an axial depth convolution layer is used to replace the original convolution in the traditional convolutional network to enhance the feature learning ability of liver tumors while reducing the number of parameters and computational complexity of the model. The axial depth convolution layer is designed based on depthwise separable convolution, replacing the cross-shaped receptive field in the vision permutator with a local receptive field. The axial depth convolution layer is implemented through depth convolution, pointwise convolution, and batch normalization. It adopts a unique pointwise convolution design without adding residual connections to adaptively adjust the number of input channels. The depth convolution performs spatial convolution independently on each input channel, while the pointwise convolution performs a linear combination in the channel dimension through a 1×1 convolution kernel.
[0053] Specifically, Figure 2 The figure shows a schematic diagram of the encoding layer. The encoding layer includes an axial depth convolution layer (Axial DW conv), a first pointwise convolution layer (PW conv), a maxpooling layer (Maxpooling), and a GeLU activation function layer connected in sequence; the output of the axial depth convolution layer is skip-connected to the corresponding decoding layer.
[0054] Specifically, Figure 3 The figure shows a schematic diagram of the decoding layer. The decoding layer includes an upsampling layer (UpSampling), a second pointwise convolution layer (PW conv), an axial depth convolution layer (Axial DW conv), a third pointwise convolution layer (PW conv), and a GeLU activation function layer connected in sequence; the output of the upsampling layer is fused with the output of the axial depth convolution layer in the encoding layer. The kernel size of the axial depth convolution layer is set to n = 7, which is used to replace the traditional convolution in the encoder and decoder. By using such a larger-sized kernel, it is possible to capture more extensive context information, reduce network parameters, lower computational complexity, and make the network more compact.
[0055] Specifically, Figure 4 A schematic diagram of the bottleneck layer is shown. The bottleneck layer includes a fourth pointwise convolution layer (PWconv), multiple axially distributed depth convolution layers (Axial DW conv) in parallel, a fifth pointwise convolution layer (PW conv), and a GeLU activation function layer connected in sequence. After the input of the bottleneck layer passes through the fourth pointwise convolution layer, it is respectively input into multiple axially distributed depth convolution layers. The outputs of each axially distributed depth convolution layer are fused, and the fused features are input into the fifth pointwise convolution layer. The kernel size of the axially distributed depth convolution layer in the bottleneck layer is set to n = 3. This design not only improves the efficiency and quality of feature extraction but also significantly reduces the storage and computational requirements of the model, thereby enhancing its applicability in resource-constrained environments while maintaining network performance.
[0056] In one embodiment, Figure 5 A schematic diagram of the interaction-aware Transformer module is shown. Refer to Figure 5 , the interaction-aware Transformer module is used to apply cross-scale attention to the input multiple feature maps F3, F4, and F5 respectively, and correspondingly obtain two query vectors Q 34 and Q 35 , key-value pairs {K4, V4}, and key-value pairs {K5, V5}, where K4 and V4 are the key vector and value vector in the key-value pair {K4, V4} respectively, and K5 and V5 are the key vector and value vector in the key-value pair {K5, V5} respectively; the resolution of the feature map F3 is greater than that of the feature map F4, and the feature map F5 is the output of the first token-aware MLP module;
[0057] C is the number of channels, H and W are the resolutions of the input image, and d is the dimension of the Transformer module. Through cross-scale attention calculation, due to the short sequence length, the computational complexity is significantly reduced. Key-value pairs of different scales correspond to different receptive fields and semantic information levels, and can fuse multi-scale feature information, enhancing the richness and diversity of feature expression.
[0058] Transpose the feature map K4 and multiply it with Q 34 to obtain a first multiplication result; multiply the first multiplication result with V4 to obtain a second multiplication result;
[0059] Transpose K5 and multiply it with Q 35 to obtain a third multiplication result; multiply the third multiplication result with V5 to obtain a fourth multiplication result;
[0060] Fuse the second multiplication result and the fourth multiplication result, and input the fused features into the feed-forward network to obtain the output of the interaction-aware Transformer module.
[0061] In one embodiment, Figure 6 A schematic diagram of the first token-aware MLP module is shown. Refer to Figure 6 , the first token-aware MLP module includes a Shifted MLP and a Norm layer. The Shifted MLP shifts the input in the width and height directions and processes it through an MLP. The output of the Shifted MLP undergoes layer normalization through the Norm layer to obtain normalized features. The normalized features are added to the input to obtain the output of the first token-aware MLP module. Figure 7 A schematic diagram of the Shifted MLP is shown.
[0062] It should be noted that the second token-aware MLP module has the same implementation function as the first token-aware MLP module.
[0063] Adopting the same inventive concept as the liver tumor segmentation method based on axial depth convolution and Transformer, this embodiment also provides a corresponding liver tumor segmentation device based on axial depth convolution and Transformer. The device includes:
[0064] A dataset acquisition module for acquiring a model training dataset. The samples in the training dataset are liver tumor images.
[0065] A model training module for training a liver tumor segmentation model based on the model training dataset to obtain a trained liver tumor segmentation model. The liver tumor segmentation model includes an encoder, a decoder, and a bottleneck layer. The encoder includes a plurality of encoding layers connected in sequence, a first token-aware MLP module, and an interaction-aware Transformer module. The outputs of two encoding layers among the plurality of encoding layers and the output of the first token-aware MLP module are all input into the interaction-aware Transformer module for integration. The decoder includes a second token-aware MLP module and a plurality of decoding layers connected in sequence. The output of the interaction-aware Transformer module is input into the bottleneck layer for feature compression, and the output of the bottleneck layer is input into the second token-aware MLP module. The plurality of encoding layers and the plurality of decoding layers correspond one by one and there are skip connections.
[0066] A prediction module for inputting the liver tumor image to be segmented into the trained liver tumor segmentation model to obtain a segmentation image of the liver tumor image to be segmented.
[0067] The liver tumor segmentation device based on axial depth convolution and Transformer in this embodiment has the same inventive concept as the above-mentioned liver tumor segmentation method based on axial depth convolution and Transformer. Therefore, the specific implementation of this device can be seen in the embodiment part of the liver tumor segmentation method based on axial depth convolution and Transformer in the previous text, and its technical effects correspond to those of the above method, which will not be elaborated here.
[0068] To further verify the effectiveness of this application, the following experimental analysis was carried out.
[0069] 1. Experimental settings
[0070] This application uses the publicly available liver tumor dataset (LiTS) for comparative experiments. This dataset contains 131 CT scan images for training and 70 CT scan images for testing, and each image is accompanied by detailed annotations of the liver and liver tumors. These CT scan images are provided by multiple medical centers and cover different types of liver tumors and their morphological changes at different stages. By slicing this dataset, a total of 6934 two-dimensional tumor images are obtained. In the experiment, this application randomly divides these images according to a ratio of 6:2:2 as the training subset, validation subset, and test subset respectively. Since data preprocessing has strict requirements on the image format, the sliced two-dimensional tumor images are converted to a lossless RGB format.
[0071] In the experiment of this application, to verify the effectiveness of liver tumor segmentation performance, the Dice similarity coefficient (Dice), volume overlap error (VOE), relative volume difference (RVD), sensitivity (Sen), and specificity (Spe) are selected as evaluation indicators; to evaluate the lightweight degree of the network, the number of model parameters (Params, M), computational complexity (FLOPs, G), and memory occupancy (Memory, M) are selected as evaluation indicators.
[0072] All comparative experiments are implemented based on the PyTorch framework and run on an NVIDIA GTX3080 GPU with 12GB of video memory, and the CUDA version is 11.2. In the experiment, the Adam optimizer with an initial learning rate of 0.0001 and a momentum of 0.9 is adopted, and the cosine annealing learning rate scheduling strategy is combined to optimize the training process. During training, the batch size is set to 8, and a total of 400 iteration cycles are carried out.
[0073] 2. Training and testing the liver tumor segmentation model
[0074] This application has conducted comparative experiments with a variety of existing advanced methods, including U-Net, Attention U-Net, nnU-Net, MALU-Net, TransAttU-Net, MS-FANet, RMAU-Net, TransU-Net, MedT, and LightM-UNet, etc. The experimental results are shown in Table 1.
[0075] Table 1: Comparative experimental results on segmentation performance
[0076]
[0077] As shown in Table 1, the results obtained using evaluation metrics such as Dice, VOE, RVD, sensitivity, and specificity are presented, and these results are mainly used to analyze the segmentation performance of this application. Among them, the Dice score of U-Net is 63.31%, the VOE is 50.13%, the RVD is 25.15%, the sensitivity is 63.73%, and the specificity is 99.93%. The Dice score of RMAU-Net based on U-Net is 71.22%, the VOE is 52.70%, the RVD is 23.24%, the sensitivity is 62.23%, and the specificity is 99.89%. The Dice score of MS FANet is 71.43%, the VOE is 39.59%, the RVD is 23.14%, the sensitivity is 60.92%, and the specificity is 98.72%. Compared with U-Net, the Dice of RMAU-Net increased by 7.91%, the VOE increased by 2.57%, the RVD decreased by 1.91%, the sensitivity decreased by 1.5%, and the specificity decreased slightly. RMAU-Net is superior to MS-FANet in terms of sensitivity and specificity, but is still lower than MS-FANet in terms of Dice coefficient, VOE, and RVD. The experimental results show that neither MS-FANet nor RMAU-Net can fully consider the impact of the fuzzy boundary and multi-scale features of liver tumors on the segmentation performance at the same time. By introducing an interactive attention Transformer module and a label-aware MLP module, this application has significantly improved the segmentation performance. Compared with UNet, the Dice, sensitivity, and specificity of this application increased by 11.96%, 4.39%, and 0.01% respectively, while the VOE and RVD decreased by 20.4% and 27.06% respectively; compared with other networks, this application achieved the best Dice score, which is 3.84% higher than the second-best indicator. The results show that this application can more effectively enhance the feature differences between the fuzzy boundary of liver tumor images and normal liver tissues, and improve the learning ability of multi-scale features, thus significantly outperforming other existing methods in terms of segmentation performance.
[0078] Table 2: Comparative experimental results on segmentation efficiency.
[0079] methods Params(M)↓ FLOPs(G)↓ Memory(M)↓ Dice(%)↑ U-Net(2015) 13.39 375.64 51.49 63.31 nnUNet(2018) 126.56 612.73 253.74 65.21 TransUNet(2021) 159.68 873.85 321.53 70.24 MedT(2021) 1.65 243.57 1.96 72.65 LightM-UNet(2024) 1.87 267.18 2.13 72.86 This application 1.58 210.54 1.51 75.27
[0080] As shown in Table 2, the present application has been comprehensively evaluated through evaluation metrics such as Params, FLOPs, Memory, and Dice, which are mainly used to analyze the performance of the present application in terms of segmentation efficiency. Given that the present application uses U-Net as the backbone network and aims to achieve the goal of lightweight, in the comparative experiment, the present application not only compares the performance with common large models, but also conducts a detailed comparative analysis with a variety of lightweight models. Among them, the Params of U-Net is 13.39M, the FLOPs is 375.64G, the Memory is 51.49M, and the Dice coefficient is 63.31%. While the number of parameters of LightM-UNet is 1.87M, the computational volume is 267.18G, the memory occupancy is 2.13M, and the Dice coefficient is 72.86%. Models such as nnU-Net and TransUNet have a large number of parameters and high complexity, while the segmentation accuracy of other lightweight models is relatively low. By introducing the axial convolution module, the present application significantly reduces the number of model parameters and complexity while maintaining excellent segmentation performance. Compared with U-Net, the Params of the present application is reduced by 11.81M, the FLOPs is reduced by 165.1G, the Memory is reduced by 49.98M, and the Dice coefficient is increased by 11.96%. Compared with LightM-UNet, the Params of the present application is reduced by 0.29M, the FLOPs is reduced by 57.27G, the Memory is reduced by 0.62M, and the Dice coefficient is increased by 2.41%. These results indicate that the present application can effectively improve the segmentation efficiency while reducing the number of parameters and complexity of the liver tumor segmentation model, and better retain the segmentation performance.
[0081] 3. Verify the effectiveness of each module in the liver tumor segmentation model of the present application.
[0082] The lightweight liver tumor segmentation model based on axial depth convolution and Transformer mainly consists of an axial depth convolution module, an interactive attention Transformer module, and a label-aware MLP module. To verify the effectiveness of this application in the lightweight research direction, a series of ablation experiments are designed in this embodiment. The goal of these experiments is to evaluate the specific contributions of each module in improving the accuracy of the segmentation task and achieving model lightweighting. As shown in Table 3, the results of the Dice, VOE, RVD, sensitivity, specificity, Params, FLOPs, and Memory evaluation metrics are presented after implementing Embodiment 1. When using U-Net as the backbone network, after adding the axial depth convolution module, the Dice increased by 3.42%, the Params decreased by 7.74M, the FLOPs decreased by 133.02G, and the Memory decreased by 41.98M. After adding the interactive attention Transformer module, the Dice increased by 7.21%, and the VOE and RVD decreased by 11.17% and 10.5% respectively. After adding the label-aware MLP module, the Dice increased by 10.18%, and the VOE and RVD decreased by 18.58% and 22.73% respectively. When the axial depth convolution module, the interactive attention Transformer module, and the label-aware MLP module are introduced simultaneously, the Dice coefficient, sensitivity, and specificity are increased by 11.96%, 4.39%, and 0.01% respectively, while the VOE and RVD are decreased by 20.4% and 27.06% respectively. At the same time, the Params, FLOPs, and Memory are decreased by 11.81M, 165.1G, and 49.98M respectively. The results show that the interactive attention Transformer module and the label-aware MLP module can significantly improve the segmentation performance of fuzzy boundaries and multi-scale features in liver tumor segmentation, but are slightly weaker in reducing the number of model parameters. The axial depth convolution module effectively reduces the number of model parameters while maintaining the accuracy of liver tumor segmentation, saving computational costs.
[0083] Table 3: Results of ablation experiments on comparison relationships and feature-level modules
[0084]
[0085] 4. Verify the effectiveness of the axial depth convolution module.
[0086] The convolution kernel size in the axial depth convolution module also has different effects on the number of model parameters and complexity. To evaluate the performance of the axial depth convolution module in this application, Figure 2The axial depth convolution operator in is replaced with a traditional depth convolution operator, and different sizes of convolution kernels are tried to compare their performance. The final results are shown in Table 4, which presents the experimental results based on the evaluation metrics of Params, FLOPs, Memory, and Dice after the execution of Example 3. The sizes of the depth convolution kernel and the axial depth convolution kernel are set to 3, 5, and 7 respectively to verify the importance of their sizes. Taking U-Net as the backbone network, when the kernel size of the depth convolution is set to 7, the Dice is 64.42%, the Params is 8.12M, the FLOPs is 257.46G, and the Memory is 10.11M; when the kernel size of the axial depth convolution is set to 7, the Dice is 66.73%, the Params is 5.74M, the FLOPs is 242.62G, and the Memory is 9.51M. The results show that although the traditional depth convolution operator can provide a larger receptive field under the same size of convolution kernel (i.e., the same kernel size n), the axial depth convolution still performs better and shows higher simplicity in terms of parameter configuration and computational complexity. This finding further proves the advantages of axial depth convolution in lightweight and performance optimization.
[0087] Table 4: Experimental results of performance comparison between depth convolution and axial depth convolution under different kernel sizes
[0088]
[0089] As mentioned above, the above are only various embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.
Claims
1. A liver tumor segmentation method based on axial deep convolution and Transformer, characterized in that: include: Obtaining a model training data set; samples in the training data set are liver tumor images; The liver tumor segmentation model is trained based on the model training data set to obtain a trained liver tumor segmentation model; the liver tumor segmentation model includes an encoder, a decoder and a bottleneck layer; the encoder includes a plurality of encoding layers, a first tag-aware MLP module and an interactively-aware Transformer module connected in sequence, and the outputs of two encoding layers in the plurality of encoding layers and the output of the first tag-aware MLP module are input into the interactively-aware Transformer module for integration; the decoder includes a second tag-aware MLP module and a plurality of decoding layers connected in sequence, the output of the interactively-aware Transformer module is input into the bottleneck layer for feature compression, and the output of the bottleneck layer is input into the second tag-aware MLP module; the plurality of encoding layers and the plurality of decoding layers correspond to each other one by one and there are jump connections; The liver tumor image to be segmented is input into the trained liver tumor segmentation model to obtain a segmented image of the liver tumor image to be segmented.
2. The method according to claim 1, characterized in that The encoding layer includes an axial depth convolution layer, a first point-by-point convolution layer, a maximum pooling layer and a GeLU activation function layer connected in sequence; the output of the axial depth convolution layer is skip-connected and output to the corresponding decoding layer.
3. The method according to claim 1, characterized in that The decoding layer includes an upsampling layer, a second point-by-point convolution layer, an axial depth convolution layer, a third point-by-point convolution layer and a GeLU activation function layer connected in sequence; the output of the upsampling layer is fused with the output of the axial depth convolution layer in the encoding layer.
4. The method according to claim 1, characterized in that The bottleneck layer includes a fourth point-by-point convolution layer, a plurality of parallel distributed axial depth convolution layers, a fifth point-by-point convolution layer and a GeLU activation function layer connected in sequence; the input of the bottleneck layer passes through the fourth point-by-point convolution layer and is respectively input into a plurality of axial depth convolution layers, the output of each axial depth convolution layer is fused, and the fused features are input into the fifth point-by-point convolution layer.
5. The method according to claim 1, characterized in that The interactive perception Transformer module is used to apply cross-scale attention to the input multiple feature maps F3, F4, and F5, and obtain two corresponding query vectors Q 34 and Q 35 , key-value pair {K4, V4}, key-value pair {K5, V5}, where K4 and V4 are the key vector and value vector in the key-value pair {K4, V4}, and K5 and V5 are the key vector and value vector in the key-value pair {K5, V5}; the resolution of the feature map F3 is greater than the resolution of the feature map F4, and the feature map F5 is the output of the first tag-aware MLP module; Transpose the feature map K4 and compare it with Q 34 Multiply them to obtain a first multiplication result; multiply the first multiplication result by V4 to obtain a second multiplication result; Transpose K5 and add it to Q 35 multiply by V5 to obtain a third multiplication result; multiply the third multiplication result by V5 to obtain a fourth multiplication result; The second multiplication result and the fourth multiplication result are fused, and the fused features are input into the feedforward network to obtain the output of the interaction perception Transformer module.
6. The method according to claim 1, characterized in that The first tag-aware MLP module includes a ShiftedMLP and a Norm layer. The Shifted MLP shifts the input in width and height directions and processes it through an MLP. The output of the Shifted MLP is normalized through a Norm layer to obtain normalized features. The normalized features are added to the input to obtain the output of the first tag-aware MLP module.
7. A liver tumor segmentation device based on axial deep convolution and Transformer, characterized in that: include: A data set acquisition module is used to acquire a model training data set; The samples in the training data set are liver tumor images; A model training module, used for training a liver tumor segmentation model based on the model training data set to obtain a trained liver tumor segmentation model; the liver tumor segmentation model includes an encoder, a decoder and a bottleneck layer; the encoder includes a plurality of encoding layers, a first tag-aware MLP module and an interactively-aware Transformer module connected in sequence, and the outputs of two encoding layers in the plurality of encoding layers and the output of the first tag-aware MLP module are input into the interactively-aware Transformer module for integration; the decoder includes a second tag-aware MLP module and a plurality of decoding layers connected in sequence, the output of the interactively-aware Transformer module is input into the bottleneck layer for feature compression, and the output of the bottleneck layer is input into the second tag-aware MLP module; the plurality of encoding layers and the plurality of decoding layers correspond to each other one-to-one and have jump connections; The prediction module is used to input the liver tumor image to be segmented into the trained liver tumor segmentation model to obtain a segmented image of the liver tumor image to be segmented.
8. The device according to claim 7, characterized in that The encoding layer includes an axial depth convolution layer, a first point-by-point convolution layer, a maximum pooling layer and a GeLU activation function layer connected in sequence; the output of the axial depth convolution layer is skip-connected and output to the corresponding decoding layer.
9. The device according to claim 7, characterized in that The interactive perception Transformer module is used to apply cross-scale attention to the input multiple feature maps F3, F4, and F5, and obtain two corresponding query vectors Q 34 and Q 35 , key-value pair {K4, V4}, key-value pair {K5, V5}, where K4 and V4 are the key vector and value vector in the key-value pair {K4, V4}, and K5 and V5 are the key vector and value vector in the key-value pair {K5, V5}; the resolution of the feature map F3 is greater than the resolution of the feature map F4, and the feature map F5 is the output of the first tag-aware MLP module; Transpose the feature map K4 and compare it with Q 34 Multiply them to obtain a first multiplication result; multiply the first multiplication result by V4 to obtain a second multiplication result; Transpose K5 and add it to Q 35 multiply by V5 to obtain a third multiplication result; multiply the third multiplication result by V5 to obtain a fourth multiplication result; The second multiplication result and the fourth multiplication result are fused, and the fused features are input into the feedforward network to obtain the output of the interaction perception Transformer module.
10. The device according to claim 7, characterized in that The first tag-aware MLP module includes a ShiftedMLP and a Norm layer. The Shifted MLP shifts the input in width and height directions and processes it through an MLP. The output of the Shifted MLP is normalized through a Norm layer to obtain normalized features. The normalized features are added to the input to obtain the output of the first tag-aware MLP module.