A Liver Tumor Segmentation Method Based on SCAF-TransUNet

By adopting the SCAF-TransUNet network in liver tumor segmentation, combining SLIC superpixel preprocessing, CECA module and FPN, the segmentation result problem caused by insufficient feature extraction in the prior art is solved, and higher segmentation accuracy and robustness are achieved.

CN119399461BActive Publication Date: 2025-05-30SHENYANG UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411445414.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2025-05-30
Estimated Expiration
2044-10-16

AI Technical Summary

Technical Problem

In the prior art, deep learning-based liver tumor segmentation algorithm is difficult to fully capture when feature extraction, resulting in missed and misdetected segmentation results, affecting the overall segmentation performance.

Method used

The liver tumor segmentation method based on SCAF-TransUNet is used to segment the image into superpixel areas through the Simple Linear Iterative Clustering (SLIC) superpixel pre-processing algorithm, reducing redundant information and retaining local details. Then, multi-scale features are extracted using convolutional neural network (CNN), and feature extraction and attention mechanism processing are performed through Transformer Layer and Channel-Efficient Convolutional Attention (CECA) modules, and finally multi-scale feature fusion is performed through Feature Pyramid Networks (FPN).

Benefits of technology

It significantly improves the accuracy and robustness of liver tumor segmentation tasks, enhances the processing ability of complex liver boundaries and small tumors, reduces calculation overhead and improves real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399461B_ABST
    Figure CN119399461B_ABST
Patent Text Reader

Abstract

The present invention discloses a liver tumor segmentation method based on SCAF-TransUNet. First, the SLIC superpixel preprocessing divides the liver image into semantically related superpixel regions, enabling the model to more accurately focus on key regions such as tumors, reducing computational redundancy, and enhancing the efficiency and accuracy of feature extraction. Second, the CECA module introduced in the skip connection simplifies the channel attention mechanism by replacing the traditional fully connected layer with one-dimensional convolution, enhances the capture of channel dependence relationships, and at the same time maintains the accuracy of feature extraction. Finally, FPN is added to the decoder to further enhance the multi-scale feature fusion ability. FPN effectively combines global semantic information and local detail features before upsampling by fusing feature maps of different resolutions layer by layer, improving the processing ability for complex liver boundaries and small tumors. Ultimately, these improvements significantly enhance the accuracy and robustness of the liver tumor segmentation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of liver tumor segmentation, and particularly relates to a liver tumor segmentation method based on SCAF-TransUNet. Background Art

[0002] With the development of the times, the algorithms for liver tumor segmentation are constantly updated and improved. From the initial traditional segmentation methods to the current deep learning-based segmentation methods, the efficiency and accuracy are getting higher and higher. Traditional image segmentation methods mainly classify based on the similarity between pixels, that is, the pixels in the image are divided into different regions to separate and extract the target region. For example, the threshold method proposed by Soler et al. in " Fully automatic anatomical,pathological, and functional segmentation from CT scans for hepatic surgery " selects an appropriate threshold according to the characteristics and requirements of the image to divide the pixels in the image into different regions or categories. Pohle et al. proposed a region growing algorithm in "Segmentation of medical images using adaptive region growing". This method first selects seed pixels or sub-regions as the target positions. If the adjacent pixels or regions meet the similarity conditions, they are merged, so that the region can grow step by step. The graph cut method proposed by Boykov et al. in "Interactive graph cuts for optimal boundary®ion segmentation of objectsin ND images" converts the image into an undirected graph, where each pixel is regarded as a node and the edges between adjacent pixels are regarded as edges. Then, the graph cut method uses the maximum flow minimum cut algorithm to process this graph to achieve image segmentation. Compared with traditional algorithms, the deep learning-based feature extraction method has better performance because it can automatically learn features from the image, does not require manual selection of feature types, and can better extract the deep features for image segmentation.

[0003] However, in the existing image segmentation technologies, although the deep learning-based methods have made significant progress, there are still some limitations, resulting in the accuracy of the segmentation results still being unsatisfactory. This is mainly reflected in that the edge segmentation of the low-contrast tumor region is not well-fitted; the existence of multi-scale tumors makes it difficult for the algorithm to fully capture due to the different sizes and shapes of the tumors, resulting in missed detections and false detections in the segmentation results, seriously affecting the overall segmentation performance.

[0004] Therefore, the existing technology still needs to be improved and enhanced. Summary of the Invention

[0005] In view of the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a liver tumor segmentation method based on SCAF-TransUNet, aiming to solve the problem that the existing algorithms in the prior art are difficult to fully capture during feature extraction, resulting in missed detections and false detections in the segmentation results.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] In the first aspect, a liver tumor segmentation method based on SCAF-TransUNet is characterized by including:

[0008] Step 1: Input the liver tumor image. First, use the Simple Linear Iterative Clustering superpixel preprocessing algorithm, which divides the image into multiple superpixel regions with similar features, reduces redundant information, and preserves the local details of the image to form a superpixel image.

[0009] Step 2: Input the superpixel image obtained after superpixel preprocessing into the encoder part. Extract the multi-scale features of the input image through a convolutional neural network CNN, and use the Transformer Layer of SCAF-TransUNet to divide the CNN feature map into small patches and use the self-attention mechanism to capture the long-range dependencies in the image.

[0010] Step 3: In each layer of the decoder stage, the upsampled feature map of the decoder is combined with the underlying encoder feature map passed through the skip connection. The encoder feature map is processed by a channel-efficient convolutional attention module during the transmission process, and finally, a feature map weighted by channel attention and spatial attention is output.

[0011] Step 4: At the last layer of the decoder, the feature map after upsampling and feature fusion compresses the number of channels to the target number of classes through a 1x1 convolution. Subsequently, through Softmax activation, a class probability distribution is generated for each pixel, and the final output is the segmentation result for each pixel, realizing the segmentation of the liver tumor image.

[0012] Further, the superpixel preprocessing of the obtained liver tumor CT image to be segmented includes:

[0013] Step 1.1: Initialize the seeds:

[0014] Assume that the image has N pixel points and it is desired to divide the image into K superpixel regions. Then the area of each superpixel is N / K, and the distance between the seed points is:

[0015]

[0016] Step 1.2, Gradient adjustment of seed points:

[0017] Within the n×n neighborhood of each seed point, usually n = 3, calculate the gradient value of the pixel, and move the seed to the pixel position with the minimum gradient;

[0018] Step 1.3, Calculate the color distance d c :

[0019]

[0020] where L, a, and b are the three channels of the Lab color space, p represents the current pixel, and c represents the superpixel seed point;

[0021] Step 1.4, Calculate the spatial distance d s :

[0022]

[0023] where x and y represent the pixel coordinates respectively;

[0024] Step 1.5, Final distance metric D′:

[0025] Combining the color distance and the spatial distance, use a weighted metric method to measure the distance from each pixel to the seed. To balance the influence of the color distance and the spatial distance, a weight parameter m is required, and this value is used to control the relative importance of color and space, usually taking the value of 10:

[0026]

[0027] where is the step size between seed points, used to standardize the spatial distance;

[0028] Step 1.6, Iterative optimization:

[0029] Optimize the division of seed points and superpixels by iterating the above process multiple times. In each iteration, the seed points will gradually adjust to the center of the superpixel region they affect and finally stabilize to form a superpixel image.

[0030] Furthermore, the connectivity between superpixel regions needs to be enhanced in Step 1. Since discontinuous superpixel regions may be generated during segmentation, the connectivity enhancement step is finally used to ensure the coherence of superpixel regions, remove isolated pixels, and optimize incoherent superpixel regions.

[0031] Furthermore, the superpixel image obtained through superpixel preprocessing is input into the encoder part, and the multi-scale features of the input image are extracted through the convolutional neural network CNN, including: in the encoder part, the convolutional neural network CNN is used to extract feature maps of different resolutions; for TransUNet, through convolutional and pooling operations, the resolution of the image decreases step by step, generating feature maps of different levels; assuming the size of the input image is H×W, the resolution of the feature map decreases step by step. During the convolutional operation, the input image will gradually shrink into different feature maps, and the shrinking process of the feature map is as follows:

[0032] C1: The size of the input image is H×W;

[0033] C2: After one pooling operation, the image resolution is reduced to to form a larger low-level feature map;

[0034] C3; Pooling again, the image resolution is further reduced to

[0035] C4: Finally, the image resolution is reduced to At this time, the feature map reflects more local information;

[0036] The calculation formula for the operation of the convolutional layer is expressed as:

[0037] F i =σ(W i *I i +b i )

[0038] where F i represents the feature map generated after the convolutional operation; W i represents the weight matrix of the convolutional kernel; I i represents the input image or the feature map of the previous layer; b i represents the bias term; * represents the convolutional operation; σ is a non-linear activation function.

[0039] Furthermore, the Transformer Layer of SCAF-TransUNet captures long-range dependencies in the image by dividing the CNN feature map into small patches and using the self-attention mechanism, including:

[0040] Step 2.1, Imageization and Embedding

[0041] First, the feature map is divided into small patches of P×P and unfolded into a one-dimensional vector; each small patch is linearly mapped to a high-dimensional space and added with position information:

[0042]

[0043] Among them, E is the linear embedding matrix; E pos represents the position encoding;

[0044] Step 2.2, Multi-Head Self-Attention Mechanism

[0045] Through the self-attention mechanism, calculate the similarity between each patch, and the formula is:

[0046]

[0047] Among them, Q, K, and V are the query, key, and value vectors respectively, calculate the similarity between each position, and obtain the final output through weighting;

[0048] Step 2.3, Multi-Layer Perceptron MLP and Output Update

[0049] After the self-attention mechanism, use the multi-layer perceptron MLP for non-linear transformation, and finally obtain the output of the Transformer layer:

[0050]

[0051] Among them, Z i represents the final output of the current layer; represents the intermediate output after the multi-head self-attention mechanism; represents layer normalization on the intermediate output layer to make the data more stable; MLP represents non-linear transformation through a multi-layer perceptron to enhance the feature representation ability; represents the residual connection, which is used to retain the input information and enhance the gradient flow.

[0052] Furthermore, the third step includes:

[0053] Process the feature map passed from the encoder part through the Channel-Efficient Convolutional Attention module, and the output is the feature map weighted by channel attention:

[0054] Step 3.1, Input Feature Map F

[0055] Assume that the input feature map is F, and H, W, and C are the height, width, and number of channels of the feature map respectively;

[0056] Step 3.2, Global Average Pooling:

[0057] Perform global average pooling on the input feature map F to generate the global representation of each channel; obtain a one-dimensional vector v ∈ R C , where each element represents the global information of a channel; the formula is as follows:

[0058]

[0059] Step 3.3, Replace the fully connected layer with a one-dimensional convolution:

[0060] Use a 1D convolution to replace the fully connected layer and generate the dependency weights between channels; the kernel size determines the range of cross-channel dependencies:

[0061] W = Conv1D(v);

[0062] Step 3.4, Activation of channel weights:

[0063] Use the Sigmoid function to map the weight W to the range of [0, 1]:

[0064] W' = σ(W)

[0065]

[0066] Step 3.5, Feature map weighting:

[0067] Use the channel weight W' to weight each channel of the input feature map to generate a weighted channel feature map:

[0068] F' = F C ·W', c = 1, 2...., C.

[0069] Furthermore, the third step includes:

[0070] Continue to process the feature map after channel attention weighting through the Channel-Efficient Convolutional Attention module, and the output is the feature map after spatial attention weighting:

[0071] Step 3.6, Input feature map F': After channel attention weighting, the feature map F' ∈ R H×W×C is fed into the spatial attention module;

[0072] Step 3.7, Pooling operation: In the spatial attention module, apply global average pooling and global max pooling to compress along the channel dimension to generate two spatial feature maps F avg ∈ R H×W×C and F max ∈ R H×W×C ;

[0073]

[0074] F max = max(F'(i, j, c)), c = 1,2...., C;

[0075] Step 3.8, Feature Map Concatenation: Concatenate the two feature maps obtained by pooling along the channel dimension to form a 2D feature map F concat ∈R H×W×2 ;

[0076] Step 3.9, Convolution Operation: Perform a 3x3 convolution on the concatenated feature map to generate a spatial attention weight map M s ∈R H×W :

[0077]

[0078] where σ is the Sigmoid activation function, which restricts the attention weights between [0, 1];

[0079] Step 3.10, Spatial Attention Weighting:

[0080] Use the spatial attention weight M s to perform spatial weighting on the previously weighted feature map F':

[0081] F″(i, j, c) = F′(i, j, c)·M s (i, j), i = 1, 2...., H, j = 1, 2...., W;

[0082] Step 3.11, Final Output: The finally output feature map F″(i, j, c) ∈ R H×W×C is the feature map after channel attention and spatial attention weighting.

[0083] Furthermore, the following is also included in Step 3:

[0084] At each layer of the decoder, the upsampled feature map is fused with the feature map passed through the Channel-Efficient Convolutional Attention module via skip connection. The fused feature map F fusion will be passed as input to the PFN module for further processing to enhance the feature expression and optimization effect:

[0085] Feature Extraction: In the encoder part, a convolutional neural network is used to extract feature maps of different resolutions; for SCAF-TransUNet, through convolution and pooling operations, the resolution of the image gradually decreases, generating feature maps of different levels; assuming the size of the input image is H×W, the resolution of the feature maps gradually decreases, and these multi-scale feature maps will be used as the input for the subsequent FPN; each feature map C i contains different levels of spatial information and semantic information;

[0086] Horizontal connection: At each resolution level, the feature map is compressed in terms of the number of channels through 1x1 convolution to obtain the downsampled feature map P i ; The role of 1x1 convolution is to reduce the number of channels of the feature map, reduce the computational complexity, and at the same time retain the most important key information:

[0087]

[0088] Top-down feature fusion: FPN adopts a top-down feature fusion strategy; the feature map of the highest layer has the smallest spatial resolution and the richest semantic information, while the feature maps of the lower layers have higher resolutions and more spatial details; during feature fusion, first the low-resolution feature map is upsampled to a higher resolution, and then added element-wise to the higher-resolution feature map:

[0089]

[0090]

[0091]

[0092] In this top-down manner, the global semantic information of the high-level feature map is gradually propagated to the low-level feature map, while retaining the low-level detailed information;

[0093] Upsampling operation: The upsampling operation enlarges the low-resolution feature map to a size matching that of the high-resolution feature map through bilinear interpolation or transposed convolution;

[0094] Fused convolution: After each upsampling and feature fusion, FPN applies a 3x3 convolution operation to further process the fused feature map; this convolution operation helps to smooth the feature map and reduce the aliasing effect, while maintaining the combination of global semantic information and local details:

[0095] P i = Conv3x3(C i );

[0096] Combination of FPN and skip connection: In the improved SCAF-TransUNet, the skip connection part has introduced the Channel-Efficient Convolutional Attention module; the feature map generated by FPN will be combined with the skip connection feature map from the encoder:

[0097] P i = P i + SkipConnection(CECA(C i ));

[0098] Further, in the SCAF-TransUNet decoder part in step three, upsampling is performed using the feature map after FPN fusion. After each upsampling, the resolution of the feature map is gradually restored until the original resolution of the input image is reached.

[0099]

[0100] Among them, P i represents the input feature map of the current layer, and the upsampled feature map will be passed to the next layer;

[0101] Finally, the resolution of the output image is restored to the size of the original image, and the segmentation result is:

[0102] Output = Upsample(P 2 ).

[0103] The technical solution adopted by the present invention has the following beneficial effects:

[0104] The network of the present invention is an improved scheme based on the TransUNet network, which combines SLIC superpixel preprocessing, the CECA (Channel-Efficient Convolutional Attention) hybrid attention module, and the multi-scale feature fusion module FPN (Feature Pyramid Networks), aiming to improve the performance of liver tumor segmentation. First, SLIC superpixel preprocessing divides the liver image into semantically related superpixel regions, enabling the model to more accurately focus on key regions such as tumors, reducing computational redundancy, and improving the efficiency and accuracy of feature extraction. Second, the CECA module introduced in the skip connection simplifies the inter-channel attention mechanism by replacing the traditional fully connected layer with one-dimensional convolution, enhances the capture of channel dependence relationships, reduces computational overhead, and at the same time maintains the accuracy of feature extraction. This makes CECA show higher real-time performance and computational efficiency in liver tumor segmentation. Finally, adding FPN to the decoder further enhances the multi-scale feature fusion ability. FPN effectively combines global semantic information and local detail features by fusing feature maps of different resolutions layer by layer, improving the ability to process complex liver boundaries and small tumors. Ultimately, these improvements significantly improve the accuracy and robustness of the liver tumor segmentation task. BRIEF DESCRIPTION OF THE DRAWINGS

[0105] Figure 1 It is the structural diagram of the SCAF-TransUNet network of the embodiment of the present invention;

[0106] Figure 2It is the FPN (Feature Pyramid Network) module of the embodiment of the present invention;

[0107] Figure 3 It is the CECA (Channel-Efficient Convolutional Attention) module of the embodiment of the present invention;

[0108] Figure 4 It is the ECA (Efficient Channel Attention) module of the embodiment of the present invention. Detailed implementation manners

[0109] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0110] In the embodiment of the present invention, please refer to Figure 1 , a liver tumor segmentation method based on SCAF-TransUNet includes the following steps:

[0111] Step 1: Input a liver tumor image. First, use the Simple Linear Iterative Clustering superpixel preprocessing algorithm. This algorithm divides the image into multiple superpixel regions with similar features, reduces redundant information and retains local details of the image, forming a superpixel image;

[0112] Step 2: Input the superpixel image obtained through superpixel preprocessing into the encoder part. Extract multi-scale features of the input image through a convolutional neural network CNN, and through the Transformer Layer of SCAF-TransUNet, divide the CNN feature map into small patches and use the self-attention mechanism to capture long-range dependencies in the image;

[0113] Step 3: In each layer of the decoder stage, the upsampled feature map of the decoder is combined with the underlying encoder feature map passed through the skip connection. The encoder feature map is processed by the channel-efficient convolutional attention module during the transmission process, and finally outputs a feature map weighted by channel attention and spatial attention;

[0114] Step 4: In the last layer of the decoder, the feature map after upsampling and feature fusion compresses the number of channels to the number of target classes through 1x1 convolution. Subsequently, through Softmax activation, a class probability distribution is generated for each pixel, and the final output is the segmentation result for each pixel, realizing the segmentation of the liver tumor image.

[0115] Step 1: Preprocessing Stage

[0116] Innovation in the pre - input stage based on the SLIC super - pixel algorithm:

[0117] Input the liver tumor image. First, adopt the SLIC (Simple Linear Iterative Clustering) super - pixel preprocessing algorithm. This algorithm divides the image into multiple super - pixel regions with similar features. By reducing redundant information and retaining local image details, the subsequent features can be obtained more efficiently and accurately. The specific steps are as follows:

[0118] Initialize seeds:

[0119] Assume the image has N pixel points and we want to divide the image into K super - pixel regions. Then the area of each super - pixel is N / K. The distance S (step size) between seed points is:

[0120]

[0121] Adjust seed points by gradient:

[0122] Within the n×n neighborhood (usually n = 3) of each seed point, calculate the gradient value of the pixel and move the seed to the pixel position with the minimum gradient. This helps to avoid seed points being located at the boundaries and improves the clustering effect.

[0123] Calculate the color distance d c :

[0124]

[0125] where L, a, and b are the three channels of the Lab color space, p represents the current pixel, and c represents the super - pixel seed point.

[0126] Calculate the spatial distance d s

[0127]

[0128] where x and y represent the pixel coordinates respectively.

[0129] Final distance metric D′:

[0130] Combining the color distance and the spatial distance, use a weighted metric method to measure the distance from each pixel to the seed. To balance the influence of the color distance and the spatial distance, a weight parameter m is required. This value is used to control the relative importance of color and space and usually takes the value of 10:

[0131]

[0132] Among them, d c and d s represent the color distance and the spatial distance respectively. is the step size between seed points, which is used to normalize the spatial distance.

[0133] Iterative optimization:

[0134] The allocation of seed points and superpixels is optimized by iterating the above process multiple times. Usually, 10 iterations are sufficient. In each iteration, the seed points are gradually adjusted to the center of the superpixel region they affect and finally stabilize.

[0135] Enhanced connectivity:

[0136] Since discontinuous superpixel regions may be generated during segmentation, the connectivity enhancement step is finally used to ensure the coherence of the superpixel regions, remove isolated pixels, and optimize the incoherent superpixel regions.

[0137] Output the processed superpixel image:

[0138] The output image will be segmented into multiple superpixel regions. These superpixel regions can be used as input data for subsequent processing by the SCAF-TransUNet model for further detailed segmentation tasks.

[0139] The main purpose of SLIC (Simple Linear Iterative Clustering) superpixel preprocessing is to segment the input image into multiple compact regions (superpixels) with semantic associations, improving the processing efficiency and accuracy of the model. Specifically, SLIC aggregates adjacent similar pixels into superpixel regions, preserving local features in the image and enabling the model to better focus on important image regions. This method is particularly effective in medical image segmentation tasks as it can significantly reduce redundant information in the image, prevent the model from wasting computational resources on irrelevant regions, and thereby improve segmentation accuracy. In the SLIC workflow, superpixel seed points are first initialized through a regular grid, and then the surrounding pixels are clustered based on spatial distance and color similarity of the pixels to gradually form semantically related superpixel regions. Next, the SLIC algorithm iteratively optimizes and continuously adjusts the boundaries of each superpixel region to ensure the minimum color and spatial differences between adjacent pixels, generating a compact superpixel segmentation result. At the same time, SLIC also pays particular attention to the preservation of boundary information, enabling the segmented superpixels to not only reflect the true boundary features in the image but also have good regional compactness. Through this preprocessing, the model simplifies the representation of the image at the input stage, making subsequent feature extraction more efficient. This characteristic of SLIC also particularly helps to enhance the model's ability to process boundaries and local details, avoiding boundary blurring or over-smoothing problems. Ultimately, SLIC superpixel preprocessing provides the model with a more concise and semantically clear input image, greatly improving the efficiency and accuracy of the segmentation task.

[0140] Step 2: Encoder stage

[0141] CNN convolutional feature extraction:

[0142] For the image input obtained through SLIC superpixel preprocessing, the encoder part first extracts multi-scale features of the input image through a convolutional neural network (CNN). The specific operations are as follows:

[0143] Convolutional layer and pooling operations:

[0144] The image first undergoes operations of convolutional and pooling layers to extract multi-scale feature maps from low-level to high-level. The pooling layer is responsible for reducing the spatial dimension of the image, while the convolutional layer extracts local features through filters. Usually, the pooling layer uses max pooling, and the convolutional layer applies a 3×3 convolutional kernel.

[0145] Multi-scale feature maps:

[0146] During the convolutional operation, the input image is gradually reduced to different feature maps. The process of feature map reduction is as follows:

[0147] C1: The size of the input image is H×W.

[0148] C2: After one pooling operation, the image resolution is reduced to to form a larger low-level feature map.

[0149] C3; Pooling again, the image resolution is further reduced to

[0150] C4: Finally, the image resolution is reduced to At this time, the feature map reflects more local information.

[0151] The calculation formula of the convolution layer operation is expressed as:

[0152] F i = σ(W i * I i + b i )

[0153] Among them, F i represents the feature map generated after the convolution operation. W i represents the weight matrix of the convolution kernel. I i represents the input image or the feature map of the previous layer. b i represents the bias term. * represents the convolution operation. σ is a non-linear activation function.

[0154] The Transformer Layer of SCAF-TransUNet divides the CNN feature map into small patches and uses the self-attention mechanism to capture long-range dependencies in the image. The calculation steps of this part are as follows:

[0155] Patch Embedding

[0156] First, the feature map is divided into P×P small patches and unfolded into one-dimensional vectors. Each small patch is linearly mapped to a high-dimensional space and added with position information:

[0157]

[0158] Among them, E is the linear embedding matrix. E pos represents the position encoding.

[0159] Multi-Head Self-Attention (MSA)

[0160] Through the self-attention mechanism, the similarity between each small patch is calculated. The formula is:

[0161]

[0162] Among them, Q, K, and V are the query, key, and value vectors respectively. The similarity between each position is calculated, and the final output is obtained through weighting.

[0163] Multi-Layer Perceptron (MLP) and output update:

[0164] After the self-attention mechanism, a multi-layer perceptron (MLP) is used for non-linear transformation, and finally the output of the Transformer layer is obtained:

[0165]

[0166] Among them, Z i represents the final output of the current layer. represents the intermediate output after the multi-head self-attention mechanism. represents layer normalization of the intermediate output layer to make the data more stable. MLP represents non-linear transformation through a multi-layer perceptron to enhance the feature representation ability. represents the residual connection, which is used to retain the input information and enhance the gradient flow.

[0167] Step 3: Decoder stage

[0168] In the embodiments of the present invention, please refer to Figure 3 , and the CECA module is introduced through skip connections to enhance the feature expression;

[0169] In each layer of the decoder stage, the upsampled feature map of the decoder is combined with the underlying encoder feature map transmitted through the skip connection. The encoder feature map is processed by the CECA module (Channel-Efficient Convolutional Attention Module) during the transmission process, enhancing the important information in the channel and spatial dimensions.

[0170] F fusion = Concat(CECA(F encoder ), F decoder )

[0171] Among them, F encoder is the feature map transmitted from the encoder part, and after being processed by the CECA module, it becomes CECA(F encoder ). F decoder is the upsampled feature map of the current layer of the decoder. F fusion represents the result of concatenating the feature map processed by the CECA module transmitted through the skip connection and the upsampled feature map of the current layer of the decoder.

[0172] CECA (Channel-Efficient Convolutional Attention) module:

[0173] Input feature map F:

[0174] Assume the input feature map is F, and H, W, and C are the height, width, and number of channels of the feature map respectively.

[0175] Global Average Pooling:

[0176] Perform global average pooling on the input feature map F to generate a global representation for each channel. Obtain a one-dimensional vector v ∈ R C , where each element represents the global information of a channel. The formula is as follows:

[0177]

[0178] Replace the fully connected layer with a one-dimensional convolution:

[0179] Use a 1D convolution to replace the fully connected layer and generate the dependency weights between channels. The size of the convolution kernel determines the range of cross-channel dependencies:

[0180] W = Conv1D(v).

[0181] Activation Function of Channel Weights:

[0182] Use the Sigmoid function to map the weight W to the range of [0, 1]:

[0183] W' = σ(W);

[0184]

[0185] Channel Reweighting:

[0186] Use the channel weight W' to weight each channel of the input feature map to generate a weighted channel feature map:

[0187] F' = F c · W', c = 1, 2...., C.

[0188] Spatial Attention Module:

[0189] 1) Input feature map F': After channel attention weighting, the feature map F' ∈ R H×W×C is passed into the spatial attention module.

[0190] 2) Pooling Operations: In the spatial attention module, apply global average pooling and global max pooling to compress along the channel dimension to generate two spatial feature maps F avg ∈ RH×W×C and F max ∈R H×W×C :

[0191]

[0192] F max = max(F′(i, j, c)), c = 1, 2...., C.

[0193] 3) Feature map concatenation: Concatenate the two feature maps obtained by pooling along the channel dimension to form a 2D feature map F concat ∈R H×W×2 .

[0194] 4) Convolution operation: Perform a 3x3 convolution on the concatenated feature map to generate a spatial attention weight map

[0195] M s = σ(Conv3x3(F concat ))

[0196] where σ is the Sigmoid activation function, which restricts the attention weights between [0, 1].

[0197] Spatial Reweighting:

[0198] Use the spatial attention weight M s to perform spatial weighting on the previously weighted feature map F′:

[0199]

[0200] Final output:

[0201] The final output feature map F″(i, j, c) ∈ R H×W×C is the feature map after channel attention and spatial attention weighting.

[0202] Among them, the Channel-Efficient Convolutional Attention module is an improvement to the channel attention mechanism in CBAM (Convolutional Block Attention Module), replacing the channel attention part therein with a more lightweight ECA (Efficient Channel Attention) module. The traditional CBAM channel attention module extracts channel features through global average pooling and global max pooling, and generates attention weights for each channel through two fully connected layers. Although this method can accurately capture the dependencies between channels, its computational complexity is relatively high, especially when dealing with high-resolution images, the computational overhead increases significantly.

[0203] In this embodiment, as Figure 4 shown, the ECA module replaces the fully connected layer with a one-dimensional convolution, simplifying the design of the channel attention mechanism. It no longer requires channel compression and expansion operations, but directly captures cross-channel dependencies through 1D convolution. This improvement greatly reduces the computational cost, especially in high-resolution scenarios, effectively reducing the computational burden. ECA retains the accurate feature extraction ability in the CBAM module through efficient convolution operations, while achieving the same or even better performance with fewer computational resources. This lightweight design is particularly suitable for tasks with high requirements for real-time performance and computational efficiency. In medical image segmentation, CECA not only improves the ability to capture important features but also speeds up the response speed of the model, ensuring the efficiency and accuracy of the segmentation task. By integrating the CECA module into the skip connection, the model further enhances the information interaction between channels, optimizes the transmission between feature layers, and finally significantly improves the accuracy and robustness of medical image segmentation.

[0204] In this embodiment, please refer to Figure 2 , and the feature fusion and information transfer optimization are achieved by introducing the FPN module:

[0205] At each layer of the decoder, the upsampled feature map is fused with the encoder feature map (processed by the CECA module) passed through the skip connection, and the fused feature map F fusion will be passed as input to the PFN module for further processing to enhance the expression and optimization effect of the features. The following is the detailed working process of the PFN module.

[0206] Feature Extraction (Encoder part)

[0207] In the encoder part, a convolutional neural network (CNN) is used to extract feature maps of different resolutions. For SCAF-TransUNet, through convolutional and pooling operations, the resolution of the image gradually decreases, generating feature maps at different levels. Assuming the size of the input image is H×W, the resolution of the feature maps gradually decreases, for example C 2 , C 3 , C 4 , C 5 , which represent different resolutions respectively:

[0208] C 2 = (H / 2, W / 2), C 3 = (H / 4, W / 4), C 4 = (H / 8, W / 8), C 5 = (H / 16, W / 16);

[0209] These multi-scale feature maps will be used as the input for the subsequent FPN. Each feature map C i contains different levels of spatial and semantic information:

[0210] Lateral Connections:

[0211] At each resolution level, the feature map is compressed in the number of channels through 1x1 convolution to obtain the feature map F i after dimensionality reduction. The role of 1x1 convolution is to reduce the number of channels of the feature map, reduce the computational complexity, and at the same time retain the most important key information:

[0212]

[0213] Top-Down Pathway:

[0214] FPN adopts a top-down feature fusion strategy. The highest-level feature map P 5 has the smallest spatial resolution and the richest semantic information, while the lower-level feature maps have higher resolutions and more spatial details. During feature fusion, first the low-resolution feature map is upsampled to a higher resolution, and then added element-wise to the higher-resolution feature map:

[0215] P 4 = P 4 + Upsample(P 5 );

[0216] P 3 = P 3 + Upsample(P 4 );

[0217] P 2 = P 2 + Upsample(P 3 ).

[0218] Through this top - down approach, the global semantic information of the high - level feature map is gradually propagated to the low - level feature map, while retaining the low - level detailed information.

[0219] Upsampling operation:

[0220] The upsampling operation enlarges the low - resolution feature map to a size matching that of the high - resolution feature map through bilinear interpolation or transposed convolution. For example, upsampling P 5 to P 4 in resolution:

[0221]

[0222] After upsampling, the sizes of the feature maps are consistent, enabling direct calculation during element - wise addition.

[0223] Fused convolution:

[0224] After each upsampling and feature fusion, FPN applies a 3x3 convolution operation to further process the fused feature map. This convolution operation helps smooth the feature map and reduce the aliasing effect, while maintaining the combination of global semantic information and local details:

[0225] P i = Conv3x3(C i ).

[0226] Combination of FPN and Skip Connections:

[0227] In the improved SCAF - TransUNet, the CECA module has been introduced in the skip connection part. The feature maps generated by FPN will be combined with the skip - connection feature maps from the encoder (through feature concatenation). These skip - connection feature maps are processed by the channel and spatial attention mechanisms of CECA, which can effectively enhance the attention to important features:

[0228] P i = P i + SkipConnection(CECA(C i )).

[0229] Upsampling operation:

[0230] In the decoder part of SCAF-TransUNet, upsampling is performed using the feature maps fused by FPN. After each upsampling, the resolution of the feature maps is gradually restored until the original resolution of the input image is reached:

[0231]

[0232] Among them, P i represents the input feature map of the current layer, and the upsampled feature map will be passed to the next layer.

[0233] Finally, the resolution of the output image is restored to the size of the original image, and the segmentation result is obtained as:

[0234] Output = Upsample(P 2 ).

[0235] Among them, FPN (Feature Pyramid Network) is introduced in the decoding stage of SCAF-TransUNet and placed before upsampling, which can significantly improve the fusion ability of multi-scale features, thereby enhancing the accuracy of medical image segmentation. FPN fuses feature maps of different resolutions (such as 1 / 8, 1 / 4, 1 / 2, etc.) layer by layer, and synthesizes these feature maps into a feature map containing global semantic information and local details before upsampling. This multi-scale feature fusion can better capture liver tumors of different sizes and shapes, especially helping to handle complex boundaries and small lesions. After completing the multi-scale feature fusion and then performing upsampling, it can ensure that the feature map already has rich details and global information before upsampling, making the upsampling process more robust and reducing the risk of boundary blurring or detail loss. Through such improvements, the model can significantly improve the segmentation accuracy and boundary processing ability when dealing with complex medical images

[0236] Step Four: Final Output

[0237] At the last layer of the decoder, the feature map after upsampling and feature fusion compresses the number of channels to the number of target classes through 1x1 convolution. Subsequently, through Softmax activation, a class probability distribution is generated for each pixel, and the final output is the segmentation result for each pixel.

[0238] The improved solution of the network of the present invention based on the TransUNet network combines SLIC superpixel preprocessing, the CECA (Channel-Efficient Convolutional Attention) hybrid attention module, and the multi-scale feature fusion module FPN (Feature Pyramid Networks), aiming to improve the performance of liver tumor segmentation. First, the SLIC superpixel preprocessing divides the liver image into semantically related superpixel regions, enabling the model to more accurately focus on key regions such as tumors, reducing computational redundancy, and improving the efficiency and accuracy of feature extraction. Secondly, the CECA module introduced in the skip connection simplifies the attention mechanism between channels by replacing the traditional fully connected layer with one-dimensional convolution, enhances the capture of channel dependence relationships, reduces computational overhead, and at the same time maintains the accuracy of feature extraction. This makes CECA show higher real-time performance and computational efficiency in liver tumor segmentation. Finally, adding FPN to the decoder further enhances the multi-scale feature fusion ability. FPN effectively combines global semantic information and local detail features before upsampling by layer-by-layer fusion of feature maps with different resolutions, improving the ability to handle complex liver boundaries and small tumors. Ultimately, these improvements significantly improve the accuracy and robustness of the liver tumor segmentation task.

[0239] After considering the specification and practicing the disclosed solutions herein, those skilled in the art will readily conceive of other embodiments of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present invention are indicated by the claims.

Claims

1. A liver tumor segmentation method based on SCAF-TransUNet, characterized in that: include: Step 1: Input the liver tumor image and first use the superpixel preprocessing algorithm, which divides the image into multiple superpixel regions with similar features, reduces redundant information and retains local details of the image to form a superpixel image; Step 2: For the super-pixel image obtained after super-pixel preprocessing, the encoder part uses the convolutional neural network CNN to extract the multi-scale features of the input image, and divides the CNN feature map into small patches through the Transformer Layer of SCAF-TransUNet, and uses the self-attention mechanism to capture the long-distance dependencies in the image; Step 3: In the decoder stage of each layer, the upsampled feature map of the decoder is combined with the underlying encoder feature map transmitted through the skip connection. The encoder feature map is processed by the channel-efficient convolutional attention module during the transmission process, and finally the feature map weighted by channel attention and spatial attention is output; Step 4: In the last layer of the decoder, the feature map after upsampling and feature fusion is compressed to the number of target categories through 1x1 convolution. Then, through Softmax activation, the category probability distribution is generated for each pixel, and the final output is the segmentation result of each pixel, realizing liver tumor image segmentation; The step three also includes: At each layer of the decoder, the upsampled feature map is fused with the feature map processed by the Channel-Efficient Convolutional Attention module passed through the jump connection. The fused feature map F fusion It is passed as input to the PFN module for further processing to improve the expression and optimization of features: Feature extraction: In the encoder part, a convolutional neural network is used to extract feature maps of different resolutions. For SCAF-TransUNet, the resolution of the image is gradually reduced through convolution and pooling operations to generate feature maps of different levels. Assuming that the size of the input image is H×W, the resolution of the feature map is gradually reduced, and these multi-scale feature maps will be used as the input of the subsequent FPN. Each feature map C i They all contain spatial and semantic information at different levels; Horizontal connection: At each resolution level, the feature map is compressed by 1x1 convolution to obtain the reduced-dimensional feature map P i ; The role of 1x1 convolution is to reduce the number of channels of the feature map, reduce the computational complexity, and retain the most important key information: P5=Conv1x1(C5), P4=Conv1x1(C4), P3=Conv1x1(C3), P2=Conv1x1(C2); Top-down feature fusion: FPN adopts a top-down feature fusion strategy; the highest-level feature map P5 has the smallest spatial resolution and the richest semantic information, while the lower-level feature maps have higher resolution and more spatial details; When fusion is performed, the low-resolution feature map is first upsampled to a higher resolution and then added element-by-element to the higher-resolution feature map: P4=P4+Upsample(P5) P5=P5+Upsample(P4) P2=P2+Upsample(P3) Through this top-down approach, the global semantic information of high-level feature maps is propagated to low-level feature maps step by step, while retaining the low-level detail information; Upsampling operation: The upsampling operation enlarges the low-resolution feature map to a size that matches the high-resolution feature map through bilinear interpolation or deconvolution; Convolution after fusion: After each upsampling and feature fusion, FPN further processes the fused feature map by applying a 3x3 convolution operation; this convolution operation helps smooth the feature map and reduce aliasing effects while maintaining the combination of global semantic information and local details: P i =Conv3x3(C i ); Combination of FPN and skip connection: In the improved SCAF-TransUNet, the skip connection part has introduced the Channel-Efficient Convolutional Attention module; the feature map generated by FPN will be combined with the skip connection feature map from the encoder: P i =P i +SkipConnection(CECA(C i ))。 2. The liver tumor segmentation method based on SCAF-TransUNet according to claim 1, characterized in that: The superpixel preprocessing of the acquired liver tumor CT image to be segmented includes: Step 1.1, initialize the seed: Suppose the image has N pixels and we want to divide the image into K superpixel regions, then the area of ​​each superpixel is N / K, and the distance between seed points is: Step 1.2, gradient adjustment seed point: In the n×n neighborhood of each seed point, usually n=3, the gradient value of the pixel is calculated and the seed is moved to the pixel position with the smallest gradient; Step 1.3, calculate the color distance d c : Among them, L, a, and b are the three channels of the Lab color space, p represents the current pixel, and c represents the superpixel seed point; Step 1.4: Calculate the spatial distance d s : Among them, x and y represent pixel coordinates respectively; Step 1.5, final distance metric D′: Combining color distance and spatial distance, a weighted metric is used to measure the distance from each pixel to the seed. In order to balance the influence of color distance and spatial distance, a weight parameter m is required. This value is used to control the relative importance of color and space, and is usually set to 10: in, is the step size between seed points, used to normalize spatial distances; Step 1.6, iterative optimization: The seed point and superpixel distribution are optimized by iterating the previous step multiple times. In each iteration, the seed point is gradually adjusted to the center of the superpixel area it affects, and finally stabilizes to form a superpixel image.

3. The liver tumor segmentation method based on SCAF-TransUNet according to claim 1, characterized in that: In the step 1, the connectivity between super-pixel regions needs to be enhanced. Since discontinuous super-pixel regions may be generated during segmentation, the connectivity enhancement step is finally performed to ensure the coherence of the super-pixel regions, remove isolated pixels, and optimize the discontinuous super-pixel regions.

4. The liver tumor segmentation method based on SCAF-TransUNet according to claim 1, characterized in that: The step of inputting the super-pixel image obtained through super-pixel preprocessing into the encoder part, extracting multi-scale features of the input image through a convolutional neural network (CNN) includes: In the encoder part, a convolutional neural network (CNN) is used to extract feature maps of different resolutions. For TransUNet, through convolution and pooling operations, the resolution of the image is gradually reduced to generate feature maps of different levels. Assuming that the size of the input image is H×W, the resolution of the feature map is gradually reduced. During the convolution operation, the input image will be gradually reduced to different feature maps. The feature map reduction process is as follows: C1: The input image size is H×W, where H represents the height of the input image, that is, the number of pixels in the vertical direction, and W represents the width of the input image, that is, the number of pixels in the horizontal direction; C2: After a pooling operation, the image resolution is reduced to Forming a low-level feature map; C3; pooling again, the image resolution is further reduced to C4: Finally, the image resolution is reduced to At this time, the feature map reflects more local information; The calculation formula of the convolutional layer operation is expressed as: F i =σ(W i *I i +b i ) Among them, F i Represents the feature map generated after the convolution operation; W i Represents the weight matrix of the convolution kernel; I i Represents the input image or the feature map of the previous layer; b i represents the bias term; * represents the convolution operation; σ is the nonlinear activation function.

5. The liver tumor segmentation method based on SCAF-TransUNet according to claim 1, characterized in that: The Transformer Layer of SCAF-TransUNet divides the CNN feature map into small patches and uses the self-attention mechanism to capture long-distance dependencies in the image, including: Step 2.1, visualization and embedding: First, the feature map is divided into P×P small blocks and expanded into a one-dimensional vector; each small block is linearly mapped to a high-dimensional space and the position information is added: Where E is the linear embedding matrix; E pos Expression position coding; Step 2.2, multi-head self-attention mechanism: Through the self-attention mechanism, the similarity between each small block is calculated, and the formula is: Among them, Q, K, V are query, key and value vectors respectively. The similarity between each position is calculated and the final output is obtained by weighting; K T represents the transpose of the key matrix, d k Represents the dimension of the key vector; Step 2.3, Multilayer Perceptron MLP and output update: After the self-attention mechanism, a multi-layer perceptron MLP is used for nonlinear changes, and finally the output of the Transformer layer is obtained: WITH l =MLP(LN(Z′ l ))+Z′ l Among them, Z l Describe the final output of the current layer; Z′ l represents the intermediate output after the multi-head self-attention mechanism; LN(Z′ l ) indicates that the intermediate output layer is normalized to make the data more stable; MLP indicates that nonlinear transformation is performed through multi-layer perceptron to improve the feature representation capability; +Z′ l Represents a residual connection, which is used to preserve input information and enhance gradient flow.

6. The liver tumor segmentation method based on SCAF-TransUNet according to claim 1, characterized in that: The step three includes: The feature map passed by the encoder is processed by the Channel-Efficient Convolutional Attention module, and the output is the feature map after channel attention weighting: Step 3.1, input feature map F: Assume that the input feature map is F, H, W, and C are the height, width, and number of channels of the feature map respectively; Step 3.2, global average pooling: Perform global average pooling on the input feature map F to generate a global representation of each channel; obtain a one-dimensional vector v∈R C , where each element represents the global information of a channel; the formula is as follows: Step 3.3, one-dimensional convolution instead of fully connected layer: Use 1D convolution instead of the fully connected layer to generate dependency weights between channels; the size of the convolution kernel determines the range of cross-channel dependency: W = Conv1D(v); Step 3.4, activation of channel weights: Use the Sigmoid function to map the weight W to the range of [0,1]: W′=σ(W) Step 3.5, feature map weighting: Use the channel weight W′ to weight each channel of the input feature map to generate a weighted channel feature map: F′=F c ·W',c=1,2...,C。 7. The liver tumor segmentation method based on SCAF-TransUNet according to claim 6, characterized in that: The step three includes: The channel-efficient convolutional attention module is used to process the feature map weighted by channel attention, and the output is the feature map weighted by spatial attention: Step 3.6, input feature map F′: After channel attention weighting, feature map F′∈R H×W×C Passed into the spatial attention module; Step 3.7, pooling operation: In the spatial attention module, global average pooling and global maximum pooling are applied to compress along the channel dimension to generate two spatial feature maps F avg ∈R H×W×C and F max ∈R H×W×C ; F max =max(F′(i,j,c)),c=1,2....,c; Step 3.8, feature map concatenation: Concatenate the two pooled feature maps along the channel dimension to form a 2D feature map F concat ∈R H×W×2 ; Step 3.9, convolution operation: perform a 3x3 convolution on the concatenated feature map to generate a spatial attention weight map M s ∈R H×W : M s =σ(Conv 3x3 (F concat )) Where σ is the Sigmoid activation function, which limits the attention weight to between [0,1]; Step 3.10, spatial attention weighting: Use the spatial attention weight M s Spatial weighting of the previously weighted feature map F′: F″(i,j,c)=F'(i,j,c)·M s (i,j),i=1,2....,H,j=1,2....,W; Step 3.11, final output: the final output feature map F″(i, j, c)∈R H×W×C It is the feature map after channel attention and spatial attention weighting.

8. The liver tumor segmentation method based on SCAF-TransUNet according to claim 1, characterized in that: In the SCAF-TransUNet decoder in step 3, upsampling is performed by using the feature map fused with FPN. After each upsampling, the resolution of the feature map is gradually restored until the original resolution of the input image is reached. Among them, P i Represents the input feature map of the current layer, the upsampled feature map The chapter will be passed to the next level; Finally, the resolution of the output image is restored to the size of the original image, and the segmentation result is: Output = Upsample (P2).

Citation Information

Patent Citations

  • Liver tumor CT image segmentation method based on AGCH-Net and multi-scale fusion

    CN117495882A

  • CT pancreatic tumor automatic segmentation method and system, terminal and storage medium

    WO2024000161A1