Liver tumor CT image segmentation method
By combining a parallel encoder and an adaptive feature fusion module with CNN and Transformer, the problems of insufficient local feature extraction and high computational complexity in existing liver tumor segmentation methods are solved, and efficient liver tumor image segmentation is achieved.
Patent Information
- Application Number
- CN202510916890.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-11-14
AI Technical Summary
In existing technologies, liver tumor segmentation methods based on convolutional neural networks suffer from insufficient extraction of local feature information and high computational resource requirements. Hybrid models fail to fully utilize the advantages of CNN and Transformer, resulting in insufficient feature fusion and high computational complexity.
A parallel encoder structure is adopted, combining CNN branches and Transformer branches. Local and global features are extracted through depthwise separable convolution and hierarchical SwingTransformer modules. An adaptive feature fusion module is designed, and a boundary-sensitive adaptive hybrid loss function is used to optimize the training process.
High-precision segmentation of liver tumor images was achieved, improving feature fusion efficiency, reducing computational complexity, and enhancing the model's segmentation performance.
Smart Images

Figure CN120953600A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image processing technology, specifically relating to a method for segmenting liver tumor CT images, which is applicable to the automated and accurate segmentation of liver tumors in CT images. Background Technology
[0002] Liver cancer, as a malignant tumor, has seriously endangered human health and posed a significant public health challenge. Computed tomography (CT) scanning technology is the primary means of liver tumor examination, clinical diagnosis, and surgical support. Traditional clinical imaging diagnosis relies on manual annotation by doctors, a time-consuming and labor-intensive process that carries the risk of misdiagnosis or missed diagnosis due to the doctor's subjective judgment.
[0003] Currently, the main technical approach relies on U-shaped encoder-decoder models based on convolutional neural networks (CNNs) and their improved variants. While these models have achieved automatic segmentation of liver tumors and improved segmentation accuracy to some extent, their performance is limited by the limited receptive field of CNN kernels, making it difficult to capture long-range relationships and global information. With the introduction of the Transformer model, its ability to effectively model long-range dependencies through its self-attention mechanism has proven outstanding in various natural language processing tasks. Given the Transformer's advantages in capturing global contextual information, some works have extended it to medical image segmentation, overcoming the limitations of convolutional networks in acquiring global information and the limitations of CNNs in handling long-range dependencies. However, this approach also suffers from a lack of local feature extraction and high computational resource requirements.
[0004] To fully leverage the advantages of CNN and Transformer networks, hybrid CNN-Transformer models have rapidly developed. However, current CNN-Transformer hybrid networks typically employ simple concatenated or parallel structures, failing to fully utilize the local modeling capabilities of CNNs and the global modeling advantages of Transformers. Existing feature fusion methods mostly use fixed weight concatenation, which cannot dynamically adapt to changes in feature importance at different scales, resulting in insufficient feature fusion. Furthermore, high computational complexity and enormous training overhead remain problems. Summary of the Invention
[0005] To fully utilize the respective advantages of CNN and Transformer and solve the problems of insufficient feature fusion and large computational load in the existing technologies, this invention proposes a liver tumor image segmentation method based on parallel encoder and feature fusion.
[0006] First, a parallel hybrid encoder structure including CNN and Transformer branches is proposed. The CNN branch uses depthwise separable convolution to extract local features to reduce computation, while the Transformer branch uses a hierarchical Swing Transformer module to model global dependencies to enhance long-distance feature association. Then, a feature fusion module is designed to fuse the output features of the CNN and Swing Transformer branches to achieve adaptive feature fusion.
[0007] 1. To achieve the above objectives, the present invention proposes a method for segmenting liver tumor CT images, characterized in that the method includes the following steps:
[0008] S1: Obtain the original dataset and perform data slicing and preprocessing operations;
[0009] S2: Divide the preprocessed data into datasets;
[0010] S3: Build a liver tumor segmentation model based on parallel encoder and feature fusion;
[0011] S4: Design Feature Fusion Module;
[0012] S5: Construct a hybrid loss function;
[0013] S6: Train the model and load the optimal model parameters for testing.
[0014] 2. The method according to claim 1, characterized in that step S1, acquiring the original dataset and performing data slicing and preprocessing operations, includes:
[0015] (1) Load the original data in nii format, slice the original image data and corresponding label data along the axis, keep only the slices containing the tumor, remove invalid data, and finally save the slices in npy format.
[0016] (2) Crop the image data slices using HU values, and set the minimum lower limit HU value for the cropping range. min and the maximum value HU max Crop to a specific area:
[0017]
[0018] Where I HU This is the original HU value, setting all values below the lower limit to the minimum value and values above the upper limit to the maximum value;
[0019] (3) Perform min-max normalization on the data after HU clipping. The mathematical expression is:
[0020]
[0021] Linearly map the original data to the interval [0,1].
[0022] 3. The method according to claim 1, characterized in that the design of the liver tumor segmentation model in step S3 includes:
[0023] This model uses the Unet network as its framework and employs a parallel encoder and a progressive upsampling decoder structure. The encoder consists of a CNN branch, a Transformer branch, and a feature fusion module. Each branch contains four stages. In the first two stages of the CNN branch, features are extracted using depthwise separable convolutions. In the second stage, residual connections are introduced to alleviate the gradient vanishing problem in deep training. Then, downsampling is performed with a stride of 2. In the third stage, dilated convolutions are introduced to expand the receptive field. Finally, global average pooling is used to obtain high-level semantic information. In the first stage, the input of the Transformer branch is first processed through patch partitioning and linear embedding, and then through multiple SwinTransformer Blocks. The remaining stages use a combination of patch merging and multiple Swin Transformer Blocks. In each stage, the output of each branch is used as the input of the feature fusion module for feature fusion. The final output of the fused features from each stage is used for subsequent skip connections.
[0024] The decoder gradually restores the spatial resolution through transposed convolution. In each upsampling stage, it sequentially fuses the skip connection features from stage three, stage two, and stage one, and finally outputs a segmentation mask through 1×1 convolution and Sigmoid activation.
[0025] 4. The method according to claim 1, characterized in that the implementation of the feature fusion module in step S4 includes:
[0026] (1) First, unify the feature channel dimension of each branch output through convolution. Then, use the cross-attention mechanism with Swin features as key / value to guide the CNN features to learn global dependencies. The CNN features are used as key / value to supplement the Swin features with local details to calculate the attention weights. The mathematical expression is:
[0027]
[0028] Finally, multiply its weight by the corresponding value to generate the attention output;
[0029] (2) In the adaptive fusion stage, the two features that have undergone cross-attention processing are concatenated along the channel dimension, then the spatial information is compressed by global average pooling, and then the channel relationship is learned by two layers of 1×1 convolutions containing ReLU. Finally, the fusion weight w is generated by Softmax. CNN and w Transformer The mathematical expression for the entire weight generation process is as follows:
[0030]
[0031] in For concatenation along the channel dimension, GAP uses global average pooling;
[0032] The output features of the adaptive fusion stage are obtained by weighted fusion using generated weights, and the mathematical expression is as follows:
[0033] F fused =w CNN ·F CNN +w Transformer ·F Transformer
[0034] In the formula, F CNN F Transformer These represent the inputs of each branch after being processed with a unified feature channel dimension;
[0035] (3) The output of the adaptive fusion stage is subjected to feature enhancement operation through depthwise separable convolution, and the final fused feature is output after GELU activation.
[0036] 5. The method according to claim 1, characterized in that step S5, which constructs a hybrid loss function, specifically includes the following:
[0037] A boundary-sensitive adaptive hybrid loss function (BSAM Loss) is proposed, and its mathematical expression is as follows:
[0038] L BSAM =σ(α)·L WBCE +(1-σ(α))·L Dice
[0039] In the formula, L WBCE It is a binary cross-entropy loss function that introduces a boundary-sensitive weighting mechanism, and its mathematical expression is:
[0040]
[0041] g i It is the true label of pixel i in the sample image; p i W is the model's prediction of pixel i in the sample image; N is the number of pixels in the sample image; Wi This is the weight of each pixel, calculated using the following formula:
[0042] W i =1+λ|(L*y) i |
[0043] Where L is the Laplacian operator, λ is the boundary enhancement coefficient, y represents the ground truth label map, and * represents the convolution operation;
[0044] L Dice This is the Dice loss function, and its mathematical expression is:
[0045]
[0046] ∑p i y i ∑p represents the intersection of the prediction obtained by summing pixel-wise products with the label. i and ∑y i These are the total number of foreground pixels for prediction and labeling, respectively, and ∈ is a smoothing term to prevent division by zero.
[0047] σ(α) represents the weight factor of the sigmoid function, which maps the unconstrained scalar parameter α to the (0,1) interval to ensure the rationality of the mixed weight value. Attached Figure Description
[0048] Figure 1 This is a flowchart of the implementation method of the present invention;
[0049] Figure 2 This is a diagram showing the CT image preprocessing results of the present invention;
[0050] Figure 3 This is a diagram of the parallel encoder structure of the segmentation model of this invention;
[0051] Figure 4 This is a diagram of the Swing Transformer Block structure;
[0052] Figure 5 This is a basic structural diagram of the feature fusion module designed in this invention. Detailed Implementation
[0053] The present invention will now be described in further detail with reference to the accompanying drawings and examples.
[0054] The implementation flowchart of the embodiments of the present invention is shown in the appendix. Figure 1 As shown, the liver tumor CT image segmentation method in this example includes the following steps:
[0055] S1: Obtain the original dataset and perform data slicing and preprocessing operations;
[0056] S101: Load the raw CT image data in nii format, slice the raw image data and corresponding label data along the cross-sectional direction of the tomographic scan to obtain 2D data slices, and finally save the data slices in npy format; the data used are all from the public dataset LiTS2017, which contains 131 cases of raw abdominal CT images and corresponding manually labeled liver regions. Each data slice obtained is a single-channel 512×512 grayscale image;
[0057] S102: Perform HU value cropping, windowing, and normalization operations on the image data slices. Set the minimum HU value to -100; all pixels below this value will be set to -100. The maximum HU value to 200; all pixels above this value will be set to 200. Finally, use the minimum-maximum-value normalization method to map all data values to the range [0,1]. The preprocessing results are shown in the appendix. Figure 2 As shown;
[0058] S103: Use the where function of the NumPy library to convert the single-channel multi-class label image into a single-channel binary label image for the corresponding label data.
[0059] S2: Randomly sample the preprocessed data slices and divide the dataset into training, validation and test sets in an 8:1:1 ratio.
[0060] S3: Build a liver tumor segmentation model with parallel encoder and feature fusion;
[0061] S301: The basic structure diagram of the encoder is attached. Figure 3 As shown. The input is a single-channel 512×512 CT image. In the first-stage CNN branch, the input image is first scaled to 256×256 using bilinear interpolation, and then passed through a 4×4 depthwise separable convolution with a stride of 1 and a 3×3 depthwise separable convolution with a stride of 2, finally outputting 64×128×128 resolution features. The Transformer branch directly divides the input image into 4×4 non-overlapping blocks using Patch Partition to reduce computation, initially establishing pixel relationships within the module, and then feeding it into two Swin Transformer Blocks after linear embedding to obtain 128×128×64 resolution features. The basic structure of the Swin Transformer Block is shown in the attached figure. Figure 4 As shown, the final dual-branch output is feature F1 after feature fusion;
[0062] S302: In the second stage, the CNN branch passes the output of the previous stage branch through two layers of 3×3 depthwise separable convolutions with a stride of 1, and then passes them through 3×3 depthwise separable convolutions with a stride of 2 to downsample to 128×64×64. Residual connections are introduced to alleviate the gradient vanishing problem in deep training. The Transformer branch passes the output of the previous stage through Patch Merging and then through two Swing Transformer Blocks to obtain 64×64×128 resolution features. Finally, the fused feature F2 is output.
[0063] S303: In the third stage, the CNN branch introduces dilated convolutions with dilation rates of 2 and 4 to expand the receptive field. Then, it is downsampled to 256×32×32 by a 3×3 depthwise separable convolution with a stride of 2. The Transformer branch uses the same structure as the previous stage to obtain 32×32×256 resolution features. After feature fusion, the output feature F3 is obtained.
[0064] S304: Finally, in the fourth stage, the input of the CNN branch first undergoes global average pooling to obtain high-level semantic information and obtains 512×16×16 resolution features. The Transformer branch is the same as the previous stage, outputting 16×16×512 resolution features. Finally, feature fusion is used to output feature F4.
[0065] S305: The decoder gradually restores the spatial resolution through transposed convolution. In each upsampling stage, the skip connection fusion features F3, F2, and F1 from stage three, stage two, and stage one are fused in sequence. Finally, the output segmentation mask is activated by 1×1 convolution and Sigmoid.
[0066] S4: Design Feature Fusion Module;
[0067] S401: The basic structure of the feature fusion module is shown in the appendix. Figure 5 As shown, the outputs of each branch are first convolved to unify the feature channel dimensions. Then, a cross-attention mechanism is used, with Swin features as keys and values, to guide the CNN features to learn global dependencies. The CNN features, as keys and values, supplement the Swin features with local details to calculate the attention weights. The mathematical expression is:
[0068]
[0069] Finally, multiply its weight by the corresponding value to generate the attention output;
[0070] S402: In the adaptive fusion stage, the two features processed by cross-attention are concatenated along the channel dimension, then global average pooling is used to compress spatial information, followed by two layers of 1×1 convolutions containing ReLU to learn the channel relationships, and finally, Softmax is used to generate the fusion weight w. CNN and w Transformer The mathematical expression for the entire weight generation process is as follows:
[0071]
[0072] in For concatenation along the channel dimension, GAP uses global average pooling; the output features of the adaptive fusion stage are obtained by weighted fusion using generated weights, and the mathematical expression is:
[0073] F fused =w CNN ·F CNN +w Transformer ·F Transformer
[0074] In the formula, F CNN F Transformer These represent the inputs of each branch after being processed with a unified feature channel dimension;
[0075] S403: The output of the adaptive fusion stage is subjected to feature enhancement operation through depthwise separable convolution, and the final fused feature is output after GELU activation.
[0076] S5: Construct a hybrid loss function;
[0077] To address the problems of existing binary cross-entropy and Dice hybrid loss functions, such as gradient conflicts leading to training oscillations, boundary insensitivity failing to effectively handle key boundary regions in medical images, and fixed-ratio hybrid weights being unsuitable for different training stages and sample characteristics, a boundary-sensitive adaptive hybrid loss function (BSAM Loss) is proposed. Its mathematical expression is as follows:
[0078] L BSAM =σ(α)·L WBCE +(1-σ(α))·L Dice
[0079] In the formula, L WBCE It is a binary cross-entropy loss function that introduces a boundary-sensitive weighting mechanism, and its mathematical expression is:
[0080]
[0081] g i p is the true label of pixel i in the sample image. iThis is the model's prediction of pixel i in the sample image, where N is the number of pixels in the sample image, and W... i This is the weight of each pixel, calculated using the following formula:
[0082] W i =1+λ|(L*y) i |
[0083] Where L is the Laplacian operator, λ is the boundary enhancement coefficient, y represents the ground truth label map, and * represents the convolution operation;
[0084] L Dice This is the Dice loss function, and its mathematical expression is:
[0085]
[0086] ∑p i y i ∑p represents the intersection of the prediction obtained by summing pixel-wise products with the label. i and ∑y i These are the total number of foreground pixels for prediction and labeling, respectively, and ∈ is a smoothing term to prevent the denominator from being zero;
[0087] σ(α) represents the weight factor of the sigmoid function, which maps the unconstrained scalar parameter α to the (0,1) interval to ensure the rationality of the mixed weight value.
[0088] S6: Train the model and load the optimal model parameters for testing;
[0089] S601: The Adam optimizer is used, with an initial learning rate of 0.003, a batch size of 2, a weight decay of 0.00005, a total of 150 training epochs, and a tolerance value of 20 for the early stopping strategy.
[0090] S602: Train the network model and optimize hyperparameters on the training and validation sets using the backpropagation algorithm until the loss function basically converges, thus completing the training.
[0091] S603: Load the optimal parameters obtained from the above training onto the constructed segmentation model, perform the same preprocessing operation on the test set data, and input it into the model to obtain the final segmentation result.
[0092] The foregoing has described specific illustrative embodiments of the present invention to enable those skilled in the art to understand the invention. It should be noted that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various modifications are readily apparent as long as they fall within the spirit and scope of the invention as defined and determined by the appended claims. All inventions utilizing the concept of this invention are protected.
Claims
1. A method for segmenting liver tumor CT images, characterized in that, The method includes the following steps: S1: Obtain the original dataset and perform data slicing and preprocessing operations; S2: Divide the preprocessed data into datasets; S3: Build a liver tumor segmentation model based on parallel encoder and feature fusion; S4: Design Feature Fusion Module; S5: Construct a hybrid loss function; S6: Train the model and load the optimal model parameters for testing.
2. The method according to claim 1, characterized in that, Step S1, which involves obtaining the original dataset and performing data slicing and preprocessing, includes: (1) Load the original data in nii format, slice the original image data and corresponding label data along the axis, keep only the slices containing the tumor, remove invalid data, and finally save the slices in npy format. (2) Crop the image data slices using HU values, and set the minimum lower limit HU value for the cropping range. min and the maximum value HU max Crop to a specific area: Where I HU This is the original HU value, setting all values below the lower limit to the minimum value and values above the upper limit to the maximum value; (3) Perform min-max normalization on the data after HU clipping. The mathematical expression is: Linearly map the original data to the interval [0,1].
3. The method according to claim 1, characterized in that, The design of the liver tumor segmentation model in step S3 includes: This model uses the Unet network as its framework and employs a parallel encoder and a progressive upsampling decoder structure. The encoder consists of a CNN branch, a Transformer branch, and a feature fusion module. Each branch contains four stages. In the first two stages of the CNN branch, features are extracted using depthwise separable convolutions. In the second stage, residual connections are introduced to alleviate the gradient vanishing problem in deep training. Then, downsampling is performed with a stride of 2. In the third stage, dilated convolutions are introduced to expand the receptive field. Finally, global average pooling is used to obtain high-level semantic information. In the first stage, the input of the Transformer branch is first processed through patch partitioning and linear embedding, followed by multiple Swin Transformer Blocks. The remaining stages use a combination of patch merging and multiple Swin Transformer Blocks stacked together. In each stage, the output of each branch is used as the input of the feature fusion module for feature fusion. The final output of the fused features from each stage is used for subsequent skip connections. The decoder gradually restores the spatial resolution through transposed convolution. In each upsampling stage, it sequentially fuses the skip connection features from stage three, stage two, and stage one, and finally outputs a segmentation mask through 1×1 convolution and Sigmoid activation.
4. The method according to claim 1, characterized in that, The implementation of the feature fusion module in step S4 includes: (1) First, unify the feature channel dimension of each branch output through convolution. Then, use the cross-attention mechanism with Swin features as key / value to guide the CNN features to learn global dependencies. The CNN features are used as key / value to supplement the Swin features with local details to calculate the attention weights. The mathematical expression is: Finally, multiply its weight by the corresponding value to generate the attention output; (2) In the adaptive fusion stage, the two features that have undergone cross-attention processing are concatenated along the channel dimension, then the spatial information is compressed by global average pooling, and then the channel relationship is learned by two layers of 1×1 convolutions containing ReLU. Finally, the fusion weight w is generated by Softmax. CNN and w Transformer The mathematical expression for the entire weight generation process is as follows: in For concatenation along the channel dimension, GAP uses global average pooling; The output features of the adaptive fusion stage are obtained by weighted fusion using generated weights, and the mathematical expression is as follows: F fused =w CNN ·F CNN +w Transformer ·F Transformer In the formula, F CNN F Transformer These represent the inputs of each branch after being processed with a unified feature channel dimension; (3) The output of the adaptive fusion stage is subjected to feature enhancement operation through depthwise separable convolution, and the final fused feature is output after GELU activation.
5. The method according to claim 1, characterized in that, Step S5 involves constructing a hybrid loss function, which includes the following: A boundary-sensitive adaptive hybrid loss function (BSAM Loss) is proposed, and its mathematical expression is as follows: L BSAM =σ(α)·L WBCE +(1-σ(α))·L Dice In the formula, L WBCE It is a binary cross-entropy loss function that introduces a boundary-sensitive weighting mechanism, and its mathematical expression is: g i It is the true label of pixel i in the sample image; p i It is the result of the model predicting pixel i in the sample image; N is the number of pixels in the sample image; W i This is the weight of each pixel, calculated using the following formula: W i =1+λ|(L*y) i | Where L is the Laplacian operator, λ is the boundary enhancement coefficient, y represents the ground truth label map, and * represents the convolution operation; L Dice This is the Dice loss function, and its mathematical expression is: ∑p i y i ∑p represents the intersection of the prediction obtained by summing pixel-wise products with the label. i and ∑y i These are the total number of foreground pixels for prediction and labeling, respectively, and ∈ is a smoothing term to prevent division by zero. σ(α) is the weighting factor passed through the sigmoid function, ensuring that the scalar parameter α is mapped to the (0,1) interval.