Liver cancer segmentation system and method based on deep learning

By introducing multifunctional modules, such as the double cavity convolution module and the polarized multi-scale feature self-attention module, the problems of loss of feature information and unconsidered context information in liver cancer segmentation are solved, and precise segmentation and diagnostic assistance of liver cancer areas are realized.

CN120125595APending Publication Date: 2025-06-10CHONGQING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510048918.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing deep learning algorithms have problems such as low segmentation accuracy, loss of feature information, and unconsideration of context information in the liver cancer segmentation task, which is difficult to meet the precise needs of liver cancer diagnosis.

Method used

A liver cancer segmentation system based on deep learning is proposed, adopting a multifunctional module combination, including a double hollow convolution module, an attention module, a large-core attention gate module, a downsampling module, an upsampling module and a polarized multi-scale feature self-attention module. Through these modules, feature information is extracted and enhanced, and the problems of feature loss and context information are not considered.

Benefits of technology

Accurate segmentation of liver cancer areas is achieved, segmentation accuracy is improved, feature information is reduced, and context information is enhanced by the model, thereby assisting doctors in more accurate diagnosis of liver cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125595A_ABST
    Figure CN120125595A_ABST
Patent Text Reader

Abstract

The invention provides a liver cancer segmentation system and method based on deep learning. A dual-cavity convolution module extracts features of an input image; the attention module introduces a channel attention mechanism, extracts context information of the feature map by using average pooling and convolution operation, and generates an attention map according to the context; the big kernel attention gate module respectively performs convolution processing on the feature maps output by the encoder and the decoder, and combines high-level features with low-level features to obtain an enhanced feature map; the down-sampling module is used for carrying out average pooling and maximum pooling on the input feature maps respectively, and splicing the feature maps; the up-sampling module is used for generating a dynamic offset from the input feature map to obtain a coordinate after dynamic offset, and sampling on dynamic interpolation is carried out; the polarized multi-scale feature self-attention module carries out resolution integration and channel attention and space attention calculation on an input image in sequence. And the hepatocellular carcinoma area can be accurately segmented, so that a doctor can be effectively assisted to diagnose and treat liver cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image segmentation in medical cancer, and particularly provides a liver cancer segmentation system and method based on deep learning. Background Art

[0002] The liver is one of the most important organs in the human body to maintain life activities, and has functions such as helping the human body's metabolism and detoxification. However, due to the increasingly rich material life, people have gradually developed bad living habits such as overeating and staying up late, which makes the burden on the human liver heavier and heavier, and it is extremely easy to suffer from various problems, such as liver cancer. Liver cancer is the most common malignant tumor, but due to its early symptoms being not obvious, the best way to prevent it currently is early detection and early treatment. Clinically, since computed tomography (CT) can clearly show the internal tissue structure of the human body and facilitate doctors to diagnose the condition, it has become one of the most important means for the prevention, diagnosis and treatment of liver cancer. However, due to the traditional manual diagnosis method being time-consuming and laborious, and being extremely susceptible to the influence of doctors' personal experience, in the face of a large number of liver CT image layers, even for a single patient's images ranging from dozens to hundreds of layers, doctors will inevitably have the phenomenon of missed diagnosis or misdiagnosis. Therefore, in modern liver cancer treatment and diagnosis, there is an urgent need for a precise and fast liver cancer image segmentation technology algorithm to achieve rapid automatic segmentation of liver CT, clearly show the location and size of liver cancer, so as to assist doctors in diagnosing the condition.

[0003] With the development of deep learning technology, it has shown broad prospects in the field of medical diagnosis, and has demonstrated extremely excellent performance in multiple medical directions, such as the detection and segmentation of lung nodules and breast cancer, and its detection and segmentation results are sufficient to reach the level of experienced doctors' manual detection and segmentation. However, for the segmentation of liver tumors, although some scholars in the industry have used networks such as UNet or its variants such as SelfReg-UNet for the segmentation of liver cancer, due to the disadvantages of low contrast, blurred edges in liver CT images, and the different shapes and sizes of liver tumor regions, as well as low contrast in the edge regions, these algorithms are not satisfactory and it is difficult to meet the standards required for liver cancer segmentation. Therefore, currently, the deep learning algorithms that can segment the liver cancer region from CT images usually have the problem of low segmentation accuracy. There are currently multiple difficulties in liver cancer segmentation based on deep learning:

[0004] 1. The liver and surrounding organs have similar shapes and colors, there is no obvious distinction between the colors and shapes of liver cancer cells and normal cells, and the shapes and sizes of liver cancer regions are different, making the segmentation accuracy of current deep learning network methods such as UNet and SelfReg-UNet algorithms relatively low;

[0005] 2. When current liver cancer segmentation algorithms such as UNet and Attention-UNet models perform feature downsampling and upsampling, the methods used are too simple, resulting in the loss of many original feature information after downsampling by the model. When performing upsampling, the context information of the input feature map cannot be considered, leading to the loss of position features;

[0006] 3. Existing liver cancer segmentation algorithms such as UNet and SAR-UNet do not fully utilize the information of all feature maps extracted by their model encoders during feature extraction, and do not consider the long-term dependence relationship between the feature maps output by encoders of different layers. Summary of the Invention

[0007] Based on this, the present invention provides a liver cancer segmentation system and method based on deep learning, which can accurately segment the hepatocellular carcinoma region, thereby effectively assisting doctors in the diagnosis and treatment of liver cancer.

[0008] To achieve the above object, in the first aspect, the present invention provides a liver cancer segmentation system based on deep learning, including multiple functional modules:

[0009] Double Atrous Convolution Module, extracting the features of the input image;

[0010] Attention Module, introducing a channel attention mechanism, using average pooling and convolution operations to extract the context information of the feature map, and then generating an attention map according to the context to automatically select more important features and enhance their expression;

[0011] Large-kernel Attention Gate Module (LGAG), convolving the output of the encoder and the feature map output by the decoder respectively, and adding them together to combine high-level features with low-level features to obtain an enhanced feature map;

[0012] Downsampling Module, performing average pooling and max pooling on the input feature map respectively, and splicing them to avoid losing detailed features;

[0013] Upsampling Module, generating a dynamic offset for the input feature map, obtaining the coordinates after dynamic offset, and performing dynamic interpolation upsampling; and

[0014] Polarized Multi-scale Feature Self-attention Module (PMFS), performing resolution integration, channel attention, and spatial attention calculations on the input image in sequence to enhance multi-scale feature fusion.

[0015] Further, the dual dilated convolution module includes a dilated convolution and an activation part, and its working process includes:

[0016] The first layer:

[0017] F 1 = LeakReLU(BatchNorm(DilationConv(X))) (1)

[0018] The second layer:

[0019] F 2 = LeakReLU(BatchNorm(DilationConv(F 1 ))) (2)

[0020] Then, a dropout layer is added after equation (2) to reduce overfitting of the model:

[0021] F 3 = Dropout(F 2 ) (3)

[0022] Then, the input X is connected with the output F3 of the convolutional layer in a residual connection to complement the original feature information that may be lost during the feature extraction process:

[0023] R = X + F 3 (4)

[0024] Finally, the output R of the residual connection is input into the context anchor attention module (CAA) to enhance the feature information of R;

[0025] Y = CAA(R) (5)

[0026] Wherein, X represents the input image features, and Y represents the output high-level image features after passing through the dual dilated convolution module.

[0027] Further, the working process of the attention module includes:

[0028] First, global average pooling is performed on the input feature X to reduce the spatial dimension:

[0029] F 1 = Avgpool(X) (6)

[0030] Then, the feature is processed through 1*1 convolution, batch normalization, and activation function:

[0031] F 2 = LeakRelu(Batchnorm(Conv 1×1 (F 1 ))) (7)

[0032] Then, perform depth convolution operations for feature extraction:

[0033] F 3 = Conv 1×1 (Conv 1×7 (Conv 7×1 (F 2 ))) (8)

[0034] Then, normalize and activate the extracted feature maps:

[0035] F 4 = LeakRelu(Batchnorm(F 3 )) (9)

[0036] Then, generate the final attention weights through the sigmoid function:

[0037] F out = Sigmoid(F 4 ) (10)

[0038] Finally, multiply the above output element-wise with the original input feature map to obtain the feature map enhanced by features;

[0039] Y = X·F out (11)

[0040] Wherein, X represents the input image features, and Y represents the attention image features extracted after passing through the context anchor attention module (CAA).

[0041] Furthermore, the working process of the large kernel attention gate module (LGAG) includes:

[0042] Perform convolution processing and batch normalization processing on the feature map output by the encoder and the upsampled feature map through two branches respectively:

[0043] F 1 = Batchnorm(Conv 3×3 (X 1 )) (12)

[0044] F 2 = Batchnorm(Conv 3×3 (X 2 )) (13)

[0045] Then, add the feature maps extracted from the two branches element-wise for feature fusion:

[0046] F 3= F 1 + F 2 (14)

[0047] Then, the fused features are processed through 1*1 convolution, normalization, and the sigmoid function (Sigmoid) to generate attention weights:

[0048] F out = Sigmoid(Batchnorm(Conv 1×1 (F 3 ))) (15)

[0049] Finally, the attention weights are applied to the original input X through element-wise multiplication to obtain the final enhanced feature map: 2 = Batchnorm(Conv 3×3 (X 2 ))

[0050] Y = X·F out (16)

[0051] Wherein, X1 is the feature map output by the encoder, X2 is the feature map obtained after passing through the upsampling module, and Y is the high-level feature map output after passing through the large kernel attention gate module (LGAG).

[0052] Furthermore, the working process of the downsampling module includes:

[0053] First, global average pooling and global max pooling operations are respectively performed on the input feature map X for feature extraction operations and to reduce the resolution of the feature map:

[0054] F 1 = Avgpool(X) (17)

[0055] F 2 = Maxpool(X) (18)

[0056] Then, 1*1 convolution, normalization, and activation function processing are respectively performed on the extracted feature maps to further extract features:

[0057] F 1out = LeakRelu(Batchnorm(Conv 1×1 (F 1 ))) (19)

[0058] F 2out = LeakRelu(Batchnorm(Conv 1×1 (F 2 ))) (20)

[0059] Finally, the features extracted from the two branches are concatenated for feature fusion:

[0060] Y = Concatenate(F 1out , F 2out ) (21)

[0061] Among them, X is the high-level feature map output after passing through the double dilated convolution module, and Y is the feature map output after passing through the downsampling module.

[0062] Furthermore, the working process of the upsampling module includes:

[0063] First, the input feature map X generates dynamic offsets through a 1*1 convolutional layer:

[0064] F 1 = Conv 1×1 (X) (22)

[0065] Then, add it to the initialized input network coordinates C to obtain the coordinates after dynamic offset:

[0066] F 2 = F 1 + C (23)

[0067] Then, perform coordinate normalization on the obtained coordinate network:

[0068] F 3 = Normalize(F 2 ) (24)

[0069] Finally, perform dynamic interpolation upsampling on the basis of the input feature map X according to the normalized coordinates:

[0070] Y = Grid_upsample(F 3 ) (25)

[0071] Among them, X is the feature map output after passing through the polarization multi-scale feature self-attention module (PMFS) or the feature map output by the decoder, Normlize represents coordinate normalization, and Gr id_upsample represents dynamic interpolation upsampling.

[0072] Furthermore, the working process of the polarization multi-scale feature self-attention module (PMFS) includes:

[0073] First, normalize the input multiple feature maps respectively through parallel paths to convert them into the same resolution and number of channels:

[0074] F i = Conv(Maxpool(X i)) i = 1, 2, ..., n (26)

[0075] Then, the transformed feature maps are concatenated:

[0076] F all = Concatenate(F i ) i = 1, 2, ..., n (27)

[0077] Then, the channel self-attention is calculated, where the key matrix K, query matrix Q, and value matrix V are calculated by the following formulas respectively:

[0078] K = Conv(F all ) Q = Conv(F all ) V = Conv(F all ) (28)

[0079] Then, the attention weights are generated using the Sigmoid function (Softmax):

[0080] Attention 1 = Softmax(K · Q T ) (29)

[0081] Then, convolution and normalization are performed:

[0082] F refined = Sigmoid(Layernorm(Conv(Attention 1 ))) (30)

[0083] Then, the features are updated to enhance the channel information of the input multi-scale feature maps:

[0084] F attention1 = F refined · V (31)

[0085] · represents element-wise multiplication.

[0086] Furthermore, the working process of the polarization multi-scale feature self-attention module (PMFS) further includes:

[0087] The spatial attention is calculated, and the feature maps enhanced in features obtained above are respectively input into the convolutional layer to obtain the key matrix K, query matrix Q, and value matrix V:

[0088] K = Conv(F attention1 ) Q = Conv(F attention1 ) V = Conv(F attention1 ) (32)

[0089] Then, global average pooling is performed and global attention weights are generated:

[0090] Attention 2 = Softmax(Avgpool(K·Q T )) (33)

[0091] Then, global feature updates of the multi-scale feature maps are performed:

[0092] F attention2 = Sigmoid(Attention 2 )·V (34)

[0093] Finally, the features output by multiple branches are weighted and fused through a convolutional layer and output:

[0094] Y = Conv(F attention1 + F attention2 ) (35)

[0095] Among them, X1, X2, and X3 are the feature maps output by each encoder of the MSA-UNet network, and K, Q, and V are the key matrix, query matrix, and value matrix respectively; Y is the high-level feature extracted through multi-scale feature fusion and self-attention after passing through the polarization multi-scale feature self-attention module (PMFS).

[0096] To achieve the above object, in a second aspect, the present invention provides a liver cancer segmentation method based on deep learning. Using the above-mentioned liver cancer segmentation system based on deep learning, the steps include:

[0097] S100. Input a medical image X with a size of H×W×3;

[0098] S200. Define the MSA-UNet model;

[0099] S300. As the encoder of the downsampling path, for the I1 to IN encoders, perform the following operations: Apply a double dilated convolution module to the input image or feature map for feature extraction, store the generated feature map Fi for skip connection in the subsequent encoder part, and use a downsampling module to reduce the spatial size of the input feature map, and end the loop;

[0100] S400. The bottleneck layer applies a double dilated convolution module to the input feature map and applies the PMFS module to the output feature map for multi-scale feature aggregation enhancement;

[0101] S500. As the decoder of the upsampling path, for the IN to I1 decoder, perform the following operations: Use the upsampling module to increase the spatial dimension of the feature map, use the large kernel attention gate module (LGAG) to optimize and enhance the features of the upsampled feature map, concatenate the stored feature map Fi with the current decoder output feature map, use the double dilated convolution module for feature extraction, and end the loop;

[0102] S600. Output layer: Apply 1×1 convolution to reduce the number of input feature map channels to the number of output classes, and apply the Softmax activation function to generate the probability map;

[0103] S700. Inference process, load the trained liver cancer segmentation algorithm model, and for each test image, perform the following operations: Through the forward propagation of the model, obtain the segmentation mask Ypred, and end the loop;

[0104] S800. Return the predicted H×W×1 segmentation mask Ypred.

[0105] Furthermore, before the inference process, train the liver cancer segmentation algorithm model, and for each training epoch, perform the following operations:

[0106] Load the training images and their corresponding ground truth segmentation masks, perform forward propagation through the liver cancer segmentation algorithm and calculate the predicted masks, calculate the total loss using the Dice loss (similarity coefficient loss) and binary cross-entropy loss, backpropagate the loss and update the network weights using the weighted decay adaptive moment estimation optimizer (AdamW optimizer), and end the loop of the training epoch.

[0107] The technical advantages of a system and method for liver cancer segmentation based on deep learning provided by the present invention are at least reflected in:

[0108] 1. A liver cancer segmentation algorithm (MSA-UNet) based on 2D CNN is proposed, which can accurately and quickly segment the hepatocellular carcinoma region from CT images;

[0109] 2. The segmentation attention module is introduced to improve the interpretability of the model, enabling the model to more effectively segment small liver cancer regions;

[0110] 3. The downsampling module (ADown, Adaptive Downsample) and the upsampling module (DY_upsample, Dynamic Upsample) are introduced, which solve the problems of previous algorithms such as UNet and SAR-UNet losing too many original image feature details during downsampling and only performing simple fixed interpolation during upsampling without considering context information, resulting in the loss of some spatial information in the upsampled feature map;

[0111] 4. Introduce the PMFS module, enabling the algorithm to encode multi-scale long-term dependencies more meticulously and precisely, and through special design, providing a more lightweight solution than the traditional self-attention mechanism. BRIEF DESCRIPTION OF THE DRAWINGS

[0112] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure. With reference to the drawings, the present disclosure can be more clearly understood according to the following detailed description, wherein:

[0113] Figure 1 Schematic diagram of the MSA-UNet structure;

[0114] Figure 2 Structural diagram of the dual dilated convolution module;

[0115] Figure 3 Schematic diagram of the CAA attention module;

[0116] Figure 4 LGAG module;

[0117] Figure 5 Structure of the ADown module;

[0118] Figure 6 Structure of the DY upsampling module;

[0119] Figure 7 Structure of the PMFS module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0120] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the drawings. The description of the exemplary embodiments is merely illustrative and in no way limits the present disclosure and its application or use. The present disclosure can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to make the present disclosure thorough and complete and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, the composition of materials, numerical expressions, and numerical values set forth in these embodiments should be construed as merely exemplary and not as limitations.

[0121] Aiming at the problems of low segmentation accuracy of the above-mentioned automatic liver cancer segmentation algorithm, loss of a lot of original feature information during upsampling and downsampling, and failure to consider the long-term dependence relationship between different feature maps, this patent proposes a liver cancer segmentation algorithm based on convolutional neural network (CNN) (Multi-scale feature Self-Attention U-shaped Network, abbreviated as MSA-UNet). It can effectively solve these difficult problems mentioned above, accurately segment the hepatocellular carcinoma region, and thus effectively assist doctors in the diagnosis and treatment of liver cancer.

[0122] The present invention relates to the field of image segmentation in medical cancer, specifically to propose a fast, accurate and robust method for segmenting liver cancer images, so as to provide more accurate and detailed information about liver anatomical features than previous algorithms such as UNet, SAR-UNet, etc., thereby assisting doctors in more effectively planning treatment plans for patients, such as tumor resection, living transplantation and other intervention measures.

[0123] The following drawings will detail the components of the present invention.

[0124] 1. Dual dilated convolution module

[0125] Through the dual dilated convolution module, MSA-UNet can extract features of the input image better than other UNet networks and their variants, and the use of dilated convolution enables the model to have a larger receptive field and obtain better performance without changing the number of parameters.

[0126] As Figure 2 shown, the module formula includes:

[0127] Dilated convolution and activation part:

[0128] The first layer:

[0129] F 1 =LeakReLU(BatchNorm(DilationConv(X))) (1)

[0130] The second layer:

[0131] F 2 =LeakReLU(BatchNorm(DilationConv(F 1 ))) (2)

[0132] Then, a Dropout (dropout) layer is added after equation (2) to reduce overfitting of the model:

[0133] F 3 =Dropout(F 2) (3)

[0134] Then, perform a residual connection between the input X and the output F3 of the convolutional layer to complement the original feature information that may be lost during the feature extraction process:

[0135] R = X + F 3 (4)

[0136] Finally, input the output R of the residual connection into the (CAA, Context Anchor Attention), that is, the context anchor attention module, to enhance the feature information of R.

[0137] Y = CAA(R) (5)

[0138] Wherein, X represents the input image features, and Y represents the output high-level image features after passing through the double dilated convolution module.

[0139] The present invention is a liver cancer segmentation algorithm based on 2D CNN. By improving the traditional Unet network, it greatly improves the performance of the model in the liver cancer segmentation task. In the encoder part, this algorithm proposes a double dilated convolution module (Double_Dilation_Conv), which replaces the convolutional layer used in the traditional double convolutional layer with a dilated convolution. The dilated convolution can increase the receptive field of the model without increasing the number of model parameters, enabling the model to more easily capture global context information and making it perform more excellently than the traditional convolution in the liver cancer segmentation task.

[0140] 2. CAA Attention Module

[0141] The context anchor attention module (CAA) extracts the context information of the feature map by introducing a channel attention mechanism, using average pooling and convolutional operations, and then generates an attention map according to the context. This map adjusts the input feature map according to the characteristics of each channel. In this way, the network can automatically select more important features and enhance their expression. For unimportant features, due to the small attention value, they play an "inhibitory" role in the feature map. This module helps the proposed model to focus more on effective information and enhances its expression ability.

[0142] As Figure 3 shown, the working process of this module includes:

[0143] First, perform global average pooling on the input feature X to reduce the spatial dimension:

[0144] F 1 = Avgpool(X) (6)

[0145] Then, the features are processed through 1*1 convolution, batch normalization, and activation function:

[0146] F 2 = LeakRelu(Batchnorm(Conv 1×1 (F 1 ))) (7)

[0147] Immediately afterwards, depth convolution operation is performed for feature extraction:

[0148] F 3 = Conv 1×1 (Conv 1×7 (Conv 7×1 (F 2 ))) (8)

[0149] Then, the extracted feature maps are normalized and activated:

[0150] F 4 = LeakRelu(Batchnorm(F 3 )) (9)

[0151] Then, the final attention weights are generated through the Sigmoid (S-shaped) function:

[0152] F out = Sigmoid(F 4 ) (10)

[0153] Finally, the above output is multiplied element-wise with the original input feature map to obtain the feature map enhanced by features.

[0154] Y = X·F out (11)

[0155] Among them, X represents the input image features, and Y represents the attention image features extracted after passing through the Context Anchor Attention (CAA) module.

[0156] This algorithm introduces the CAA (Context Anchor Attention) channel and spatial attention mechanism in the double convolutional layer, improving the interpretability of the model for tumor regions of different sizes and enhancing the segmentation and extraction ability of the model.

[0157] 3. LAGA Module

[0158] Compared with the crude method of the original UNet that simply adds the extracted high-level semantic information and low-level feature information directly by simple skipping, the large kernel attention gate module (LGAG) combines the output of the encoder and the feature maps output by the decoder through two convolutional layers respectively, and then adds them together and passes through a 1*1 convolutional layer and an activation function layer. This enables the model to effectively combine local and global information, organically combines high-level features with low-level features, and enables the model to use high-level features to effectively guide the extraction of low-level features. By using the attention mechanism, the expressive ability of the input feature maps is effectively enhanced, and the attention of the model to important features is improved.

[0159] As Figure 4 shown, the working process of this module includes:

[0160] The feature maps output by the encoder and the upsampled feature maps are respectively subjected to convolutional processing and batch normalization through two branches:

[0161] F 1 = Batchnorm(Conv 3×3 (X 1 )) (12)

[0162] F 2 = Batchnorm(Conv 3×3 (X 2 )) (13)

[0163] Then, the feature maps extracted from the two branches are added element-wise for feature fusion:

[0164] F 3 = F 1 + F 2 (14)

[0165] Then, the fused features are processed through 1*1 convolution, normalization, and Sigmoid (S-shaped function) to generate attention weights:

[0166] F out = Sigmoid(Batchnorm(Conv 1×1 (F 3 ))) (15)

[0167] Finally, the attention weights are applied to the original input X through element-wise multiplication to obtain the final enhanced feature map:

[0168] Y = X·F out (16)

[0169] Among them, X1 is the feature map output by the encoder, X2 is the feature map obtained after passing through the upsampling module, and Y is the high-level feature map output after passing through the large-kernel attention gate module (LGAG).

[0170] Compared with the original UNet's rough method of simply adding the extracted high-level semantic information and low-level feature information directly by the model through simple skipping, this algorithm uses the LGAG module (Large-kernel Attention Gate), which combines local and global information and uses the attention mechanism to enhance the input feature map, improving the model's attention to important features.

[0171] 4. ADown Module

[0172] Traditional UNet networks and their variants often simply use the max-pooling layer for downsampling during the encoder's downsampling process. However, this often leads to the loss of detailed information in the input image, especially local features such as edges and textures. This loss of information has an adverse impact on tasks like liver cancer segmentation that require detailed information, and max-pooling is an irreversible operation. After downsampling, the lost spatial information (such as pixel positions) cannot be recovered through backpropagation. This means that during the upsampling process, it may be difficult to recover all the spatial details in the original image. Therefore, this paper uses the adaptive downsampling module (ADown), which effectively avoids the occurrence of the above problems. By separately performing average pooling and max-pooling on the input feature map and then concatenating them, it effectively reduces the required number of parameters while avoiding the loss of detailed features.

[0173] As Figure 5 shown, the working process of the ADown (adaptive downsampling) module is as follows:

[0174] First, perform global average pooling and global max-pooling operations on the input feature map X respectively for feature extraction and to reduce the resolution of the feature map:

[0175] F 1 = Avgpool(X) (17)

[0176] F 2 = Maxpool(X) (18)

[0177] Then, perform 1*1 convolution, normalization, and activation function processing on the extracted feature maps respectively to further extract features:

[0178] F 1out = LeakRelu(Batchnorm(Conv 1×1 (F 1 ))) (19)

[0179] F 2out = LeakRelu(Batchnorm(Conv 1×1 (F 2 ))) (20)

[0180] Finally, the features extracted from the two branches are concatenated for feature fusion:

[0181] Y = Concatenate(F 1out , F 2out ) (21)

[0182] Among them, X is the high-level feature map output after passing through the double dilated convolution module, and Y is the feature map output after passing through the ADown (adaptive downsampling) module.

[0183] During the upsampling process, the UNet network and its variants generally perform upsampling through simple linear interpolation. Although this can reduce the computational cost required for model training, this method will also lead to the problem of losing spatial position features because the image context information cannot be considered. Therefore, MSA-UNet introduces a dynamic upsampling module (DY), which combines convolution operations, displacement offsets, and pixel rearrangement operations through a dynamic sampling upsampling method, that is, dynamically adjusts the sampling position to improve the upsampling accuracy. And it takes into account the balance between the amount of computation and the accuracy.

[0184] 5. DY Upsampling Module

[0185] As Figure 6 shown, the DY upsampling module can be summarized as follows:

[0186] First, generate dynamic offsets for the input feature map X through a 1*1 convolutional layer:

[0187] F 1 = Conv 1×1 (X) (22)

[0188] Then, add it to the initialized input network coordinates C to obtain the coordinates after dynamic offset:

[0189] F 2 = F 1 + C (23)

[0190] Then, perform coordinate normalization on the obtained coordinate network:

[0191] F 3 = Normalize(F 2 ) (24)

[0192] Finally, dynamic interpolation upsampling is performed based on the normalized coordinates on the basis of the input feature map X:

[0193] Y = Grid_upsample(F 3 ) (25)

[0194] where X is the feature map output after the PMFS module or the feature map output by the decoder, Normlize represents coordinate normalization, and Grid_upsample represents dynamic interpolation upsampling.

[0195] By introducing the ADown (Adaptive Downsampling) module and the DY (Dynamic) upsampling module, this algorithm greatly improves the problem that a large amount of original image feature information is lost during downsampling of the feature map and the upsampling does not consider the image context information, resulting in the loss of spatial position features, compared with the original Unet model.

[0196] 6. PMFS Module

[0197] Traditional multi-scale feature fusion modules often have the characteristic of large computational complexity. Especially after introducing the attention mechanism, the computational cost increases significantly. The Polarized Multi-scale Feature Self-Attention (PMFS) module takes a different approach and avoids the problem of computational redundancy in the model through special structural settings. The Polarized Multi-scale Feature Self-Attention module (PMFS) effectively realizes the function of multi-scale feature fusion enhancement by performing resolution integration on the input image and calculating channel attention and spatial attention successively.

[0198] As Figure 7 shown, the working process of the Polarized Multi-scale Feature Self-Attention (PMFS) module is as follows:

[0199] First, the input multiple feature maps are respectively normalized through parallel paths to be converted into the same resolution and number of channels:

[0200] F i = Conv(Maxpool(X i )) i = 1, 2,..., n (26)

[0201] Then, the converted feature maps are concatenated:

[0202] F all = Concatenate(F i ) i = 1, 2,..., n (27)

[0203] Then, channel self-attention is calculated. Among them, the key matrix K, the query matrix Q, and the value matrix V are calculated by the following formulas respectively:

[0204] K = Conv(F all ) Q = Conv(F all ) V = Conv(F all ) (28)

[0205] Then, use the sigmoid function (Softmax) to generate attention weights:

[0206] Attention 1 = Softmax(K·Q T ) (29)

[0207] Then, perform convolution processing and normalization:

[0208] F refined = Sigmoid(Layernorm(Conv(Attention 1 ))) (30)

[0209] Then, update the features to enhance the channel information of the input multi-scale feature map:

[0210] F attention1 = F refined ·V (31)

[0211] · represents element-wise multiplication

[0212] Then, calculate the spatial attention. Input the feature maps with enhanced features obtained above into the convolutional layer respectively to obtain the key matrix K, query matrix Q, and value matrix V:

[0213] K = Conv(F attention1 ) Q = Conv(F attention1 ) V = Conv(F attention1 ) (32)

[0214] Then, calculate the global average pooling and generate the global attention weights:

[0215] Attention 2 = Softmax(Avgpool(K·Q T )) (33)

[0216] Then, perform the global feature update of the multi-scale feature map:

[0217] F attention2 = Sigmoid(Attention 2 )·V (34)

[0218] Finally, weight and fuse the features output by multiple branches through the convolutional layer and output:

[0219] Y = Conv(F attention1 + F attention2 )(35)

[0220] Wherein, X1, X2, and X3 are the feature maps output by each encoder of the MSA-UNet network, and K, Q, and V are the key matrix, query matrix, and value matrix respectively. Y is the high-level feature extracted through multi-scale feature fusion and self-attention after passing through the Polarized Multi-Scale Feature Self-Attention (PMFS) module.

[0221] In addition, the introduction of the Polarized Multi-Scale Feature Self-Attention module (PMFS) makes the model more robust. By encoding the long-term multi-scale dependencies of the feature maps more carefully and precisely, the feature extraction ability of the model is stronger than previous models. And the special module design makes PMFS more lightweight than the traditional self-attention mechanism, reducing the computational cost required by the model.

[0222] Finally, the overall implementation pseudo-code of the MSA-UNet liver cancer segmentation algorithm is as follows.

[0223]

[0224]

[0225]

[0226] The entire segmentation process is as Figure 1 shown: The picture is input into MSA-UNet, and after the feature extraction and decoding output of the model, the final output effect is obtained.

[0227] To verify the segmentation performance of the network, the following evaluation metrics are used for measurement.

[0228] Dice Similarity Coefficient, which is used to measure the similarity between the segmentation result and the ground truth label. The larger the value, the better the effect.

[0229] IOU (Intersection over Union), which is an index used to measure the overlapping part between the segmentation result and the ground truth label. The larger the value, the better the segmentation effect.

[0230] HD (Hausdorff Distance), which is used to evaluate the accuracy of the segmentation edge. The smaller the value, the better the effect.

[0231] On the LITS public dataset, MSA-UNet achieved the highest Dice coefficient, IOU score, and the lowest HD score.

[0232] Table 1 Performance comparison of the MSA-UNet liver cancer segmentation system on the LITS dataset

[0233] Dice(%) IOU(%) HD UNet 81.94 74.73 2.8084 Attention_UNet 83.76 77.33 2.6389 SAR_UNet 83.22 76.14 2.7077 SelfReg_UNet 82.75 76.08 2.72 Trans_Unet 78.28 70.36 2.9882 MSA-UNet(ours) 87.24 81.18 2.4899

[0234] As can be seen from the above table, in the liver cancer segmentation problem of the LITS dataset, MSA-UNet achieved the best performance and ranked first in segmentation accuracy among all segmentation networks. Compared with the second-ranked segmentation network, the Dice score of MSA-UNet increased by 3.48%, the IOU score increased by 3.85%, and the HD score was 0.149 lower. This fully demonstrates the superiority of MSA-UNet in the liver cancer segmentation task and its excellent segmentation accuracy.

[0235] By using the MSA-UNet segmentation network, the segmentation of the liver cancer area can be accurately achieved, greatly reducing the workload of radiologists and providing effective assistance for doctors to diagnose and treat liver cancer.

[0236] So far, the embodiments of the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description. Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments can be modified or some technical features can be equivalently replaced without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A liver cancer segmentation system based on deep learning, characterized in that: include: Double dilated convolution module to extract features of input images; The attention module introduces the channel attention mechanism, uses average pooling and convolution operations to extract the context information of the feature map, and then generates an attention map based on the context to automatically select more important features and enhance their expression; The large core attention gate module convolves the feature maps of the encoder output and the decoder output respectively, and adds them together to combine high-level features with low-level features to obtain enhanced feature maps; The downsampling module performs average pooling and maximum pooling on the input feature maps and concatenates them to avoid losing detail features. The upsampling module generates a dynamic offset from the input feature map, obtains the coordinates after dynamic offset, and performs sampling on dynamic interpolation; as well as The polarized multi-scale feature self-attention module performs resolution integration, channel attention and spatial attention calculations on the input image in sequence, thereby enhancing the fusion of multi-scale features.

2. The liver cancer segmentation system based on deep learning according to claim 1, characterized in that: The dual dilated convolution module includes dilated convolution and activation parts, and the working process includes: First layer: F1=LeakReLU(BatchNorm(DilationConv(X))) (1) Second layer: F2=LeakReLU(BatchNorm(DilationConv(F1))) (2) Then, a dropout layer is added after equation (2) to reduce the overfitting of the model: F3=Dropout(F2) (3) Then, the input X is residually connected to the convolutional layer output F3 to complete the original feature information that may be lost during the feature extraction process: R=X+F3 (4) Finally, the output R of the residual connection is input into the attention module to enhance the feature information of R; Y=CAA(R) (5) Among them, X represents the input image features, and Y represents the output high-level image features after passing through the double hole convolution module.

3. The liver cancer segmentation system based on deep learning according to claim 1, characterized in that: The working process of the attention module includes: First, the input feature X is globally averaged pooled to reduce the spatial dimension: F1=Avgpool(X) (6) Then, the features are processed by 1*1 convolution, batch normalization and activation function: <h2 style=";text-align:left;direction:ltr">F2=LeakRelu(Batchnorm(Conv<h2 style=";text-align:left;direction:ltr"> 1×1 <h2 style=";text-align:left;direction:ltr"> (F1))) (7) Then, feature extraction is performed by performing a depthwise convolution operation: F3=Conv 1×1 (Conv 1×7 (Conv 7×1 (F2))) (8) Then, the extracted feature map is normalized and activated: F4=LeakRelu(Batchnorm(F3)) (9) Then, the final attention weights are generated through the S-shaped function: F out =Sigmoid(F4) (10) Finally, the above output is multiplied element-wise with the original input feature map to obtain the feature-enhanced feature map; Y=X·F out (11) Among them, X represents the input image features, and Y represents the attention image features extracted after the attention module.

4. The liver cancer segmentation system based on deep learning according to claim 1, characterized in that: The working process of the large core attention gate module includes: The feature map output by the encoder and the upsampled feature map are convolved and batch normalized through two branches respectively: F1=Batchnorm(Conv 3×3 (X1)) (12) <h2 style=";text-align:left;direction:ltr">F2 = Batchnorm(Conv<h2 style=";text-align:left;direction:ltr"> 3×3 <h2 style=";text-align:left;direction:ltr"> (X2)) (13) Then, the feature maps extracted by the two branches are added element by element for feature fusion: F3=F1+F2 (14) Then, the fused features are processed by 1*1 convolution, normalization and Sigmoid function to generate attention weights: F out =Sigmoid(Batchnorm(Conv 1×1 (F3))) (15) Finally, the attention weights are applied to the original input X by element-wise multiplication to obtain the final enhanced feature map: Y=X·F out (16) Among them, X1 is the feature map output by the encoder, X2 is the feature map obtained after the upsampling module, and Y is the high-level feature map output after the large-core attention gate module.

5. The liver cancer segmentation system based on deep learning according to claim 1, characterized in that: The working process of the downsampling module includes: First, global average pooling and global maximum pooling operations are performed on the input feature map X to perform feature extraction operations and reduce the resolution of the feature map: F1=Avgpool(X) (17) F2=Maxpool(X) (18) Then, the extracted feature maps are subjected to 1*1 convolution, normalization and activation function processing to further extract features: F 1out =LeakRelu(Batchnorm(Conv 1×1 (F1))) (19) <h2 style=";text-align:left;direction:ltr">F<h2 style=";text-align:left;direction:ltr"> 2out <h2 style=";text-align:left;direction:ltr"> =LeakRelu(Batchnorm(Conv<h2 style=";text-align:left;direction:ltr"> 1×1 <h2 style=";text-align:left;direction:ltr"> (F2)) (20) Finally, the features extracted from the two branches are concatenated for feature fusion: Y=Concatenate(F 1out ,F 2out ) (21) Among them, X is the high-level feature map output after the double hole convolution module, and Y is the feature map output after the downsampling module.

6. The liver cancer segmentation system based on deep learning according to claim 1, characterized in that: The working process of the upsampling module includes: First, the input feature map X is passed through a 1*1 convolutional layer to generate a dynamic offset: F1=Conv 1×1 (X) (22) Then, add it to the initial input network coordinate C to get the coordinate after dynamic offset: F2=F1+C (23) Then, the obtained coordinate network is normalized: F3=Normalize(F2) (24) Finally, dynamic interpolation upsampling is performed based on the input feature map X according to the normalized coordinates: Y=Grid_upsample(F3) (25) Among them, X is the feature map output after the polarized multi-scale feature self-attention module or the feature map output by the decoder, Normalize represents coordinate normalization, and Grid_upsample represents dynamic interpolation upsampling.

7. The liver cancer segmentation system based on deep learning according to claim 1, characterized in that: The working process of the polarized multi-scale feature self-attention module includes: First, multiple input feature maps are normalized through parallel paths and converted to the same resolution and number of channels: F i =Conv(Maxpool(X i )) i=1,2,...,n (26) Then, the transformed feature maps are spliced: F all =Concatenate(F i ) i=1,2,...,n (27) Then, the channel self-attention is calculated, where the key matrix K, query matrix Q, and value matrix V are calculated by the following formulas: K=Conv(F all ) Q=Conv(F all ) V=Conv(F all ) (28) Then, the attention weights are generated using the S-shaped function (Softmax): Attention1=Softmax(K·Q T ) (29) Then, convolution and normalization are performed: F refined =Sigmoid(Layernorm(Conv(Attention1))) (30) Then, the features are updated to enhance the channel information of the input multi-scale feature map: F attention1 =F refined ·V (31) · Represents element-by-element multiplication.

8. The deep learning-based liver cancer segmentation system according to claim 7, characterized in that: The working process of the polarized multi-scale feature self-attention module also includes: Perform spatial attention calculations and input the feature maps obtained above into the convolutional layers to obtain the key matrix K, query matrix Q, and value matrix V: K=Conv(F attention1 ) Q=Conv(F attention1 ) V=Conv(F attention1 ) (32) Then, the global average pooling and global attention weight generation are calculated: Attention2=Softmax(Avgpool(K·Q T )) (33) Then, the global feature update of the multi-scale feature map is performed: F attention2 =Sigmoid(Attention2)·V (34) Finally, the features output by multiple branches are weighted and fused through the convolution layer and output: Y=Conv(F attention1 +F attention2 ) (35) Among them, X1, X2, and X3 are the feature maps output by each encoder of the MSA-UNet network, K, Q, and V are the key matrix, query matrix, and value matrix respectively; Y is the high-level feature extracted by multi-scale feature fusion and self-attention after the polarized multi-scale feature self-attention module.

9. A liver cancer segmentation method based on deep learning, characterized in that: Using the liver cancer segmentation system based on deep learning according to any one of claims 1 to 8, the steps include: S100. Input a medical image X of size H×W×3; S200. Define liver cancer segmentation algorithm model; S300. As an encoder of the downsampling path, for the I1 to IN encoder, perform the following operations: apply a double hole convolution module to the input image or feature map for feature extraction, store the generated feature map Fi for use in the jump connection of the subsequent encoder part, use a downsampling module to reduce the spatial size of the input feature map, and end the loop; S400. The bottleneck layer applies a double dilated convolution module to the input feature map, and applies a PMFS module to the output feature map for multi-scale feature aggregation enhancement; S500. As a decoder of the upsampling path, for the IN to I1 decoder, perform the following operations: use an upsampling module to increase the spatial size of the feature map, use a large core attention gate module to optimize and enhance the upsampled feature map, concatenate the stored feature map with the current decoder output feature map, use a double hole convolution module to extract features, and end the loop; S600. Output layer: Apply 1×1 convolution to reduce the number of input feature map channels to the number of output categories, and apply S-shaped activation function to generate probability map; S700. Inference process: load the trained liver cancer segmentation algorithm model, and for each test image, perform the following operations: obtain the segmentation mask Ypred through forward propagation of the model, and end the loop; S800. Return the predicted H×W×1 segmentation mask Ypred.

10. The liver cancer segmentation method based on deep learning according to claim 9, characterized in that: Before the inference process, the liver cancer segmentation algorithm model is trained. For each training round, the following operations are performed: Load the training image and its corresponding true segmentation mask, perform forward propagation through the liver cancer segmentation algorithm and calculate the predicted mask, calculate the total loss using the similarity coefficient loss and the binary cross entropy loss, backpropagate the loss and use the weighted decay adaptive moment estimation optimizer to update the network weights, and end the training cycle.

Citation Information

Cited By

  • Pavement crack segmentation algorithm based on YOLO11

    CN121236375A