Liver tumor segmentation method fusing three kinds of attention
By integrating three attention mechanisms in liver tumor segmentation and building a TAF-Net model, the shortcomings of traditional methods in boundary recognition and multi-scale information processing are solved, and higher segmentation accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510089749.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The prior art is difficult to accurately identify boundaries in liver tumor segmentation, and traditional methods have shortcomings in processing multi-scale information and enhancing boundary perception capabilities, resulting in inaccurate identification of small tumors by models or complex morphological tumors.
A liver tumor segmentation method that fuses three attentions is adopted. By constructing a dual attention dynamic convolution and edge fusion attention module, combined with multi-scale edge segmentation loss, a TripleAttentionFusionNetwork (TAF-Net) model is constructed to improve the accuracy and robustness of liver tumor segmentation.
It significantly improves the accuracy and robustness of liver tumor segmentation, can more accurately identify liver tumor areas, and is efficient and adaptable, providing strong technical support for the precise segmentation of liver tumors.
Smart Images

Figure CN120031900A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and in particular to a liver tumor segmentation method integrating three types of attention. Background Art
[0002] In recent years, medical image segmentation has played a vital role in the diagnosis, treatment plan specification, and postoperative evaluation of liver tumors, especially in terms of automation and accurate segmentation. However, due to the significant diversity of liver tumors in terms of morphology, size, and boundary fuzziness, the segmentation task faces great challenges. Especially in abdominal computed tomography (CT) images, the contrast between tumors and normal liver tissues is low, and the images are easily disturbed by factors such as noise, artifacts, and physiological differences between different patients, making it difficult for the accuracy and robustness of traditional segmentation methods to meet actual needs. With the rapid development of computer technology, deep learning technology has provided a new research direction for achieving automated and high-precision liver tumor segmentation.
[0003] Existing segmentation methods based on Convolutional Neural Networks (CNN), such as U-Net, have shown remarkable results in liver tumor segmentation tasks. However, these methods still have deficiencies in capturing subtle features and multi-scale information and enhancing the ability to perceive boundaries. In particular, for liver tumors with blurred boundaries, accurate identification of their contours is more difficult, which may cause the model to miss small tumors or inaccurately identify complex morphological tumors. In addition, models such as U-Net also have certain limitations in their modeling capabilities for processing global information, and the lack of receptive field may limit their segmentation performance for large-scale lesions.
[0004] Dynamic convolution, as a mechanism that can dynamically adjust the weights of the convolution kernel according to the input, provides a new idea for enhancing the feature extraction ability of the model. However, the existing dynamic convolution method still has certain limitations in some complex features and it is difficult to fully express the detailed characteristics of the tumor. On the other hand, the attention mechanism has been widely used in medical image segmentation in recent years. However, the existing attention mechanism usually focuses on the salient feature areas, while the boundary information is less utilized. For the liver tumor segmentation task, the modeling of boundary information is crucial to accurately identify the tumor area.
[0005] In order to solve these problems, it is necessary to explore new technologies and methods to improve the accuracy and robustness of liver tumor segmentation. Therefore, how to use the characteristics of dynamic convolution and attention mechanism and the advantages of U-Net architecture to effectively improve the accuracy and reliability of liver tumor segmentation has become a key issue to be solved in the field of medical image segmentation. Summary of the invention
[0006] In order to solve the above technical problems, the present invention provides a liver tumor segmentation method integrating three types of attention, which can provide users with a more accurate, efficient and highly adaptable liver tumor segmentation solution.
[0007] Specifically, the method comprises the following steps:
[0008] S1: Preprocess abdominal CT data and divide them into training set, validation set and test set in proportion;
[0009] S2: Combine dual attention and convolution operations to construct dual attention dynamic convolution;
[0010] S3: Fuse edge information and attention mechanism to build edge fusion attention module;
[0011] S4: Fusion edge supervision loss and segmentation loss to construct multi-scale edge segmentation loss;
[0012] S5: Fusing dual attention dynamic convolution, edge fusion attention module and multi-scale edge segmentation loss to build TripleAttentionFusionNetwork (TAF-Net) model;
[0013] S6: Use the experimental dataset to complete model training, validation optimization and performance evaluation.
[0014] Preferably, S1 comprises the following steps:
[0015] S1.1: Preprocessing of abdominal CT data, including data format conversion, cropping, resampling, denoising, and normalization operations;
[0016] S1.2: The preprocessed abdominal CT data are divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0017] Preferably, S2 comprises the following steps:
[0018] S2.1: Enhance features using spatial attention and self-attention mechanisms;
[0019] S2.2: Fusion convolution operations to build a dual-attention dynamic convolution (DADC) module.
[0020] Preferably, S2.1 comprises the following steps:
[0021] First, the original image is convolved to obtain the input feature map F, and F is reduced in dimension using a 1×1 convolution. Then the output tensor is normalized using the sigmoid activation function to obtain the spatial attention weight A.s (F), finally, the obtained spatial attention weight is multiplied pixel by pixel by the input feature map to obtain the output feature map F s The calculation process is as follows:
[0022] A s (F) = σ(Conv 1×1 (F)),
[0023]
[0024] In the above formula, F is the input feature map, A s (F) is the spatial attention weight, represents the element-by-element multiplication operation, σ represents the sigmoid activation function, and F s is the output feature map.
[0025] Output feature map F of the spatial attention module s , the self-attention mechanism is used to capture the global dependency and interaction between channels, and the channel attention weight A is obtained c , and finally obtain the channel weight W after being enhanced by the spatial attention module and the self-attention mechanism c The calculation process is as follows:
[0026] F c =AvgPool(F s ),
[0027] Q=W Q ⊙F c ,
[0028] K=W K ⊙F c ,
[0029] V=W V ⊙F c ,
[0030]
[0031] W c =A c ⊙V,
[0032] In the above formula, F c is the channel feature after global average pooling (AvgPool), W Q ,W K ,W V is a learnable weight matrix, Q, K, V represent the query, key and value of the channel dimension respectively, A c represents the weight on the channel dimension, ⊙ represents matrix multiplication, C represents the channel dimension of the input feature map, W cIt represents the channel weight after being enhanced by the dual attention mechanism.
[0033] Preferably, S2.2 comprises the following steps:
[0034] The channel weights are converted into dynamic convolution weights μ and convolution operations are performed. The constructed dual attention dynamic convolution (DADC) module can be described as:
[0035] Convert the channel weights after dual attention enhancement to dynamic convolution weights:
[0036] μ=Softmax(W p ⊙W c ),
[0037] In the above formula, W p It is a linear transformation matrix that realizes the transformation from channel dimension to convolution kernel dimension, and ⊙ represents matrix multiplication.
[0038] Then, based on the preset 4 convolution kernels, the final dynamic convolution kernel k is generated by dynamic weighting:
[0039]
[0040] In the above formula, μ i (i=1,2,3,4) are the dynamic weights of the four convolution kernels, conv i (i=1,2,3,4) represents the convolution kernel.
[0041] Finally, the output F of the spatial attention module is s As the convolution object, the dynamic convolution kernel k is used to perform convolution operation on it to obtain the output F of the DADC module out :
[0042] F out =Conv(k(F s )),
[0043] Preferably, S3 comprises the following steps:
[0044] S3.1: Extract edge information and generate boundary features, background features and high-frequency features;
[0045] S3.2: Fusion attention mechanism, constructing the Edge-Fusion Attention (EFA) module.
[0046] Preferably, S3.1 comprises the following steps:
[0047] In order to enhance the network's perception of edge information, the feature expression capability is further optimized by extracting edge information and generating multiple features. in, first generate the predicted segmentation result P. Then process P and get the edge weight A edge , the edge weight and the input feature are multiplied to obtain the boundary feature F edge ; Then the background weight A is obtained by calculating the reverse attention bg , the background weight and the input feature are multiplied to obtain the background feature F bg ; Finally, the high-frequency edge weight A is extracted through the Laplacian pyramid hf , the high-frequency edge weight and the input feature are multiplied to obtain the high-frequency feature F hf The overall generation process is as follows:
[0048] By input feature F in Generate predicted segmentation result P:
[0049] P = σ(Conv 1×1 (F in )),
[0050] Extract boundary features F from the prediction result P edge :
[0051] P filtered =ConvGauss(P),
[0052] P up =Upsample(Downsample(P filtered )),
[0053]
[0054]
[0055] The background feature F is obtained by calculating the background focus area of the prediction result P. bg :
[0056]
[0057]
[0058] Extract high-frequency features F through Laplacian pyramid hf :
[0059] A hf =Laplacian pyramid(F in ),
[0060]
[0061] In the above formula, σ(·) represents the Sigmoid function, Laplacian pyramid(·) represents the Laplacian pyramid method, represents XOR operation, P represents prediction result, ConvGauss(·) represents Gaussian convolution, Downsample represents downsampling, Upsample represents upsampling, Represents an element-wise multiplication operation.
[0062] Preferably, S3.2 comprises the following steps:
[0063] For boundary features F edge 、Background features F bg And the high frequency feature F hf These three features are fused, and the attention mechanism is introduced to highlight the significant information of the target area to construct the Edge-Fusion Attention (EFA) module. The specific implementation is as follows:
[0064] The three features are fused to obtain the fused feature F fusion :
[0065] F fusion =Conv(Concat(F edge ,F bg ,F hf )),
[0066] Generate the attention weight A of the fused feature fusion , and get the enhanced feature F enhanced :
[0067] A fusion =σ(Conv(F fusion )),
[0068]
[0069] Add the enhanced fusion feature to the residual input to get the intermediate feature F temp :
[0070]
[0071] Finally, the feature representation capability is further enhanced by channel attention and spatial attention to obtain the final output F of the EFA module. EFA :
[0072]
[0073] M s (F temp )=σ(f 7×7 (Concat(AvgPool(Ftemp ),MaxPool(F temp )))),
[0074]
[0075]
[0076] In the above formula, Concat(·) represents concatenation of features in the channel dimension, Conv(·) represents convolution operation, represents an element-wise multiplication operation, represents the element-by-element addition operation, σ represents the sigmoid activation function, and f 7×7 represents a filter with a kernel size of 7, MLP represents a multi-layer perceptron, and M c , M s They represent the channel attention layer and the spatial attention layer respectively.
[0077] Preferably, S4 comprises the following steps:
[0078] S4.1: Calculate marginal supervision loss;
[0079] S4.2: Calculate segmentation loss;
[0080] S4.3: Construct multi-scale edge segmentation loss.
[0081] Preferably, S4.1 comprises the following steps:
[0082] For feature maps of different scales, edge features are extracted by Laplacian operator Guide the model to learn the boundary area more accurately. Edge supervision loss L ES The definition is as follows:
[0083]
[0084]
[0085] In the above formula, represents the predicted edge feature, e represents the true edge feature, represents the segmentation prediction result, Laplace(·) represents the Laplace operator operation, and N represents the total number of pixels.
[0086] Preferably, S4.2 comprises the following steps:
[0087] The segmentation loss combines Dice loss and cross entropy loss. SEG The definition is as follows:
[0088]
[0089]
[0090] L SEG =λ 1 L DICE +λ 2 L CE ,
[0091] Among them, y i represents the actual segmentation result, represents the predicted segmentation result, N represents the total number of pixels, and λ 1 and λ 2 is the weight hyperparameter.
[0092] Preferably, S4.3 comprises the following steps:
[0093] Use multi-scale supervision and integrate edge supervision loss L ES and segmentation loss L SEG , we get the multi-scale edge segmentation loss L MES As shown below:
[0094]
[0095] Among them, S is set to 4, α s represents the loss weight of the s-th scale, represents the marginal supervision loss of the s-th scale, represents the segmentation loss of the s-th scale.
[0096] Preferably, S5 comprises the following steps:
[0097] Taking the U-shaped network as the basic architecture, the DADC module is first used to replace the second, third, and fourth convolutional layers of the encoder and all convolutional layers of the decoder in the U-shaped network. Then, the EFA module is introduced into the bottleneck layer of the U-shaped network. Finally, the multi-scale edge segmentation loss function L is combined. MES Optimize the overall learning process of the model to build the TAF-Net model.
[0098] Preferably, S6 comprises the following steps:
[0099] Use the divided experimental data set to complete the model training and verification optimization, save the optimal model weights, and then use the test set to evaluate the model performance and obtain various model indicators.
[0100] The beneficial effects of the method of the present invention are: first, in the data preprocessing stage, by performing data format conversion, cropping, resampling, denoising and normalization operations on the collected abdominal CT images, the data quality is comprehensively improved, the interference of irrelevant information is effectively reduced, and the data is proportionally divided into training set, validation set and test set, which lays a solid foundation for subsequent model training and testing. In the process of model construction, the dual attention mechanism is innovatively combined with the convolution operation to design the dual attention dynamic convolution (DADC) to realize the dynamic weight allocation of feature information. In addition, by fusing the edge information with the attention mechanism, an edge fusion attention (EFA) module is constructed. Furthermore, for feature information at different scales, the edge supervision loss and segmentation loss are fused to construct a multi-scale edge segmentation loss L MES . Finally, the TAF-Net model was constructed based on the U-shaped network. Through multi-dimensional optimization design, the model demonstrated excellent ability in capturing edge features and multi-scale information, significantly improving the accuracy of liver tumor segmentation. At the same time, it also adaptively adjusted the convolution kernel weights to effectively avoid the drag on model performance caused by feature redundancy and invalid features. The overall design not only accurately identifies liver tumor areas, but also has both high efficiency and robustness, providing strong technical support for the accurate segmentation of liver tumors. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the prior art and the drawings required for use in the embodiments. The following drawings are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0102] Figure 1 It is a flowchart of a liver tumor segmentation method integrating three types of attention in the present invention;
[0103] Figure 2 This is a TAF-Net model architecture diagram of a liver tumor segmentation method integrating three types of attention in the present invention;
[0104] Figure 3 It is a schematic diagram of the DADC structure of a liver tumor segmentation method integrating three types of attention in the present invention;
[0105] Figure 4 It is a schematic diagram of the EFA structure of a liver tumor segmentation method integrating three types of attention in the present invention; Specific implementation plan
[0106] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all other embodiments obtained by ordinary technicians in this field without creative work based on the embodiments of the present invention are within the scope of protection of the present invention.
[0107] The embodiment of the present invention provides a liver tumor segmentation method integrating three types of attention, which is used to achieve accurate segmentation of liver tumors, assist clinical diagnosis and treatment decision-making, and improve the accuracy and efficiency of medical image processing.
[0108] Reference Figure 1 , the method comprises the following steps:
[0109] S1: Preprocess abdominal CT data and divide them into training set, validation set and test set in proportion;
[0110] S2: Combine dual attention and convolution operations to construct dual attention dynamic convolution;
[0111] S3: Fuse edge information and attention mechanism to build edge fusion attention module;
[0112] S4: Fusion edge supervision loss and segmentation loss to construct multi-scale edge segmentation loss;
[0113] S5: Fusing dual attention dynamic convolution, edge fusion attention module and multi-scale edge segmentation loss to build TripleAttentionFusionNetwork (TAF-Net) model;
[0114] S6: Use the experimental dataset to complete model training, validation optimization and performance evaluation.
[0115] Further, S1 comprises the following steps:
[0116] S1.1: Preprocessing of abdominal CT data, including data format conversion, cropping, resampling, denoising, and normalization operations;
[0117] Further, the preprocessing of the abdominal CT data includes: data format conversion, cropping, resampling, denoising and normalization operations, which specifically include:
[0118] The collected abdominal CT data, including public data sets and data sets provided by cooperative hospitals, each case image completely covers the entire liver area. First, the labeled abdominal CT data stored in nii format is converted to npy matrix format, and then the abdominal CT data is cropped to remove irrelevant areas and resampled to a uniform size of 512×512. Then, the abdominal CT data is denoised using Gaussian filtering, and the pixel values of each data are normalized to a uniform range (0-1). The mask corresponding to the liver and background areas is set to 0, and the mask corresponding to the liver tumor area is set to 1.
[0119] S1.2: The preprocessed abdominal CT data are divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0120] Furthermore, the step of dividing the pre-processed abdominal CT data into a training set, a validation set, and a test set according to a ratio of 8:1:1 specifically includes:
[0121] The abdominal CT data after preprocessing were divided into training set, validation set and test set in a ratio of 8:1:1.
[0122] Further, refer to Figure 3 , S2 includes the following steps:
[0123] S2.1: Enhance features using spatial attention and self-attention mechanisms;
[0124] S2.2: Fusion convolution operations to build a dual-attention dynamic convolution (DADC) module.
[0125] Further, S2.1 includes the following steps:
[0126] First, the original image is convolved to obtain the input feature map F, and F is reduced in dimension using a 1×1 convolution. Then the output tensor is normalized using the sigmoid activation function to obtain the spatial attention weight A. s (F), finally, the obtained spatial attention weight is multiplied pixel by pixel by the input feature map to obtain the output feature map F s The calculation process is as follows:
[0127] A s (F) = σ(Conv 1×1 (F)),
[0128]
[0129] In the above formula, F is the input feature map, A s (F) is the spatial attention weight, represents the element-by-element multiplication operation, σ represents the sigmoid activation function, and F s is the output feature map.
[0130] Output feature map F of the spatial attention module s , the self-attention mechanism is used to capture the global dependency and interaction between channels, and the channel attention weight A is obtained c , and finally obtain the channel weight W after being enhanced by the spatial attention module and the self-attention mechanism c The calculation process is as follows:
[0131] F c =AvgPool(F s ),
[0132] Q=W Q ⊙F c ,
[0133] K=W K ⊙F c ,
[0134] V=W V ⊙F c ,
[0135]
[0136] W c =A c ⊙V,
[0137] In the above formula, F c is the channel feature after global average pooling (AvgPool), W Q ,W K ,W V is a learnable weight matrix, Q, K, V represent the query, key and value of the channel dimension respectively, A c represents the weight on the channel dimension, ⊙ represents matrix multiplication, C represents the channel dimension of the input feature map, W c It represents the channel weight after being enhanced by the dual attention mechanism.
[0138] Further, S2.2 includes the following steps:
[0139] The channel weights are converted into dynamic convolution weights μ and convolution operations are performed. The constructed dual attention dynamic convolution (DADC) module can be described as:
[0140] Convert the channel weights after dual attention enhancement to dynamic convolution weights:
[0141] μ=Softmax(W p ⊙Wc ),
[0142] In the above formula, W p It is a linear transformation matrix that realizes the transformation from channel dimension to convolution kernel dimension, and ⊙ represents matrix multiplication.
[0143] Then, based on the preset 4 convolution kernels, the final dynamic convolution kernel k is generated by dynamic weighting:
[0144]
[0145] In the above formula, μ i (i=1,2,3,4) are the dynamic weights of the four convolution kernels, conv i (i=1,2,3,4) represents the convolution kernel.
[0146] Finally, the output F of the spatial attention module is s As the convolution object, the dynamic convolution kernel k is used to perform convolution operation on it to obtain the output F of the DADC module out :
[0147] F out =Conv(k(F s )),
[0148] Further, refer to Figure 4 , S3 includes the following steps:
[0149] S3.1: Extract edge information and generate boundary features, background features and high-frequency features;
[0150] S3.2: Fusion attention mechanism, constructing the Edge-Fusion Attention (EFA) module.
[0151] Further, S3.1 includes the following steps:
[0152] In order to enhance the network's perception of edge information, the feature expression capability is further optimized by extracting edge information and generating multiple features. in , first generate the predicted segmentation result P. Then process P and get the edge weight A edge , the edge weight and the input feature are multiplied to obtain the boundary feature F edge ; Then the background weight A is obtained by calculating the reverse attention bg , the background weight and the input feature are multiplied to obtain the background feature F bg ; Finally, the high-frequency edge weight A is extracted through the Laplacian pyramid hf , the high-frequency edge weight and the input feature are multiplied to obtain the high-frequency feature F hfThe overall generation process is as follows:
[0153] By input feature F in Generate predicted segmentation result P:
[0154] P = σ(Conv 1×1 (F in )),
[0155] Extract boundary features F from the prediction result P edge :
[0156] P filtered =ConvGauss(P),
[0157] P up =Upsample(Downsample(P filtered )),
[0158]
[0159]
[0160] The background feature F is obtained by calculating the background focus area of the prediction result P. bg :
[0161]
[0162]
[0163] Extract high-frequency features F through Laplacian pyramid hf :
[0164] A hf =Laplacian pyramid(F in ),
[0165]
[0166] In the above formula, σ(·) represents the Sigmoid function, Laplacian pyramid(·) represents the Laplacian pyramid method, represents XOR operation, P represents prediction result, ConvGauss(·) represents Gaussian convolution, Downsample represents downsampling, Upsample represents upsampling, Represents an element-wise multiplication operation.
[0167] Further, S3.2 includes the following steps:
[0168] For boundary features F edge 、Background features F bg And the high frequency feature Fhf These three features are fused, and the attention mechanism is introduced to highlight the significant information of the target area to construct the Edge-Fusion Attention (EFA) module. The specific implementation is as follows:
[0169] The three features are fused to obtain the fused feature F fusion :
[0170] F fusion =Conv(Concat(F edge ,F bg ,F hf )),
[0171] Generate the attention weight A of the fused feature fusion , and get the enhanced feature F enhanced :
[0172] A fusion =σ(Conv(F fusion )),
[0173]
[0174] Add the enhanced fusion feature to the residual input to get the intermediate feature F temp :
[0175]
[0176] Finally, the feature representation capability is further enhanced by channel attention and spatial attention to obtain the final output F of the EFA module. EFA :
[0177] M c (F temp )=σ(MLP(AvgPool(F temp ))⊕MLP(MaxPool(F temp ))),
[0178] M s (F temp )=σ(f 7×7 (Concat(AvgPool(F temp ),MaxPool(F temp )))),
[0179]
[0180]
[0181] In the above formula, Concat(·) represents concatenation of features in the channel dimension, Conv(·) represents convolution operation, represents an element-wise multiplication operation, represents the element-by-element addition operation, σ represents the sigmoid activation function, and f 7×7 represents a filter with a kernel size of 7, MLP represents a multi-layer perceptron, and M c , M s They represent the channel attention layer and the spatial attention layer respectively.
[0182] Further, S4 comprises the following steps:
[0183] S4.1: Calculate marginal supervision loss;
[0184] S4.2: Calculate segmentation loss;
[0185] S4.3: Construct multi-scale edge segmentation loss.
[0186] Further, S4.1 includes the following steps:
[0187] For feature maps of different scales, edge features are extracted by Laplacian operator Guide the model to learn the boundary area more accurately. Edge supervision loss L ES The definition is as follows:
[0188]
[0189]
[0190] In the above formula, represents the predicted edge feature, e represents the true edge feature, represents the segmentation prediction result, Laplace(·) represents the Laplace operator operation, and N represents the total number of pixels.
[0191] Further, S4.2 includes the following steps:
[0192] The segmentation loss combines Dice loss and cross entropy loss. SEG The definition is as follows:
[0193]
[0194]
[0195] L SEG =λ 1 L DICE +λ 2 L CE ,
[0196] Among them, y i represents the actual segmentation result, represents the predicted segmentation result, N represents the total number of pixels, and λ 1 and λ 2 is the weight hyperparameter.
[0197] Further, S4.3 includes the following steps:
[0198] Use multi-scale supervision and integrate edge supervision loss L ES and segmentation loss L SEG , we get the multi-scale edge segmentation loss L MES As shown below:
[0199]
[0200] Among them, S is set to 4, α s represents the loss weight of the s-th scale, represents the marginal supervision loss of the s-th scale, represents the segmentation loss of the s-th scale.
[0201] Further, refer to Figure 2 , S5 comprises the following steps:
[0202] Taking the U-shaped network as the basic architecture, the DADC module is first used to replace the second, third, and fourth convolutional layers of the encoder and all convolutional layers of the decoder in the U-shaped network. Then, the EFA module is introduced into the bottleneck layer of the U-shaped network. Finally, the multi-scale edge segmentation loss function L is combined. MES Optimize the overall learning process of the model to build the TAF-Net model.
[0203] Further, S6 comprises the following steps:
[0204] Use the divided experimental data set to complete the model training and verification optimization, save the optimal model weights, and then use the test set to evaluate the model performance and obtain various model indicators.
[0205] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art may make other equivalent modifications or substitutions without violating the spirit of the invention, and these equivalent modifications or substitutions are included in the scope defined by the application claims.
Claims
1. A liver tumor segmentation method integrating three types of attention, characterized in that: The method comprises the following steps: S1: Preprocess abdominal CT data and divide them into training set, validation set and test set in proportion; S2: Combine dual attention and convolution operations to construct dual attention dynamic convolution; S3: Fuse edge information with attention mechanism to build edge fusion attention module; S4: Fusion edge supervision loss and segmentation loss to construct multi-scale edge segmentation loss; S5: Fusing dual attention dynamic convolution, edge fusion attention module and multi-scale edge segmentation loss to build TripleAttentionFusionNetwork (TAF-Net) model; S6: Use the experimental dataset to complete model training, validation optimization and performance evaluation.
2. The method for liver tumor segmentation integrating three types of attention according to claim 1, characterized in that: S1 includes the following steps: S1.1: Preprocessing of abdominal CT data, including data format conversion, cropping, resampling, denoising, and normalization operations; S1.2: The preprocessed abdominal CT data are divided into a training set, a validation set, and a test set according to the ratio of 8:1:
1.
3. The method for liver tumor segmentation integrating three types of attention according to claim 1, characterized in that: S2 includes the following steps: S2.1: Enhance features using spatial attention and self-attention mechanisms; S2.2: Fuse convolution operations and build a Dual-Attention Dynamic Convolution (DADC) module; Among them, S2.1 includes the following steps: First, the original image is convolved to obtain the input feature map F, and F is reduced in dimension using a 1×1 convolution. Then the output tensor is normalized using the sigmoid activation function to obtain the spatial attention weight A. s (F), finally, the obtained spatial attention weight is multiplied pixel by pixel by the input feature map to obtain the output feature map F s The calculation process is as follows: A s (F)=σ(Conv 1×1 (F)), In the above formula, F is the input feature map, A s (F) is the spatial attention weight, represents the element-by-element multiplication operation, σ represents the sigmoid activation function, and F s is the output feature map; Output feature map F of the spatial attention module s , the self-attention mechanism is used to capture the global dependency and interaction between channels, and the channel attention weight A is obtained c , and finally obtain the channel weight W after being enhanced by the spatial attention module and the self-attention mechanism c The calculation process is as follows: F c =AvgPool(F s ), Q=W Q ⊙F c , K=W K ⊙F c , V=W V ⊙F c , W c =A c ⊙V, In the above formula, F c is the channel feature after global average pooling (AvgPool), W Q ,W K ,W V is a learnable weight matrix, Q, K, V represent the query, key and value of the channel dimension respectively, A c represents the weight on the channel dimension, ⊙ represents matrix multiplication, C represents the channel dimension of the input feature map, W c It represents the channel weight after being enhanced by the dual attention mechanism; Among them, S2.2 includes the following steps: The channel weights are converted into dynamic convolution weights μ and convolution operations are performed. The constructed dual attention dynamic convolution (DADC) module can be described as: Convert the channel weights after dual attention enhancement to dynamic convolution weights: μ=Softmax(W p ⊙W c ), In the above formula, W p It is a linear transformation matrix that realizes the transformation from channel dimension to convolution kernel dimension, and ⊙ represents matrix multiplication. Then, based on the preset 4 convolution kernels, the final dynamic convolution kernel k is generated by dynamic weighting: In the above formula, μ i (i=1,2,3,4) are the dynamic weights of the four convolution kernels, conv i (i=1,2,3,4) represents the convolution kernel; Finally, the output F of the spatial attention module is s As the convolution object, the dynamic convolution kernel k is used to perform convolution operation on it to obtain the output F of the DADC module out : F out =Conv(k(F s ))。 4. The method for liver tumor segmentation integrating three types of attention according to claim 1, characterized in that: S3 includes the following steps: S3.1: Extract edge information and generate boundary features, background features and high-frequency features; S3.2: Fusion attention mechanism, constructing the Edge-Fusion Attention (EFA) module; Among them, S3.1 includes the following steps: In order to enhance the network's perception of edge information, the feature expression capability is further optimized by extracting edge information and generating multiple features. in , first generate the predicted segmentation result P. Then process P and get the edge weight A edge , the edge weight and the input feature are multiplied to obtain the boundary feature A edge ; Then the background weight A is obtained by calculating the reverse attention bg , the background weight and the input feature are multiplied to obtain the background feature F bg ; Finally, the high-frequency edge weight A is extracted through the Laplacian pyramid hf , the high-frequency edge weight and the input feature are multiplied to obtain the high-frequency feature F hf The overall generation process is as follows: By input feature F in Generate predicted segmentation result P: P=σ(Conv 1×1 (F in )), Extract boundary features F from the prediction result P edge : P filtered =ConvGauss(P), P up =Upsample(Downsample(P filtered )), The background feature F is obtained by calculating the background focus area of the prediction result P. bg : Extract high-frequency features F through Laplacian pyramid hf : A hf =Laplacian pyramid(F in ), In the above formula, σ(·) represents the Sigmoid function, Laplacian pyramid(·) represents the Laplacian pyramid method, represents XOR operation, P represents prediction result, ConvGaus(·) represents Gaussian convolution, Downsample represents downsampling, Upsample represents upsampling, Represents an element-by-element multiplication operation; Among them, S3.2 includes the following steps: For boundary features F edge 、Background features F bg And the high frequency feature F hf These three features are fused, and the attention mechanism is introduced to highlight the significant information of the target area to construct the Edge-Fusion Attention (EFA) module. The specific implementation is as follows: The three features are fused to obtain the fused feature F fusion : F fusion =Conv(Concat(F edge ,F bg ,F hf )), Generate the attention weight A of the fusion feature fusion , and get the enhanced feature F enhanced : A fusion =σ(Conv(F fusion )), Add the enhanced fusion feature to the residual input to get the intermediate feature F temp : F temp =F enhanced ⊕F in , Finally, the feature representation capability is further enhanced by channel attention and spatial attention to obtain the final output F of the EFA module. EFA : M c (F temp )=σ(MLP(AvgPool(F temp ))⊕MLP(MaxPool(F temp ))), M s (F temp )=σ(f 7×7 (Concat(AvgPool(F temp ),MaxPool(F temp )))), In the above formula, Concat(·) represents concatenation of features in the channel dimension, Conv(·) represents convolution operation, represents the element-by-element multiplication operation, ⊕ represents the element-by-element addition operation, σ represents the sigmoid activation function, and f 7×7 represents a filter with a kernel size of 7, MLP represents a multi-layer perceptron, and M c , M s They represent the channel attention layer and the spatial attention layer respectively.
5. The method for liver tumor segmentation integrating three types of attention according to claim 1, characterized in that: S4 includes the following steps: S4.1: Calculate marginal supervision loss; S4.2: Calculate segmentation loss; S4.3: Construct multi-scale edge segmentation loss; Among them, S4.1 includes the following steps: For feature maps of different scales, edge features are extracted by Laplacian operator Guide the model to learn the boundary area more accurately. Edge supervision loss L ES The definition is as follows: In the above formula, represents the predicted edge feature, e represents the true edge feature, represents the segmentation prediction result, Laplace(·) represents the Laplace operator operation, and N represents the total number of pixels. Among them, S4.2 includes the following steps: The segmentation loss combines Dice loss and cross entropy loss. SEG The definition is as follows: L SEG =λ1L DICE +λ2L CE , In the above formula, y i represents the actual segmentation result, Represents the predicted segmentation result, N represents the total number of pixels, and λ1 and λ2 are weight hyperparameters. Among them, S4.3 includes the following steps: Use multi-scale supervision and integrate edge supervision loss L ES and segmentation loss L SEG , we get the multi-scale edge segmentation loss L MES As shown below: Among them, S is set to 4, α s represents the loss weight of the s-th scale, represents the marginal supervision loss of the s-th scale, represents the segmentation loss of the s-th scale.
6. The method for liver tumor segmentation integrating three types of attention according to claim 1, characterized in that: S5 includes the following steps: Taking the U-shaped network as the basic architecture, the DADC module is first used to replace the second, third, and fourth convolutional layers of the encoder and all convolutional layers of the decoder in the U-shaped network. Then, the EFA module is introduced into the bottleneck layer of the U-shaped network. Finally, the multi-scale edge segmentation loss function L is combined. MES Optimize the overall learning process of the model to build the TAF-Net model.
7. The method for liver tumor segmentation integrating three types of attention according to claim 1, characterized in that: S6 includes the following steps: Use the divided experimental data set to complete the model training and verification optimization, save the optimal model weights, and then use the test set to evaluate the model performance and obtain various model indicators.
Citation Information
Patent Citations
Lung CT image segmentation method based on transfer learning and attention mechanism
CN115457049A
Endoscopic polyp segmentation method based on boundary supervision and time sequence association
CN116824139A
Kidney tumor image segmentation method and system based on boundary extraction
CN118115738A
MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on multi-scale attention mechanism and secondary segmentation strategy
CN119131051A
Thangka image segmentation method based on edge feature guidance and detail feature denoising
CN119169296A