Medical image segmentation system and method combining position and channel double attention
By combining the medical image segmentation method with position and channel dual attention, using residual network and Transformer encoder, the feature representation and spatial information are enhanced, and the problems of spatial information loss and insufficient information utilization in the prior art are solved, and the accuracy and robustness of medical image segmentation are improved.
Patent Information
- Application Number
- CN202510204091.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing medical image segmentation method leads to the loss of spatial information during the downsampling process and lacks the utilization of context and global information, resulting in the model paying attention to the uninterested area during the training process, and the segmentation effect is poor.
Using a medical image segmentation method combining position and channel dual attention, a residual network module, a Transformer encoder and a dual attention block (DRA-Block) is used to enhance feature representation and retain spatial information, while using position and channel attention block (CPAM and CCAM) to extract sparsely encoded features.
It effectively alleviates the loss of spatial information caused by pooling operations, enhances the utilization of global information, improves the robustness and segmentation accuracy of the model, reduces the sensitivity of overfitting, and improves the overall performance and generalization capabilities of the model.
Smart Images

Figure CN120279032A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of neural networks and image processing, and particularly to a medical image segmentation system and method combining position and channel dual attention. Background Art
[0002] Medical image segmentation is one of the core tasks of medical image analysis and plays a crucial role in disease diagnosis, treatment plan formulation, and patient prognosis evaluation. With the rapid development of deep learning technology, its application in the field of medical image segmentation has also made remarkable progress. However, the analysis and diagnosis of medical images often require accurate image segmentation to identify and locate various abnormalities, such as tumors, lesions, etc. Traditional medical image segmentation methods often consume a large amount of time and manpower and have limited accuracy.
[0003] For example, the existing patent CN111681252B discloses a medical image segmentation method and system based on deep learning. In the downsampling process, two adjacent feature layers with different resolutions in any one of the historical magnetic resonance imaging (MRI) modality images in the training set are input into a neural network model for multi-level feature re-extraction and aggregation to determine the segmented MRI modality image. The two adjacent feature layers with different resolutions include a low-resolution feature layer and a high-resolution feature layer; the two adjacent feature layers with different resolutions pass through a residual convolution unit, a resolution fusion unit, and an aggregation unit in sequence to determine the segmented MRI modality image. However, the continuous pooling operation in the downsampling process of this patented technology results in the loss of spatial information.
[0004] Another example is the existing patent CN111681252B, which discloses a medical image automatic segmentation method based on multi-path attention fusion. The input image is input into a multi-path encoder, which consists of 4 paths respectively composed of 1, 2, 3, and 4 residual blocks. Then, the outputs of each path are fused through an attention mechanism, and then decoded by a decoder to output the segmented image. Although this patent uses attention to fuse the features of different path outputs, it lacks the utilization of context information and global information.
[0005] Therefore, it is of great significance to develop an artificial intelligence-based medical image segmentation method. Summary of the Invention
[0006] The purpose of the present invention is to provide a medical image segmentation system and method combining position and channel dual attention, which effectively solves the problems of the loss of spatial information caused by pooling operations in the circuit downsampling process and the lack of utilization of context information and global information. At the same time, it overcomes the situation where the model pays more attention to uninteresting regions during training due to pixel class imbalance in medical images, resulting in poor segmentation effects of the model.
[0007] The technical solution adopted by the present invention is as follows:
[0008] A medical image segmentation method combining position and channel dual attention, which comprises the following steps:
[0009] Step 1: Input the medical image to be segmented into a feature encoder, and obtain a first feature map with an image size half of the input medical image and 64 channels through a standard convolutional layer, a group normalization layer, and a ReLU activation function;
[0010] Step 2: After the first feature map passes through max pooling and 3 residual blocks, a second feature map with a halved image size and 256 channels is obtained (as Figure 7 shown);
[0011] Step 3: The second feature map passes through 4 residual blocks to obtain a third feature map with a halved image size and 512 channels;
[0012] Step 4: The third feature map passes through 9 residual blocks to obtain a fourth feature map with a halved image size and 768 channels;
[0013] Step 5: The fourth feature map passes through a dual attention block (DRA-Block) to enhance the depth of the feature representation while retaining the inherent features of the input map to obtain a fifth feature map;
[0014] Step 6: After the fifth feature map is subjected to image serialization processing, a Transformer encoder is executed. After passing through 12 encoding blocks Encoder Block (as Figure 5 shown), a sixth feature map is obtained. The Transformer encoder includes a number of encoding blocks, and each encoding block uses a spatial reduction attention (SRA) layer to replace the traditional multi-head self-attention (MHA) layer in the encoder, so as to form a block with a multi-layer perceptron (MLP) block;
[0015] Step 7: After shaping the sixth feature map, a convolution with a convolution kernel size of 3x3 and a ReLU activation function are performed to obtain the output features of the final feature encoder;
[0016] Step 8: The output features of the feature encoder are processed by a feature decoder to obtain a seventh feature map,
[0017] Step 9: After the seventh feature map is upsampled once, it is concatenated with the output features of the corresponding feature encoder, and then passes through a decoding block and a convolution with a convolution kernel size of 1x1 to obtain an eighth feature map;
[0018] Step 10: Repeat the operations of Step 8 and Step 9 twice to obtain a ninth feature map with 64 channels;
[0019] Step 11, perform upsampling on the ninth feature map and then perform convolutions with convolution kernels of size 3x3 and 1x1 to obtain the final segmentation map.
[0020] Further, in Step 1, the convolution kernel of the standard convolution layer has a size of 7x7 and a stride of 2.
[0021] Further, in Step 5, the position attention module CPAM of the dual attention block performs the following steps:
[0022] Step 5-1-1, pass the input feature through a multi-scale pooling layer, and then input it into a 1x1 convolution layer to obtain pooled features of different sizes, with sizes of 1x1, 2x2, 3x3, and 6x6 respectively; regard each position of the pooled feature as an aggregation center, and reshape the pooled feature into a size of LxL Concatenate all the pooled features to obtain the position set center F ∈ R C×M , where M is the sum of the number of grid points of all pooled features;
[0023] Step 5-1-2, merge the aggregation center to each pixel according to semantic relevance, that is, input the input feature and the position set center F into a 1x1 convolution layer and a fully connected layer to obtain features B and C respectively, where is the number of feature channels after convolution, B represents the feature obtained through the 1x1 convolution layer; C represents the feature obtained through the fully connected layer; then use matrix multiplication and the softmax layer to obtain the spatial attention map S ∈ R N×M , where N = H×W is the number of pixels; then input the position set center F into the fully connected layer to obtain the feature D ∈ R C×M , D represents the feature obtained after processing the position set center F through the fully connected layer;
[0024] Step 5-1-3, obtain the output feature E of the position attention module composed of the spatial attention map S and the feature D, and the corresponding expression is as follows:
[0025]
[0026] where, α represents a learnable parameter of the position attention module CPAM, initialized to 0, and gradually learns to assign more weights to control the influence of the feature from the aggregation center in the output of the attention module; D i represents the center of the i-th channel of the feature D; A j represents the j-th channel of the input feature A;
[0027] Further, in Step 5, the channel attention module CCAM of the dual attention block performs the following steps:
[0028] Step 5-2-1: Use a 1×1 convolutional layer to reduce the channel dimension of the input feature A to obtain the dimension-reduced feature F′∈R K×H×W ; Consider each channel mapping of the dimension-reduced feature F′ as a set center, and calculate the channel attention map X∈R C×K ; where the channel attention map X∈R C×K has the following specific expression:
[0029]
[0030] where, x ji represents the influence of the i-th center of the measured input feature A on the j-th channel; K represents the number of channels of the dimension-reduced feature F′, that is, the number of aggregation centers; H and W respectively represent the height and width of the feature map; C represents the original number of channels of the input feature; A j represents the j-th channel of the input feature A, and F i represents the i-th channel center of the dimension-reduced feature F′.
[0031] Step 5-2-2: Use the channel attention map to selectively integrate the channel aggregation centers into each channel of the feature A to obtain the output feature E′; the expression of the output feature E′ is as follows:
[0032]
[0033] where, β represents a learnable parameter of the channel attention module CCAM, which is used to gradually adjust the influence of the channel aggregation centers on the final output feature, and the initial value is 0.
[0034] Further, the convolution with a kernel size of 1x1 in Step 9 is used for dimension reduction.
[0035] A medical image segmentation system that combines position and channel dual attention, which includes a feature encoder and a feature decoder. The feature encoder uses a CNN-Transformer Hybrid coding unit. The CNN-Transformer Hybrid coding unit includes a ResNet50 part and a Transformer encoder part; the ResNet50 part includes a standard convolutional layer, a group normalization layer, a ReLU activation function, a max pooling layer, and three different stage blocks arranged in sequence. Different stage blocks use residual blocks with different numbers of layers; the Transformer encoder part includes several coding blocks, and the feature decoder uses a cascaded upsampling unit.
[0036] Among them, the output features of the ReLU activation function in the ResNet50 part and the output features of the first two stage blocks are respectively connected to the cascaded upsampling units at different levels of the feature decoder through dual attention blocks, so as to help the decoder reconstruct the feature map more accurately; the output features of the third stage block in the ResNet50 part are connected to the Transformer encoder part through a dual attention block; the dual attention block includes a position attention module CPAM and a channel attention module CCAM. The dual attention block extracts the features of sparse coding from the perspectives of position and channel, extracting more valuable information while reducing redundancy; the encoding block uses a spatial reduction attention layer to replace the traditional multi-head self-attention layer in the encoder, so that the spatial reduction attention layer and the multi-layer perceptron block form the encoding block; the cascaded upsampling unit includes an upsampling module and a convolutional module.
[0037] Furthermore, the three stage blocks respectively adopt 3-layer residual blocks, 4-layer residual blocks and 9-layer residual blocks.
[0038] Furthermore, the medical image passes through a standard convolutional layer, a group normalization layer and a ReLU activation function to obtain a first feature map with an image size half of the input medical image and 64 channels; the first feature map passes through a max pooling and 3-layer residual blocks to obtain a second feature map with a halved image size and 256 channels; the second feature map passes through 4-layer residual blocks to obtain a third feature map with a halved image size and 512 channels; the third feature map passes through 9-layer residual blocks to obtain a fourth feature map with a halved image size and 768 channels; the fourth feature map passes through a dual attention block (DRA-Block) to enhance the depth of feature representation while retaining the inherent features of the input mapping to obtain a fifth feature map; the fifth feature map is serially processed and then performs Transformer encoding, and passes through 12-layer encoding blocks to obtain a sixth feature map; after shaping the sixth feature map, convolution with a convolution kernel size of 3x3 and a ReLU activation function are performed to obtain the output features of the final feature encoder.
[0039] Furthermore, the input layer of the position attention module CPAM is divided into three-way outputs. One output of the input layer passes through a 1×1 convolution to obtain feature B; another output of the input layer passes through a multi-scale pooling layer and a 1×1 convolutional layer in sequence to form the center of the position set, that is, the output of the input layer passes through the multi-scale pooling layer and the 1×1 convolutional layer to obtain pooling features of different sizes, and the pooling features of different sizes are cascaded to obtain the center of the position set; one output feature of the center of the position set passes through a fully connected layer to obtain feature C, and another output feature of the center of the position set passes through a fully connected layer to obtain feature D; after feature B and feature C are concatenated, they are connected to the softmax layer; after the output end of the softmax layer is concatenated with feature D, and then concatenated with the third output of the input layer and connected to the output layer of the position attention module CPAM.
[0040] Specifically, the position attention module CPAM passes the input features through a multi-scale pooling layer, and then inputs them into a 1×1 convolutional layer to obtain pooled features of different sizes, with sizes of 1×1, 2×2, 3×3, and 6×6 respectively; for simplicity, Figure 3 the features of 6×6 are not drawn; then, each position of the pooled features is regarded as an aggregation center, and the pooled features are reshaped into Finally, all the pooled features are concatenated to obtain the position set center F∈R C×M , where M is the total number of grid points of all the pooled features;
[0041] The aggregation centers are adaptively merged onto each pixel according to semantic relevance. Specifically, the input feature A and the position set center F are input into a 1×1 convolutional layer and a fully connected layer to obtain features B and C respectively, where is the number of feature channels after convolution, B represents the feature obtained through the 1×1 convolutional layer; C represents the feature obtained through the fully connected layer; then matrix multiplication and a softmax layer are used to obtain the spatial attention map S∈R N×M , where N = H×W is the number of pixels;
[0042]
[0043] where, s ji represents the relationship between the i-th center and the j-th pixel. Then the set center F is input into a fully connected layer to obtain the feature D∈R C×M ; after obtaining the features S and D, the output feature E is then output, and the corresponding expression is as follows:
[0044]
[0045] where, α represents a learnable parameter of the position attention module CPAM, initialized to 0 and gradually learning to assign more weights, used to control the influence of the features from the aggregation center in the output of the attention module; D i represents the i-th channel center of the feature D; A j represents the j-th channel of the input feature A.
[0046] Furthermore, the input layer of the channel attention module CCAM is divided into three-way outputs. One output of the input layer passes through a 1×1 convolution to obtain a dimensionality-reduced feature; one output of the dimensionality-reduced feature is concatenated with another output of the input layer and then output to the softmax layer; the output of the softmax layer is concatenated with another output of the dimensionality-reduced feature, and then concatenated with the third output of the input layer and connected to the output layer of the channel attention module CCA.
[0047] Specifically, the channel attention module CCAM uses a 1×1 convolutional layer to reduce the channel dimension of the input feature A to obtain the feature F'∈R after dimensionality reduction K×H×W ; regarding each channel map of the feature F' after dimensionality reduction as a set center, calculate the channel attention map X∈R C×K ; selectively integrate the channel aggregation center into each channel of the feature A using the channel attention map to obtain the output feature E; where the channel attention map X∈R C×K The specific expression is as follows:
[0048]
[0049] where, x ji represents the influence of the i-th center of the measured input feature A on the j-th channel; K represents the number of channels of the feature F' after dimensionality reduction, that is, the number of aggregation centers; H and W respectively represent the height and width of the feature map; C represents the original number of channels of the input feature; A j represents the j-th channel of the input feature A, and F i represents the i-th channel center of the feature F' after dimensionality reduction.
[0050] The expression of the output feature E' is as follows:
[0051]
[0052] where, β represents a learnable parameter of the channel attention module CCAM, used to gradually adjust the influence of the channel aggregation center on the final output feature, and the initial value is 0.
[0053] The present invention adopts the above technical solutions and has the following technical advantages compared with the prior art: 1. By adopting the residual network module, the loss of spatial information caused by the pooling operation is alleviated, and the feature extraction ability of the network is further improved. 2. Transformers encode long-range dependencies to make up for the lack of capture of global information by CNN due to the limitation of the convolutional kernel size, strengthen the utilization of global information, and enhance the feature expression ability of the network. And a DRA-Block is included before the Transformer layer to perform specialized image processing on the post-convolutional features, enhancing the feature extraction ability of the Transformer for image content. A spatial reduction attention (SRA) layer is used to replace the traditional multi-head self-attention (MHA) layer in the encoder, reducing the computational / memory cost. 3. In the DRA_TransUnet model of the present invention, dual attention blocks (DRA-Blocks) are introduced in three skip connection layers respectively. Integrating the DRA-Block into the skip connection can refine the sparsely encoded features from the perspectives of position and channel, extract more valuable information while reducing redundancy. By doing so, the DRA-Block can help the decoder reconstruct the feature map more accurately. In addition, the addition of the DRA-Block not only enhances the robustness of the model, but also effectively reduces the sensitivity of the model to overfitting, improving the overall performance and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The following further describes the present invention in detail with reference to the drawings and specific embodiments;
[0055] Figure 1 It is a schematic structural diagram of a medical image segmentation system combining position and channel dual attention according to the present invention;
[0056] Figure 2 It is a schematic structural diagram of a dual attention block (DRA-Block) according to the present invention;
[0057] Figure 3 It is a schematic structural diagram of a position attention module (CPAM) according to the present invention;
[0058] Figure 4 It is a schematic structural diagram of a channel attention module (CCAM) according to the present invention;
[0059] Figure 5 It is a schematic structural diagram of an encoder block according to the present invention;
[0060] Figure 6 It is a schematic structural diagram of an attention (SRA) layer according to the present invention;
[0061] Figure 7This is a schematic diagram of the structure of the residual block of the present invention. Detailed implementation manners
[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application.
[0063] As Figures 1 to 7 shown in one of them, the present invention discloses a medical image segmentation method combining position and channel dual attention, which includes the following steps:
[0064] Step 1: Input the medical image to be segmented into a feature encoder, and obtain a first feature map with an image size half of the input medical image and 64 channels through a standard convolutional layer, a group normalization layer, and a ReLU activation function;
[0065] Step 2: After the first feature map passes through max pooling and 3 residual blocks, a second feature map with a halved image size and 256 channels is obtained (as Figure 7 shown);
[0066] Step 3: The second feature map passes through 4 residual blocks to obtain a third feature map with a halved image size and 512 channels;
[0067] Step 4: The third feature map passes through 9 residual blocks to obtain a fourth feature map with a halved image size and 768 channels;
[0068] Step 5: The fourth feature map passes through a dual attention block (DRA-Block) to enhance the depth of the feature representation while retaining the inherent features of the input map to obtain a fifth feature map;
[0069] Step 6: After the fifth feature map is serially processed, a Transformer encoder is executed. After passing through 12 encoding blocks Encoder Block (as Figure 5 shown), a sixth feature map is obtained. The Transformer encoder includes a number of encoding blocks, and each encoding block uses a spatial reduction attention (SRA) layer to replace the traditional multi-head self-attention (MHA) layer in the encoder, so as to form a block with a multi-layer perceptron (MLP) block;
[0070] Step 7: After the sixth feature map is reshaped, a convolution with a convolution kernel size of 3x3 and a ReLU activation function are performed to obtain the output features of the final feature encoder;
[0071] Step 8: The output features of the feature encoder are processed by a feature decoder to obtain a seventh feature map,
[0072] Step 9, the seventh feature map is upsampled once and concatenated with the output features of the corresponding feature encoder, and then passes through a decoding block and a convolution with a kernel size of 1x1 to obtain the eighth feature map;
[0073] Step 10, repeat the operations of Step 8 and Step 9 twice to obtain a ninth feature map with 64 channels;
[0074] Step 11, upsample the ninth feature map and then perform convolutions with kernel sizes of 3x3 and 1x1 to obtain the final segmentation map.
[0075] Further, in Step 1, the convolution kernel size of the standard convolution layer is 7x7 and the stride is 2.
[0076] Further, in Step 5, the position attention module CPAM of the dual attention block performs the following steps:
[0077] Step 5-1-1, pass the input features through a multi-scale pooling layer, and then input them into a 1×1 convolution layer to obtain pooled features of different sizes, with sizes of 1×1, 2×2, 3×3, and 6×6 respectively; treat each position of the pooled features as an aggregation center, and reshape the pooled features with a size of L×L into Concatenate all the pooled features to obtain the position set center F∈R C×M , where M is the sum of the grid points of all the pooled features;
[0078] Step 5-1-2, merge the aggregation center to each pixel according to semantic relevance, that is, input the input features and the position set center F into a 1×1 convolution layer and a fully connected layer respectively to obtain features B and C, where is the number of feature channels after convolution, B represents the feature obtained through the 1×1 convolution layer; C represents the feature obtained through the fully connected layer; then use matrix multiplication and the softmax layer to obtain the spatial attention map S∈R N×M , where N = H×W is the number of pixels; then input the position set center F into the fully connected layer to obtain the feature D∈R C×M , D represents the feature obtained after processing the position set center F through the fully connected layer;
[0079] Step 5-1-3, obtain the output feature E of the position attention module composed of the spatial attention map S and the feature D, and the corresponding expression is as follows:
[0080]
[0081] Among them, α represents a learnable parameter of the position attention module CPAM, which is initialized to 0 and gradually learns to assign more weights to control the influence of the features from the aggregation center in the output of the attention module; Di Represents the center of the i-th channel of feature D; A j Represents the j-th channel of the input feature A.
[0082] Furthermore, the channel attention module CCAM of the dual attention block in step 5 performs the following steps:
[0083] Step 5-2-1, use a 1×1 convolutional layer to reduce the channel dimension of the input feature A to obtain the reduced-dimensional feature F′∈R K×H×W ; regard each channel map of the reduced-dimensional feature F′ as a set center, and calculate the channel attention map X∈R C×K ; where the channel attention map X∈R C×K The specific expression is as follows:
[0084]
[0085] where, x ji represents the influence of the i-th center of the measured input feature A on the j-th channel; K represents the number of channels of the reduced-dimensional feature F′, that is, the number of aggregation centers; H and W respectively represent the height and width of the feature map; C represents the original number of channels of the input feature; A j represents the j-th channel of the input feature A, F i represents the center of the i-th channel of the reduced-dimensional feature F′.
[0086] Step 5-2-2, use the channel attention map to selectively integrate the channel aggregation centers into each channel of the feature A to obtain the output feature E′; the expression of the output feature E′ is as follows:
[0087]
[0088] where, β represents a learnable parameter of the channel attention module CCAM, which is used to gradually adjust the influence of the channel aggregation center on the final output feature, and the initial value is 0.
[0089] Furthermore, the convolution with a kernel size of 1x1 in step 9 is used for dimensionality reduction.
[0090] A medical image segmentation system that combines position and channel dual attention, which includes a feature encoder and a feature decoder. The feature encoder uses a CNN-Transformer Hybrid encoding unit, and the CNN-Transformer Hybrid encoding unit includes a ResNet50 part and a Transformer encoder part. The ResNet50 part includes a standard convolutional layer, a group normalization layer, a ReLU activation function, a max pooling layer, and three different stage blocks arranged in sequence. Different stage blocks use residual blocks with different numbers of layers. The Transformer encoder part includes several encoding blocks, and the feature decoder uses a cascaded upsampling unit.
[0091] Among them, the output features of the ReLU activation function in the ResNet50 part and the output features of the first two stage blocks are respectively connected to the cascaded upsampling units at different levels of the feature decoder through dual attention blocks, so as to help the decoder more accurately reconstruct the feature map. The output features of the third stage block in the ResNet50 part are connected to the Transformer encoder part through a dual attention block. The dual attention block includes a position attention module CPAM and a channel attention module CCAM. The dual attention block refines the sparsely encoded features from the perspectives of position and channel, extracting more valuable information while reducing redundancy. The encoding block uses a spatial reduction attention layer to replace the traditional multi-head self-attention layer in the encoder, so that the spatial reduction attention layer and the multi-layer perceptron block form the encoding block. The cascaded upsampling unit includes an upsampling module and a convolutional module.
[0092] Further, the three stage blocks respectively use 3-layer residual blocks, 4-layer residual blocks, and 9-layer residual blocks.
[0093] Further, the medical image passes through a standard convolutional layer, a group normalization layer, and a ReLU activation function to obtain a first feature map with an image size half of the input medical image and 64 channels. The first feature map passes through max pooling and 3-layer residual blocks to obtain a second feature map with a halved image size and 256 channels. The second feature map passes through 4-layer residual blocks to obtain a third feature map with a halved image size and 512 channels. The third feature map passes through 9-layer residual blocks to obtain a fourth feature map with a halved image size and 768 channels. The fourth feature map passes through a dual attention block (DRA-Block) to enhance the depth of feature representation while retaining the inherent features of the input mapping to obtain a fifth feature map. The fifth feature map undergoes image serialization processing and then performs Transformer encoding, and passes through 12-layer encoding blocks to obtain a sixth feature map. After shaping the sixth feature map, convolution with a convolution kernel size of 3x3 and a ReLU activation function are performed to obtain the output features of the final feature encoder.
[0094] Furthermore, the input layer of the position attention module CPAM is divided into three outputs. One output of the input layer passes through a 1×1 convolution to obtain feature B; another output of the input layer passes through a multi-scale pooling layer and a 1×1 convolution layer in sequence to form the center of the position set, that is, the output of the input layer passes through the multi-scale pooling layer and the 1×1 convolution layer to obtain pooling features of different sizes, and the pooling features of different sizes are cascaded to obtain the center of the position set; one output feature of the center of the position set passes through a fully connected layer to obtain feature C, and another output feature of the center of the position set passes through a fully connected layer to obtain feature D; feature B and feature C are concatenated and then connected to the softmax layer; after the output end of the softmax layer is concatenated with feature D, it is concatenated with the third output of the input layer and then connected to the output layer of the position attention module CPAM.
[0095] Furthermore, the input layer of the channel attention module CCAM is divided into three outputs. One output of the input layer passes through a 1×1 convolution to obtain a dimensionality-reduced feature; one output of the dimensionality-reduced feature is concatenated with another output of the input layer and then output to the softmax layer; after the output of the softmax layer is concatenated with another output of the dimensionality-reduced feature, it is concatenated with the third output of the input layer and then connected to the output layer of the channel attention module CCA.
[0096] Furthermore, the feature decoder includes feature fusion, a segmentation head, and three upsampling convolutional blocks; feature fusion integrates the feature map transmitted through the skip connection with the existing feature map, thereby helping the decoder to reconstruct the original feature map; the segmentation head restores the finally output feature map to its original size; the three upsampling convolutional blocks incrementally double the size of the input feature map at each step, effectively restoring the resolution of the image.
[0097] The specific principle of the present invention will be described in detail below:
[0098] The present invention is applicable to the segmentation of medical images and helps doctors accurately judge lesions clinically. For example Figure 1As shown in the figure, the DRA_TransUnet model of the present invention mainly includes three main parts: DRA-Block, CNN-Transformer Hybrid encoding, and skip connections. The CNN-Transformer Hybrid combines CNN feature extraction and Transformer and is further enriched by the DRA-Block specifically introduced in this model architecture. The cascaded upsampling consists of upsampling and convolution. In contrast, the decoder mainly uses traditional convolution mechanisms. To optimize the skip connections, the DRA-Block is a key component in the DRA_TransUnet architecture. The DRA-Block consists of dual attention. The DRA-Block filters out irrelevant information in the skip connections and improves the accuracy of image reconstruction. In summary, compared with traditional convolution methods and the widespread use of Transformer, DRA_TransUnet uniquely uses the DRA-Block to extract and utilize image-specific location and channel features. This strategic combination significantly improves the overall performance of the model.
[0099] As Figure 2 shown, as a feature extraction module, the DRA-Block integrates image-specific location and channel features. In this way, feature extraction can be performed according to the unique attributes of the image. Especially in the U-Net type architecture, the special feature extraction ability of the DRA-Block is crucial. Although Transforms are good at using attention mechanisms to extract global features, they are not specifically tailored for image-specific attributes. In contrast, the DRA-Block performs well in both location-based and channel-based feature extraction and can obtain a more detailed and accurate feature set. Therefore, the present invention incorporates it into the encoder and skip connections to improve the segmentation performance of the model. The DRA-Block consists of two main components: one with a Compact Position Attention Module (CPAM), and the other containing a Channel Attention Module (CCAM).
[0100] CPAM (Compact Position Attention Module): Since the present invention needs to obtain the relationship between any two pixels through the inner product of vectors, if the number of pixels is large, expensive GPU memory storage and computational costs are required. To alleviate this problem, the present invention proposes a CPAM method, as Figure 3 shown. This method constructs the relationship between each pixel and several acquisition centers. By collecting feature vectors from a subset of pixels in the input tensor, the collection centers are formally defined as compact feature vectors. They are implemented through a spatial pyramid pooling scheme, which provides context information from different spatial scales.
[0101] First, the features are input into the multi-scale pooling layer, and then into the 1×1 convolutional layer to obtain several pooled features with sizes of 1×1, 2×2, 3×3, and 6×6 respectively. For simplicity, Figure 3 the 6×6 feature is not drawn. Then, each position of the pooled features is regarded as an aggregation center, and these features are reshaped into Finally, all the pooled features are concatenated to obtain the set center F∈R C×M , where M is the total number of grid points of all the pooled features.
[0102] Next, the aggregation centers are adaptively merged into each pixel according to semantic relevance. Specifically, the input feature A and the position set center F are input into the 1×1 convolutional layer and the fully connected layer to obtain features B and C respectively, where is the number of feature channels after convolution, B represents the feature obtained through the 1×1 convolutional layer; C represents the feature obtained through the fully connected layer; then matrix multiplication and the softmax layer are used to obtain the spatial attention map S∈R N×M , where N = H×W is the number of pixels;
[0103]
[0104] where, s ji represents the relationship between the i-th center and the j-th pixel. Then the set center F is input into the fully connected layer to obtain the feature D∈R C×M ; after obtaining the features S and D, the output feature E is obtained, and the corresponding expression is as follows:
[0105]
[0106] where, α represents a learnable parameter of the position attention module CPAM, initialized to 0, and gradually learns to assign more weights to control the influence of the features from the aggregation center in the output of the attention module; D i represents the i-th channel center of the feature D; A j represents the j-th channel of the input feature A.
[0107] CCAM (Compact Channel Attention Module): As Figure 4 shown, this is CCAM, which is good at extracting channel features. The channels of the input tensor are aggregated to obtain the channel set center. Specifically, the present invention uses a 1×1 convolutional layer to reduce the channel dimension of the input feature A to obtain the reduced-dimensional feature F′∈R K×H×W , and each channel map of the reduced-dimensional feature F′ is regarded as a set center. Then the channel attention map X∈R is calculatedC×K As follows:
[0108]
[0109] Among them, x ji represents the influence of the i-th center of the measured input feature A on the j-th channel; K represents the number of channels of the feature F' after dimensionality reduction, that is, the number of aggregation centers; H and W respectively represent the height and width of the feature map; C represents the original number of channels of the input feature; A j represents the j-th channel of the input feature A, and F i represents the i-th channel center of the feature F' after dimensionality reduction.
[0110] Then the present invention selectively integrates the channel aggregation centers into each channel of the feature A to obtain the expression of the output feature E' as follows:
[0111]
[0112] Among them, β represents a learnable parameter of the channel attention module CCAM, which is used to gradually adjust the influence of the channel aggregation center on the final output feature, and the initial value is 0.
[0113] The feature encoder adopts CNN-Transformer Hybrid encoding, which combines ResNet50 and Transformer. The ResNet50 here is a bit different from the traditional ResNet50. The initial convolutional layer of the ResNet50 here uses StdConv2d instead of the traditional Conv2d, and then all BatchNorm layers are replaced with GroupNorm layers. The number of stages and the number of repeated stacks in each stage are also different. Particularly importantly, a DRA-Block is included before the Transformer layer. This design aims to perform specialized image processing on the post-convolutional features and enhance the feature extraction ability of the Transformer for image content.
[0114] The Transformer layer was originally composed of alternating multi-head self-attention layers (MHA) and multi-layer perceptron (MLP) blocks. The MLP contains layers with GELU (activation function) non-linearity. To reduce the computational / memory cost, a spatial reduction attention (SRA) layer is used to replace the traditional multi-head self-attention (MHA) layer in the encoder, as Figure 5 shown. Before the input data enters the SRA and MLP modules, it is normalized by a normalization layer (Layer Normalization, LN), and the result is connected to the input in a residual manner. The dimension and size of the feature map remain unchanged during the entire encoding process.
[0115] Similar to the multi-head self-attention (MHA) layer, the added spatial reduction attention (SRA) layer receives a query Q, a key K, and a value V as inputs and outputs a refined feature. The difference is that the added SRA layer reduces the spatial scale of K and V before the attention operation (see Figure 6 ), which greatly reduces the computational / memory overhead. The details of the SRA in the i-th stage can be formulated as follows:
[0116]
[0117] where Concat(·) is the concatenation operation, and are the linear projection parameters. S i is the number of heads of the attention layer in the i-th stage. Therefore, the dimension of each head (i.e., dhead) is equal to C i is the number of channels of the output in the i-th stage. SR(·) is the operation of reducing the spatial dimension of the input sequence (i.e., K or V), which can be written as:
[0118] SR(X) = Norm(Reshape(X, R i )W A ) (7)
[0119] where represents an input sequence, R i represents the reduction rate of the attention layer in the i-th stage, Reshape(X, R i ) is the operation of reshaping the input sequence X into a sequence of size . is the linear projection of reducing an input sequence to C i . Norm(·) is layer normalization. Similar to the original Transformer, the attention operation Attention(·) is calculated as follows:
[0120]
[0121] Through these formulas, it can be found that the computational / memory cost of the attention operation of the present invention is lower than that of MHA by times. Therefore, the SRA of the present invention can process larger input feature maps / sequences with limited resources.
[0122] Skip connections with dual attention: Similar to other u-structure models, skip connections are added between the encoder and the decoder to bridge the semantic gap existing between them. To further reduce this semantic gap, we introduce dual attention blocks (DRA-Block) in three skip connection layers, as shown in Figure 1As shown. This decision is based on our observation that traditional skip connections usually transmit redundant features, while DRA-Block can effectively filter these features. Integrating DRA-Block into skip connections can refine the features encoded sparsely from both the positional and channel perspectives, extracting more valuable information while reducing redundancy. By doing so, DRA-Block can help the decoder reconstruct the feature map more accurately. In addition, the addition of DRA-Block not only enhances the robustness of the model but also effectively reduces the model's sensitivity to overfitting, improving the overall performance and generalization ability of the model.
[0123] The feature decoder part is a cascaded upsampling composed of upsampling and convolution. The components of the decoder include feature fusion, a segmentation head, and three upsampling convolutional blocks. The first component: Feature fusion needs to integrate the feature map transmitted through the skip connection with the existing feature map, thus helping the decoder reconstruct the original feature map. The second component: The segmentation head is responsible for restoring the finally output feature map to its original size. The third component: The three upsampling convolutional blocks incrementally double the size of the input feature map at each step, effectively restoring the resolution of the image. The specific process is as follows: (1) After one upsampling of the output of the decoding part, it is concatenated with the output of the corresponding encoding layer, and then passed through a convolution with a kernel size of 1x1, where the 1x1 convolution is used for dimensionality reduction. (2) Repeat the operation in (1) twice to obtain a feature map with 64 channels. (3) The output is upsampled and then passed through convolutions with kernel sizes of 3x3 and 1x1 to obtain the final segmentation map.
[0124] The technical features of the present invention are: (1) Using DRA-Block composed of dual attention, the position attention module constructs the relationship between each pixel and several acquisition centers. By collecting feature vectors from a subset of pixels in the input tensor, the collection center is formally defined as a compact feature vector. It is implemented through a spatial pyramid pooling scheme, which provides context information from different spatial scales. The channel attention module uses a 1×1 convolutional layer to reduce the channel dimension of the input features. (2) To reduce the computational / memory cost, a spatial reduction attention (SRA) layer is used to replace the traditional multi-head self-attention (MHA) layer in the encoder. (3) The encoder-decoder structure model presents a U-shaped structure. The feature encoder adopts CNN-Transformer Hybrid, and DRA-Block is added before the Transformer layer in the encoder. Enhancing the feature extraction ability of Transformer for image content. And to reduce the semantic gap between the encoder and the decoder, DRA-Block is introduced in each of the three skip connection layers.
[0125] The proposed DRA_TransUnet is applied to different 2D medical image segmentation tasks. As shown in Tables 1, 2, and 3, there are three different 2D medical image segmentation tasks: breast tumor segmentation, skin lesion segmentation, and polyp segmentation. To evaluate the model and for expansion, the present invention uses five commonly used segmentation task metrics. These metrics are Pixel Accuracy (PA), Class Pixel Accuracy (CPA), IoU, Recall, and Dice Score. The present invention demonstrates its effectiveness and advancement by comparing the segmentation task metric values with other segmentation methods.
[0126] Table 1 Comparison of Breast Malignant Tumor Segmentation Performance
[0127]
[0128] Table 2 Comparison of Skin Lesion Segmentation Performance
[0129]
[0130] Table 3 Comparison of Polyp Segmentation Performance
[0131]
[0132] The present invention adopts the above technical solutions and has the following technical advantages compared with the prior art: 1. By adopting the residual network module, the loss of spatial information caused by the pooling operation is alleviated, and the feature extraction ability of the network is further improved. 2. Transformers encode long-range dependencies to make up for the lack of capture of global information by CNN due to the limitation of the convolutional kernel size, strengthen the utilization of global information, and enhance the feature expression ability of the network. And a DRA-Block is included before the Transformer layer to perform specialized image processing on the post-convolutional features, enhancing the feature extraction ability of the Transformer for image content. A Spatial Reduction Attention (SRA) layer is used to replace the traditional multi-head self-attention (MHA) layer in the encoder, reducing the computational / memory cost. 3. In the DRA_TransUnet model of the present invention, dual attention blocks (DRA-Blocks) are introduced into the three skip connection layers respectively. Integrating the DRA-Block into the skip connection can refine the sparsely encoded features from the perspectives of position and channel, extract more valuable information while reducing redundancy. By doing so, the DRA-Block can help the decoder reconstruct the feature map more accurately. In addition, the addition of the DRA-Block not only enhances the robustness of the model, but also effectively reduces the sensitivity of the model to overfitting, improving the overall performance and generalization ability of the model.
[0133] Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. Generally, the components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
Claims
1. A medical image segmentation system combining location and channel dual attention, which includes a feature encoder and a feature decoder, the feature encoder It includes a ResNet50 part and a Transformer encoder part; the ResNet50 part includes a standard convolutional layer, a group normalization layer, a ReLU activation function, a max pooling layer, and three different stage blocks arranged in sequence, and different stage blocks adopt residual blocks with different numbers of layers; the Transformer encoder part includes several encoding blocks, and the feature decoder adopts a cascaded upsampling unit. It is characterized in that: the output features of the ReLU activation function in the ResNet50 part and the output features of the first two stage blocks are respectively connected to the cascaded upsampling units at different levels of the feature decoder through dual attention blocks, so as to help the decoder reconstruct the feature map more accurately; the output features of the third stage block in the ResNet50 part are connected to the Transformer encoder part through a dual attention block; the dual attention block includes a position attention module CPAM and a channel attention module CCAM, and the dual attention block refines the features of the sparse coding from the perspectives of position and channel, extracting more valuable information while reducing redundancy; the encoding block adopts a spatial reduction attention layer to replace the traditional multi-head self-attention layer in the encoder, so that the spatial reduction attention layer and the multi-layer perceptron block form the encoding block; the cascaded upsampling unit of the feature decoder includes an upsampling module and a convolutional module.
2. The medical image segmentation system with combined location and channel dual attention according to claim 1, wherein: The three stage blocks respectively adopt 3-layer residual blocks, 4-layer residual blocks, and 9-layer residual blocks; in the ResNet50 part, the medical image passes through the standard convolutional layer, the group normalization layer, and the ReLU activation function to obtain a first feature map with an image size half of the input medical image and a channel number of 64; the first feature map passes through the max pooling and 3-layer residual blocks to obtain a second feature map with a halved image size and a channel number of 256; the second feature map passes through 4-layer residual blocks to obtain a third feature map with a halved image size and a channel number of 512; the third feature map passes through 9-layer residual blocks to obtain a fourth feature map with a halved image size and a channel number of 768; the fourth feature map passes through a dual attention block to enhance the depth of the feature representation while retaining the inherent features of the input mapping to obtain a fifth feature map. The fifth feature map is processed by image serialization and then performs Transformer encoding, and passes through 12-layer encoding blocks to obtain a sixth feature map; after shaping the sixth feature map, it is convolved with a convolution kernel of size 3x3 and passes through the ReLU activation function to obtain the output features of the final feature encoder.
3. The medical image segmentation system with combined location and channel dual attention according to claim 1, characterized in that: The input layer of the position attention module CPAM is divided into three outputs. One output of the input layer passes through a 1×1 convolution to obtain feature B; another output of the input layer passes through a multi-scale pooling layer and a 1×1 convolution layer in sequence to form the center of the position set, that is, the output of the input layer passes through the multi-scale pooling layer and the 1×1 convolution layer to obtain pooling features of different sizes, and the pooling features of different sizes are concatenated to obtain the center of the position set; one output feature of the center of the position set passes through a fully connected layer to obtain feature C, and another output feature of the center of the position set passes through a fully connected layer to obtain feature D; feature B and feature C are concatenated and then connected to the softmax layer; after the output end of the softmax layer is concatenated with feature D, it is concatenated with the third output of the input layer and then connected to the output layer of the position attention module CPAM.
4. The medical image segmentation system with combined location and channel dual attention according to claim 1, characterized in that: The input layer of the channel attention module CCAM is divided into three outputs. One output of the input layer passes through a 1×1 convolution to obtain a dimensionality-reduced feature; one output of the dimensionality-reduced feature is concatenated with another output of the input layer and then output to the softmax layer; after the output of the softmax layer is concatenated with another output of the dimensionality-reduced feature, it is concatenated with the third output of the input layer and then connected to the output layer of the channel attention module CCA.
5. The medical image segmentation system with combined location and channel dual attention according to claim 1, characterized in that: The feature decoder includes feature fusion , a segmentation head, and three upsampling convolutional blocks; Feature fusion integrates the feature maps transmitted through skip connections with the existing feature maps, thereby helping the decoder to reconstruct the original feature maps; the segmentation head restores the finally output feature maps to their original size; the three upsampling convolutional blocks incrementally double the size of the input feature maps at each step, effectively restoring the resolution of the image.
6. The medical image segmentation method combining position and channel dual attention adopts the medical image segmentation system combining position and channel dual attention according to any one of claims 1 to 5, and is characterized in that: The method includes the following steps: Step 1, input the medical image to be segmented into the feature encoder, and obtain the first feature map with an image size half of the input medical image and 64 channels through the standard convolutional layer, group normalization layer, and ReLU activation function; Step 2, the first feature map passes through max pooling and 3 residual blocks to obtain the second feature map with the image size halved and 256 channels; Step 3, the second feature map passes through 4 residual blocks to obtain the third feature map with the image size halved and 512 channels; Step 4, the third feature map passes through 9 residual blocks to obtain the fourth feature map with the image size halved and 768 channels; Step 5, the fourth feature map passes through the dual attention block to enhance the depth of the feature representation while retaining the inherent features of the input map to obtain the fifth feature map; Step 6, the fifth feature map is processed by image serialization and then the Transformer encoder is executed. After passing through 12 encoding blocks, the sixth feature map is obtained. The Transformer encoder includes several encoding blocks, and each encoding block uses a spatial reduction attention layer to replace the traditional multi-head self-attention layer in the encoder, so as to form a block with the multi-layer perceptron block; Step 7, after shaping the sixth feature map, perform convolution with a convolution kernel size of 3x3 and the ReLU activation function to obtain the output feature of the final feature encoder; Step 8, the output feature of the feature encoder is processed by the feature decoder to obtain the seventh feature map, Step 9: After one upsampling, the seventh feature map is concatenated with the output features of the corresponding feature encoder, and then passes through a decoding block and a convolution with a convolution kernel size of 1x1 to obtain the eighth feature map; Step 10: Repeat the operations of Step 8 and Step 9 twice to obtain a ninth feature map with 64 channels; Step 11: After upsampling the ninth feature map, it passes through convolutions with convolution kernel sizes of 3x3 and 1x1 to obtain the final segmentation map.
7. The medical image segmentation method combining location and channel dual attention according to claim 6, characterized in that: In Step 5, the position attention module CPAM of the dual attention block performs the following steps: Step 5-1-1, pass the input features through the multi-scale pooling layer, and then input them into the 1×1 convolutional layer to obtain pooled features of different sizes, with sizes of 1×1, 2×2, 3×3, and 6×6 respectively; regard each position of the pooled features as an aggregation center, and reshape the pooled features into Concatenate all the pooled features to obtain the center of the position set F∈R C ×M , where M is the sum of the number of grid points of all the pooled features; Step 5-1-2, merge the aggregation centers to each pixel according to semantic relevance, that is, input the input feature and the center F of the position set to a 1×1 convolutional layer and a fully connected layer to obtain features B and C respectively, where is the number of feature channels after convolution, B represents the feature obtained through the 1×1 convolutional layer; C represents the feature obtained through the fully connected layer; then use matrix multiplication and the softmax layer to obtain the spatial attention map S∈R N×M , where N = H×W is the number of pixels; then input the center F of the position set to the fully connected layer to obtain the feature D∈R C×M , D represents the feature obtained after processing the center F of the position set through the fully connected layer; Step 5-1-3: Obtain the spatial attention map S and the feature D to form the output feature E of the position attention module. The corresponding expression is as follows: Among them, α represents a learnable parameter of the position attention module CPAM, which is initialized to 0 and gradually learns to assign more weights to control the influence of the aggregated center features in the output of the attention module; D i represents the center of the i-th channel of the feature D; A j represents the j-th channel of the input feature A.
8. The medical image segmentation method combining location and channel dual attention according to claim 6, characterized in that: In Step 5, the channel attention module CCAM of the dual attention block performs the following steps: Step 5-2-1, use a 1×1 convolutional layer to reduce the channel dimension of the input feature A to obtain the dimension-reduced feature F′∈R K ×H×W ; regard each channel mapping of the dimension-reduced feature F′ as a set center, and calculate the channel attention map X∈R C×K ; where the specific expression of the channel attention map X∈R C×K is as follows: Among them, x ji represents the influence of the i-th center of the measured input feature A on the j-th channel; K represents the number of channels of the feature F' after dimensionality reduction, that is, the number of aggregation centers; H and W respectively represent the height and width of the feature map; C represents the original number of channels of the input feature; A j represents the j-th channel of the input feature A, and F i represents the i-th channel center of the feature F' after dimensionality reduction; Step 5-2-2: Selectively integrate the channel aggregation center into each channel of the feature A using the channel attention map to obtain the output feature E'; the expression of the output feature E' is as follows: Among them, β represents a learnable parameter of the channel attention module CCAM, which is used to gradually adjust the influence of the channel aggregation center on the final output feature, and the initial value is 0.
9. The medical image segmentation method combining location and channel dual attention according to claim 6, wherein: In Step 6, before the input data of the Transformer encoder enters the attention layer and the multi-head self-attention layer, it is normalized by the normalization layer, and the result is connected with the input in a residual connection. The dimension and size of the feature map remain unchanged during the entire encoding process.
Citation Information
Cited By
Hard-tipped-pen character stroke segmentation and extraction method, system and device and storage medium
CN122024256A