Medical image segmentation method and system based on efficient double attention
By introducing an efficient dual attention mechanism and feature fusion module in medical image segmentation technology, the shortcomings of local detail processing and multi-scale information fusion in the existing technology are solved, and higher segmentation accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510482139.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing medical image segmentation technology has shortcomings in processing local details and multi-scale information fusion, resulting in low segmentation accuracy and high computational complexity.
Using a medical image segmentation method based on efficient dual attention, the Dual Transformer module, encoder module, decoder module, multi-scale normalized channel attention module (MNCA) and feature fusion residual module (FFRM), is used to capture local and global features, and optimize the segmentation results through jump connection and comprehensive loss functions.
The accuracy and efficiency of medical image segmentation are improved, especially in the target positioning and multi-scale information fusion in complex backgrounds, which significantly improves the quality of segmentation results.
Smart Images

Figure CN119991709A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image segmentation, and in particular to a medical image segmentation method and system based on efficient dual attention, aiming to improve the segmentation accuracy and processing efficiency of medical images. Background Art
[0002] With the development of deep learning, especially convolutional neural networks (CNN), traditional CNN-based image segmentation methods, such as U-Net, are widely used in medical image segmentation, but their local receptive field and inductive bias limit the modeling of long-range dependencies. At the same time, Transformer-based methods perform well in capturing global semantics, but are weak in handling local details. Although hybrid models combining Transformer and CNN, such as TransUNet, have alleviated these problems to a certain extent, they still face challenges such as high computational cost and large training data requirements. Therefore, although medical image segmentation has made significant progress driven by deep learning, especially convolutional neural networks and Transformer architectures, its shortcomings in local details, computational complexity, and dependency on labeled data are still issues that need to be urgently addressed in current research. Especially in handling local details and multi-scale information fusion. Summary of the invention
[0003] In order to overcome the shortcomings of the prior art, the present invention provides a medical image segmentation method based on efficient dual attention, which aims to solve the problems of poor local details and insufficient segmentation accuracy in image segmentation processing in the prior art.
[0004] The present invention also provides a medical image segmentation system based on efficient dual attention.
[0005] The technical solution of the present invention is: Medical image segmentation method based on efficient dual attention, including: 1) Obtain a medical image dataset, perform data preprocessing, and divide the obtained dataset into a training set and a test set; 2) Construct a medical image segmentation model based on efficient dual attention; 3) Input the training set into the medical image segmentation model based on efficient dual attention and output the feature map; 4) Training a medical image segmentation model based on efficient dual attention to obtain an optimized medical image segmentation model based on efficient dual attention; 5) The test set is input into the optimized medical image segmentation model based on efficient dual attention, and the final medical image segmentation result is output.
[0006] Furthermore, step 1) includes the following steps: 1-1) Data collection and preprocessing: Use the Synapse dataset; apply different data enhancement methods to the Synapse dataset, including horizontal flip, vertical flip, Gaussian noise, blur, and random brightness contrast, increase the number of samples, and obtain a dataset named Data; 1-2) Data division: Divide the data set Data into training set D1 and test set D2.
[0007] Furthermore, the medical image segmentation model based on efficient dual attention includes an efficient dual attention module, namely a DualTransformer module, an encoder module, a decoder module, a multi-scale normalized channel attention module, namely an MNCA bottleneck module, and a feature fusion residual module, namely an FFRM module; The Dual Transformer module includes spatial attention and channel attention; In the Dual Transformer module, spatial attention is performed first, and then channel attention. By combining spatial attention and channel attention, local and global features are captured at the same time, and at the same time, residual connections and normalization operations are performed.
[0008] Furthermore, in spatial attention, the input feature matrix X has a dimension of n×d, where n is the sequence length and d is the feature dimension; the query matrix Q (Query), key matrix K (Key) and value matrix V (Value) are generated from the input feature matrix X, and the dimension of each matrix is n×di, where di is the dimension; the key matrix K is transposed to obtain K T , the dimension becomes dk*n; calculate V and K T The dot product of is obtained to obtain the attention weight matrix, and then the attention weight matrix is multiplied by Q to obtain the spatial attention output; the spatial attention output is residually connected with the input X, and then processed by normalization (Norm) and feed-forward network (FFN) to obtain the final spatial attention output; In channel attention, the input feature matrix X is n×d; the query matrix Q1, key matrix K1 and value matrix V1 are generated from the input X, each of which has a dimension of n×di, where di is the dimension; the key matrix K1 is transposed to obtain K1 T / β, where β is the scaling factor, calculate Q1 and K1 T The dot product of / β is used to obtain the attention weight matrix, and then the attention weight matrix is multiplied by V1 to obtain the channel attention output; the channel attention output is residually connected with the input X, and then processed by normalization (Norm) and feed-forward network (FFN) to obtain the final channel attention output.
[0009] Furthermore, the encoder module includes a Patch Partition layer, a Linear Embedding layer, a first downsampling module, a second downsampling module, and a third downsampling module; The first downsampling module, the second downsampling module and the third downsampling module each include two consecutive DualTransformer blocks and a Patch Merging layer; First, through the Patch Partition layer and the Linear Embedding layer, the Patch Partition layer divides the image into multiple non-overlapping image blocks (patches); the Linear Embedding layer maps the feature vector of each image block to an embedding vector of a fixed dimension through a linear transformation (fully connected layer), divides the input image into non-overlapping blocks of size 4×4, and maps the non-overlapping blocks to the new feature dimension C (W / 4 × H / 4 × C); Next, representation learning is performed through two consecutive Dual Transformer blocks; Subsequently, a 2x downsampling Patch Merging layer is performed to reduce the resolution while increasing the feature dimension. This process is repeated three times in the encoder to gradually refine the feature representation of the image.
[0010] According to the preferred embodiment of the present invention, the process performed in the MNCA bottleneck module is as follows: First, a deep separable convolution module (DASConv) with expansion ratios of 1, 3, 5, and 7 is used to expand the receptive field and extract rich multi-scale feature information. DASConv is a depth-wise separable convolution, including depth-wise convolution and point-wise convolution. Depth-wise convolution performs convolution operation on each input channel separately, while point-wise convolution performs linear combination between channels on the result of depth-wise convolution. In addition, DASConv is fused with the residual channel Fin to extract the features of objects of various sizes; Secondly, we further adopt normalized channel attention to enhance the correlation between feature channels; we use the BN function to batch normalize DASout to suppress unimportant weights, and then multiply the channel weight Wr and NAMout to force the dependency between feature channels; Finally, the final output of the MNCA bottleneck module is obtained after the Sigmoid activation function.
[0011] Preferably, according to the present invention, the decoder module comprises, from top to bottom, a first upsampling module, a second upsampling module and a third upsampling module, each of which comprises two consecutive Dual Transformer blocks and a Patch Expanding layer, i.e., a block expansion layer; The Patch Expanding layer upsamples the extracted deep features by a factor of 2 and reshapes the adjacent dimension feature maps into higher resolution feature maps; accordingly, the feature dimension is reduced to half of the original dimension, gradually restoring the feature map to the original image size; The encoder module and the decoder module are connected through the FFRM module.
[0012] Preferably, according to the present invention, the FFRM module includes a deep feature mapping branch m and a shallow feature mapping branch n, which perform the same feature processing operation in parallel. Before fusion, the FFRM module performs an inverse residual operation on the low and high-level feature maps respectively, and then sums them to finally obtain the output feature map m+n.
[0013] Furthermore, in step 3), the training set is input into a medical image segmentation model based on efficient dual attention, and a feature map is output; the steps include: 3-1) The i-th preprocessed image data in the training set D1 is referred to as The image is input into the medical image segmentation model based on efficient dual attention, and downsampled through the Patch Partition layer and the Liear Enbedding layer to reduce the image size to H / 4*W / 4*C, where H and W are the height and width of the original image, and C is the number of channels. The feature map is output. ; 3-2) Input to the first downsampling module; First, two Dual Transformer blocks are used for representation learning, and the output is the feature map ; Then, by performing a 2x downsampling Patch Merging layer, the feature dimension is increased while the resolution is reduced. The image size is reduced from W / 4*H / 4*C to W / 8*H / 8*2C, and the feature map is output. ; 3-3) Input to the second downsampling module; First, two Dual Transformer blocks are used for representation learning, and the output is the feature map ; Then, by performing a 2x downsampling Patch Merging layer, the feature dimension is increased while the resolution is reduced. The image size is reduced from W / 8*H / 8*2C to W / 16*H / 16*4C, and the feature map is output. ; 3-4) Input to the third down-sampling module; First, two Dual Transformer blocks are used for representation learning, and the output is the feature map ; Then, by performing a 2x downsampling Patch Merging layer, the feature dimension is increased while the resolution is reduced. The image size is reduced from W / 16*H / 16*4C to W / 32*H / 32*8C, and the feature map is output. ; 3-5) Input to the MNCA bottleneck module, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-6) Input to the third upsampling module; First, the extracted deep features are upsampled by 2 times through the Patch Expanding layer, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 32*W / 32*8C to H / 16*W / 16*4C, and the feature map is output. ; Secondly, After the FFRM module, the Add to reduce the spatial information loss caused by downsampling, and output the feature map ;at last, After two Dual Transformer blocks for representation learning, during this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-7) Input to the second upsampling module; First, the extracted deep features are upsampled by 2 times through the Patch Expanding layer, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 16*W / 16*4C to H / 8*W / 8*2C, and the feature map is output. ; Secondly, After the FFRM module, the Add to reduce the spatial information loss caused by downsampling, and output the feature map ;at last, After two Dual Transformer blocks for representation learning, during this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-8) Input to the first upsampling module; First, the extracted deep features are upsampled by 2 times through the Patch Expanding layer, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 8*W / 8*2C to H / 4*W / 4*C, and the feature map is output. ; Secondly, After the FFRM module, the (Italics) Addition reduces the spatial information loss caused by downsampling, and the output is the feature map ;at last, After two Dual Transformer blocks for representation learning, during this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-9) The extracted deep features are input to the Patch Expanding layer for 2x upsampling, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 4*W / 4*C to H / *W / *C, and the feature map is output. ; 3-10) Input to the linear projection layer (Linear Projection) for pixel-level segmentation prediction. The feature map is converted to H*W*CLASS, where CLASS is the number of categories, and the final output is obtained and named .
[0014] Preferably, according to the present invention, a comprehensive loss function method is used in training, which integrates the Dice loss function and the cross entropy loss function; the Dice loss function alleviates the problem of data imbalance in binary classification; the Dice loss function The formula is: = ; Among them, |M| and |N| represent the area of the segmentation result and the label, |M∩N| represents the area of the overlapping part of the segmentation result and the label, and the cross entropy loss function optimizes pixel-level classification. The formula is: ; Among them, y and p are the true value and predicted value when y = 0|1, p∈(0,1) respectively; The total loss function of the medical image segmentation model based on efficient dual attention It is expressed as: ; in, is the balance coefficient, by adjusting the parameter To optimize the trade-off between precision and recall.
[0015] Furthermore, step 5) includes the following steps: The test set D2 is input into the optimized medical image segmentation model based on efficient dual attention; the final medical image segmentation result graph is obtained and named .
[0016] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a medical image segmentation method based on efficient dual attention when executing the computer program.
[0017] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a medical image segmentation method based on efficient dual attention.
[0018] Medical image segmentation system based on efficient dual attention, including: The medical image data set acquisition module is configured to: perform data preprocessing and divide the obtained data set into a training set and a test set; The medical image segmentation model building module is configured to: build a medical image segmentation model based on efficient dual attention; The medical image segmentation model training module is configured to: input the training set into the medical image segmentation model based on efficient dual attention, and output a feature map; train the medical image segmentation model based on efficient dual attention, and obtain an optimized medical image segmentation model based on efficient dual attention; The medical image segmentation module is configured to: input the test set into the optimized medical image segmentation model based on efficient dual attention, and output the final medical image segmentation result.
[0019] The beneficial effects of the present invention are: 1) The present invention designs the FFRM module as a skip connection in the network to more accurately locate the target in a complex background and effectively fuse the semantic features between high-level and low-level layers, thereby extracting more significant features and preserving the semantic spatial information.
[0020] 2) The present invention designs the MNCA bottleneck module, which combines attribute convolution, normalized channel attention mechanism and depthwise separable convolution (DSConv) to enhance the inter-channel interdependence, and it fully utilizes the complementarity of global area and local edge information to maintain the global context information.
[0021] 3) This paper proposes a new symmetrical dual-branch encoder-decoder network architecture. The network model is built with the Dual Transformer block as the basic unit. By combining spatial attention and channel attention, it can capture both local and global features, thereby improving the feature extraction ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of the structure of the medical image segmentation model based on efficient dual attention; Figure 2 It is a schematic diagram of the structure of the Dual Transformer module; Figure 3 It is a schematic diagram of the structure of the MNCA bottleneck module; Figure 4 It is a structural diagram of the FFRM module; Figure 5 It is a flowchart of the medical image segmentation method based on efficient dual attention of the present invention; Figure 6 The figure is a schematic diagram of visualization of the results after being processed by the method of the present invention and the existing model. DETAILED DESCRIPTION
[0023] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.
[0024] Example 1 Medical image segmentation method based on efficient dual attention, that is, a medical image segmentation method based on efficient dual attention, such as Figure 5 As shown, including: 1) Obtain a medical image dataset, perform data preprocessing, and divide the obtained dataset into a training set and a test set; 2) Construct a medical image segmentation model based on efficient dual attention; 3) Input the training set into the medical image segmentation model based on efficient dual attention and output the feature map; 4) Training a medical image segmentation model based on efficient dual attention to obtain an optimized medical image segmentation model based on efficient dual attention; 5) The test set is input into the optimized medical image segmentation model based on efficient dual attention, and the final medical image segmentation result is output.
[0025] Example 2 The difference between the medical image segmentation method based on efficient dual attention described in Example 1 is that: Step 1) includes the following steps: 1-1) Data collection and preprocessing: The publicly available Synapse dataset was used; different data enhancement methods were applied to the Synapse dataset, including horizontal flipping, vertical flipping, Gaussian noise, blurring, and random brightness contrast, to increase the number of samples, and the dataset was obtained and named Data; The Synapse dataset, namely the Synapse multi-organ segmentation dataset, contains 3779 axial abdominal clinical CT images of 30 patients. The dataset contains 8 abdominal organs (aorta, gallbladder, left kidney, right kidney, liver, pancreas, spleen, and stomach). The present invention adjusts the resolution of all images to 512×512.
[0026] 1-2) Data division: Divide the data set Data into training set D1 and test set D2 in a ratio of 9:1.
[0027] The medical image segmentation model based on efficient dual attention includes an efficient dual attention module, namely a Dual Transformer module, an encoder module, a decoder module, a multi-scale normalized channel attention module, namely an MNCA bottleneck module, and a feature fusion residual module, namely an FFRM module; Figure 1 As shown; like Figure 2 As shown, the Dual Transformer module includes spatial attention and channel attention; In the Dual Transformer module, spatial attention is performed first, and then channel attention. By combining spatial attention and channel attention, the dual attention block can capture local features and global features at the same time, thereby improving the model's feature extraction ability. At the same time, through residual connection and normalization operations, the model can be trained more stably.
[0028] Figure 2 (a) in the figure is spatial attention. In spatial attention, the dimension of the input feature matrix X is n×d, where n is the sequence length and d is the feature dimension. The query matrix Q (Query), key matrix K (Key), and value matrix V (Value) are generated from the input feature matrix X. The dimension of each matrix is n×di, where di is the dimension. The key matrix K is transposed to obtain K T , the dimension becomes dk*n; calculate V and K TThe dot product of is obtained to obtain the attention weight matrix, and then the attention weight matrix is multiplied by Q to obtain the spatial attention output; the spatial attention output is residually connected with the input X, and then processed by normalization (Norm) and feedforward network (FFN) to obtain the final spatial attention output; the feedforward network, FFN, is a key component in the Transformer architecture. It receives the feature vector output by the self-attention mechanism, performs nonlinear transformation and feature extraction, and enhances the model's expression and learning capabilities. In the Dual Transformer module, the FFN input and output dimensions are the same, and feature interaction and mapping are achieved internally through two layers of linear transformation and activation functions, helping the model to better understand the data semantics and contextual relationships.
[0029] Figure 2 (b) in the figure is channel attention. In channel attention, the input feature matrix X has a dimension of n×d. The query matrix Q1, key matrix K1 and value matrix V1 are generated from the input X. The dimension of each matrix is n×di, where di is the dimension. The key matrix K1 is transposed to obtain K1. T / β, where β is a scaling factor used to stabilize training; calculate Q1 and K1 T The dot product of / β is used to obtain the attention weight matrix, and then the attention weight matrix is multiplied by V1 to obtain the channel attention output; the channel attention output is residually connected with the input X, and then processed by normalization (Norm) and feed-forward network (FFN) to obtain the final channel attention output.
[0030] The decoder module includes a first upsampling module, a second upsampling module and a third upsampling module from top to bottom, each of which includes two consecutive Dual Transformer blocks and a Patch Expanding layer, i.e., a block expansion layer; The first downsampling module, the second downsampling module and the third downsampling module each include two consecutive DualTransformer blocks and a Patch Merging layer; First, through the Patch Partition layer and the Linear Embedding layer, the Patch Partition layer divides the image into multiple non-overlapping image blocks (patches); the Linear Embedding layer maps the feature vector of each image block to an embedding vector of a fixed dimension through a linear transformation (fully connected layer), divides the input image into non-overlapping blocks of size 4×4, and maps the non-overlapping blocks to the new feature dimension C (W / 4 × H / 4 × C); Next, representation learning is performed through two consecutive Dual Transformer blocks; at this time, two consecutive Dual Transformer blocks are used in the encoder for representation learning. The feature matrix processed by the first Dual Transformer block is still in the shape of N×d, but the feature representation is richer and more abstract. The feature matrix processed by the second Dual Transformer block is still in the shape of N×d, but the feature representation is deeper and more abstract, providing higher quality feature input for the subsequent PatchMerging or Patch Expanding layer.
[0031] The internal processing of each Dual Transformer block: The input is the embedded feature matrix from the previous layer, with a shape of N×d, where N is the number of image blocks and d is the embedding dimension. The processing process is as follows: self-attention is calculated on the input features to capture the global dependencies between features. Each attention head is calculated independently, and then the results are concatenated and linearly transformed to obtain a new feature representation; the output of the self-attention mechanism is residually connected to the input features and normalized through a normalization layer to enhance the stability and convergence speed of the model; the normalized features are nonlinearly transformed to further extract and enhance features. FFN usually consists of two linear layers and an activation function to increase the expressive power of the model; the output of FFN is residually connected to the input features and normalized again through a normalization layer to obtain the final feature representation. After the feature matrix processed by the Dual Transformer block, the shape is still N×d, but the feature representation is richer and more abstract. In this process, the feature dimension and resolution remain unchanged.
[0032] Subsequently, a 2x downsampling Patch Merging layer is performed to reduce the resolution while increasing the feature dimension. This process is repeated three times in the encoder to gradually refine the feature representation of the image.
[0033] The processing of the Patch Merging layer is as follows: The input is the feature matrix output by the Dual Transformer block, with a shape of N×d. Processing process: Group multiple adjacent image blocks; fuse the image block features in each group, usually through linear transformation and dimensionality reduction operations, to merge the features of multiple blocks into a feature vector; through the above fusion operation, the resolution of the feature map is reduced. The feature matrix after the PatchMerging layer is processed, with a shape of N'×2d, where N' is the number of merged image blocks, and the feature dimension is increased to 2d.
[0034] In the MNCA bottleneck module, the dimension and resolution of the feature map remain consistent. The MNCA bottleneck module is divided into two main steps. They include:
[0035] First, if Figure 3 (a) shows a multi-residual multi-scale strategy, which uses a deep separable convolution module (DASConv) with expansion ratios of 1, 3, 5, and 7 to expand the receptive field and extract rich multi-scale feature information. DASConv is a depth-wise separable convolution, including depth-wise convolution and point-wise convolution. Depth-wise convolution performs convolution operation on each input channel separately, while point-wise convolution performs linear combination between channels on the result of depth-wise convolution. Figure 3 (a) shows four DASConvs with different dilation rates: 1, 3, 5, and 7. The dilation rate determines the sampling interval of the convolution kernel on the feature map. The larger the dilation rate, the wider the actual coverage of the convolution kernel and the larger the receptive field.
[0036] Dilation rate 1: The convolution kernel is sampled at an interval of 1 on the feature map, covering a smaller local area, which is suitable for capturing fine-grained features.
[0037] Dilation rate 3: The convolution kernel is sampled at an interval of 3 on the feature map, covering a larger area and suitable for capturing medium-scale features.
[0038] Dilation rate 5: The convolution kernel is sampled at an interval of 5 on the feature map, covering a larger area and suitable for capturing features of larger scales.
[0039] Dilation rate 7: The convolution kernel is sampled at an interval of 7 on the feature map, covering the largest area and suitable for capturing global features.
[0040] The feature maps output by four DASConv modules with different expansion rates are concatenated to obtain a feature map that integrates multi-scale features.
[0041] In addition, DASConv is fused with the residual channel Fin to extract the features of objects of various sizes; the fused feature map is fused with the input feature map Fin through a residual connection to enhance the feature expression ability and the convergence performance of the network.
[0042] Secondly, if Figure 3 As shown in (b), the normalized channel attention is further adopted to enhance the correlation between feature channels; the BN function is used to batch normalize DASout to suppress unimportant weights, so that NAMout contains more important feature weight information. Then, the channel weight Wr is multiplied by NAMout to force the dependency between feature channels;
[0043] Finally, the final output of the MNCA bottleneck module is obtained after the Sigmoid activation function.
[0044] The decoder module includes a first upsampling module, a second upsampling module and a third upsampling module from top to bottom, each of which includes two consecutive Dual Transformer blocks and a Patch Expanding layer, i.e., a block expansion layer; Inspired by U-Net and Swin-Unet, a decoder based on a symmetric encoder is designed. The decoder structure is similar to the encoder and is also built based on two consecutive Dual Transformer blocks. The Patch Expanding layer upsamples the extracted deep features by 2 times and reshapes the adjacent dimensional feature maps into higher resolution feature maps; accordingly, the feature dimension is reduced to half of the original dimension, gradually restoring the feature map to the original image size;
[0045] The encoder module and the decoder module are connected by the FFRM module. The skip connection fuses the encoder features with the deep features recovered from upsampling, thereby alleviating the loss of spatial data caused by downsampling.
[0046] like Figure 4 As shown in the figure, the FFRM module includes a deep feature map branch m and a shallow feature map branch n, which perform the same feature processing operations in parallel. Before fusion, the FFRM module performs an inverse residual operation on the low and high-level feature maps respectively, and then sums them up to finally obtain the output feature map m+n. It also retains more semantic information and restores more spatial details.
[0047] The FFRM module aims to fuse the deep feature map branch m and the shallow feature map branch n to achieve feature fusion and refinement. The following is the specific structure of the two branches and the fusion implementation process:
[0048] Deep feature map branch m: 1×1 convolution: The input deep feature map m is adjusted for channel number and feature transformation, using a 1×1 convolution kernel followed by a ReLU6 activation function to reduce the amount of computation and maintain feature richness.
[0049] 3×3 Depthwise Separable Convolution (Dwise 3×3): Further extracts local features and enhances feature expression capabilities. It uses a 3×3 depthwise separable convolution kernel followed by a ReLU6 activation function to reduce the amount of computation while maintaining the local perception of features.
[0050] 1×1 convolution: The features are adjusted again in terms of number of channels and linearly transformed, using a 1×1 convolution kernel without using an activation function to maintain the linear combination capability of the features.
[0051] Shallow feature map branch (n): 1×1 convolution: The input shallow feature map n is adjusted for channel number and feature transformation, using a 1×1 convolution kernel followed by a ReLU6 activation function to reduce the amount of computation and maintain feature richness.
[0052] 3×3 Depthwise Separable Convolution (Dwise 3×3): Further extracts local features and enhances feature expression capabilities. It uses a 3×3 depthwise separable convolution kernel followed by a ReLU6 activation function to reduce the amount of computation while maintaining the local perception of features.
[0053] 1×1 convolution: The features are adjusted again in terms of number of channels and linearly transformed, using a 1×1 convolution kernel without using an activation function to maintain the linear combination capability of the features.
[0054] The two branches perform the same feature processing operation in parallel. The processed high-level features are directly added to the low-level features. The high-level features provide semantic information (such as the object outline), and the low-level features supplement the details (such as texture). The two complement each other. The inverse residual structure enhances the diversity and nonlinear expression ability of features through the channel expansion-compression strategy.
[0055] In step 3), the training set is input into the medical image segmentation model based on efficient dual attention, and the feature map is output; the steps include: 3-1) The i-th preprocessed image data in the training set D1 is referred to as The image is input into the medical image segmentation model based on efficient dual attention, and downsampled through the Patch Partition layer and the Liear Enbedding layer to reduce the image size to H / 4*W / 4*C, where H and W are the height and width of the original image, and C is the number of channels. The feature map is output. ; 3-2) Input to the first downsampling module; First, two Dual Transformer blocks are used for representation learning. During this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; Then, by performing a 2x downsampling Patch Merging layer, the feature dimension is increased while the resolution is reduced. The image size is reduced from W / 4*H / 4*C to W / 8*H / 8*2C, and the feature map is output. ; 3-3) Input to the second downsampling module; First, two Dual Transformer blocks are used for representation learning. During this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; Then, by performing a 2x downsampling Patch Merging layer, the feature dimension is increased while the resolution is reduced. The image size is reduced from W / 8*H / 8*2C to W / 16*H / 16*4C, and the feature map is output. ; 3-4) Input to the third down-sampling module; First, two Dual Transformer blocks are used for representation learning. During this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; Then, by performing a 2x downsampling Patch Merging layer, the feature dimension is increased while the resolution is reduced. The image size is reduced from W / 16*H / 16*4C to W / 32*H / 32*8C, and the feature map is output. ; 3-5) Input to the MNCA bottleneck module, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-6) Input to the third upsampling module; First, the extracted deep features are upsampled by 2 times through the Patch Expanding layer, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 32*W / 32*8C to H / 16*W / 16*4C, and the feature map is output. ; Secondly, After the FFRM module, the Add to reduce the spatial information loss caused by downsampling, and output the feature map ;at last, After two Dual Transformer blocks for representation learning, during this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-7) Input to the second upsampling module; First, the extracted deep features are upsampled by 2 times through the Patch Expanding layer, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 16*W / 16*4C to H / 8*W / 8*2C, and the feature map is output. ; Secondly, After the FFRM module, the Add to reduce the spatial information loss caused by downsampling, and output the feature map ;at last, After two Dual Transformer blocks for representation learning, during this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-8) Input to the first upsampling module; First, the extracted deep features are upsampled by 2 times through the Patch Expanding layer, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 8*W / 8*2C to H / 4*W / 4*C, and the feature map is output. ; Secondly, After the FFRM module, the Add to reduce the spatial information loss caused by downsampling, and output the feature map ;at last, After two Dual Transformer blocks for representation learning, during this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-9) The extracted deep features are input to the Patch Expanding layer for 2x upsampling, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 4*W / 4*C to H / *W / *C, and the feature map is output. ; 3-10) Input to the linear projection layer (Linear Projection) for pixel-level segmentation prediction. The feature map is converted to H*W*CLASS, where CLASS is the number of categories, and the final output is obtained and named .
[0056] Due to the limited size of medical image datasets and the fact that the area to be segmented in the image only occupies a small part of the entire image, these problems may cause the model to overfit during training. A comprehensive loss function method is used in training, which combines the Dice loss function and the cross entropy loss function; the Dice loss function alleviates the problem of data imbalance in binary classification; the Dice loss function The formula is:
[0057] = ; Among them, |M| and |N| represent the area of the segmentation result and the label, and |M∩N| represents the area of the overlapping part of the segmentation result and the label. The closer the intersection is to 1, the more accurate the prediction result is. The cross entropy loss function optimizes pixel-level classification. The cross entropy loss function The formula is:
[0058] ; Among them, y and p are the true value and predicted value when y = 0|1, p∈(0,1) respectively; In order to fully utilize the complementary advantages of the Dice loss function and the cross entropy loss, the two loss functions are linearly combined with the coefficient Balanced with the value of the loss function. The total loss function of the medical image segmentation model based on efficient dual attention It is expressed as:
[0059] ; in, is the balance coefficient, by adjusting the parameter To optimize the trade-off between precision and recall.
[0060] In the training process of the efficient dual-attention medical image segmentation model, the programming language of the efficient dual-attention medical image segmentation model is Python 3.9, the deep learning framework is PyTorch, and it is trained on four NVIDIA A100 GPUs with 40GB of memory. The training configuration includes a batch size of 24, an SGD optimizer momentum of 0.9, a basic learning rate of 0.1, a weight decay of 0.0001, and the use of cross entropy loss and Dice loss ( = 0.5 + 0.5 ) was trained for 400 rounds. In order to evaluate the performance of the proposed medical image segmentation model based on efficient dual attention, the average probability similarity coefficient (DSC) and the average Hausdorff distance (HD) were used as evaluation indicators. The DSC value ranges from 0 to 1, and the larger the value, the better the performance, and the smaller the HD value, the better the performance. Finally, the optimized medical image segmentation model based on efficient dual attention was obtained.
[0061] Step 5) includes the following steps: The test set D2 is input into the optimized medical image segmentation model based on efficient dual attention; the final medical image segmentation result graph is obtained and named .
[0062] Figure 6It is a visualization diagram of the results after processing by the method of the present invention and the existing model; by comparison, it can be seen that the segmentation performance of the model of the present invention is improved. On the Synapse dataset, especially for the gallbladder, kidney, liver and spleen, the overall segmentation map is smoother and more natural.
[0063] Example 3 A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the medical image segmentation method based on efficient dual attention described in embodiment 1 or 2 are implemented.
[0064] Example 4 A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the medical image segmentation method based on efficient dual attention described in embodiment 1 or 2 are implemented.
[0065] Example 5 Medical image segmentation system based on efficient dual attention, including: The medical image data set acquisition module is configured to: perform data preprocessing and divide the obtained data set into a training set and a test set; The medical image segmentation model building module is configured to: build a medical image segmentation model based on efficient dual attention; The medical image segmentation model training module is configured to: input the training set into the medical image segmentation model based on efficient dual attention, and output a feature map; train the medical image segmentation model based on efficient dual attention, and obtain an optimized medical image segmentation model based on efficient dual attention; The medical image segmentation module is configured to: input the test set into the optimized medical image segmentation model based on efficient dual attention, and output the final medical image segmentation result.
[0066] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A medical image segmentation method based on efficient dual attention, characterized in that: include: 1) Obtain a medical image dataset, perform data preprocessing, and divide the obtained dataset into a training set and a test set; 2) Construct a medical image segmentation model based on efficient dual attention; 3) Input the training set into the medical image segmentation model based on efficient dual attention and output the feature map; 4) Training a medical image segmentation model based on efficient dual attention to obtain an optimized medical image segmentation model based on efficient dual attention; 5) The test set is input into the optimized medical image segmentation model based on efficient dual attention, and the final medical image segmentation result is output.
2. The medical image segmentation method based on efficient dual attention according to claim 1, characterized in that: Step 1) includes the following steps: 1-1) Data collection and preprocessing: Use the Synapse dataset; apply different data enhancement methods to the Synapse dataset, including horizontal flip, vertical flip, Gaussian noise, blur, and random brightness contrast, increase the number of samples, and obtain a dataset named Data; 1-2) Data division: Divide the data set Data into training set D1 and test set D2.
3. The medical image segmentation method based on efficient dual attention according to claim 1, characterized in that: The medical image segmentation model based on efficient dual attention includes an efficient dual attention module, namely a Dual Transformer module, an encoder module, a decoder module, a multi-scale normalized channel attention module, namely an MNCA bottleneck module, and a feature fusion residual module, namely an FFRM module; The Dual Transformer module includes spatial attention and channel attention; In the Dual Transformer module, spatial attention is performed first, and then channel attention. By combining spatial attention and channel attention, local and global features are captured at the same time, and at the same time, residual connections and normalization operations are performed.
4. The medical image segmentation method based on efficient dual attention according to claim 1, characterized in that: In spatial attention, the input feature matrix X is of dimension n×d, where n is the sequence length and d is the feature dimension; the query matrix Q, key matrix K, and value matrix V are generated from the input feature matrix X, and the dimension of each matrix is n×di, where di is the dimension; Transpose the key matrix K to get K T , the dimension becomes dk*n; calculate V and K T The dot product of is used to obtain the attention weight matrix, and then the attention weight matrix is multiplied by Q to obtain the spatial attention output; the spatial attention output is residually connected with the input X, and then processed by normalization and feedforward network to obtain the final spatial attention output; In channel attention, the input feature matrix X is of dimension n×d; the query matrix Q1, key matrix K1 and value matrix V1 are generated from the input X, and the dimension of each matrix is n×di, where di is the dimension; Transpose the key matrix K1 to obtain K1 T / β, where β is the scaling factor, calculate Q1 and K1 T / β to get the attention weight matrix, and then multiply the attention weight matrix with V1 to get the channel attention output; the channel attention output is residually connected with the input X, and then processed by normalization and feedforward network to get the final channel attention output; The encoder module includes a Patch Partition layer, a Linear Embedding layer, a first downsampling module, a second downsampling module, and a third downsampling module; The first downsampling module, the second downsampling module and the third downsampling module each include two consecutive DualTransformer blocks and a Patch Merging layer; First, through the Patch Partition layer and the Linear Embedding layer, the Patch Partition layer divides the image into multiple non-overlapping image blocks; the Linear Embedding layer maps the feature vector of each image block to an embedding vector of a fixed dimension through a linear transformation, divides the input image into non-overlapping blocks of size 4 × 4, and maps the non-overlapping blocks to the new feature dimension C; Next, representation learning is performed through two consecutive Dual Transformer blocks; Subsequently, a 2x downsampling Patch Merging layer is performed to reduce the resolution while increasing the feature dimension. This process is repeated three times in the encoder to gradually refine the feature representation of the image.
5. The medical image segmentation method based on efficient dual attention according to claim 1, characterized in that: The execution process in the MNCA bottleneck module is as follows: First, a deep separable convolution module (DASConv) with expansion ratios of 1, 3, 5, and 7 is used to expand the receptive field and extract rich multi-scale feature information. DASConv is a depth-wise separable convolution, including depth-wise convolution and point-wise convolution. Depth-wise convolution performs convolution operation on each input channel separately, while point-wise convolution performs linear combination between channels on the result of depth-wise convolution. In addition, DASConv is fused with the residual channel Fin to extract the features of objects of various sizes; Secondly, we further adopt normalized channel attention to enhance the correlation between feature channels; we use the BN function to batch normalize DASout to suppress unimportant weights, and then multiply the channel weight Wr and NAMout to force the dependency between feature channels; Finally, the final output of the MNCA bottleneck module is obtained after the Sigmoid activation function.
6. The medical image segmentation method based on efficient dual attention according to claim 1, characterized in that: The decoder module includes a first upsampling module, a second upsampling module and a third upsampling module from top to bottom, each of which includes two consecutive Dual Transformer blocks and a Patch Expanding layer, i.e., a block expansion layer; The Patch Expanding layer upsamples the extracted deep features by a factor of 2 and reshapes the adjacent dimension feature maps into higher resolution feature maps; accordingly, the feature dimension is reduced to half of the original dimension, gradually restoring the feature map to the original image size; The encoder module and the decoder module are connected through the FFRM module; The FFRM module includes a deep feature mapping branch m and a shallow feature mapping branch n, which perform the same feature processing operations in parallel. Before fusion, the FFRM module performs an inverse residual operation on the low and high-level feature maps respectively, and then sums them up to finally obtain the output feature map m+n.
7. The medical image segmentation method based on efficient dual attention according to claim 1, characterized in that: In step 3), the training set is input into the medical image segmentation model based on efficient dual attention, and the feature map is output; the steps include: 3-1) The i-th preprocessed image data in the training set D1 is referred to as The image is input into the medical image segmentation model based on efficient dual attention, and downsampled through the Patch Partition layer and the Liear Enbedding layer to reduce the image size to H / 4*W / 4*C, where H and W are the height and width of the original image, and C is the number of channels. The feature map is output. ; 3-2) Input to the first downsampling module; First, two Dual Transformer blocks are used for representation learning, and the output is the feature map ; Then, by performing a 2x downsampling Patch Merging layer, the feature dimension is increased while the resolution is reduced. The image size is reduced from W / 4*H / 4*C to W / 8*H / 8*2C, and the feature map is output. ; 3-3) Input to the second downsampling module; First, two Dual Transformer blocks are used for representation learning, and the output is the feature map ; Then, by performing a 2x downsampling Patch Merging layer, the feature dimension is increased while the resolution is reduced. The image size is reduced from W / 8*H / 8*2C to W / 16*H / 16*4C, and the feature map is output. ; 3-4) Input to the third downsampling module; First, two Dual Transformer blocks are used for representation learning, and the output is the feature map ; Then, by performing a 2x downsampling Patch Merging layer, the feature dimension is increased while the resolution is reduced. The image size is reduced from W / 16*H / 16*4C to W / 32*H / 32*8C, and the feature map is output. ; 3-5) Input to the MNCA bottleneck module, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-6) Input to the third upsampling module; First, the extracted deep features are upsampled by 2 times through the Patch Expanding layer, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 32*W / 32*8C to H / 16*W / 16*4C, and the feature map is output. ; Secondly, After the FFRM module, the Add to reduce the spatial information loss caused by downsampling, and output the feature map ;at last, After two Dual Transformer blocks for representation learning, during this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-7) Input to the second upsampling module; First, the extracted deep features are upsampled by 2 times through the Patch Expanding layer, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 16*W / 16*4C to H / 8*W / 8*2C, and the feature map is output. ; Secondly, After the FFRM module, the Add to reduce the spatial information loss caused by downsampling, and output the feature map ;at last, After two Dual Transformer blocks for representation learning, during this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-8) Input to the first upsampling module; First, the extracted deep features are upsampled by 2 times through the Patch Expanding layer, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 8*W / 8*2C to H / 4*W / 4*C, and the feature map is output. ; Secondly, After the FFRM module, the Add to reduce the spatial information loss caused by downsampling, and output the feature map ;at last, After two Dual Transformer blocks for representation learning, during this process, the feature dimension and resolution remain unchanged, and the output is the feature map ; 3-9) The extracted deep features are input to the Patch Expanding layer for 2x upsampling, and the adjacent dimension feature maps are reshaped into higher resolution feature maps. The image size is expanded from H / 4*W / 4*C to H / *W / *C, and the feature map is output. ; 3-10) Input to the linear projection layer for pixel-level segmentation prediction; transform the feature map into H*W*CLASS, where CLASS is the number of categories, and get the final output and name it .
8. The medical image segmentation method based on efficient dual attention according to any one of claims 1 to 6, characterized in that: A comprehensive loss function method is used in training, which combines the Dice loss function and the cross entropy loss function; the Dice loss function alleviates the problem of data imbalance in binary classification; the Dice loss function The formula is: = ; Among them, |M| and |N| represent the area of the segmentation result and the label, |M∩N| represents the area of the overlapping part of the segmentation result and the label, and the cross entropy loss function optimizes pixel-level classification. The formula is: ; Among them, y and p are the true value and predicted value when y = 0|1, p∈(0,1) respectively; The total loss function of the medical image segmentation model based on efficient dual attention It is expressed as: ; in, is the balance coefficient, by adjusting the parameter To optimize the trade-off between precision and recall.
9. The medical image segmentation method based on efficient dual attention according to any one of claims 1 to 6, characterized in that: Step 5) includes the following steps: The test set D2 is input into the optimized medical image segmentation model based on efficient dual attention; the final medical image segmentation result graph is obtained and named .
10. A medical image segmentation system based on efficient dual attention, characterized in that, include: The medical image data set acquisition module is configured to: perform data preprocessing and divide the obtained data set into a training set and a test set; The medical image segmentation model building module is configured to: build a medical image segmentation model based on efficient dual attention; The medical image segmentation model training module is configured to: input the training set into the medical image segmentation model based on efficient dual attention, and output a feature map; Training a medical image segmentation model based on efficient dual attention to obtain an optimized medical image segmentation model based on efficient dual attention; The medical image segmentation module is configured to: input the test set into the optimized medical image segmentation model based on efficient dual attention, and output the final medical image segmentation result.
Citation Information
Patent Citations
Medical image segmentation model construction method based on multi-attention fusion
CN116309648A
Medical image segmentation method and system based on double-branch embedded attention mechanism
CN116309650A
Medical image segmentation method based on CNN and Transform fusion network
CN117173412A
Medical image multi-organ segmentation method based on global context interaction Transform
CN118570222A
Bone tumor medical image segmentation method based on improved mixed attention mechanism and TransUnet model
CN119579897A
Cited By
Three-dimensional medical image segmentation method and device based on deep learning and medium
CN120451195A
Three-dimensional medical image segmentation method, device and medium based on deep learning
CN120451195B
Medical image segmentation method and imaging method based on Mama network
CN120689296A
Medical image segmentation method based on mamba network and imaging method
CN120689296B