Medical image segmentation method based on multi-attention optimization network
By introducing multi-attention optimization network methods with EMA, CB and UFO modules based on TransUNet, the shortcomings of existing medical image segmentation models in capturing complex morphology and grayscale characteristics of COVID-19 lesions are solved, and higher segmentation accuracy and computing efficiency are achieved.
Patent Information
- Application Number
- CN202510144337.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-16
AI Technical Summary
Existing medical image segmentation models are difficult to effectively capture the complex morphological patterns of COVID-19 lesions and grayscale or texture features similar to healthy lung tissues, resulting in unsatisfactory segmentation results.
A medical image segmentation method based on multi-attention optimization network is proposed, called AO-TransUNet. By introducing multiple attention mechanisms based on TransUNet, including efficient multi-scale attention (EMA) module, context broadcast (CB) module and unit force operation (UFO) module, to enhance the model's attention to lesion morphology details and global context.
Through experimental evaluation, AO-TransUNet significantly improves the accuracy of medical image segmentation, especially when dealing with complex and variable COVID-19 lesions, it can more effectively retain the morphological details and characteristic information of the lesions, improve the recognition ability of the lesion boundaries, and reduce the computational complexity.
Smart Images

Figure CN120014412A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to technical fields such as medical image segmentation and deep learning, and in particular to a medical image segmentation method based on a multi-attention optimization network. Background Art
[0002] Deep learning methods are currently widely used in the field of COVID-19 lesion segmentation. These include U-Net, V-Net, and UNet++. Recently proposed segmentation models that combine U-Net and Transformer architectures include TransUNet and SwinUnet. TransUNet enhances the model's ability to model long-range dependencies by integrating a Transformer module. Seamlessly integrated with the U-shaped TransUNet architecture, the Transformer module extracts global information from medical images and enhances semantic representation capabilities. This effectively addresses the challenges posed by large-scale, high-resolution medical images.
[0003] In previous studies, various variants of U-networks have successfully addressed many challenges in the field of medical image segmentation, as well as many problems in COVID-19 lesion image segmentation, such as suppressing irrelevant feature responses and highlighting key features to improve segmentation accuracy, optimizing the mapping relationship between spatial information, etc. However, there are still some research gaps. First, medical imaging related to COVID-19 faces unique challenges due to the highly variable size, shape, and distribution of lesions, such as ground glass opacity and consolidation. Existing segmentation models often have difficulty effectively capturing these complex and diverse morphological patterns, resulting in suboptimal depictions of affected areas. Second, these lesions often have similar grayscale or texture features to healthy lung tissue, making it difficult to establish clear boundaries between lesions and healthy areas. Summary of the Invention
[0004] Considering medical images related to COVID-19 presents unique challenges. For example, lesions often appear in various forms (such as ground glass shadows and consolidation shadows) with great differences in size, shape, and distribution. In addition, these lesions may have similar grayscale or texture characteristics to normal lung tissue, making it difficult to delineate clear boundaries between affected and healthy areas.
[0005] This paper presents a medical image segmentation method based on a multi-attention optimization network, the Attention Optimized TransUNet (AO-TransUNet), which is built on the foundation of TransUNet. By integrating multiple attention mechanisms, the method aims to minimize the loss of critical information during the segmentation and dimensionality reduction phase. AO-TransUNet enhances the dense interactions between all pixels, ensuring the preservation of morphological details and characteristic information of lesions, improving the model's ability to detect subtle structural differences and effectively segment complex COVID-19 lesions.
[0006] The performance of the proposed AO-TransUNet was verified through experimental evaluation on the dataset. The results show that AO-TransUNet outperforms the existing state-of-the-art networks, demonstrating its effectiveness in medical image segmentation. The proposed solution has the potential for further promotion and application in the field of medical image segmentation by addressing the challenges of complex and variable lesions (such as COVID-19). The ability of this method to preserve morphological details and improve pixel-level interactions suggests that it has broader applicability to other medical image analysis challenges.
[0007] The proposed TransUNet-centric multi-attention optimization network, called AO-TransUNet, is specifically designed for medical image segmentation, particularly in the context of COVID-19. The proposed AO-TransUNet optimizes the attention in the encoder, decoder, and skip connections separately. The present invention uses the integrated weighted binary cross-entropy loss function (BCE) and the dice coefficient loss function (dice) as the objective function. In addition, the present invention utilizes a pre-trained model and designs its loss function by balancing the contributions of the cross-entropy loss and the dice loss. To address the loss of key information in the dimensionality reduction process caused by the complex and variable shapes of COVID-19 lesions and retain the morphological details and feature information of the lesions, the present invention proposes to use efficient multi-scale attention (EMA) for AO-TransUNet. This method involves shaping the channel subset into a batch dimension and grouping the channel dimension into multiple sub-features. This helps to evenly distribute spatial semantic features in each feature group, thereby achieving multi-scale attention. Therefore, this method helps to retain detailed information within the lesion area, thereby improving the accuracy of the segmentation model. Furthermore, while studying the grayscale or texture features that resemble COVID-19 lesions and normal lung tissue in images, recognizing the complexity of dense attention maps and their importance in ViTs, the present invention introduces uniform attention via the context broadcast (CB) module in AO-TransUNet. This distributed form of attention is achieved through average pooling of all tokens, supplementing each individual token in each intermediate layer. This inclusion improves dense interactions and overall attention performance, enhancing the model's ability to understand dense interactions between all pixels in the image and helping it identify weak or implicit tissue differences at the boundaries of COVID-19 lesions. To address computational complexity, particularly in the field of medical image segmentation, which deals with extensive and complex datasets such as COVID-19 images, the Unit Force Operation (UFO) module is employed in AO-TransUNet. This module reduces computational complexity from quadratic to linear by decomposing the self-attention operation into force and operator sub-operations. This not only improves the model's inference speed but also ensures practicality without compromising performance. Rigorous experiments on various datasets, including the COVID-19 image set, demonstrate the state-of-the-art performance of the proposed model. The research results highlight the effectiveness of the method of the present invention and its significant contribution to the advancement of the field of medical image segmentation. In summary, the main contributions of the present invention are as follows: 1) For COVID-19 image segmentation, AO-TransUNe is proposed to optimize the attention-related components in the encoder-decoder and skip connections respectively.
[0008] 2) The complex and variable shapes of COVID-19 lesions can lead to loss of critical information during dimensionality reduction during segmentation. To maximize the preservation of morphological details and feature information of the lesions, EMA was introduced. This method ensures a fair distribution of spatial semantic features across each feature group, facilitating multi-scale attention. This effectively preserves detailed information about the lesion region, thereby improving the overall accuracy of the lesion segmentation model.
[0009] 3) COVID-19 lesions often appear gray or textured in images similar to normal lung tissue, making it difficult to accurately define the boundary between lesions and normal tissue. Therefore, this paper adds a CB module to the MLP layer of the Transformer to enhance the model's focus on dense spatial interactions. This helps the model understand the dense interactions between all pixels in the image and facilitates the identification of weak or hidden tissue structural differences in the boundaries of COVID-19 lesions.
[0010] 4) To address computational complexity, especially in the context of large-scale datasets such as medical image segmentation and COVID-19 images, our model utilizes the UFO module. Replacing the traditional multi-head self-attention in the Transformer layer with the UFO module improves inference speed while maintaining model performance, making the model more suitable for practical applications.
[0011] The specific technical solutions are as follows: A medical image segmentation method based on a multi-attention optimization network: medical image segmentation is performed through an attention-optimized TransUNet; the attention-optimized TransUNet is based on TransUNet: the SA in the Transformer is replaced by the unit force operation UFO module; the context broadcast CB module is introduced into the multi-layer perceptron MLP of the Transformer; in the skip connection of the U-Net, the feature map of the encoder is processed by the multi-scale attention EMA module and then channel-wise spliced with the upsampled features of the decoder. In the decoder, the EMA module is located after each upsampling layer to enhance the multi-scale semantic information of the upsampled features.
[0012] Furthermore, the attention-optimized TransUNet adjusts the loss function by balancing the contributions of cross entropy loss and dice loss as the objective function.
[0013] Furthermore, in the attention-optimized TransUNet, features of three different scales and partial upsampling are purified through the EMA module, the purified skip connections are fused with the purified decoder, and the channels are restored to the same resolution as the input image to obtain image prediction results.
[0014] Furthermore, the EMA module adopts three parallel routes to obtain attention weight descriptors from the grouped feature maps, of which two routes are located in the 1x1 branch and the remaining one is located in the 3x3 branch; cross-channel information interaction is modeled in the channel direction; the G group is reshaped and arranged into batch dimensions, and the input tensor is redefined; two encoded features are combined along the height direction of the image, sharing the same 1x1 convolution without reducing the dimension of the 1x1 branch; after decomposing the output of the 1x1 convolution into two vectors, two nonlinear Sigmoid functions are used to model the two-dimensional binomial distribution through linear convolution; the two channel attention maps in each group are combined by multiplication; the 3x3 branch captures local cross-channel interactions through 3x3 convolution.
[0015] Furthermore, the CB module performs average pooling on all tokens to merge the average pooled tokens into each individual token in each intermediate layer.
[0016] Furthermore, the UFO module utilizes the associative law, multiplies the key and value through the cross-normalized XNorm, and then performs another multiplication operation using the XNorm query, and obtains a linear process to n by decomposing the self-attention operation into two different sub-operations; wherein the features of the force sub-operation input are projected into a new feature space, and the operator sub-operation performs self-attention based on the projected features.
[0017] And, a medical image segmentation system based on a multi-attention optimization network, based on a computer system, including: attention-optimized TransUNet; the attention-optimized TransUNet is based on TransUNet: the SA in the Transformer is replaced by the unit force operation UFO module; the context broadcast CB module is introduced into the multi-layer perceptron MLP of the Transformer; in the jump connection of the U-Net, the feature map of the encoder is processed by the multi-scale attention EMA module and then channel-wise spliced with the upsampled features of the decoder. In the decoder, the EMA module is located after each upsampling layer to enhance the multi-scale semantic information of the upsampled features.
[0018] And, an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.
[0019] A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that the computer program implements the steps of the above method when executed by a processor.
[0020] Existing multi-attention mechanisms have been widely used to improve the performance of models such as TransUNet and UNet. However, the architecture proposed in this invention and its preferred embodiments not only follows these mechanisms but also introduces new strategies, particularly for more effectively optimizing the allocation of attention to COVID-19 lesions, focusing on lesion regions with greater accuracy. First, the EMA module partitions the channel dimension of the input feature map into subgroups, each of which is processed independently; this ensures that features corresponding to both small and large lesions are effectively captured. Unlike traditional uniform channel compression, this approach avoids the loss of critical information during dimensionality reduction by distributing local and global information across different feature groups. Furthermore, EMA dynamically reallocates spatial weights using cross-spatial information aggregation, emphasizing lesion boundaries while suppressing background interference, thereby preserving morphological details of the lesion. Second, the CB module enhances global context understanding by using global average pooling to summarize features across all pixels and broadcast this global context back to each token. This mechanism reduces the reliance of individual tokens on regions with similar grayscale or texture characteristics, such as COVID-19 lesions and healthy lung tissue. By incorporating global image context into the attention computation, CB helps the model detect subtle differences and hidden structural variations, which is particularly important for distinguishing weak or blurred lesion boundaries. Third, the combined use of EMA and CB modules ensures a balanced approach to attention optimization, addressing the challenges of local and global feature extraction. While EMA focuses on preserving multi-scale morphological details, CB enhances global contrast and context awareness. Together, these modules enhance the model's ability to accurately segment complex COVID-19 lesions, demonstrating a powerful and optimized attention mechanism tailored for this challenging task.
[0021] At the same time, the UFO module effectively reduces the computational burden by optimizing the computational complexity of MSA from O(N²) to O(N). Experimental results show that FLOPs are reduced from 24.727G to 24.711G, while the number of parameters remains unchanged. This optimization improves computational efficiency and allows AO-TransUNet to perform inference at a faster speed on ordinary GPUs (e.g., NVIDIA RTX 2080Ti, 11GB memory). BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments: Figure 1 This is a framework diagram of AO-TransUNet according to an embodiment of the present invention.
[0023] Figure 2 Schematic diagram of the EMA module according to an embodiment of the present invention; wherein, "g" represents grouping, "X Avg Pool" represents one-dimensional horizontal global pooling, and "Y Avg Pool" represents one-dimensional vertical global pooling; Figure 3 This is a schematic diagram of a CB module according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a UFO module according to an embodiment of the present invention; Figure 5 PR and ROC curves of the model proposed in this embodiment of the present invention and the baseline model TransUNet on the COVID-19 dataset, where (a) is CT data and (b) is CXR data.
[0024] Figure 6 This is a qualitative comparison result diagram of AO-TransUNet, an embodiment of the present invention, and the comparison model on the datasets (COVID-19 CT, COVID-19CXRs, Kvasir-Instrument). DETAILED DESCRIPTION
[0025] To make the features and advantages of this patent more clearly understood, the following embodiments are specifically described in detail as follows: It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the art to which this application belongs.
[0026] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0027] This paper proposes an AO-TransUNet segmentation model, which is primarily used to achieve more accurate segmentation results for COVID-19 lesion images and can also be applied to the segmentation of other medical images. This improvement is achieved primarily by introducing three modules: EMA, UFO, and CB.
[0028] To mitigate the loss of critical information about complex and variable COVID-19 lesions caused by channel dimensionality reduction modeling within the attention mechanism, EMA was employed to preserve the morphological details and characteristic information of the lesions. To address the difficulty in accurately distinguishing COVID-19 lesions from normal lung tissue images due to similar grayscale or texture features, the CB module was incorporated to enhance dense interaction and improve attention performance. Finally, the UFO module was integrated to reduce computational complexity, thereby accelerating inference when processing complex data such as COVID-19 images.
[0029] Figure 1 The overall structure of the segmentation model is presented. This segmentation network is based on an encoder-decoder architecture. The input medical image is fed into the encoder of a Transformer with UFO and CB modules. Features at three different scales and partial upsampling are then purified through an EMA module. Finally, the purified features are fused with skip connections to the decoder, restoring the channels to the same resolution as the input image, resulting in the final image prediction result.
[0030] 1) In the process of designing the scheme, it is considered that the neurons in the efficient multi-scale attention module (THE EFFICIENT MULTI-SCALE ATTENTION) benefit from a large local receptive field, which enables it to collect multi-scale spatial information. Therefore, the EMA (THE EFFICIENT MULTI-SCALE ATTENTION) module proposed in this embodiment uses three parallel routes to obtain attention weight descriptors from the grouped feature maps, such as Figure 2 As shown, two routes are located in the 1x1 branch, and the remaining one is located in the 3x3 branch. To capture all channel dependencies and manage the computational budget, cross-channel information interactions are modeled in the channel direction. EMA reshapes and arranges the G group into the batch dimension, redefining the input tensor. EMA combines two encoded features along the height direction of the image, sharing the same 1x1 convolution, without reducing the dimensionality of the 1x1 branch. After decomposing the output of the 1x1 convolution into two vectors, two nonlinear sigmoid functions are used to model a two-dimensional binomial distribution through linear convolution. To achieve diverse cross-channel interaction features in the two parallel routes of the 1x1 branch, the two channel attention maps in each group are combined through a simple multiplication. The 3x3 branch captures local cross-channel interactions through 3x3 convolutions, expanding the feature space. This scheme allows EMA to encode inter-channel information to adjust the importance of different channels while accurately preserving the spatial structure within the channels. EMA enhances feature fusion by establishing interdependencies between channels and spatial positions and applying cross-spatial information aggregation techniques across different spatial dimensions. Utilizing 2D global average pooling, EMA encodes global spatial information in the outputs of the 1x1 and 3x3 branches. These outputs are directly converted to the corresponding dimensional shapes before the channel-wise feature joint activation mechanism. Ultimately, the output feature map within each group is calculated by aggregating the two resulting spatial attention weights, followed by a sigmoid function. This process captures pixel-level pairwise relationships and emphasizes the global context of all pixels.
[0031] EMA is incorporated into U-Net as follows: In the U-Net's skip connections, the encoder's feature maps are processed by the EMA module and then channel-wise concatenated with the decoder's upsampled features (feature concatenation). For example, in the decoder's layer II, the EMA module performs multi-scale attention optimization on the encoder's layer II output before fusing it with the decoder's upsampled results.
[0032] 2) In this embodiment, the CB (Context Broadcast) module introduces the most intensive form of unified attention by performing average pooling on all tokens. Figure 3 As shown in Figure 2, this module merges average pooled tokens into each individual token at each intermediate layer. Adding CB to the ViT (Vision Transformer) model reduces the density of attention maps across all layers, simplifying overall optimization and improving generalization. Essentially, the CB module helps the ViT model reallocate resources from learning dense attention maps to capturing other information signals.
[0033] like Figure 3 As shown in the figure, in the specific scheme, the CB module is integrated into the MLP layer of the Transformer, immediately after the self-attention calculation, to enhance the global perception of intermediate features. It merges the average pooled tokens into each individual token of each intermediate layer. This approach reduces the density of the attention maps of each layer, simplifies the optimization process and enhances generalization ability. It is particularly useful for dealing with the problem of blurred boundaries between COVID-19 lesions and normal lung tissue, because it helps the model understand the dense interactions between all pixels in the image and helps to identify weak or implicit tissue structure differences in the lesion boundaries. AO-TransUNet significantly improves the model's ability to distinguish blurred boundaries and overall segmentation accuracy by first using the CB module to provide global context information and then performing multi-scale feature extraction and optimization through the EMA module. This design takes full advantage of the advantages of both mechanisms and provides a more powerful solution for medical image segmentation tasks.
[0034] 3) In this embodiment, the UFO (Unit Force Operation) module introduces a self-attention mechanism to solve the quadratic complexity of the original self-attention mechanism, such as Figure 4As shown. Cross normalization (XNorm) is a common L2 norm applied in two dimensions. Exploiting the associative law, the key and value are multiplied by XNorm, and then another multiplication is performed using the XNorm query, resulting in a linear process up to n. This is achieved by decomposing the self-attention operation into two different sub-operations: force and operation. The force operation projects the input features into a new feature space, while the operation operation performs self-attention based on the projected features. This decomposition allows UFO to reduce the computational complexity from quadratic to linear, resulting in faster inference speed and lower memory requirements.
[0035] Empirical results confirm that the UFO-ViT model has faster inference speed and lower GPU memory requirements compared to traditional Transformer.
[0036] exist Figure 1 In
[15] , the AO-TransUNet segmentation model proposed in this paper is described. Its overall architecture is similar to TransUNet and is divided into three core components: encoder, decoder and skip connection. The encoder seamlessly integrates traditional convolutional neural network (CNN) and Transformer layers, leveraging the combined advantages of these two architectures to enhance feature extraction. Convolutional neural networks effectively capture local features, while Transformer promotes global feature extraction through its self-attention mechanism. The decoder uses a traditional convolution mechanism, coupled with multiple upsampling layers, to generate segmentation results. The skip connection between the encoder and decoder facilitates upper-layer feature extraction and accurate restoration of the input image.
[0037] The design principles of AO-TransUNet take into account the limitations and inherent advantages of the Transformer and U-Net architectures in feature extraction. While the Transformer excels in global feature extraction, the original self-attention mechanism faces challenges with O(N²) time and storage complexity. The self-attention mechanism introduces the UFO module, replacing the multi-head self-attention module in TransUNet, reducing floating-point operations (FLOPs) and improving overall performance.
[0038] When studying how to accurately identify weak tissue at the boundaries of COVID-19 lesions or structural differences in crypts, the present invention focused on dense attention maps. By studying the ViT model, the present invention discovered the difficulty of learning dense attention maps via gradient descent. This challenge was extended to TransUNet, prompting the introduction of the CB module. CB reallocates resources within the TransUNet model from learning dense attention maps to other information signals, facilitating overall optimization and improving generalization.
[0039] In the decoding and skip connection components of the CNN, this paper utilizes deep convolutional layers to enhance the learned feature representations. In an encoder-decoder architecture, the decoder typically incorporates semantic information from the encoder output. However, this transmission method can potentially lead to loss of semantic information. Therefore, this paper introduces an attention mechanism, enabling the decoder to focus more on key details in the encoder output, thereby mitigating the potential loss of semantic information. However, due to channel-wise dimensionality reduction modeling, traditional attention mechanisms can still suffer from information loss. In particular, COVID-19 lesions present in various forms in medical images, such as ground-glass opacities and consolidation. These morphologies vary significantly in size, shape, and distribution. The loss of key information caused by the dimensionality reduction process hinders the preservation of morphological details and characteristic information of the lesions. To address this issue, the present invention incorporates an attention mechanism, specifically the EMA module, which outperforms traditional channel-wise attention. The EMA module excels at enhancing the learning of discriminative feature representations by aggregating cross-spatial information and effectively models long-range dependencies with precise location information. In summary, the AO-TransUNet architecture combines the strengths of convolutional neural networks and the Transformer, addressing their respective limitations through innovative modules such as UFO, CB, and EMA. These enhancements help improve the segmentation accuracy of COVID-19 lesion images. Below is a detailed description of the model architecture: 1) ENCODER: Figure 1In the left half of the Transformer layer, the UFO module is introduced to replace the traditional self-attention (SA) mechanism in the Transformer. This module solves the quadratic computational complexity problem associated with self-attention and provides a linear complexity alternative. Notably, the UFO module effectively scales the model to larger input sizes without incurring prohibitive computational costs. By eliminating the need for softmax functions, it alleviates their limitations, such as over-emphasis on local information and the challenge of modeling long-range dependencies. The UFO module obtains new features by performing matrix multiplication operations with learnable matrices. Subsequently, a spatial attention mechanism is applied to weigh the importance of different spatial positions, and the final matrix multiplication operation produces the final output features. Integrating the UFO module in TransUNet to replace the multi-head self-attention module reduces FLOPs and improves overall performance.
[0040] Furthermore, a context broadcast (CB) module is integrated into the multi-layer perceptron (MLP) within the Transformer layer. This module enhances the model's ability to capture contextual relationships between input tokens. The CB module projects the input tokens into a continuous feature space via a linear transformation layer, applies a nonlinear activation function, and maps the features back to the original input token space. This process allows the Transformer model to capture dependencies between distant tokens in the input sequence, contributing to improved contextual information and better representation. The output of the CB module is then fed into the subsequent MLP layer to further refine the token representation and generate the model's output. The CB module enhances information sharing between tokens across the model's layers through a unified attention mechanism, thereby improving the ability to capture global context. Compared to simply using an MLP layer, it facilitates global fusion of features. It enables the trans model to learn better representations and understand complex relationships, thereby improving the performance of the TransUNet task.
[0041] 2) DECODER AND skip-connections: The main function of the decoder is to reconstruct the original feature map using the features obtained from the encoder and the features received through the skip connection, using operations such as upsampling. This reconstruction process is crucial for preserving semantic information while restoring spatial details in the image. To achieve this, skip connections and decoders are strategically used to restore the original feature map.
[0042] To further reduce the semantic gap, we choose to incorporate EMA modules in the upsampling portion and skip connections of the decoder. With the addition of EMA modules, the skip connection component of TransUNet performs channel-wise and spatial-wise attention on the feature maps obtained by the self-attention layer, respectively. These two attention mechanisms are then combined through channel-wise multiplication to generate a new feature map, which is then added to the original feature map to obtain more accurate high-level features. This approach effectively mitigates potential side effects associated with modeling cross-channel relationships when extracting deep visual representations through channel-wise dimensionality reduction.
[0043] The EMA module plays a key role in the upsampling stage of the decoder. After each transposed convolutional layer, an EMA module is introduced to capture multi-scale contextual information and enhance the feature representation capabilities of the upsampling module. This enhancement enables the decoder to more effectively model local cross-channel interactions and utilize contextual information, ultimately producing higher-quality upsampled images. The integration of the EMA module significantly improves the performance of both TransUNet and the decoder by strengthening their feature representation capabilities and enhancing their ability to capture contextual information.
[0044] In the skip connection part, the EMA module extracts multi-scale features from the feature map output by the encoder, performs channel and spatial attention fusion, and then splices them with the features of the decoder.
[0045] In the decoder, the EMA module is located after each upsampling layer to enhance the multi-scale semantic information of the upsampled features.
[0046] like Figure 1 As shown, in the preferred embodiment of this invention, three EMA modules are added to the skip connections at three different scales, and an EMA module is also used in the upsampling stage of the decoder to capture multi-scale contextual information and improve the feature representation capability of the upsampling module. Each EMA module independently processes features at the corresponding scale, ensuring a balance between local and global information.
[0047] 3) LOSS FUNCTION and PRE-TRAINING: Since the closer the Dice Coefficient is to 1, the more similar the predicted value is to the true labeled value, and since the loss function generally expects a smaller value to indicate better performance, the Dice Coefficient is converted to the Dice LOSS formula as follows: Dice Loss is suitable for dealing with category imbalance and is more sensitive to errors in boundary areas. It reduces the negative impact of foreground-background imbalance in samples and focuses more on mining foreground areas during training, but it also suffers from loss saturation. The cross-entropy loss algorithm calculates the loss of each pixel on average, and the loss of the current point is only related to the distance between the current predicted value and the true label value, which may cause some problems. Using dice loss or cross-entropy loss alone often does not produce better results in medical image segmentation and needs to be used in combination. The present invention adjusts its loss function by balancing the contributions of cross-entropy loss and dice loss. The loss function is defined as follows: .
[0048] The following test examples demonstrate the effectiveness and superiority of the embodiments of the present invention: This embodiment uses the PyTorch framework and the NVIDIA RTX 2080ti GPU. In all models, the pre-trained model "R50-ViT" is used, the input image size is 224±224, the batch size is 8, and the patch size is 16. The present invention uses the Adam optimizer, configures the learning rate to be 1e-2, the momentum to be 0.9, and the weight decay to be 1e-4. For the COVID-19 CT dataset, the COVID-19CXRs dataset, and the Kvasir-Instrument dataset, the datasets are randomly divided into training and test sets at a ratio of 70%:30%. For the Synapse dataset, the dataset is randomly divided into training and test sets at a ratio of 75%:25%.
[0049] A. Ablation Studies To evaluate the effectiveness of the proposed AO-TransUNet model for medical image segmentation, we conducted an ablation experiment on the Synapse dataset. The experimental results are shown in Table 1. The purpose was to investigate whether including EMA blocks in the encoding layer and skip connections can help improve segmentation performance.
[0050] Table 1: Effect of multiple modules on encoder and jumper connections The present invention introduces EMA modules into the three skip-connected components of the baseline model, TransUNet. Results show only marginal improvements in segmentation performance, with no significant enhancement observed. Subsequently, the EMA modules are incorporated into the upsampling and skip-connected components of the decoder. While this adjustment results in a two-percentage-point decrease in the Dice Similarity Coefficient (DSC), the Hausdorff distance at the 95th percentile (HD95) metric shows improvement. Further exploration focuses on the Transformer's multi-layer perceptron (MLP) layer, incorporating CB modules into both the MLP layer and the encoder. This modification leads to further improvements in the HD95 index while increasing the DSC index. Finally, the present invention replaces the traditional multi-head self-attention (SA) module in the Transformer layer with a UFO model. The present invention investigates the impact of including EMA blocks and skip connections in the encoder layer on the segmentation performance of the newly formed Transformer layer, replacing the UFO attention module and the CB module, respectively. Notably, the HD95 of the new Transformer model, which integrates the EMA module only into the skip-connected components, further decreases, while slightly increasing the DSC. Notably, using EMA modules in both the encoding layer and the skip connection significantly improves both DSC and HD95 metrics. Importantly, these enhancements are achieved without increasing the number of parameters, and floating-point operations (FLOPs) are optimized.
[0051] The experimental evaluation of the present invention uses receiver operating characteristic (ROC) and precision recall (PR) curves to evaluate the performance of the proposed model and TransUNet model on the COVID-19 CT and chest X-ray datasets, as shown in Figure 2. Figure 5 As shown. The ROC curve is a widely used indicator for evaluating the performance of a model, plotting the false positive rate and true positive rate at different thresholds. The PR curve illustrates the precision and recall rate at different thresholds. The high curves on both graphs indicate the superior performance of the model proposed in this invention. Figure 5 In the figure, the blue curve represents the model proposed by the present invention, while the orange curve represents the TransUNet model. The model of the present invention has improved performance on both the ROC curve and the PR curve, indicating stronger classification ability and higher prediction accuracy on pixel-level COVID-19 lesion images. The segmentation results of the model of the present invention enable excellent performance in identifying true positive samples while maintaining high accuracy, which helps healthcare professionals achieve more accurate diagnosis and treatment strategies for COVID-19 patients. The evaluation results show that the multi-attention optimization network based on TransUNet performs well in COVID-19 lesion segmentation compared with the baseline TransUNet model.
[0052] As can be seen from the above results, the integration of EMA addresses the issues inherent in the channel dimensionality reduction inherent in traditional attention mechanisms. This enhancement ensures that the network captures the complex characteristics of lesions at different scales, thereby improving segmentation accuracy. Furthermore, the introduction of the UFO module and the CB module addresses the challenges associated with learning dense attention maps and the resource-intensive nature of traditional MSA algorithms. This not only enhances the learning capabilities of the Transformer but also amplifies overall network performance, enabling the model to effectively process complex lesion images.
[0053] We selected TransUNet and SwinUNet, two representative medical image segmentation models from the past two years, and compared them with our proposed approach. We compared seven segmentation performance metrics across four different datasets. The results are shown in Table 2. These results demonstrate the effectiveness of our proposed model.
[0054] Table 2: Experimental results on datasets (COVID-19 CT, COVID-19 CXRs, Synapse, Kvasir-Instrument) On the COVID-19 CT and CXRs datasets, the model of the present invention showed significant advantages in Dice coefficient and IoU value, as shown in Table 2. Specifically, the model of the present invention achieved high Dice coefficient scores of 0.74022 and 0.77966, surpassing TransUNet with scores of 0.72946 and 0.75443 and SwinUNet with scores of 0.67780 and 0.75102. The Dice coefficient of the model proposed in the present invention on COVID-19 CT and CXRs images is at least 2 percentage points higher than that of other models, and is even about 7 percentage points higher than SwinUNet on COVID-19 CT images. This shows that the model of the present invention has an enhanced ability to accurately identify and segment key areas of lesions with complex and irregular shapes and different densities in COVID-19 lung CT and chest X-ray images.
[0055] In addition, the IoU scores obtained by the model of the present invention are 0.62783 and 0.67810, which are about 2 percentage points higher than TransUNet and about 3 percentage points higher than SwinUNet. This shows that the model of the present invention has advantages in considering the consistency between the predicted segmentation results and the true labels of the COVID-19 lesion area. In terms of the HD index, the model proposed by the present invention is always better than TransUNet and SwinUNet, which shows that the boundaries generated by the segmentation model of the present invention are very close to the actual boundaries in the COVID-19 images. In addition, the model of the present invention shows obvious advantages in indicators such as sensitivity, further confirming its advantages in segmenting complex shapes and structures in COVID-19 images.
[0056] On the Synapse dataset and the Kvasir-Instrument dataset, the model of the present invention consistently outperformed TransUNet and SwinUNet in Dice coefficient, with improvements of 1.3%, 3%, 0.5%, and 3%, respectively. Similarly, the IoU index also increased by 2%, 0.4%, 4%, and 5%, respectively. These results show that the model of the present invention has been comprehensively improved compared to the current state-of-the-art models, especially compared to the latest model SwinUNet. It is worth noting that the model of the present invention has achieved substantial improvements in most indicators (such as Hausdorff distance and sensitivity).
[0057] In order to further verify the segmentation performance of the model of the present invention, the segmentation results of the model of the present invention and several other models were intuitively compared on the COVID-19 CT dataset, COVID-19CXRs dataset and Kvasir-Instrument dataset, as shown in Figure 2. Figure 6 The compared models include PSPnet, DeeplabV3+, U-Net, SwinUNet and TransUNet models.
[0058] The first column shows the original image, and the last column shows the segmentation labels for pixel values 0 and 255. Experimental results show that the compared methods only identify a portion of the lesion area and provide rough boundaries. In contrast, the proposed network captures the vast majority of the lesion area and produces high-quality segmentation results with smooth boundaries compared to other methods.
[0059] The proposed model was compared with several state-of-the-art segmentation models, including V-Net, DARR, R50-UNet, UNet, at-UNet, UNet++, R50-AttnUNet, UNet3+, R50-ViT, ViT, TransUNet, UCTransNet, TransNorm, and SwinUNet. The results show that the proposed model outperforms TransUNet in terms of overall segmentation results and organ edge prediction. Specifically, the proposed model achieved an average Dice Similarity Coefficient (DSC) of 78.52% and an average HD of 29.54 mm, which are improvements of 1.3% and 3.62 mm, respectively, over TransUNet. This demonstrates that the proposed model has enhanced segmentation capabilities in capturing both overall structure and fine organ edges. For the Dice coefficients of specific organ segmentations, including aorta, gallbladder, kidney (L), kidney (R), liver, pancreas, spleen, and stomach, the proposed model consistently outperformed TransUNet. The magnitude of improvement ranged from 0.05% for kidney (L) to 3.64% for gallbladder. It is worth noting that our model achieves the highest DSC index, indicating a higher region overlap between the ground truth and the segmentation results.
[0060] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.
[0061] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, performs the above-described method. The storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0062] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present disclosure. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
[0064] The present invention is not limited to the above-mentioned optimal implementation mode. Anyone can derive various other forms of medical image segmentation methods based on multi-attention optimization networks under the inspiration of the present invention. All equal changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.
Claims
1. A medical image segmentation method based on a multi-attention optimization network, characterized in that: Medical image segmentation is performed by attention-optimized TransUNet; the attention-optimized TransUNet is based on TransUNet: the SA in Transformer is replaced by the unit force operation UFO module; the context broadcast CB module is introduced into the multi-layer perceptron MLP of Transformer; in the jump connection of U-Net, the feature map of the encoder is processed by the multi-scale attention EMA module and then channel-wise spliced with the upsampled features of the decoder. In the decoder, the EMA module is located after each upsampling layer to enhance the multi-scale semantic information of the upsampled features.
2. The medical image segmentation method based on multi-attention optimization network according to claim 1, characterized in that: The attention-optimized TransUNet adjusts the loss function by balancing the contributions of the cross entropy loss and the dice loss as the objective function.
3. The medical image segmentation method based on multi-attention optimization network according to claim 1, characterized in that: In the attention-optimized TransUNet, features of three different scales and partial upsampling are purified through the EMA module, the purified skip connections are fused with the purified decoder, and the channels are restored to the same resolution as the input image to obtain image prediction results.
4. The medical image segmentation method based on multi-attention optimization network according to claim 1, characterized in that: The EMA module uses three parallel routes to obtain attention weight descriptors from the grouped feature maps, two of which are located in the 1x1 branch and the remaining one is located in the 3x3 branch; Cross-channel information interaction is modeled in the channel direction; the G group is reshaped and arranged into batch dimensions, and the input tensor is redefined; the two encoded features are combined along the height direction of the image, sharing the same 1x1 convolution without reducing the dimension of the 1x1 branch; after decomposing the output of the 1x1 convolution into two vectors, two nonlinear Sigmoid functions are used to model the two-dimensional binomial distribution through linear convolution; the two channel attention maps in each group are combined by multiplication; the 3x3 branch captures local cross-channel interactions through 3x3 convolution.
5. The medical image segmentation method based on multi-attention optimization network according to claim 1, characterized in that: The CB module performs average pooling on all tokens to merge the average pooled tokens into each individual token in each intermediate layer.
6. The medical image segmentation method based on multi-attention optimization network according to claim 1, characterized in that: The UFO module uses the associative law, and the key and value are multiplied by cross-normalized XNorm, and then another multiplication operation is performed using the XNorm query. The self-attention operation is decomposed into two different sub-operations to obtain a linear process to n; the features of the force sub-operation input are projected into a new feature space, and the operation sub-operation performs self-attention based on the projected features.
7. A medical image segmentation system based on a multi-attention optimization network, based on a computer system, characterized in that: include: Attention optimized TransUNet; the attention optimized TransUNet is based on TransUNet: the SA in Transformer is replaced by the unit force operation UFO module; the context broadcast CB module is introduced in the multi-layer perceptron MLP of Transformer; in the jump connection of U-Net, the feature map of the encoder is processed by the multi-scale attention EMA module and then channel-wise spliced with the upsampled features of the decoder. In the decoder, the EMA module is located after each upsampling layer to enhance the multi-scale semantic information of the upsampled features.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the medical image segmentation method based on a multi-attention optimization network as described in any one of claims 1 to 6 are implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the medical image segmentation method based on a multi-attention optimization network as described in any one of claims 1 to 6 are implemented.