Cross-shaped window self-attention and edge perception medical image segmentation method based on channel enhancement
By combining cross-shaped window self-attention with edge perception in medical image segmentation, this method solves the problems of insufficient accuracy in global dependency modeling and boundary segmentation in existing technologies, achieving efficient and accurate medical image segmentation, applicable to organ segmentation in CT and MRI images.
Patent Information
- Application Number
- CN202511440865.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-01-09
AI Technical Summary
Existing medical image segmentation techniques cannot simultaneously meet the requirements of efficient global dependency modeling, accurate channel feature identification, fine boundary preservation, and strong generalization ability, especially when dealing with blurred organ boundaries and complex backgrounds, where there is insufficient segmentation accuracy.
A medical image segmentation method based on channel-enhanced cross-shaped window self-attention and edge perception is adopted. Global dependency modeling and adaptive calibration of channel features are achieved through the ECCA block of the encoder and decoder. Fine-grained boundary preservation is performed by combining the edge perception module, and a composite loss function is used for training.
It improves the accuracy and efficiency of medical image segmentation, especially the segmentation accuracy of organ boundaries in CT and MRI images, enhances the adaptability to different modalities and anatomical structures, and achieves efficient fine-grained boundary preservation and semantic discrimination.
Smart Images

Figure CN121304585A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image processing and computer-aided diagnosis (CAD) technology, and relates to medical image segmentation methods, specifically a medical image segmentation method based on channel enhancement, cross-shaped window self-attention, and edge perception. Background Technology
[0002] Medical image segmentation is a key technology for extracting target anatomical structures (such as organs and lesion areas) from medical images such as CT and MRI, and its accuracy directly affects the formulation of clinical diagnosis and treatment plans. However, medical images have problems such as variable organ shape, size, height, low boundary contrast, and complex backgrounds, which make automatic segmentation still a huge challenge.
[0003] Existing medical image segmentation techniques are mainly divided into two categories:
[0004] 1. Convolutional Neural Network (CNN) based methods: Represented by U-Net and its variants (such as U-Net++ and UNet3+), these methods capture multi-scale local features through encoder-decoder architecture and skip connections, performing well in simple scenes. However, the local operation characteristics of CNNs make it difficult to model "long-range pixel dependencies" in images, and they cannot effectively distinguish adjacent organs with similar intensity distributions but different anatomical structures (such as liver and stomach, pancreas and intestines).
[0005] 2. Visual Transformer (ViT) based methods: Represented by TransUNet, SwinUNet, and TransBTS, these methods model the global context through self-attention mechanisms, solving the long-range dependency problem of CNNs. However, these methods have significant drawbacks:
[0006] TransUNet loses local details when fusing global and local information;
[0007] While SwinUNet reduces computational complexity, it tends to overlook key details when dealing with complex backgrounds and small targets.
[0008] TransBTS enhances multimodal fusion, but significantly increases computational resource consumption;
[0009] The cross-shaped window (CSWin) Transformer decomposes global attention into local computations of horizontal and vertical stripes (reducing complexity from O(N)). 2 While the resolution has been reduced to O(N), problems such as loss of local fine details and lack of channel feature modulation mechanism still exist, making it impossible to highlight key semantic features.
[0010] In addition, existing methods generally suffer from the problem of "insufficient boundary segmentation accuracy": organ boundaries in medical images are blurred, and traditional segmentation models have difficulty accurately preserving fine-grained boundary structures; some edge enhancement methods (such as multi-scale feature fusion and edge-guided attention) can improve the boundaries, but they increase network complexity and computational cost, or cause the segmentation performance of other regions to decline due to excessive focus on edges.
[0011] In summary, existing technologies cannot simultaneously meet the medical image segmentation requirements of "efficient global dependency modeling, accurate channel feature identification, and fine boundary preservation," and a new solution that balances accuracy, efficiency, and robustness is urgently needed. Summary of the Invention
[0012] To overcome the shortcomings of the prior art, the purpose of this invention is to provide a medical image segmentation method based on channel-enhanced cross-shaped window self-attention and edge perception, solving the following problems existing in the prior art:
[0013] This invention aims to solve the following core problems existing in current medical image segmentation technology:
[0014] 1. The contradiction between global dependency modeling and computational efficiency: Existing ViT methods have strong global modeling capabilities but high computational costs, while improved methods such as CSWin still lack channel modulation mechanisms;
[0015] 2. Insufficient discriminative power of channel features: It cannot adaptively enhance key semantic features and suppress redundant information, making it difficult to distinguish organs with similar intensity distributions;
[0016] 3. Low boundary segmentation accuracy: It lacks a lightweight and effective edge supervision mechanism and cannot accurately preserve fine-grained anatomical boundaries;
[0017] 4. Limited generalization ability: It is difficult to maintain stable performance on images of different modalities (CT / MRI) and different anatomical structures (abdominal organs / heart).
[0018] It is suitable for the precise segmentation of anatomical structures (such as abdominal organs and heart structures) in medical images such as CT and MRI, and can provide support for clinical decision-making such as tumor volume measurement, organ function assessment, and radiotherapy planning.
[0019] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0020] A medical image segmentation method based on channel-enhanced cross-shaped window self-attention and edge perception, characterized by the following steps:
[0021] S1. Perform convolutional token embedding on the input medical image (CT / MRI), and convert the image into block token features through 7×7 convolution (stride 4);
[0022] S2. Multi-scale features are extracted using an encoder. The encoder consists of four stages, each consisting of an ECCA block and a downsampling layer. The ECCA block integrates cross-shaped window self-attention (CSWin) and channel attention (SE) to achieve global dependency modeling and adaptive calibration of channel features.
[0023] S3. Features are reconstructed using the decoder. The decoder is symmetrical with the encoder and contains four stages. Each stage recovers spatial details through CARAFE upsampling and fuses the features of the corresponding stage of the encoder through U-shaped jump connections.
[0024] S4. Output results through dual prediction heads: The segmentation head generates a semantic segmentation map through 1×1 convolution, and the edge perception module generates an edge prediction map through 3×3 convolution, batch normalization, activation function (ReLU) and 1×1 convolution;
[0025] S5. The model is trained using a composite loss function, which includes Dice loss, cross-entropy loss and edge loss, and edge supervision is activated in the later stage of training (the last 10% of epochs).
[0026] The workflow of the ECCA block is as follows:
[0027] A1. Input features are normalized using LayerNorm(LN);
[0028] A2. Applying CSWin self-attention: Divide the features into horizontal and vertical stripes, generate Q, K, and V, calculate the self-attention weights, and aggregate global features;
[0029] A3. Residual Connection: The CSWin self-attention output is added to the residual of the original input features;
[0030] A4. After LN normalization, the data is input into a multilayer perceptron (MLP) to refine local features;
[0031] A5. Apply SE channel attention: Generate channel weights through global average pooling, fully connected layers and sigmoid, reweight the features channel by channel and sum the residuals.
[0032] The cross-shaped window self-attention will input the feature map Divided into horizontal and vertical stripes, for each stripe X i The query, key, and value are obtained through linear projection:
[0033] Q i =W Q X i ,K i =W K Xi V i =W V X i (1)
[0034] in, It is a learnable matrix, d n =C / K is the dimension of each head, and there are a total of K heads. Then, the self-attention of the nth head is calculated as follows:
[0035]
[0036] Generate the attention output for the i-th stripe. The combined output of the horizontal and vertical stripes is:
[0037] H Attn (X)=[γ1,...,γ M ],V Attn (X)=[γ′1,...,γ′ S (3)
[0038] The final result for the bulls was:
[0039]
[0040] Among them W o This indicates the output projection. This design effectively captures remote spatial interactions.
[0041] To further refine the feature representation, SE attention adaptively recalibrates the channels. A global descriptor for each channel is obtained through average pooling.
[0042]
[0043] Then through two fully connected layers:
[0044] z′ c =ReLU(W1z) c ),a c =σ(W2z′) c (6)
[0045] Where a c The channel attention weights are generated by the sigmoid activation function, and the input is ultimately reweighted using residual enhancement.
[0046]
[0047] This mechanism enhances effective information channels while suppressing less relevant ones.
[0048] By combining spatial self-attention with adaptive channel attention recalibration, the ECCA block improves context modeling and feature selectivity, thereby enhancing segmentation performance for anatomically complex structures.
[0049] The edge labels of the edge perception module are generated by applying the Canny edge detector to the ground-truth mask of the medical image segmentation to generate a real edge map offline.
[0050] The formula for the composite loss function is as follows:
[0051] Among them α=0.54, β=0.36, γ=0.10, For Dice's loss, For cross-entropy loss, This is an edge loss based on BCE loss.
[0052] The number of ECCA blocks in the encoder / decoder is configured as [1,2,9,1], that is, encoder stages 1 to 4 contain 1, 2, 9 and 1 ECCA blocks respectively.
[0053] The beneficial effects of this invention are:
[0054] 1. High segmentation accuracy: Validated on two public benchmark datasets, its performance outperforms mainstream methods.
[0055] Synapse abdominal CT dataset: The average Dice similarity coefficient (DSC) reached 81.90%, and the average Hausdorff distance (HD) reached 20.05 mm. Compared with TransUNet, SwinUNet, and CSWin-UNet, the DSC was improved by 4.42%, 2.27%, and 2.02%, respectively, and the HD was reduced by 11.64%, 1.5%, and 7.93%, respectively.
[0056] ACDC cardiac MRI dataset: The average DSC reaches 91.10%, and the segmentation accuracy of the right ventricle (RV), myocardium (MYO), and left ventricle (LV) is superior to existing methods;
[0057] 2. Superior computational efficiency: By using CSWin self-attention, the complexity of global attention is reduced from O(N^2) to O(N^2). 2 The computational efficiency is reduced to O(N), while the edge sensing module is designed to be lightweight and does not require a large amount of additional computing resources.
[0058] 3. Good boundary preservation: The edge perception module, combined with a delayed supervision strategy, accurately extracts the boundaries of fuzzy organs, solving the problem of inaccurate segmentation of fine-grained structures;
[0059] 4. Strong generalization ability: It performs stably in both CT (abdomen) and MRI (heart) modalities and on multiple anatomical structures, and can adapt to the diverse medical image segmentation needs in clinical practice.
[0060] 5. Accurate semantic discrimination: The SE channel in the ECCA block adaptively enhances key semantic features, effectively distinguishing adjacent organs with similar intensity distributions (such as pancreas and intestine, left kidney and right kidney). Attached Figure Description
[0061] Figure 1 A graph showing organ segmentation performance on the Synapse dataset;
[0062] Figure 2 Visualization of edge prediction;
[0063] Figure 3 Comparison chart of segmentation results for the Synapse dataset;
[0064] Figure 4 A graph showing the performance of heart structure segmentation in the ACDC dataset;
[0065] Figure 5 A performance comparison chart for different ECCA block counts;
[0066] Figure 6 Visualization of segmentation results for different numbers of ECCA blocks;
[0067] Figure 7 This is a diagram showing the influence of the weights on the composite loss function. Detailed Implementation
[0068] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0069] A medical image segmentation method based on channel-enhanced cross-shaped window self-attention and edge perception, characterized by the following steps:
[0070] S1. Perform convolutional token embedding on the input medical image (CT / MRI), and convert the image into block token features through 7×7 convolution (stride 4);
[0071] S2. Multi-scale features are extracted using an encoder. The encoder consists of four stages, each consisting of an ECCA block and a downsampling layer. The ECCA block integrates cross-shaped window self-attention (CSWin) and channel attention (SE) to achieve global dependency modeling and adaptive calibration of channel features.
[0072] S3. Features are reconstructed using the decoder. The decoder is symmetrical with the encoder and contains four stages. Each stage recovers spatial details through CARAFE upsampling and fuses the features of the corresponding stage of the encoder through U-shaped jump connections.
[0073] S4. Output results through dual prediction heads: The segmentation head generates a semantic segmentation map through 1×1 convolution, and the edge perception module generates an edge prediction map through 3×3 convolution, batch normalization, activation function (ReLU) and 1×1 convolution;
[0074] S5. The model is trained using a composite loss function, which includes Dice loss, cross-entropy loss and edge loss, and edge supervision is activated in the later stage of training (the last 10% of epochs).
[0075] Based on the above method, a medical image segmentation system based on cross-shaped window self-attention and edge perception is adopted, including:
[0076] Image preprocessing module: used for normalization, data augmentation, and convolutional token embedding of medical images;
[0077] Feature encoding module: contains an encoder with four stages, each stage consisting of the ECCA block and downsampling layer as described in claim 2, to extract multi-scale global features;
[0078] Feature decoding module: contains a 4-stage decoder, each stage consisting of a CARAFE upsampling layer and a U-shaped jump connection to reconstruct spatial details;
[0079] Dual prediction module: includes a segmentation head and an edge perception module, which output semantic segmentation map and edge prediction map respectively;
[0080] Model training module: Implements the calculation of the composite loss function and parameter optimization as described in claim 4.
[0081] The enhancement operations of the image preprocessing module include horizontal and vertical flipping, random rotation (-15° to 15°), window width and window level adjustment of CT images, and Gaussian noise addition.
[0082] The model training module uses a stochastic gradient descent (SGD) optimizer with an initial learning rate of 0.05, a batch size of 24, and a training duration of 100 epochs. The marginal loss weight γ = 0 for the first 90 epochs and γ = 0.10 for the last 10 epochs.
[0083] Example:
[0084] ECCA-UNet is implemented using Python and the PyTorch framework. Model training and evaluation are performed on a system with 24GB of VRAM. GeForce RTX TMThe training was performed on a 3090 GPU. We initialized the CSWinTransformer block with pre-trained weights to leverage prior knowledge and accelerate convergence during training. To enhance the diversity of the training dataset, data augmentation techniques such as flipping and rotating were applied, which helps improve the model's generalization ability to unseen data. During training, the batch size was set to 24 and the learning rate to 0.05. The optimization method used was stochastic gradient descent (SGD) with a momentum of 0.9 and a weight decay of 10. -4 This approach is chosen to optimize the balance between fast learning and stable convergence.
[0085] Dataset:
[0086] Synapse dataset:
[0087] The Synapse dataset originates from the MICCAI 2015 Multi-Image Atlas Abdominal Organ Segmentation Challenge. It is a widely used public dataset for abdominal CT image segmentation. The Synapse dataset contains 30 abdominal CT scans, totaling 3779 CT images. Each CT scan consists of 85 to 198 slices, each with a resolution of 512×512 pixels. Following the settings in CSWin-UNet, we selected 18 datasets for training and 12 datasets for evaluation. In the field of medical image segmentation, many studies employ similar segmentation ratios (e.g., 60% for training and 40% for testing) to ensure the reproducibility and comparability of experimental results. This segmentation method has become standard practice for the Synapse dataset.
[0088] Eighteen training subsets (2211 images in total) provide sufficient sample size to train deep learning models, especially complex models like ECCA-UNet. Ample training data helps the model learn the diversity and complexity of abdominal organs, thereby improving generalization ability. Twelve test subsets (1568 images in total) are used to fully evaluate the model's performance on unseen data. The test set includes CT scans from different cases, reflecting the model's performance in real-world applications, particularly its robustness in handling different anatomical structures and lesions.
[0089] We used the mean Dice similarity coefficient (DSC) and the mean Hausdorff distance (HD) as evaluation metrics to assess the model’s segmentation performance on eight abdominal organs (aorta, gallbladder, left kidney, right kidney, liver, pancreas, spleen, and stomach).
[0090] ACDC dataset:
[0091] The Automated Cardiac Diagnosis Challenge (ACDC) dataset, released during the 2017 ACDC Challenge, is a multi-class cardiac 3D MRI dataset consisting of 100 short-axis MR images. These images were acquired using a cine MRI scanner at 1.5T and 3T. Expert annotations were provided for three cardiac structures: right ventricle (RV), myocardium (MYO), and left ventricle (LV). For training purposes, 70 MR images were randomly selected, with 10 used for validation and 20 for evaluation. The ACDC dataset uses the mean Dice similarity coefficient (DSC) as the evaluation metric to assess the segmentation performance of the three cardiac structures.
[0092] Experimental results:
[0093] Results from the Synapse dataset:
[0094] As shown in Table 1, the proposed method improves the mean DSC and HD for each organ on the Synapse dataset. Meanwhile, Figure 1 The error bars (95% confidence intervals) for mean DSC, mean HD, and DSC for each organ on the Synapse dataset are shown. Compared to TransUNet, SwinUNet, and CSWin-UNet, our network improves mean DSC by 4.42%, 2.27%, and 2.02%, and means mean HD by 11.64%, 1.5%, and 7.93%, respectively.
[0095] Table 1: Test results on the Synapse dataset, including the Dice coefficient for each organ and the final average Dice and HD values. The first and second best values are highlighted in red and blue, respectively.
[0096]
[0097] The segmentation of organs such as the kidneys, liver, pancreas, spleen, and stomach is challenging due to variations in their shape, location, and relative relationship with surrounding organs. Larger organs, in particular, with relatively regular shapes, still adjacent to other organs (such as the stomach and spleen), may exhibit complex background noise. Smaller organs, such as the pancreas and spleen, often face problems with blurred boundaries and shape variations. The proposed method demonstrates excellent segmentation performance for organs such as the kidneys, liver, pancreas, spleen, and stomach, proving its strong robustness and generalization ability. It can effectively handle organs with different shapes, locations, and anatomical structures. By combining global information with local detail enhancement, the model can not only accurately segment large organs (such as the liver and stomach) and small organs (such as the pancreas and spleen), but also adapt to challenges such as complex backgrounds and blurred boundaries, exhibiting excellent performance in multi-organ segmentation tasks.
[0098] This demonstrates that the ECCA-UNet method exhibits high accuracy and broad application potential, particularly in medical image segmentation tasks such as abdominal CT image segmentation.
[0099] To enhance the interpretability of our method and demonstrate the effectiveness of edge-aware learning, we propose a comprehensive visualization of the learned edge predictions. For example... Figure 2 As shown, the first column displays the original input image, showing the raw anatomical structures without any processing. The second column presents an overlay visualization, combining the original image with ground truth edge maps (shown in blue) for corresponding labels, providing a clear reference for edge localization accuracy. The third column shows the predicted edge maps generated by our model, highlighting the learned edge features that contribute to improved segmentation performance.
[0100] To visually demonstrate the segmentation results of our method, we selected some slices for visualization, such as... Figure 3 As shown in the image, the first column displays the original image, the second column displays the labels, and the third column displays the segmentation results obtained by our method. It can be observed that the second and third columns are highly consistent.
[0101] Results on the ACDC dataset:
[0102] As shown in Table 2, the proposed method improves the average DSC for each organ in the ACDC dataset. Figure 4 Error bars representing the 95% confidence interval of the mean DSC value for each cardiac anatomical structure are shown. The anatomical regions are represented as follows: RV (right ventricle), MYO (myocardium), and LV (left ventricle). The results show that the proposed ECCA-UNet framework can effectively identify and segment the target: Table 2 compares it with state-of-the-art methods on the ACDC dataset. The first and second best values are highlighted in red and blue, respectively.
[0103]
[0104] These organs achieved an overall accuracy of 91.10%, and demonstrated strong generalization ability and computational stability. Model configuration analysis:
[0105] In this section, we provide a comprehensive analysis of the ECCA-UNet configuration on the Synapse dataset. Specifically, we evaluate the impact of different network architectures and loss function parameters on model performance.
[0106] Network architecture:
[0107] The depth of a neural network directly impacts its feature extraction capability and overall performance. Insufficient layers may lead to inadequate feature representation, while excessive layers increase computational complexity and may even cause non-convergence during training. Therefore, network design needs to balance depth and computational resources. To avoid convergence issues in overly deep networks, we set the number of modules in the final stage to 1. Through comparative analysis with other Transformer models, we selected [1,2,6,1], [1,2,9,1], and [1,2,12,1] as the module configurations for the encoder and decoder, aiming to balance network depth and performance through experiments. These configurations ensure that the model extracts sufficient features in the first two stages, while using more blocks (such as 6, 9, or 12 blocks) in the middle stages to further improve performance, while avoiding the problems associated with overly deep networks.
[0108] Experimental results:
[0109] Table 3: Ablation study of the number of ECCA blocks at each stage on the Synapse dataset. The first and second best values are highlighted in red and blue, respectively.
[0110]
[0111] The configuration [1,2,9,1] was confirmed to provide the best performance, as shown in Table 3.
[0112] At the same time, we Figure 5 (a) and Figure 5 (b) presents the mean DSC, HD, and mean DSC for each organ when the number of ECCA blocks is set to [1,2,6,1] and [1,2,12,1], respectively.
[0113] Visualization of slices, such as Figure 6 As shown in the image, we selected a slice for comparison. From top to bottom, the block configurations are [1,2,6,1], [1,2,9,1], and [1,2,12,1]. From left to right, the images are the original image, the ground truth label, and the predicted segmentation, respectively. This is a multi-organ segmentation task for eight abdominal organs, aiming to simultaneously identify and segment these eight organs from CT images. The prediction results in the right image show that the model performs well in segmenting these organs, with the configuration [1,2,9,1] showing the best performance.
[0114] Parameters of the loss function:
[0115] To investigate the contribution of each component in the proposed loss function, we systematically varied the weights α, β, and γ. In the first stage, γ was set to 0 to exclude the edge-aware term, thereby isolating the effects of the Dice loss and cross-entropy loss. This design allows for a focused evaluation of region-level overlap and pixel-level classification without being affected by the confusion caused by boundary supervision.
[0116] Five representative weight combinations—[1,0], [0,1], [0.5,0.5], [0.4,0.6], and [0.6,0.4]—were selected to examine the sensitivity of segmentation performance. [1,0] and [0,1] represent dedicated uses of the Dice loss and cross-entropy loss, respectively, serving as baseline extremes. The [0.5,0.5] setting provides a neutral balance, while [0.4,0.6] and [0.6,0.4] introduce a moderate bias towards one component. These configurations help assess whether a slight bias towards a particular component can improve optimization without allowing any single loss to dominate.
[0117] like Figure 7 As shown, all hybrid loss settings outperform Dice or cross-entropy loss alone. The best performance is observed at α = 0.6 and β = 0.4 without edge supervision.
[0118] Building on these results, we then evaluated the effectiveness of edge-aware supervision by introducing a non-zero γ, while maintaining the ratio α:β = 3:2. Specifically, we tested γ ∈ {0.05, 0.10, 0.20} and scaled α and β to preserve their relative influence. This strategy ensures that region-based supervision remains balanced while allowing for boundary-guided controlled integration. The selected γ values represent weak, moderate, and relatively strong edge emphasis, which helps in analyzing the model's sensitivity to boundary information.
[0119] Among them, the configuration with α=0.54, β=0.36, and γ=0.10 had the best overall segmentation effect, with an average DSC of 81.90% and an average HD of 20.05mm.
[0120] Experimental results demonstrate the complementary nature of the three loss function components: Dice loss optimizes overall consistency at the shape level, cross-entropy loss improves pixel-level classification accuracy, and edge-aware loss refines boundary space localization accuracy. The synergistic effect of these three components enables the segmentation model of this invention to simultaneously maintain global structural integrity and local detail accuracy, thereby achieving anatomically coherent and robust segmentation results in various organ segmentation tasks.
[0121] like Figure 7 As shown, the blue bars represent the DSC score (the higher the value, the better the segmentation performance), and the orange bars represent the HD value (the lower the value, the more accurate the boundary positioning). The HD index is displayed using an inverted coordinate axis to facilitate a visual comparison of the changing trends of the two indicators.
[0122] Ablation studies:
[0123] To evaluate the effectiveness of each component in our proposed ECCA-UNet, we conducted a series of ablation experiments on the Synapse dataset. As shown in Table 4, we progressively integrated the SE attention module and the edge awareness module into the baseline network (CSWin-UNet) and evaluated the performance gains.
[0124] Experimental results show that the baseline network using only the CSWin encoder achieved an average DSC similarity coefficient of 79.88% and an average HD of 27.98. When the edge-aware module is added, the performance improves to 80.67% DSC and 24.36 HD, demonstrating the benefits of explicit boundary modeling for enhanced spatial localization.
[0125] In contrast, combining only the SE attention module yields a more significant improvement (mean DSC: 81.20%, mean HD: 21.39), indicating that enhancing the channel feature response significantly improves semantic representation and suppresses background noise.
[0126] When the SE module and the edge-aware module work together, the complete ECCA-UNet architecture achieves optimal performance, with an average Dice similarity coefficient of 81.90% and an average Hausdorff distance reduced to 20.05. Experimental data validates the complementarity between the two modules: the SE attention mechanism enhances semantic feature recognition capabilities, while the edge-aware branch improves boundary localization accuracy. Their synergistic effect significantly improves both the accuracy and robustness of the segmentation results.
[0127] Table 4: Ablation studies of each module on the Synapse dataset. The first and second best values are highlighted in red and blue, respectively.
[0128]
[0129] In summary:
[0130] This invention proposes a medical image segmentation network, ECCA-UNet, based on a cross-shaped window self-attention architecture, which effectively solves the technical challenges of organ morphology diversity, boundary localization ambiguity, and interference from complex anatomical backgrounds in existing technologies. This invention constructs a complete technical solution by integrating a cross-shaped window self-attention mechanism (achieving efficient global context modeling), a squeeze excitation module (achieving adaptive channel feature enhancement), and an edge-aware auxiliary branch (achieving boundary representation refinement).
[0131] Validation experiments on multiple standard datasets demonstrate that the present invention achieves advanced segmentation accuracy across different anatomical structures and imaging modalities (including CT and MRI), and exhibits excellent generalization performance. In particular, the present invention excels in fine-grained boundary preservation and low-contrast region discrimination, making it valuable for practical clinical deployment.
Claims
1. A medical image segmentation method based on channel-enhanced cross-shaped window self-attention and edge perception, characterized in that, Includes the following steps: S1. Perform convolutional token embedding on the input medical image (CT / MRI), and convert the image into block token features through 7×7 convolution (stride 4); S2. Multi-scale features are extracted using an encoder. The encoder consists of four stages, each consisting of an ECCA block and a downsampling layer. The ECCA block integrates cross-shaped window self-attention (CSWin) and channel attention (SE) to achieve global dependency modeling and adaptive calibration of channel features. S3. Features are reconstructed using the decoder. The decoder is symmetrical with the encoder and contains four stages. Each stage recovers spatial details through CARAFE upsampling and fuses the features of the corresponding stage of the encoder through U-shaped jump connections. S4. Output results through dual prediction heads: The segmentation head generates a semantic segmentation map through 1×1 convolution, and the edge perception module generates an edge prediction map through 3×3 convolution, batch normalization, activation function (ReLU) and 1×1 convolution; S5. The model is trained using a composite loss function, which includes Dice loss, cross-entropy loss and edge loss, and edge supervision is activated in the later stage of training (the last 10% of epochs).
2. The medical image segmentation method based on channel enhancement and cross-shaped window self-attention and edge perception according to claim 1, characterized in that, The workflow of the ECCA block is as follows: A1. Input features are normalized using LayerNorm(LN); A2. Applying CSWin self-attention: Divide the features into horizontal and vertical stripes, generate Q, K, V and calculate self-attention weights, and aggregate global features; A3. Residual Connection: The output of CSWin self-attention is added to the residual of the original input features; A4. After LN normalization, the data is input into a multilayer perceptron (MLP) to refine local features; A5. Apply SE attention: Generate channel weights through global average pooling, fully connected layers and sigmoid, reweight the features channel by channel and sum the residuals.
3. The medical image segmentation method based on channel enhancement and cross-shaped window self-attention and edge perception according to claim 1, characterized in that, The cross-shaped window self-attention will input the feature map Divided into horizontal and vertical stripes, for each stripe X i The query, key, and value are obtained through linear projection: Q i =W Q X i ,K i =W K X i ,V i =W V X i (1) in, It is a learnable matrix, d n =C / K is the dimension of each head, there are a total of K heads, and then the self-attention of the nth head is calculated as follows: Generate the attention output for the i-th stripe. The combined output of the horizontal and vertical stripes is: H Attn (X)=[γ1,...,γ M ],V Attn (X)=[γ1,...,γ′ S ] (3) The final result for the bulls was: Among them W o This indicates the output projection. This design effectively captures remote spatial interactions. To further refine the feature representation, SE attention adaptively recalibrates the channels. A global descriptor for each channel is obtained through average pooling. Then through two fully connected layers: With c =ReLU(W1z c ),and c =σ(W2z′ c ) (6) Where a c The channel attention weights are generated by the sigmoid activation function, and the input is ultimately reweighted using residual enhancement. This mechanism enhances effective information channels while suppressing less relevant ones. By combining spatial self-attention with adaptive channel attention recalibration, the ECCA block improves context modeling and feature selectivity, thereby enhancing segmentation performance for anatomically complex structures.
4. The medical image segmentation method based on channel enhancement and cross-shaped window self-attention and edge perception according to claim 1, characterized in that, The edge labels of the edge perception module are generated by applying the Canny edge detector to a mask of the real image of the medical image to generate a real edge map offline.
5. The medical image segmentation method based on channel enhancement and cross-shaped window self-attention and edge perception according to claim 1, characterized in that, The formula for the composite loss function is as follows: Among them α=0.54, β=0.36, γ=0.10, For Dice's loss, For cross-entropy loss, This is the edge loss.
6. The medical image segmentation method based on channel enhancement and cross-shaped window self-attention and edge perception according to claim 1, characterized in that, The number of ECCA blocks in the encoder and decoder is configured as [1,2,9,1], that is, encoder stages 1 to 4 contain 1, 2, 9 and 1 ECCA blocks respectively.