Asymmetry-based lightweight medical image segmentation network (ABUNet) and implementation method thereof
By employing the asymmetric design of ABUNet and optimizing feature processing using FSCB, FACB, and MSDB modules, the problems of feature redundancy, multi-scale fusion, and context modeling in medical image segmentation are solved, achieving efficient and lightweight medical image segmentation.
Patent Information
- Application Number
- CN202510502020.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-11-18
AI Technical Summary
Existing medical image segmentation technologies struggle to balance lightweight deployment with sub-pixel segmentation accuracy, exhibiting issues such as feature redundancy and computational efficiency imbalance, coarse multi-scale feature fusion, and insufficient granularity in context modeling. Consequently, they fail to effectively capture complex intra-sample feature relationships and the advantages of multi-scale features.
An asymmetric lightweight medical image segmentation network (ABUNet) is adopted. In the encoding stage, feature subtraction convolutional blocks (FSCB) are used to reduce redundancy, and feature addition convolutional blocks (FACB) are used to fuse multi-path features in the decoding stage. Multi-scale deep convolutional blocks (MSDB) are introduced to perform multi-scale feature fusion. The feature representation is optimized by combining SE modules and residual connections.
It achieves efficient feature representation and detail restoration, improving the accuracy and efficiency of medical image segmentation, and significantly outperforming existing methods in segmentation performance on the ISIC2017 and ISIC2018 datasets.
Smart Images

Figure CN120976229A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical image segmentation, in particular to a lightweight medical image segmentation network based on asymmetry and an implementation method thereof. BACKGROUND
[0002] As a core task of modern medical image analysis, medical image segmentation aims to accurately locate lesion regions and anatomical structures through pixel-level classification. With the rapid development of imaging technologies such as CT, MRI, and ultrasound, this technology has become a key support for disease diagnosis, surgical planning, and efficacy evaluation.
[0003] As a key enabling technology for precision medicine, medical image segmentation faces dual challenges in algorithm design: on the one hand, it must meet the real-time clinical requirements of lightweight deployment; on the other hand, it needs to maintain sub-pixel-level segmentation accuracy in complex anatomical structures. Currently, mainstream methods are mainly based on symmetric UNet architecture. However, limited by the symmetry between the encoder and the decoder, these methods often increase the complexity of modules to improve performance, resulting in a trade-off between accuracy and efficiency. The existing methods have three major limitations: (1) imbalance between feature redundancy and computational efficiency, the encoding stage usually adopts a feature stacking strategy, resulting in deep feature maps containing a large number of low-information channels. For example, TransUNet introduces a Transformer, resulting in a 7% increase in parameters but only a 12% increase in effective features, indicating low efficiency in the use of computing resources; (2) rough multi-scale feature fusion, the traditional skip connection in the decoding stage makes it difficult to achieve fine registration of multi-scale features. For example, SwinUNet uses cross-window attention to enhance global modeling capabilities, but its computational complexity grows quadratically, limiting its practicality; (3) insufficient context modeling granularity, the bridging stage relies on single-scale convolution, making it difficult to capture different morphological features of lesions. For example, UNet++ enhances information flow through dense connections, but introduces 42% redundant parameters, greatly increasing computational overhead.
[0004] Medical image data exhibits unique distribution characteristics, with minimal inter-sample feature transformation between lesion regions and normal tissues, but significant intra-sample feature variation. Although mainstream self-attention (SA) based enhancement modules have shown good prospects in feature extraction, they struggle to effectively capture complex intra-sample feature relationships, limiting their ability to fully utilize rich details in medical images. To address this issue, the FACB module proposed in the present application focuses on the decoding stage and uses additive operations to effectively aggregate features from multiple branches. This approach enables deeper exploration of intra-sample feature information, significantly enhancing feature compression and decoder feature expansion, thereby overcoming the limitations of existing methods.
[0005] In medical imaging, target objects vary greatly in size, shape and position, so effective use of multi-stage and multi-scale information is crucial to achieve accurate segmentation. However, many existing models fail to fully utilize the complementary advantages of multi-scale features in the information fusion process, thus limiting the accuracy of segmentation. To solve this problem, the present invention introduces the MSDB module, which utilizes parallel multi-core convolution operations to process grouped and cross-scale features, playing a key role in the bridging stage. This greatly enhances the model's multi-scale feature representation of the lesion area, improving the accuracy of information transmission and spatial context awareness. By effectively addressing the challenges of multi-scale information utilization, this innovative design provides a more efficient solution for advancing medical image segmentation. SUMMARY
[0006] In order to achieve the above-mentioned purpose of the present invention, the present invention provides an asymmetric-based lightweight medical image segmentation network (ABUNet), comprising the following steps:
[0007] S1, encoding stage: through the feature subtraction convolution block (FSCB), the input medical image is extracted by feature subtraction to obtain high-frequency information, effectively enhancing edge details and textures while suppressing redundant low-frequency components, thereby improving feature discriminability. Its core formula is:
[0008] X sub =GELU(X f1 )-X f2
[0009] Where X f1 and X f2 are the outputs of parallel 1x1 convolution layers, X sub is the result of subtraction operation, and X f1 is first passed through the GELU activation function, effectively introducing nonlinearity, allowing the model to better capture complex feature relationships. Subsequently, combined with the SE module, the expression ability of key features is further enhanced, allowing the model to focus more on important feature information, thereby improving the overall feature representation effect.
[0010] S2, bridging stage: adopt multi-scale deep convolution block (MSDB), capture point, local and global features through parallel 1x1, 3x3, 5x5 deep convolution, sum the outputs after batch normalization (BN) and GELU activation, the core formula is:
[0011]
[0012] Where DWConv kdenotes a depth convolution with kernel size k x k, and the MSDB module receives features from different layers of the encoder and feedback signals from the decoder, facilitating information exchange and multi-scale feature fusion between the encoder and the decoder.
[0013] S3, decoding stage: The feature addition convolution block (FACB) adopts an addition-based fusion mechanism, which integrates multi-path feature representations to enhance the expression of target region information and promote the accurate reconstruction of boundaries and structural details. The core formula is:
[0014] Y add = GELU(Y f1 ) + Y f2
[0015] where Y f1 and Y f2 are the outputs of parallel 1 x 1 convolution layers, and X add is the result of performing addition operation. In this way, the FACB module can fuse feature information from different branches, fully utilize complementary multi-scale features, and thus enhance semantic consistency, helping the model more accurately recover the details of the target region. In addition, the FACB module also combines residual connection and SE module. The residual connection helps to optimize the gradient flow, prevent gradient vanishing, and speed up the model convergence. The SE module dynamically adjusts the weights of feature channels through channel attention mechanism, further optimizes the feature reconstruction process, so that the model can better focus on key features and improve overall performance.
[0016] S4, network integration: The FSCB module is integrated into the encoder part of the U-shaped architecture, the FACB module is integrated into the decoder part, and the MSDB module is used to replace the processing operation in each group to construct an asymmetric model ABUNet. This architecture extracts high-frequency information and reduces redundant low-frequency information in the encoding stage, recovers details in the decoding stage, and fuses multi-scale features in the bridge stage, achieving a balance between high segmentation accuracy and lightweight efficiency.
[0017] Further, the network described in the present application adopts an asymmetric channel configuration with six layers of encoder and decoder. Each level of the encoder and the decoder is provided with different channel numbers, specifically {8, 16, 24, 32, 48, 64}. This asymmetric design enables the network to gradually extract and compress features in the encoding stage, while gradually recovering feature details in the decoding stage, effectively balancing feature expression capability and computational complexity.
[0018] Further, the FSCB module and the FACB module further comprise a residual connection and an SE attention module. The residual connection ensures efficient transmission of information within the module, prevents gradient disappearance problem, and enables the network to perform deep learning more stably. The SE attention module dynamically adjusts the weight of the feature channel through the channel attention mechanism, enhances the expression of key features, and suppresses the interference of redundant features. In the FSCB module, the SE attention module helps to highlight the key features of the lesion area and improve the distinguishability of the features; in the FACB module, the SE attention module helps to retain important detail information during feature fusion and improve the feature reconstruction quality in the decoding stage.
[0019] Further, the present application adopts deep supervision to calculate the loss function at different stages, which is a weighted combination of cross-entropy loss and Dice loss, aiming to solve the class imbalance problem and improve the segmentation performance. The core formula is as follows:
[0020]
[0021] In summary, due to the adoption of the above technical solutions, the present application has the following advantages:
[0022] (1) Three key modules are innovatively designed. FSCB effectively improves the distinguishability of features by retaining edge information and texture details while reducing redundant low-frequency information. FACB can effectively aggregate multi-path features, optimize feature reconstruction and detail recovery in the decoding process. In addition, MSDB enhances multi-scale feature modeling of the lesion area, improves information transmission accuracy and spatial context understanding ability in the bridging stage, thereby improving the overall segmentation performance.
[0023] (2) An innovative asymmetric lightweight medical image segmentation network ABUNet is proposed.
[0024] (3) Comprehensive experiments were conducted on the ISIC2017 and ISIC2018 datasets. The results show that ABUNet significantly outperforms existing lightweight networks in terms of segmentation accuracy, fully verifying the efficiency of our method.
[0025] Additional aspects and advantages of the application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS
[0026] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description, taken in conjunction with the accompanying drawings, in which:
[0027] Figure 1 is a schematic diagram of the framework structure of the medical image segmentation of the present application based on deep learning.
[0028] Figure 2 is the comparison result between the parameter quantity (M) and the average intersection over union (mIoU, %) of different medical image segmentation methods and the method of the present application on the ISIC2017 dataset.
[0029] Figure 3 is the comparison result between the parameter quantity (M) and the average intersection over union (mIoU, %) of different medical image segmentation methods and the method of the present application on the ISIC2018 dataset.
[0030] Figure 4 is the visualization comparison chart of the segmentation results of different medical image segmentation methods and the method of the present application.
[0031] Figure 5 is the feature visualization atlas of the six encoding layer outputs of the baseline method and the method of the present application. DETAILED DESCRIPTION
[0032] Embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, in which the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the drawings are exemplary only, and are intended to explain the present application, and cannot be understood as a limitation of the present application.
[0033] Medical image segmentation faces dual challenges in algorithm design: one is to meet the real-time clinical needs of lightweight deployment, and the other is to achieve sub-pixel level segmentation accuracy in complex anatomical structures. Existing methods mostly increase the complexity of modules to improve performance, but this leads to an imbalance between precision and efficiency. Specifically, the following three points: 1. Feature redundancy and computational efficiency imbalance: the feature stacking strategy in the encoding stage is prone to produce a large number of low information channels, resulting in reduced computational resource utilization; 2. Coarse multi-scale feature fusion: traditional skip connection in the decoding stage makes it difficult to achieve fine registration of multi-scale features, limiting the practicality of the model; 3. Insufficient context modeling granularity: the bridging stage relies on single-scale convolution, making it difficult to capture the diverse morphological features of lesions, limiting the accuracy of information transmission.
[0034] To solve these problems, we propose a lightweight medical image segmentation network based on asymmetry and its implementation method.
[0035] 1. RELATED WORK
[0036] 1.1 Symmetric segmentation architecture
[0037] The classic UNet introduces a symmetric encoder-decoder structure, laying the foundation for medical image segmentation. However, this rigid symmetry leads to two fundamental limitations.
[0038] First, there is an inherent conflict between the information loss caused by down-sampling in the encoder and the feature recovery based on linear interpolation in the decoder. Although UNet++ attempts to alleviate this issue through nested skip connections, its parameters increase by 9.7 times compared to the original UNet, significantly increasing the computational complexity.
[0039] Second, the equal-channel design in the symmetric structure produces a large number of inefficient intermediate features. Although AttentionUNet improves feature quality by introducing a channel attention mechanism, it does not fundamentally break the symmetry of the architecture. This structural symmetry constraint introduces existing improvements into the "local optimization, global inefficiency" dilemma, making it difficult to achieve both computational efficiency and superior segmentation performance.
[0040] 1.2 Lightweight Design
[0041] In recent years, lightweight segmentation networks have developed along two main technical routes. The first is the innovation of design principles, such as MobileNets, ShuffleNets, and RepVGG. These models introduce various efficient lightweight design strategies, including depthwise separable convolution, inverted residual bottleneck, channel shuffle, and structure reparameterization. For example, UNeXt uses depthwise separable convolution to compress the parameter quantity to 0.1M, but at the cost of losing key spatial details. The second route focuses on dynamic routing mechanisms, such as DCNet, which uses deformable convolution to adaptively adjust the receptive field. However, this dynamic computation introduces an additional 21% delay, reducing real-time efficiency. MALUNet is designed under limited computational resources and integrates innovative attention modules to improve lightweight medical image segmentation. However, the combination of multiple modules increases the difficulty of coordination, leading to unstable segmentation performance on complex and irregular medical images. To address the high parameter quantity and computational burden of Transformer-based architectures, EGE-UNet introduces GHPA and GAB modules, significantly reducing model complexity. However, its reliance on feature grouping and attention mechanisms limits its effectiveness in homogeneous medical images while minimizing feature contrast.
[0042] In contrast, ABUNet redefines lightweight segmentation from the perspective of feature operations, introducing a new asymmetric mechanism that more accurately allocates computational resources. In the encoding phase, FSCB performs subtraction operations to actively compress the feature space, reducing redundancy and improving computational efficiency. In the decoding phase, FACB uses addition operations to selectively enhance key features, improving feature reconstruction. This asymmetric architecture optimizes the allocation of computational resources, achieving high segmentation accuracy and lightweight efficiency in medical image analysis.
[0043] 1.3 Multi-scale Modeling
[0044] The quality of multi-scale feature fusion directly affects the accuracy of lesion boundary localization. TransFuse enhances global perception by adopting a dual-branch Transformer-CNN hybrid architecture, but this multi-modal computation increases GPU memory consumption by 72%, significantly limiting scalability. EGE-UNet introduces gated hierarchical attention, improving information flow modeling through feature grouping. However, its grouping strategy leads to insufficient feature interaction, reducing segmentation accuracy.
[0045] To address these limitations, ABUNet introduces MSDB, utilizing a parallel multi-core deep convolutional architecture. This design enables the simultaneous extraction of 1x1, 3x3, and 5x5 multi-scale features in a single computational stream, eliminating redundant steps in traditional serial processing. Theoretical analysis shows that this approach improves computational efficiency by times compared to traditional serial architectures (where n represents the number of convolutional kernels). This advancement significantly improves multi-scale feature extraction efficiency while maintaining low computational overhead, effectively breaking the bottleneck in multi-scale modeling and providing more accurate lesion boundary definition in medical image segmentation.
[0046] 2. Method
[0047] 2.1 Feature Subtraction Convolutional Block
[0048] In medical image segmentation tasks, feature processing in the encoding stage plays a crucial role in determining the final segmentation accuracy of the model. Traditional feature processing methods often rely on feature stacking, which can lead to excessive redundant information, increase computational burden, and hinder the extraction of key features, ultimately affecting segmentation accuracy. To address this issue, the present invention proposes a feature subtraction convolutional block (FSCB).
[0049] The workflow of FSCB begins with a deep convolution operation. When the input medical image feature map enters the module, the deep convolution extracts local spatial features while maintaining the number of channels, laying a solid foundation for subsequent feature selection and enhancement. Next, two parallel 1x1 convolutional layers operate simultaneously, with the primary task being to expand the feature dimension. In FSCB, these 1x1 convolutional layers expand the feature dimension to twice the original size, effectively increasing the model's ability to learn complex features. This design enables the model to capture a more comprehensive range of feature information, significantly improving its representation ability.
[0050] After obtaining multi-dimensional features, FSCB enters the core feature processing stage. The output of one 1x1 convolutional layer is passed through a GELU[8] activation function, which effectively introduces nonlinearity and enhances the model's ability to capture complex feature relationships. Then, the activated features are subjected to a subtraction operation with the output of the other 1x1 convolutional layer, formally defined as:
[0051] Xsub = GELU(X f1 ) - X f2
[0052] where X f1 and X f2 are the outputs of parallel 1×1 convolution layers, and X sub is the result of subtraction operation. This subtraction operation is crucial as it selectively suppresses redundant features, significantly improving feature purity. This enables the model to focus more effectively on key segmentation-related features, resulting in a more compact and efficient feature representation.
[0053] To ensure training stability and efficiency, FSCB introduces residual connections. Specifically, the original input feature maps are converted through a 1×1 convolution and added to the main feature output, preserving the original information and preventing the loss of important details during complex processing; optimizing gradient flow, accelerating model convergence.
[0054] Additionally, at the end of the module, an SE (Squeeze-and-Excitation) block is integrated to further enhance feature expression capabilities. The SE block performs global average pooling to extract global feature information, then learns the inter-channel dependency through a fully connected layer, generating adaptive channel attention weights. This process dynamically optimizes the feature map, thereby improving feature quality and representation ability.
[0055] 2.2 Feature Addition Convolution Block
[0056] In the decoding stage of medical image segmentation, accurate fusion and reconstruction of features are crucial for restoring the fine details of the target region. However, traditional skip connection methods face limitations in aligning multi-scale features, leading to loss of detail information and insufficient accuracy of segmentation results. To address these challenges, we propose the Feature Addition Convolution Block (FACB) to achieve more effective feature fusion and reconstruction.
[0057] The structure of FACB is similar to FSCB, also starting with deep convolution. This step extracts local spatial features, laying the foundation for subsequent processing. Next, two parallel 1×1 convolution layers operate to expand the feature dimension. Unlike FSCB, FACB expands the feature dimension by 3 times. In addition, the scaling factor in the encoding stage is different from that in the decoding stage to adapt to the feature fusion and reconstruction needs of the decoder. This asymmetric design ensures that the decoder can effectively merge encoder features while accurately restoring lesion details, improving segmentation accuracy.
[0058] The core operation of the FACB module is addition. A 1×1 convolution output is introduced through GELU activation to introduce nonlinearity, then added to the output of another 1×1 convolution layer. This process is mathematically represented as:
[0059] Y add = GELU(Y f1 ) + Y f2
[0060] where Y f1 and Y f2 are the outputs of two 1x1 convolutional layers, and Y add is the result of the addition operation. This addition process effectively fuses multi-path feature information, fully utilizes complementary multi-scale features, enhances semantic consistency, and helps the model more accurately recover the details of the target region.
[0061] Similar to FSCB, the processed feature map passes through a deep convolutional layer and a 1x1 convolutional layer. The deep convolution further refines the features, while the 1x1 convolution adjusts the channel dimension to adapt to the processing flow of the model and improves the computational efficiency. In addition, FACB also integrates residual connections, and the original input feature map is added to the main feature output after being converted by a 1x1 convolution. At the end of the module, FACB also integrates an SE block to analyze global feature information to adaptively adjust channel weights, enhance important features, and suppress less relevant features. This further improves the quality and representation ability of the features, providing strong support for accurate segmentation results.
[0062] The subtraction operation in FSCB and the addition operation in FACB form a complementary mechanism. FSCB enhances feature purity by selectively suppressing redundant features, while FACB improves semantic consistency by integrating multi-path features. Through collaborative work, these two modules optimize feature processing in the encoding and decoding stages, significantly improving overall segmentation performance.
[0063] 2.3 Multi-scale deep convolutional block
[0064] Lesion regions in medical images exhibit diverse morphologies, with significant differences in size, shape, and texture features. In the segmentation process, single-scale convolutional operations often cannot fully capture these complex features, limiting the segmentation accuracy. To enhance the model's ability to capture multi-scale features of lesion regions, the invention proposes a multi-scale deep convolutional block (MSDB). This module plays a key role in the bridging part of ABUNet, utilizing parallel multi-scale convolutional operations to improve lesion region modeling. By providing rich multi-scale feature representations, MSDB significantly enhances the decoder's ability to reconstruct fine details, improving overall segmentation performance.
[0065] The bridging component of ABUNet sits between the encoder and decoder, its main task being to facilitate information exchange and multi-scale feature fusion between them. The MSDB module receives features from different layers of the encoder, as well as feedback information from the decoder. These features contain information at different scales, but directly using them for information transfer and fusion often yields poor results. To address this issue, MSDB employs a parallel multi-core deep convolutional architecture, combining 1×1, 3×3, and 5×5 deep convolutional layers, enabling simultaneous processing of input feature maps within the same computational flow. Each convolutional layer has a unique role: the 1×1 deep convolution captures point-like features and facilitates inter-channel information interaction, adjusting channel information without altering spatial dimensions; the 3×3 deep convolution captures local features, its moderate receptive field effectively detecting edges and small-scale patterns in the feature map; the 5×5 deep convolution has a wider receptive field, capturing larger-scale features, including the overall shape of the target region and global patterns.
[0066] Following the convolutional operation, each deep convolutional layer undergoes batch normalization (BN) and GELU activation. BN normalization normalizes the convolutional outputs, helping to stabilize the training process and enhance the model's generalization ability. The GELU activation function introduces non-linearity, enhancing the model's ability to represent complex features and better adapt to the diversity of medical images. Finally, the outputs of all deep convolutional layers are summed element-wise, enabling the model to integrate multi-scale feature information, including point-like, local, and global features. This fusion strategy significantly enhances the model's ability to capture diverse morphological features of lesion regions and improves the accuracy of information transfer between the encoder and decoder. By providing the decoder with rich multi-scale contextual information, the MSDB module optimizes spatial context capture, playing a crucial role in improving the accuracy of medical image segmentation.
[0067] 2.4ABUNet
[0068] The FSCB and FACB modules of this invention are integrated into the encoder and decoder parts of the U-shaped architecture, respectively, and the processing operations in each group are replaced with the MSDB module of this invention, resulting in the asymmetric model ABUNet, as shown below. Figure 1 As shown.
[0069] The proposed ABUNet employs a six-layer encoder and decoder, forming an asymmetric U-shaped architecture. Given an input image... Where C, H, and W represent the number of channels, height, and width of a 2D medical image, respectively. ABUNet first applies 3×3 convolutions for shallow feature extraction, generating shallow feature maps. where 8 represents the number of channels in the first convolutional block. It directly extracts key features and reduces redundant information accumulation. ABUNet integrates five consecutive FSCB modules in the encoding stage. In the first three encoding blocks, the number of channels increases by 8 per layer while the resolution is halved. Therefore, in the ith encoding block, the deep feature representation extracted by ABUNet is denoted as where i ∈ {1, 2, 3}. In the last two encoding blocks, the number of channels increases by 12 per layer while the resolution continues to be halved. However, in the final encoding block, the resolution remains unchanged, and the deepest feature map is finally generated In the decoding stage, the main goal is feature reconstruction and detail recovery. ABUNet adopts an asymmetric design because the redundant features have been significantly reduced in the encoding process through feature differentiation processing. Therefore, the decoding stage directly utilizes the FACB module for feature reconstruction while adding the features extracted by the multi-scale bridging module to the corresponding decoder features. This approach reduces the semantic gap between the encoder and decoder features while alleviating the information loss caused by downsampling, ultimately improving the segmentation accuracy.
[0070] Notably, ABUNet does not employ any complex attention mechanism. Instead, it only integrates a simple SE attention mechanism in FSCB and FACB to optimize channel dependency. Despite this, ABUNet still achieves state-of-the-art segmentation performance on the ISIC2017 and ISIC2018 datasets, demonstrating the efficiency and effectiveness of its architectural design.
[0071] 3. Experiment
[0072] 3.1 Dataset and Experimental Settings
[0073] To evaluate the effectiveness of the proposed method, we conducted method validation based on ISIC2017 and ISIC2018 benchmark dermoscopy image datasets, which were released by the International Skin Imaging Collaboration (ISIC). These two authoritative datasets are widely used in segmentation studies of skin lesion images and provide high-quality expert-annotated segmentation masks. Specifically, the ISIC2017 dataset contains 2150 annotated images, while the ISIC2018 dataset contains 2694 expert-annotated skin lesion segmentation masks. To ensure the rigor and comparability of the experimental design, we followed the data partitioning strategy of previous studies and adopted a random partitioning method to divide each dataset into training and testing sets at a 7:3 ratio. Therefore, the ISIC2017 dataset was divided into 1500 training samples and 650 testing samples, while the ISIC2018 dataset consisted of 1886 training samples and 808 testing samples. It is emphasized that comprehensive comparative experiments were conducted on both datasets to ensure the robustness and generalization ability of the model. In addition, we also selected the larger ISIC2018 dataset for ablation studies to analyze the effectiveness of each model component and its contribution to the overall segmentation performance.
[0074] 3.2 Implementation details
[0075] ABUNet was implemented using the PyTorch framework with a six-layer encoder and decoder forming an asymmetric network structure. The number of channels at each level was {8, 16, 24, 32, 48, 64}. All experiments were conducted on a single NVIDIA RTX 4090 GPU to ensure efficient training and inference. The model used AdamW as the optimizer with an initial learning rate of 0.001, and CosineAnnealLR as the learning rate scheduler with a maximum number of iterations of 50 and a minimum learning rate reduced to 1e-5. The model was trained for a total of 300 epochs with a batch size of 8. During the experiment, all input images were normalized and resized to 256x256 resolution. In addition, various data augmentation techniques were applied, including random flipping, horizontal flipping, and vertical rotation, to enhance the generalization ability of the model. In this study, we adopted deep supervision to calculate the loss function at different stages, which is a weighted combination of cross-entropy loss and Dice loss, aiming to address class imbalance issues and improve segmentation performance.
[0076]
[0077] where λ i is the weight of each stage, i is the stage index, and the values of λ i are set to {1, 0.5, 0.4, 0.3, 0.2, 0.1} for i ∈ {0, 1, 2, 3, 4, 5}.
[0078] To comprehensively evaluate the segmentation performance of ABUNet, we compare it with several commonly used medical image segmentation models that are widely used in medical image segmentation competitions. In addition, we adopt five key evaluation metrics to measure the segmentation quality, including mean intersection over union (mIoU), Dice similarity coefficient (DSC), accuracy (Acc), sensitivity (Sen), and specificity (Spe). Furthermore, to evaluate the computational cost, we report the model parameters (in millions) and the use of floating-point operations (GFLOPs) as a measure of computational complexity. It is important to note that all model parameter sizes and computational complexities are measured at an input resolution of 256x256 to ensure a fair comparison.
[0079] 3.3 Comparison Results
[0080] The experimental results table (1) shows that the proposed ABUNet achieves excellent segmentation performance and computational efficiency on both ISIC2017 and ISIC2018 datasets. As shown in Table 1, ABUNet requires only 0.157M parameters (98.0% less than UNet and 99.4% less than TransFuse), with a computational cost of only 0.115 GFLOPs, but achieves mIoU scores of 80.85% and 82.11% on ISIC2017 and ISIC2018, respectively, outperforming the second-best model TransFuse by 1.64% and 1.48%. In addition, ABUNet achieves a DSC score of 89.41% on ISIC2017 and 90.18% on ISIC2018, further verifying the effectiveness of its integrated segmentation module. It is worth noting that in the lightweight model comparison, ABUNet is comparable in parameter size to MALUNet (0.177M) and EGE-UNet (0.053M), while significantly improving segmentation accuracy. For example, on the ISIC2018 dataset, ABUNet surpasses MALUNet by 1.86% in mIoU, highlighting the advantages of its lightweight feature extraction asymmetric complementary mechanism. In addition, ABUNet performs well in balancing sensitivity (Sen) and specificity (Spe): on ISIC2017, its sensitivity reaches 89.70% (2.56% higher than TransFuse), while on ISIC2018, its specificity reaches 97.48% (0.76% higher than UNetxS), demonstrating its dual optimization in lesion boundary sensitivity and background false positive suppression.
[0081] As Figure 5As shown, we demonstrate the feature extraction results of different encoder layers for the baseline model (EGUNet) and the proposed model (ABUNet). The first row shows the original features of the input image, and each subsequent column shows the feature maps extracted from different encoder layers. The feature maps reflect the ability of the model to extract features at each layer. By comparing the feature maps of our method (ABUNet) with the baseline model, we can clearly observe the following feature representation capabilities: At deeper encoder layers (e.g., encoder 3 and beyond), the feature maps extracted by ABUNet show more detailed structural information, especially in the lesion area. In contrast, the feature maps of the baseline model are more blurred, losing some important local feature information. Detail capture: ABUNet more effectively captures higher-level semantic features at each layer, with more accurate identification of edges and details. In particular, in encoder 3, encoder 4, and encoder 5, ABUNet retains more structural information, while the feature maps of the baseline model appear sparse and not fine enough.
[0082] 3.4 Ablation Study
[0083] We conducted extensive ablation experiments to verify the effectiveness of the proposed modules and asymmetric architecture. In our experiments, the baseline model is based on EGUNet, which follows a six-stage U-shaped architecture with symmetric encoder-decoder components and utilizes deep supervision for layer-wise loss calculation. Within this framework, we analyze the contributions of three key modules, FSCB, FACB, and MSDB. As shown in Table 2, we first integrate MSDB into the bridge module, replacing the standard convolution in each group of the baseline model. MSDB captures cross-scale features using multi-scale fusion, providing rich multi-scale contextual information to the decoder, significantly enhancing the decoder's ability to reconstruct fine details and improving overall segmentation performance.
[0084] Next, we replace the 2nd to 6th stages of the baseline encoder with FSCB and the 2nd to 6th stages of the baseline decoder with FACB, forming an asymmetric network architecture. FSCB introduces a feature subtraction operation, emphasizing contrast and difference feature enhancement in the encoding stage, improving feature discrimination ability while selectively suppressing redundant features to enhance feature purity. On the other hand, FACB adopts a feature addition operation, enhancing semantic consistency by fusing multi-path features, which is crucial for restoring fine details in target regions. Through collaborative work, FSCB and FACB achieve more effective feature processing in the encoding and decoding stages, significantly improving overall segmentation performance.
[0085] Based on the ablation experiments on the ISIC2018 dataset, we determined the best architecture of ABUNet. This study aims to achieve the best balance between the parameter size, computational complexity, and segmentation performance, and design a lightweight medical image segmentation model to achieve the best results. Specifically, according to the ablation results in Table 3, the Dropout rate in the encoder-decoder module is set to 0.2. In addition, in the FSCB and FACB, the expansion size of 2 and 3 is selected respectively, and the SE attention mechanism is combined in both modules to achieve the best performance. The experimental results in Table 4 further show that when the MSDB expansion size is set to 1 and the multi-scale convolution kernel (1, 3, 5) is selected, the model achieves higher segmentation accuracy. In addition, the ablation results in Table 5 show that the highest performance can be achieved by using an asymmetric architecture, in which FSCB is used for the encoding stage and FACB is used for the decoding stage.
[0086] Through these experiments, we have determined the best module parameters and architecture of ABUNet. To further verify the effectiveness of the asymmetric "shrink-expand" mechanism in lightweight models, we present the performance comparison of ABUNet under different channel configurations in Table 6. The experimental results on two datasets show that as the channel configuration decreases, the model achieves better performance, further confirming the superiority of ABUNet in lightweight medical image segmentation tasks.
[0087] Table 1 Experimental results of ABUNet and other seven different methods on ISIC2017 and ISIC2018 datasets
[0088]
[0089] Table 2 Influence analysis of model configuration on segmentation performance
[0090]
[0091] Table 3 Detailed ablation analysis of FCSB and FACB modules: (a) Dropout regularization parameter adjustment, where D-I does not apply Dropout, D-II sets the Dropout rate of FSCB and FACB modules to 0.1, D-III to 0.2, and D-IV to 0.3; (b) MLP ratio adjustment, where M-I sets the MLP ratio of FSCB and FACB to 1, M-II to 2, M-III to 3, M-IV sets the MLP ratio of the encoder to 1 and the decoder to 2, and M-V sets the MLP ratio of the encoder to 2 and the decoder to 3.
[0092]
[0093] Table 4 Detailed ablation analysis of MSDB module, including (a) multi-scale kernel parameter adjustment, where K-I uses 1,3 scale kernel, K-II uses 3,5 scale kernel, K-III uses 5,7 scale kernel, K-IV uses 1,3,5 scale kernel, K-V uses 3,5,7 scale kernel, K-VI uses 1,3,5,7 scale kernel; and (b) expansion factor parameter adjustment, where E-I sets the expansion factor to 1, E-II sets the expansion factor to 2, E-III sets the expansion factor to 3, to evaluate the impact of different expansion strategies on the performance of MSDB module.
[0094]
[0095] Table 5 Ablation analysis of symmetric and asymmetric encoder-decoder architecture configurations: (a) symmetric configuration: S-I encoder and decoder both use FSCB module; S-II encoder and decoder both use FACB module; (b) asymmetric configuration: A-I encoder uses FSCB module, decoder uses FACB module; A-II encoder uses FACB module, decoder uses FSCB module; A-III encoder uses FACB module, decoder keeps baseline design; A-IV decoder uses FACB module, encoder keeps baseline design.
[0096]
[0097]
[0098] Table 6 Performance comparison of ABU Net under different channel configurations. Among them, S corresponds to channel configuration {8, 16, 24, 32, 48, 64}, M corresponds to channel configuration {16, 32, 64, 128, 160, 256}, L corresponds to channel configuration {16, 32, 64, 128, 256, 512}, aiming to evaluate the impact of different channel sizes on segmentation performance.
[0099]
Claims
1. A lightweight medical image segmentation network based on asymmetry (ABUNet) and its implementation method, characterized in that, Includes the following steps: S1, Encoding Stage: The input medical image is processed by Feature Subtraction Convolutional Block (FSCB) to extract high-frequency information through feature subtraction. This effectively enhances edge details and texture while suppressing redundant low-frequency components, thereby improving feature discriminability. Its core formula is: X sub =GELU(X f1 )-X f2 Among them, X f1 and X f2 X is the output of a parallel 1×1 convolutional layer. sub X is the result of the subtraction operation. f1 First, the model is activated by the GELU function, which effectively introduces non-linearity, enabling it to better capture complex feature relationships. Then, the SE module is used to further enhance the expressive power of key features, allowing the model to focus more on important feature information, thereby improving the overall feature representation performance. S2, Bridging Stage: Multi-scale deep convolutional blocks (MSDB) are used to capture point-like, local, and global features through parallel 1×1, 3×3, and 5×5 depthwise convolutions. The output is summed after batch normalization (BN) and GELU activation. The core formula is: Among them, DWConv k This represents a depthwise convolution with a kernel size of k×k. The MSDB module receives features from different layers of the encoder and feedback signals from the decoder, facilitating information exchange and multi-scale feature fusion between the encoder and decoder. S3, Decoding Stage: The Feature Additive Convolutional Block (FACB) employs an additive fusion mechanism. By integrating multi-path feature representations, it enhances the expression of target region information and facilitates the accurate reconstruction of boundaries and structural details. Its core formula is: AND add =GELU(Y f1 )+Y f2 Among them, Y f1 and Y f2 X is the output of a parallel 1×1 convolutional layer. add This is the result of the scoring operation. In this way, the FACB module can fuse feature information from different branches, fully utilizing complementary multi-scale features to enhance semantic consistency and help the model more accurately recover details of the target region. Furthermore, the FACB module combines residual connections and the SE module. Residual connections help optimize gradient flow, prevent gradient vanishing, and accelerate model convergence. The SE module, through a channel attention mechanism, dynamically adjusts the weights of feature channels to further optimize the feature reconstruction process, enabling the model to better focus on key features and improve overall performance. S4, Network Integration: The FSCB module is integrated into the encoder part of the U-shaped architecture, the FACB module is integrated into the decoder part, and the processing operations in each group are replaced with the MSDB module to construct the asymmetric model ABUNet. This architecture compresses features and reduces redundancy in the encoding stage, restores details in the decoding stage, and fuses multi-scale features in the bridging stage, achieving a balance between high segmentation accuracy and lightweight efficiency.
2. The lightweight medical image segmentation network (ABUNet) based on asymmetry and its implementation method according to claim 1, characterized in that, The FSCB module further includes: • Introducing residual connections preserves the original information by adding the input features after a 1×1 convolution to the main path output; • Integrates the SE attention module, which generates channel weights through global average pooling and fully connected layers, and dynamically adjusts the weights of feature channels.
3. The lightweight medical image segmentation network (ABUNet) based on asymmetry and its implementation method according to claim 1, characterized in that, During the bridging phase, the MSDB module receives multi-level features from the encoder and feedback information from the decoder. It extracts point-like, local, and global features through parallel multi-kernel convolution and fuses them into a multi-scale contextual representation.
4. The lightweight medical image segmentation network (ABUNet) based on asymmetry and its implementation method according to claim 1, characterized in that, The encoder and decoder of the network adopt an asymmetric channel configuration. The number of encoder channels is {8, 16, 24, 32, 48, 64}, and the decoder expands the feature dimension through the FACB module with an expansion factor of 3.
5. The lightweight medical image segmentation network (ABUNet) based on asymmetry and its implementation method according to claim 1, characterized in that, The loss function of the network is a weighted combination of cross-entropy loss and Dice loss, as shown in the formula: Where, λ i Here, λ represents the weight of each stage, where i is the stage index and λ is the weight. i The values are set to {1, 0.5, 0.4, 0.3, 0.2, 0.1}, for i ∈ {0, 1, 2, 3, 4, 5}.
6. The network and method according to claim 1, characterized in that, The network has a total of 0.157M parameters, a computational complexity of 0.115GFLOPs, an input image resolution of 256×256, an optimizer of AdamW, an initial learning rate of 0.001, uses CosineAnnealLR as the learning rate scheduler, sets the maximum number of iterations to 50, trains the model for a total of 300 epochs, and has a batch size of 8.
7. A medical image segmentation system, characterized in that, include: · Encoder: Composed of multiple FSCB modules, used to enhance edge details and textures while suppressing redundant low-frequency information; • Bridging module: Composed of MSDB modules, used for effective feature fusion and reconstruction; • Decoder: Composed of multiple FACB modules, used to facilitate accurate reconstruction of boundaries and structural details; • Output layer: Generates the final segmentation mask.
8. The system according to claim 7, characterized in that, The encoder and decoder transmit multi-level features through skip connections. In the skip connections, the feature maps are adjusted in channel dimension by 1×1 convolution and then input to the corresponding stage of the decoder.
9. The method according to claim 1, characterized in that, The method supports deployment on low-power devices and is suitable for real-time medical image segmentation of skin lesions and tumors.
Citation Information
Cited By
Esophageal CT image segmentation method based on asymmetric perception enhancement
CN121304704A
Real-time lung CT segmentation method and device based on spatial frequency guided network
CN121391892A
A real-time lung CT segmentation method and device based on a spatial frequency guide network
CN121391892B