Skin lesion image segmentation method and device based on boundary enhancement dual optimization
By introducing a boundary-segment bidirectional enhancement framework and a boundary-aware dynamic sparse attention module in the skin lesion segmentation model, BES-UNet solves the balance between boundary accuracy and regional consistency of lightweight models, and has achieved significant improvements in computing efficiency and segmentation accuracy.
Patent Information
- Application Number
- CN202510543657.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
Existing lightweight skin lesion segmentation models have challenges in balancing boundary accuracy and overall regional consistency, while it is difficult to make a reasonable trade-off between computational efficiency and segmentation accuracy.
A skin lesion image segmentation method based on the boundary-enhanced sparse attention network (BES-UNet) is proposed. By constructing a boundary-segment bidirectional enhancement framework and a boundary-perceived dynamic sparse attention module, the precise segmentation and calculation efficiency of the skin lesion area are achieved.
Through the innovative design of the boundary-segment bidirectional enhancement framework and the boundary-perceived dynamic sparse attention module, BES-UNet significantly improves the segmentation accuracy of skin lesions, while reducing model parameters and calculation complexity, achieving efficient skin lesions segmentation.
Smart Images

Figure CN120070478A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a skin lesion image segmentation method and device based on boundary enhancement double optimization, belonging to the technical fields of medical image processing and computer vision. Background Art
[0002] In recent years, the incidence of skin cancer has been showing a rapid upward trend and has become a major global public health problem. Early diagnosis is crucial for improving the survival rate of patients, and accurate skin lesion segmentation is a key link in computer-aided diagnosis systems, laying the foundation for subsequent classification and treatment planning.
[0003] Deep learning models have made significant progress in the field of medical image segmentation. U-Net and its variants have become the mainstream architectures in medical image segmentation with their effective encoder-decoder structure and skip connection design. With the rise of Vision Transformer (ViT), TransUNet combines CNN and ViT, and TransFuse adopts a dual-path structure to capture local and global features simultaneously. However, these Transformer-based models have a large number of parameters and high computational complexity, making it difficult to be deployed on resource-constrained devices such as intelligent dermoscopes.
[0004] The academic community has proposed a variety of lightweight skin lesion segmentation models to solve the above problems. UNeXt significantly reduces the number of parameters by combining MLP blocks with U-Net; MALUNet further reduces the model size by reducing the model channels and introducing an attention mechanism; EGE-UNet proposes a group multi-axis Hadamard product attention module and a group aggregation bridging module, controlling the parameters at about 50KB; LB-UNet enhances the model's perception ability of lesion boundaries through a boundary assistance module.
[0005] However, lightweight segmentation models still face two key challenges: (1) how to balance boundary accuracy and overall region consistency; (2) how to make a reasonable trade-off between model computational efficiency and segmentation accuracy. These problems are particularly prominent in skin lesion segmentation because lesions often have irregular edges, blurred contours, and diverse texture features. In clinical applications, accurately identifying the boundaries of skin lesions is crucial for subsequent diagnosis and decision-making; at the same time, real-time segmentation on resource-constrained platforms such as mobile devices requires a model with extremely high computational efficiency.
[0006] The prior art generally has the following deficiencies: First, traditional segmentation methods adopt the same processing strategy for boundaries and internal regions, resulting in insufficient boundary accuracy; Second, existing lightweight strategies often sacrifice segmentation performance in exchange for improved computational efficiency; Third, the quadratic computational complexity of the standard self-attention mechanism limits its application in high-resolution medical images. These technical problems severely restrict the application effect of skin lesion segmentation algorithms in clinical practice. Summary of the Invention
[0007] In order to improve the accuracy of skin lesion segmentation and ensure computational efficiency at the same time, the present invention provides a skin lesion image segmentation method and device based on boundary enhancement dual optimization. The technical solutions are as follows: The first object of the present invention is to provide a skin lesion image segmentation method. The method realizes precise segmentation of skin lesion regions by constructing a boundary-enhanced sparse attention network. The boundary-enhanced sparse attention network includes: an encoder module, a decoder module, and a boundary-segmentation bidirectional enhancement framework; The encoder module includes encoders in 6 stages. Among them, the first 3 stages adopt standard convolutional blocks, the 4th and 5th stages adopt boundary-aware dynamic sparse attention modules, and the 6th stage adopts the boundary-aware dynamic sparse attention module and a max pooling layer; The boundary-segmentation bidirectional enhancement framework is connected to the output end of the encoder in each stage, and is also connected to the decoder at the corresponding level; The decoder module includes decoders in 6 stages, and adopts a symmetric structure to be connected to the corresponding layer of the encoder module through the boundary-segmentation bidirectional enhancement framework; The boundary-segmentation bidirectional enhancement framework includes a representation space adaptive transformation and boundary tolerance perception loss architecture, a multi-scale depth supervision architecture based on an adaptive training process, and a boundary-region bidirectional feature interaction and decoupling architecture; the representation space adaptive transformation and boundary tolerance perception loss architecture transforms boundary labels and predictions through a dynamic soft binary function and constructs a boundary tolerance perception loss function; the multi-scale depth supervision architecture based on an adaptive training process dynamically adjusts the supervision weights according to the training stage; the boundary-region bidirectional feature interaction and decoupling architecture realizes the collaborative optimization of boundary and region features through feature decoupling and an attention mechanism to obtain comprehensively optimized features; The boundary-aware dynamic sparse attention module includes: an adaptive dynamic sparsification strategy, a boundary-aware regularization and feature enhancement module, and a multi-scale shared memory space processing module; the adaptive dynamic sparsification strategy generates a dynamic sparsity rate according to feature complexity, the boundary-aware regularization and feature enhancement module maintains the priority of boundary information through a boundary detection branch, and the multi-scale shared memory space processing module captures multi-resolution context information through multi-scale memory units.
[0008] Optionally, the boundary tolerance-aware loss function is:
[0009] where represents the predicted boundary map, represents the ground-truth boundary label, is the Dice loss for the boundary region, is the balance coefficient, represents the gradient weight field, expressed as:
[0010] where is the boundary core region, is the boundary peripheral region, and are the weight coefficients for the boundary core region and the boundary peripheral region, respectively.
[0011] Optionally, the multi-scale deep supervision architecture based on the adaptive training process dynamically adjusts the supervision weights of each layer according to the current training step t :
[0012] where t is the current training step, , and represent the step nodes in the early, middle, and late stages of training, respectively, is the base weight; Combined with the dynamic weight adjustment strategy, the adaptive weighted composite loss is finally obtained:
[0013] where is the hierarchical weight coefficient, N is the number of supervision layers, is the boundary tolerance-aware loss of the i -th layer, represents the hierarchical segmentation loss.
[0014] Optionally, the calculation process of the boundary-region bidirectional feature interaction and decoupling architecture includes: First, separate the boundary and internal region information through the feature decoupling operation:
[0015] where is the segmentation prediction map of the current layer, is the boundary prediction map, Erode and Dilate are morphological erosion and dilation operations respectively, k is the kernel size, is based on P b the adaptive threshold based on statistical characteristics; Generate the boundary attention map:
[0016] Among them, is the boundary feature extracted from the segmentation prediction map through the Sobel operator Concat represents the concatenation in the channel dimension, is a 3×3 convolution, is the Sigmoid activation function; Calculate the comprehensive optimization feature:
[0017] Among them, F seg is the segmentation feature map of the current layer, is the boundary enhancement coefficient, is the regional consistency preservation coefficient.
[0018] Optionally, the calculation process of the adaptive dynamic sparsification strategy includes: Dynamically generate the optimal sparsity rate according to the image content characteristics:
[0019] Among them is the global feature summary, generated by the lightweight predictor, and are the minimum and maximum sparsity rates respectively; Generate the statistics-aware dynamic threshold:
[0020] Among them, and are the mean and standard deviation of the importance map respectively, and are learnable parameters; Construct a boundary-aware importance estimator, and process the input features in parallel through the main feature branch and the boundary detection branch: Among them, the main feature branch uses depthwise separable convolution to extract feature representations:
[0021] The boundary detection branch uses the Sobel operator and lightweight convolution to detect the boundary:
[0022] Fuse the features of the two branches to generate an importance map:
[0023] Generate the final sparse mask:
[0024] where is the boundary detection map, is the boundary importance threshold, m is the activation steepness parameter.
[0025] Optionally, the calculation process of the multi-scale shared memory space processing module includes: Group the input feature F into 3 groups evenly according to the channel dimension:
[0026] Input the first group and the third group into the multi-scale storage space. The multi-scale storage space is divided into two scales of 4×4 and 8×8, each focusing on the context information of a specific scale, and jointly constituting a multi-resolution feature representation system:
[0027] where represents the j th shared memory space, represents the feature transformation function:
[0028] Q , K , V respectively represent the query, key-value, and value transformation functions; Process the second group of features using depth convolution:
[0029] Apply the sparse mask to each feature:
[0030] Retain the boundary information through residual connection:
[0031] Feature recombination and channel shuffle recombine the processed three groups of features:
[0032] Finally, process through depthwise separable convolution:
[0033] Among them, represents the features finally output by the multi-scale shared memory space processing module.
[0034] Optionally, the boundary-aware regularization and feature enhancement module constructs a boundary importance regularization loss:
[0035] Where is the entropy of the attention mask, is the boundary region importance metric, and are weight coefficients; The boundary region importance metric is expressed as:
[0036] is the average importance of the boundary center region, is the boundary detection intensity, is the boundary strength threshold.
[0037] Optionally, the overall loss function for training the boundary-enhanced sparse attention network is:
[0038] Where represents the segmentation loss, represents the boundary loss, represents the region consistency loss, represents the importance regularization loss, represents the auxiliary supervision loss, , , respectively represent the weights of the boundary loss, region consistency loss, and importance regularization loss.
[0039] The second object of the present invention is to provide a skin lesion image segmentation device, including a memory and a processor; The memory is used to store a computer program; The processor is used to implement the skin lesion image segmentation method as described in any one of the above when executing the computer program.
[0040] The third object of the present invention is to provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the skin lesion image segmentation method as described in any one of the above is implemented.
[0041] The beneficial effects of the present invention are as follows: The present invention realizes the precise segmentation of skin lesion areas by constructing a boundary-enhanced sparse attention network (BES-UNet). The advantages of the BES-UNet network are as follows: First, a boundary-segmentation bidirectional enhancement framework (BF) is proposed. This framework constructs a dynamic adaptive soft binarization function, breaking through the limitations of traditional fixed-threshold binarization. It realizes the precise conversion of boundary representation through adaptively calculated thresholds and temperature parameters. A boundary loss mechanism with a gradient weight field is constructed, fundamentally solving the problem of boundary uncertainty in medical images caused by differences in expert annotations. A theoretical framework for decoupling and recombining boundary and region features is proposed. Through the ingenious combination of boundary attention maps and internal region masks, a bidirectional promotion mechanism between boundary features and region features is established, effectively improving the accuracy of medical image boundary segmentation.
[0042] Second, the present invention proposes a boundary-aware dynamic sparse attention module (BDAM), constructs a dynamic computing resource allocation theory, and realizes the collaborative optimization of boundary enhancement and computing efficiency. This module generates an adaptive sparsity rate based on the complexity of image content, breaking through the limitations of traditional fixed sparsity rates. A dynamic threshold algorithm based on the statistical characteristics of importance distribution is designed to solve the problem of inconsistent sparsification effects of fixed thresholds on different images. A sparsification strategy for preferential protection of boundary regions is proposed, ensuring the complete retention of boundary information during the sparse calculation process at the algorithm level. The proposed BDAM module not only ensures the high accuracy of the attention mechanism at the algorithm level but also balances the computing efficiency, effectively improving the segmentation accuracy while reducing the computational complexity.
[0043] In an implementation manner of the present invention, a multi-component loss function and an adaptive supervision mechanism can be constructed to realize the multi-objective collaborative optimization of boundary accuracy, region consistency, and computing efficiency. Through the organic combination of segmentation loss, boundary loss, region consistency loss, importance regularization loss, and hierarchical auxiliary supervision loss, a complete model optimization theory system is established, ensuring the excellent performance of the model in all dimensions.
[0044] The experimental results prove that while maintaining high segmentation accuracy, the present invention achieves extreme compression of model parameters and computational complexity. Compared with LB-UNet, the number of parameters of BES-UNet is reduced from 0.038M to 0.032M (a decrease of about 15.8%), the computational volume is reduced from 0.098 GFLOPs to 0.089 GFLOPs (a reduction of about 9.2%), and the training time is shortened from 4 hours to about 2.5 hours (an acceleration of about 60%). This extremely lightweight design makes BES-UNet one of the most efficient skin lesion segmentation models at present, provides a technological innovation with great theoretical significance and application value for the field of medical image segmentation, and is particularly suitable for deployment in resource-constrained environments. Description of the Drawings
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0046] Figure 1 It is the overall structure diagram of the boundary-enhanced sparse attention network (BES-UNet) framework of the present invention.
[0047] Figure 2 It is the structure diagram of the boundary-segmentation bidirectional enhancement framework (BF) of the present invention.
[0048] Figure 3 It is the architecture diagram of the boundary-aware dynamic sparse attention module (BDAM) of the present invention.
[0049] Figure 4 It is the flowchart for constructing the boundary-enhanced sparse attention network (BES-UNet) framework of the present invention.
[0050] Figure 5 It is the visualization effect diagram of the segmentation results of different models on the ISIC2018 dataset.
[0051] Figure 6 It is the performance comparison result of different methods on the ISIC2017 and ISIC2018 datasets.
[0052] Figure 7 It is the analysis result of the segmentation accuracy of the boundary regions of different methods.
[0053] Figure 8 It is the ablation experiment result of the contributions of different components of the present invention.
[0054] Figure 9 It is the ablation experiment result of the boundary-aware dynamic sparse attention mechanism. Specific Embodiments
[0055] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the accompanying drawings.
[0056] Embodiment 1: This embodiment provides a precise skin lesion segmentation method based on boundary enhancement and dual optimization. By constructing a Boundary Enhancement Sparse Attention Network (BES-UNet) framework, precise segmentation of skin lesion regions is achieved. A boundary-segmentation bidirectional enhancement adaptive learning framework is designed, and through a boundary tolerance perception loss and a bidirectional feature interaction strategy, collaborative optimization of boundary precise positioning and region segmentation is realized. A boundary-aware dynamic sparse attention mechanism is designed. Through a content-adaptive dynamic sparsification strategy and a multi-scale shared memory space, the computational complexity of the model is significantly reduced, and at the same time, the recognition ability of lesion edges is enhanced.
[0057] The construction and training of the BES-UNet framework in this embodiment include the following processes: Step 1: Dataset preparation and preprocessing.
[0058] First, collect and organize the publicly available ISIC2017 and ISIC2018 skin lesion datasets required for the experiment. ISIC2017 contains 2,150 dermoscopic images, and ISIC2018 contains 2,694 dermoscopic images. Each image has a segmentation mask annotated by experts. Randomly divide the training set and the test set according to a ratio of 7:3 to ensure the balance of data distribution. Standardize all images, including pixel value normalization (subtract the mean and divide by the standard deviation) and size adjustment (unify to a resolution of 256×256 pixels). To enhance the generalization ability of the model, design a data augmentation strategy, including random horizontal and vertical flipping (probabilities are both 0.5), random rotation (±15 degrees), random brightness and contrast adjustment (±0.2), etc., to enrich the diversity of training samples.
[0059] Step 2: Construct the overall architecture of BES-UNet and initialize it.
[0060] As Figure 1 shown, BES-UNet includes: an encoder, a decoder, a boundary-segmentation bidirectional enhancement framework, and a boundary-aware dynamic sparse attention module.
[0061] The encoder consists of 6 stages, and the number of channels is successively {8, 16, 24, 32, 48, 64}. The first three stages adopt the standard Convolutional Block Chain (CBC), and the last three stages adopt the Boundary-aware Dynamic Sparse Attention Module (BDAM). Among them, the 4th and 5th stages adopt the BDAM module + the maximum pooling layer (Boundary-aware Dynamic Attention Pooling, BDAP). As Figure 1 shown, the input image is first processed by the CBC module in stage 1 to generate the feature image F1, and the boundary feature B1 and the region feature R1 are extracted through the Boundary-segmentation Bidirectional Enhanced Framework (BF); then the feature image F1 is downsampled by the ResidualBoundary Pooling (RBP) and input into the CBC module in stage 2 to generate the feature image F2, and at the same time the BF module extracts the boundary feature B2 and the region feature R2; similarly, the feature image F2 is downsampled by RBP and enters the CBC module in stage 3 to generate the feature image F3, and the BF framework extracts the boundary feature B3 and the region feature R3; the feature image F3 is then input into the BDAP module in stage 4 through RBP to generate the feature image F4, and the BF module extracts the boundary feature B4 and the region feature R4; similarly, the feature image F4 passes through the BDAP module in stage 5 to generate the feature image F5, and the feature image F5 finally passes through the BDAM module in stage 6 to generate the feature image F6.
[0062] The BF module is connected to the output end of the encoder in each stage, and is also connected to the decoder at the corresponding level. The output of the BF module includes the boundary features (B1 - B6) and the region features (R1 - R6), and these features are transmitted to the corresponding levels in the decoder for feature fusion.
[0063] The decoder adopts a symmetric structure and fuses with the corresponding layer of the encoder through skip connections. The top layer of the decoder upsamples the feature image F6 through the Upsampling Module (UP) and fuses it with the feature image F5, and at the same time combines the boundary feature B5 and the region feature R5 for boundary enhancement; each subsequent layer adopts a similar upsampling fusion structure to generate the D5, D4, D3, D2, and D1 features in turn. Finally, D1 generates the prediction mask P through a 1×1 convolution and the Sigmoid function.
[0064] When initializing the model parameters, He initialization is used for the convolutional layer, and Xavier initialization is used for the attention module and the linear layer to ensure a good starting point for training.
[0065] Step 3: Construct a boundary-segmentation bidirectional enhancement framework (BF module).
[0066] The structure of the boundary-segmentation bidirectional enhancement framework is as Figure 2 shown, mainly including a structure representing spatial adaptive transformation and boundary tolerance-aware loss, a multi-scale deep supervision architecture based on an adaptive training process, and a boundary-region bidirectional feature interaction and decoupling architecture.
[0067] Among them, the structure representing spatial adaptive transformation and boundary tolerance-aware loss obtains multi-scale features extracted by the BES-UNet encoder , and first converts them into boundary features through a 1×1 convolutional layer . The spatial adaptive transformation mechanism seamlessly transforms boundary labels and predictions through a dynamic adaptive soft binary function:
[0068] where represents the boundary probability value (0~1), τ represents the adaptive threshold, represents the temperature parameter; The adaptive threshold and the temperature parameter are represented by and respectively, and are dynamically adjusted according to the statistical characteristics of the input image :
[0069] where and represent the mean and standard deviation of the input image features respectively, and are learnable parameters, and this design ensures the complete transmission of the gradient flow and training stability.
[0070] The boundary filtering mechanism selectively retains significant boundary features through an adaptive gating unit, and its calculation process is as follows:
[0071] where and are the learning parameters of the gating unit, is the Sigmoid activation function. This mechanism can effectively filter out noise and weak boundaries, and only retain strong boundary information that is meaningful for the segmentation task.
[0072] The multi-scale boundary fusion module captures boundary information through dilated convolutions of different scales:
[0073] Among them, 、 and are 3×3 dilated convolutions with different dilation rates (1, 2, 4) respectively. Concat represents the concatenation operation in the channel dimension, is a 1×1 convolution. This design can expand the receptive field without increasing the computational complexity and capture boundary features of different scales.
[0074] To solve the problem of boundary uncertainty caused by differences in expert annotations in medical images, the present invention introduces a boundary tolerance-aware loss function:
[0075] Among them, represents the predicted boundary map, represents the ground truth boundary label, is the Dice loss of the boundary region, is the balance coefficient.
[0076] This loss uses a gradient weight field with a fine transition from the boundary core to the periphery :
[0077] Among them, is the boundary core region, is the boundary periphery region, and are the weight coefficients of the boundary core region and the periphery region respectively.
[0078] The multi-scale deep supervision architecture based on the adaptive training process obtains the feature maps of the outputs of each layer of the encoder and the features of each layer of the decoder , and first generates multi-scale segmentation predictions through the deep supervision layer on the decoder side:
[0079] Among them, is a 1×1 convolution operation, is the Sigmoid activation function. These multi-scale predictions are compared with the ground truth segmentation annotation to calculate the hierarchical segmentation loss:
[0080] Among them, BCE is the binary cross-entropy loss, Dice is the Dice coefficient loss, is the balance coefficient.
[0081] The training process adaptive weight regulator dynamically adjusts the supervision weights of each layer according to the current training step t as follows:
[0082] where t is the current training step, , and represent the time nodes in the early, middle, and late stages of training (set to 20%, 50%, and 80% of the total training steps), respectively, is the base weight (set to 0.8). This mechanism enables the model to pay more attention to the learning of overall regional features in the early stage of training and gradually increase the attention to boundary details as the training process progresses.
[0083] Combined with the multi-level weight allocation strategy, the adaptive weighted composite loss is finally obtained:
[0084] where is the hierarchical weight coefficient, which increases from 0.1 to 0.5 as the feature level is refined (for example = 0.1, = 0.2,..., = 0.5), N is the number of supervision layers (usually 5), is the i -th layer boundary tolerance perception loss application.
[0085] The boundary-region bidirectional feature interaction and decoupling architecture obtains the boundary-aware adaptive loss and the adaptive weighted composite loss , as well as the current feature and . First, the boundary and internal region information are separated through the feature decoupling operation:
[0086] where is the segmentation prediction map of the current layer, is the boundary prediction map, Erode and Dilate are morphological erosion and dilation operations respectively, k is the kernel size, is the adaptive threshold based on the statistical characteristics of Pb.
[0087] Then, the boundary attention map is generated:
[0088] where The boundary features extracted from the segmentation prediction map through the Sobel operator, Concat represents concatenation in the channel dimension, is a 3×3 convolution, and is the Sigmoid activation function.
[0089] Next, perform boundary-region bidirectional feature interaction calculation:
[0090] Among them, F seg is the segmentation feature map of the current layer, is the boundary enhancement coefficient, and
[0091] is the region consistency preservation coefficient. This mechanism allows boundary features and region features to promote each other rather than compete with each other.
[0092] Among them, and represent the mean and variance of the region respectively, and represent the internal regions of the predicted region and the ground truth region respectively, and are the weight coefficients for mean and variance consistency.
[0093] The architecture finally outputs the enhanced comprehensive optimized features , while extracting the boundary features B and the region features R :
[0094] Step 4: Construct a Boundary-Aware Dynamic Sparse Attention Module (BDAM).
[0095] The structure of the Boundary-Aware Dynamic Sparse Attention Module is as shown in Figure 3 , including an adaptive dynamic sparsification strategy, a boundary-aware regularization and feature enhancement module, and a multi-scale shared memory space processing module. The adaptive dynamic sparsification strategy is used to dynamically allocate computing resources according to the content complexity, the boundary-aware regularization and feature enhancement module is used to maintain the priority of boundary information, and the multi-scale shared memory space processing module is used to efficiently process context information at different scales.
[0096] First, the adaptive dynamic sparsification strategy dynamically generates the optimal sparsity rate according to the image content characteristics:
[0097] Among them is the global feature summary, generated by a lightweight predictor, and are the minimum and maximum sparsity rates respectively. This mechanism can automatically adjust the processing density according to the image complexity, using a high sparsity for simple regions and reserving more computing resources for complex regions.
[0098] Secondly, generate a statistic-aware dynamic threshold:
[0099] Among them, and are the mean and standard deviation of the importance map respectively, and are learnable parameters. This design enables the threshold to adapt to both the central tendency and the dispersion of importance, significantly improving the sparsification accuracy.
[0100] The boundary-aware importance estimator processes the input features in parallel through the main feature branch and the boundary detection branch. The main feature branch uses depthwise separable convolutions to extract feature representations:
[0101] The boundary detection branch uses the Sobel operator and lightweight convolutions to detect boundaries:
[0102] The feature fusion layer fuses the features of the two branches to generate an importance map:
[0103] The boundary enhancement mask process takes the dynamic threshold and the importance map output by the boundary-aware importance estimator and the boundary detection map to generate the final sparse mask:
[0104] Among them is the boundary detection map, is the boundary importance threshold, m is the activation steepness parameter.
[0105] The multi-scale shared memory space processing module first groups the input features, evenly dividing them into 3 groups along the channel dimension:
[0106] Input the first group and the third group into the multi-scale storage space. The multi-scale storage space is divided into two scales of 4×4 and 8×8, each focusing on the context information of a specific scale, and jointly constituting a multi-resolution feature representation system:
[0107] Among them represents the j th shared memory space, represents the feature transformation function:
[0108] Q , K , V represent the query, key-value, and value transformation functions respectively. This design reduces the number of parameters by about 60% and captures the structural information at different semantic levels.
[0109] Perform deep convolution on the second group of features:
[0110] In the mask application stage, apply the sparse mask to each feature:
[0111] The boundary enhancement process retains the boundary information through residual connections:
[0112] Feature recombination and channel shuffling recombine the three groups of processed features:
[0113] Finally, perform processing through depthwise separable convolution:
[0114] The boundary-aware regularization and feature enhancement module includes the boundary importance regularization loss:
[0115] Among them is the entropy of the attention mask, is the boundary region importance metric:
[0116] is the average importance of the boundary center region, is the boundary detection intensity, is the boundary strength threshold. This regularization mechanism ensures that the model does not ignore key boundary information while maintaining computational efficiency.
[0117] Step 5: Design an adaptive training strategy.
[0118] Design a gradient flow control strategy with an adaptive training process to dynamically adjust the optimization weights of boundary detection and region segmentation according to the training stages (early, middle, and late). Implement a multi-scale deep supervision mechanism, introduce boundary-aware auxiliary supervision signals at different decoding levels, and form a progressive optimization framework from coarse to fine. Construct a hierarchical weight allocation strategy to gradually increase the supervision weight as the feature hierarchy is refined (from 0.1 to 0.5). Design a comprehensive loss function, including segmentation loss, boundary loss, region consistency loss, importance regularization loss, and auxiliary supervision loss, to achieve multi-objective collaborative optimization. The overall loss function is:
[0119] Step 6: Model training configuration and optimization.
[0120] Configure the training environment, including hardware settings (AMD EPYC 9654 96-core processor, NVIDIA GeForce RTX3090 GPU) and software environment (PyTorch framework, Ubuntu 20.04 system). Set training hyperparameters, including optimizer (AdamW, initial learning rate 0.001, weight decay 0.01), learning rate policy (CosineAnnealingLR, maximum number of iterations 50, minimum learning rate 0.00001), batch size (8), and number of training epochs (300 epochs). Implement a training process monitoring mechanism, including real-time tracking of the loss function curve, validation set performance metrics, and resource utilization, to ensure the stability and effectiveness of the training process.
[0121] Step 7: Model training.
[0122] Load the preprocessed training data and apply data augmentation strategies to generate diverse training samples. Initialize the model parameters, including parameter settings for the encoder, decoder, boundary-segmentation bidirectional enhancement framework, and boundary-aware dynamic sparse attention module. Execute the iterative training process, including forward propagation, loss calculation, backward propagation, and parameter update. Apply the adaptive training strategy to dynamically adjust the weights of each loss component according to the training process to ensure the stability and effectiveness of model optimization. Regularly evaluate the model performance on the validation set, monitor the changing trends of boundary accuracy and overall segmentation accuracy, and adjust the training strategy in a timely manner.
[0123] Step 8: Performance evaluation and comparison.
[0124] Comprehensively evaluate the performance of BES-UNet on the test set using multiple evaluation metrics: mean intersection over union (mIoU), Dice similarity coefficient (DSC), number of parameters (Params), and floating-point operations (GFLOPs). Pay special attention to the segmentation accuracy of the boundary region, and calculate the F1 score, precision, and recall rate of the boundary region. Conduct a comprehensive comparison with existing lightweight skin lesion segmentation models, including methods such as UNeXt, MALUNet, EGE-UNet, and LB-UNet. Generate visualization graphs of the segmentation results to intuitively demonstrate the advantages of BES-UNet in terms of boundary accuracy and overall region consistency.
[0125] Step 9: Ablation experiments and mechanism verification.
[0126] Design and execute detailed ablation experiments to verify the effectiveness of each key component: Remove the boundary-aware dynamic sparse attention mechanism (BDAM) and the boundary-segmentation bidirectional enhancement framework (BEBIF) respectively, and evaluate their impact on the model performance. Conduct an in-depth analysis of the boundary-aware dynamic sparse attention mechanism, and evaluate the contributions of the three sub-components of dynamic sparsity rate, boundary enhancement, and shared memory space respectively. Explore the impact of different hyperparameter configurations on the model performance, including the range of sparsity rate, gradient weight field parameters, and loss function weights. Verify the scientificity and effectiveness of the proposed technical solution through detailed experimental data and analysis.
[0127] Step 10: Model optimization and deployment.
[0128] Based on the evaluation results and ablation experiments, further optimize the BES-UNet model, including structural fine-tuning, parameter configuration adjustment, and training strategy improvement. Optimize the inference efficiency of the model, including techniques such as model quantization, knowledge distillation, and computation graph optimization, to further reduce the computational complexity and storage overhead of the model. Convert the optimized model into a format suitable for mobile deployment (such as ONNX, TFLite, or Core ML), and conduct actual tests on mobile devices to verify its real-time performance in resource-constrained environments. Summarize the technical characteristics and application value of the model to provide technical support for subsequent practical applications in intelligent dermoscopes and mobile medical devices.
[0129] To illustrate the beneficial effects of the skin lesion precise segmentation method based on BES-UNet of the present invention, the following experimental data are provided, and the following evaluation metrics are used: Mean intersection over union (mIoU): Used to evaluate the segmentation accuracy, defined as the ratio of the intersection to the union of the predicted region and the ground truth region, and a higher value indicates better segmentation results.
[0130] Dice Similarity Coefficient (DSC): Measures the similarity between two sets. The calculation formula is 2×|X∩Y| / (|X|+|Y|), where X is the predicted region and Y is the ground truth region. The value range is [0,1], and a higher value indicates better segmentation results.
[0131] Number of Parameters (Params): The total number of parameters of the model, used to evaluate the model size, with the unit of million (M).
[0132] GigaFloating Point Operations per Second (GFLOPs): The number of floating point operations per second, used to evaluate the computational complexity of the model.
[0133] Boundary Region F1 Score: Specifically calculates the F1 score of the boundary region, used to evaluate the segmentation accuracy of the model in the boundary region.
[0134] The experimental environment is shown in Table 1: Table 1 Detailed Parameters for Network Model Training
[0135] To comprehensively evaluate the performance of BES-UNet, multiple representative methods were selected as comparison algorithms, mainly including UNet, MobileViTv2, MobileNetv3, UNeXt-S, MALUNet, EGE-UNet, and LB-UNet. All experiments were conducted under the same hardware environment and training settings to ensure the fairness of the evaluation. The experimental results are as Figure 6 shown.
[0136] From Figure 6 it can be seen that BES-UNet achieved 81.98% mIoU and 89.65% DSC on the ISIC2018 dataset. Compared with the current best method LB-UNet, the mIoU increased by 0.61 percentage points and the DSC increased by 0.11 percentage points, while the number of parameters decreased by 15.8%. Compared with UNeXt-S, the mIoU increased by 2.89 percentage points, the DSC increased by 1.32 percentage points, and the number of parameters decreased by 90%.
[0137] The visualization of the segmentation results of different models on the ISIC2018 dataset is as Figure 5 shown. From Figure 5 it can be clearly seen that the BES-UNet model performs significantly better than other comparison methods in the segmentation of the skin lesion boundary region.
[0138] In the first row of examples, BES-UNet can accurately capture the bifurcated structure of the lesion, while LB-UNet, EGE-UNet, and MALUNet have boundary missing phenomena in the regions marked by the red circles.
[0139] In the second row example, BES-UNet successfully retained the concave features of the lesion, while other models had inaccurate predictions at the upper left corner edge (the area marked by the red circle), and MALUNet even had significant segmentation errors in the right area (the area marked by the oval red circle).
[0140] In the third row example, BES-UNet was able to completely segment the irregular edges of the lesion, while LB-UNet and EGE-UNet had inaccurate boundary problems at the bottom edge (the area marked by the red circle), and MALUNet had serious boundary missing problems on both the left and right sides (the area marked by the red circle).
[0141] In the fourth row example, BES-UNet accurately segmented the complex boundaries at the bottom of the lesion, while other models had obvious boundary missing or misdetection problems in the bottom and right areas (the areas marked by the red circles). These visualization results fully demonstrate that the boundary-segmentation bidirectional enhancement framework and the boundary-aware dynamic sparse attention mechanism proposed in the present invention can significantly improve the model's perception ability of lesion boundaries, especially having obvious advantages when dealing with fuzzy, irregular, and complex boundary regions. Through the boundary tolerance perception loss and the adaptive feature interaction strategy, BES-UNet has successfully achieved the balance between boundary accuracy and region consistency, providing reliable technical support for the precise segmentation of skin lesions.
[0142] The accuracy analysis of the boundary region segmentation of different models is as Figure 7 shown. In the accuracy analysis of the boundary region segmentation, the BES-UNet of the present invention reached 60.17% in the F1 score of the boundary region, 1.05 percentage points higher than LB-UNet and 2.22 percentage points higher than EGE-UNet. This proves that the boundary-segmentation bidirectional enhancement framework and the boundary-aware dynamic sparse attention mechanism proposed in the present invention can significantly improve the segmentation accuracy of the model in the lesion boundary region.
[0143] The ablation experiment results of the contributions of different components are as Figure 8 shown. The results show that after adding the boundary-aware dynamic sparse attention mechanism (BDAM), the mIoU increased by 0.41 percentage points, the DSC increased by 0.07 percentage points, and the number of parameters decreased by 7.89%; after adding the boundary-segmentation bidirectional enhancement framework (BF), the mIoU and DSC increased by 0.20 and 0.04 percentage points respectively; compared with the baseline model, the complete BES-UNet model had the mIoU increased by 0.61 percentage points, the DSC increased by 0.11 percentage points, while the number of parameters decreased by 15.79% and the computational cost decreased by 9.18%.
[0144] The ablation experiment results of the boundary-aware dynamic sparse attention mechanism are asFigure 9 As shown, the dynamic sparsity rate increases the mIoU by 0.22 percentage points, verifying the effectiveness of the adaptive sparsity strategy; the boundary enhancement mechanism further increases the mIoU by 0.20 percentage points; the multi-scale shared memory space not only reduces the number of parameters from 0.044M to 0.032M and the computational volume from 0.105 to 0.089 GFLOPs, but also increases the mIoU by 0.19 percentage points, demonstrating that each component makes a significant contribution to performance improvement and model lightweighting.
[0145] In summary, the experimental results fully verify the superior performance of BES-UNet. Through the innovative design of the boundary-segmentation bidirectional enhancement adaptive learning framework and the boundary-aware dynamic sparse attention mechanism, the present invention realizes the dual optimization of boundary accuracy and computational efficiency, providing a new technical solution for the field of skin lesion segmentation.
[0146] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.
[0147] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A skin lesion image segmentation method based on boundary enhancement and dual optimization, characterized in that: The method realizes accurate segmentation of skin lesion area by constructing a boundary-enhanced sparse attention network, wherein the boundary-enhanced sparse attention network comprises: an encoder module, a decoder module, and a boundary-segmentation bidirectional enhancement framework; The encoder module includes a 6-stage encoder, wherein the first 3 stages use standard convolution blocks, the 4th and 5th stages use boundary-aware dynamic sparse attention modules, and the 6th stage uses the boundary-aware dynamic sparse attention module and a maximum pooling layer; The boundary-segmentation bidirectional enhancement framework is connected to the encoder output of each stage and is also connected to the decoder of the corresponding level; The decoder module includes a 6-stage decoder, which adopts a symmetrical structure and is connected to the corresponding layers of the encoder module through the boundary-segmentation bidirectional enhancement framework; The boundary-segmentation bidirectional enhancement framework includes a representation space adaptive conversion and boundary tolerance perception loss architecture, a multi-scale deep supervision architecture based on an adaptive training process, and a boundary-region bidirectional feature interaction and decoupling architecture; the representation space adaptive conversion and boundary tolerance perception loss architecture converts boundary labels and predictions through a dynamic soft binarization function and constructs a boundary tolerance perception loss function; the multi-scale deep supervision architecture based on an adaptive training process dynamically adjusts the supervision weight according to the training stage; the boundary-region bidirectional feature interaction and decoupling architecture realizes the coordinated optimization of boundary and region features through feature decoupling and attention mechanism to obtain comprehensive optimization features; The boundary-aware dynamic sparse attention module includes: an adaptive dynamic sparsification strategy, a boundary-aware regularization and feature enhancement module, and a multi-scale shared memory space processing module.
2. The skin lesion image segmentation method based on boundary enhancement and dual optimization according to claim 1 is characterized in that: The adaptive dynamic sparsification strategy generates a dynamic sparsification rate according to feature complexity, the boundary-aware regularization and feature enhancement module maintains boundary information priority through a boundary detection branch, and the multi-scale shared memory space processing module captures multi-resolution context information through multi-scale memory units.
3. The skin lesion image segmentation method based on boundary enhancement and dual optimization according to claim 2 is characterized in that: The boundary tolerance perception loss function is: in, represents the predicted boundary map, represents the true boundary label, is the Dice loss in the boundary area, is the balance coefficient, represents the gradient weight field, expressed as: in, The core area of the boundary, For the area around the border, and are the weight coefficients of the boundary core area and the boundary peripheral area respectively.
4. The skin lesion image segmentation method based on boundary enhancement and dual optimization according to claim 2 is characterized in that: The multi-scale deep supervision architecture based on the adaptive training process is based on the current training step number. t Dynamically adjust the supervision weights of each layer: in, t is the current training step number, , and Represent the step nodes in the early, middle and late stages of training respectively. is the basic weight; Combined with the dynamic weight adjustment strategy, we finally get the adaptive weighted composite loss: in, is the level weight coefficient, N is the number of supervision layers, For the i Boundary tolerance aware loss for layers, represents the layer-level segmentation loss.
5. The skin lesion image segmentation method based on boundary enhancement and dual optimization according to claim 2 is characterized in that: The calculation process of the boundary-region bidirectional feature interaction and decoupling architecture includes: First, the boundary and internal area information are separated by feature decoupling operation: in, is the segmentation prediction map of the current layer, is the boundary prediction map, Erode and Dilate are morphological corrosion and expansion operations respectively. k is the kernel size, Based on P b Adaptive thresholding of statistical properties; Generate boundary attention map: in, To predict the image from the segmentation through the Sobel operator The extracted boundary features, Concat represents the concatenation of channel dimensions, is a 3×3 convolution, is the Sigmoid activation function; Calculate the comprehensive optimization feature: in, F seg is the segmentation feature map of the current layer, is the boundary enhancement coefficient, is the regional consistency retention coefficient.
6. The skin lesion image segmentation method based on boundary enhancement and dual optimization according to claim 2 is characterized in that: The calculation process of the adaptive dynamic sparsification strategy includes: Dynamically generate the optimal sparsity rate based on image content characteristics: in is a global feature summary generated by a lightweight predictor, and are the minimum and maximum sparsity rates, respectively; Generate statistics-aware dynamic thresholds: in, and are the mean and standard deviation of the importance map, and is a learnable parameter; Construct a boundary-aware importance estimator to process input features in parallel through the main feature branch and the boundary detection branch: The main feature branch uses deep separable convolution to extract feature representation: The boundary detection branch uses the Sobel operator and lightweight convolution to detect boundaries: The features of the two branches are combined to generate an importance map: Generate the final sparse mask: in is the boundary detection map, is the boundary importance threshold, m is the activation steepness parameter.
7. The skin lesion image segmentation method based on boundary enhancement and dual optimization according to claim 6 is characterized in that: The calculation process of the multi-scale shared memory space processing module includes: For input features F Group them evenly into 3 groups according to the channel dimension: The first and third groups are input into a multi-scale storage space, which is divided into two scales: 4×4 and 8×8. Each of them focuses on the context information of a specific scale, and together they constitute a multi-resolution feature representation system: in Indicates j A shared memory space, Represents the feature conversion function: Q , K , V Represent query, key value, and value transformation functions respectively; Use deep convolution to process the second set of features: Apply a sparse mask to each feature: Preserve boundary information through residual connections: Feature reorganization and channel shuffling recombines the three sets of processed features: Finally, it is processed by depth-wise separable convolution: in, Represents the features finally output by the multi-scale shared memory space processing module.
8. The skin lesion image segmentation method based on boundary enhancement and dual optimization according to claim 7 is characterized in that: The boundary-aware regularization and feature enhancement module constructs a boundary importance regularization loss: in is the entropy of the attention mask, is the boundary region importance measure, and is the weight coefficient; Boundary Region Importance Metrics It is expressed as: is the average importance of the central region of the boundary, is the boundary detection strength, is the boundary strength threshold.
9. The skin lesion image segmentation method based on boundary enhancement and dual optimization according to claim 2, characterized in that: The overall loss function of the boundary-enhanced sparse attention network training is: in, represents the segmentation loss, represents the boundary loss, represents the regional consistency loss, represents the importance regularization loss, represents the auxiliary supervision loss, , , They represent the weights of boundary loss, regional consistency loss, and importance regularization loss respectively.
10. A skin lesion image segmentation device based on boundary enhancement and dual optimization, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the skin lesion image segmentation method based on boundary enhancement and double optimization as described in any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
Medical image segmentation method based on boundary perception and attention mechanism
CN117078930A
Medical image segmentation method based on boundary optimization
CN118657956A
Skin lesion segmentation method based on boundary direction connectivity
CN119445103A
Cited By
Medical image segmentation method and device, equipment and medium
CN120747135A
Medical image segmentation method based on boundary enhancement and double decoders
CN122115484A