Skin Lesion Image Segmentation Method and Device Based on Boundary Enhancement and Dual Optimization

By building a boundary-enhanced sparse attention network and optimizing the coordinated segmentation of boundary and regional characteristics, the problem of balance between boundary accuracy and computational efficiency of the skin lesion segmentation model is solved, and efficient skin lesion segmentation is achieved.

CN120070478BActive Publication Date: 2025-07-22JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510543657.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-22
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing skin lesion segmentation model is difficult to balance between boundary accuracy and overall regional consistency, and the trade-off between computing efficiency and segmentation accuracy is unreasonable, making it difficult to achieve real-time segmentation on resource-constrained devices.

Method used

Build a boundary-enhanced sparse attention network (BES-UNet), including an encoder module, a decoder module and a boundary-segmented bidirectional enhancement framework, adopts a boundary-aware dynamic sparse attention module and a multi-scale deep supervision architecture, and optimizes the coordinated segmentation of boundary and regional features through adaptive sparse strategies and feature interaction mechanisms.

Benefits of technology

It realizes that while maintaining high segmentation accuracy, it significantly reduces model parameters and calculation complexity, making BES-UNet an efficient skin lesion segmentation model, suitable for deployment in resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070478B_ABST
    Figure CN120070478B_ABST
Patent Text Reader

Abstract

The present invention relates to a skin lesion image segmentation method and device based on boundary enhancement dual optimization, belonging to the technical fields of medical image processing and computer vision. The method realizes precise segmentation of skin lesion regions by constructing a boundary-enhanced sparse attention network, and the boundary-enhanced sparse attention network includes: an encoder module, a decoder module, and a boundary-segmentation bidirectional enhancement framework; the encoder module includes multiple stages and respectively uses a standard convolutional block and a boundary-aware dynamic sparse attention module for feature extraction. Experimental results prove that by constructing a bidirectional promotion mechanism of boundary features and region features and a boundary-aware dynamic sparse attention module, the present invention effectively compresses model parameters and computational complexity while maintaining high segmentation accuracy, provides a technical innovation with great theoretical significance and application value for the field of medical image segmentation, and is particularly suitable for deployment in resource-constrained environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a skin lesion image segmentation method and device based on boundary enhancement double optimization, belonging to the technical fields of medical image processing and computer vision. Background Art

[0002] In recent years, the incidence rate of skin cancer has shown a rapid upward trend and has become a major global public health problem. Early diagnosis is crucial for improving the survival rate of patients, and accurate skin lesion segmentation is a key link in computer-aided diagnosis systems, laying a foundation for subsequent classification and treatment planning.

[0003] Deep learning models have made remarkable progress in the field of medical image segmentation. U-Net and its variants have become the mainstream architectures in medical image segmentation with their effective encoder-decoder structure and skip connection design. With the rise of Vision Transformer (ViT), TransUNet combines CNN and ViT, and TransFuse adopts a dual-path structure to capture local and global features simultaneously. However, these Transformer-based models have a large number of parameters and high computational complexity, making it difficult to be deployed on resource-constrained devices such as intelligent dermoscopes.

[0004] The academic community has proposed a variety of lightweight skin lesion segmentation models to solve the above problems. UNeXt significantly reduces the number of parameters by combining MLP blocks with U-Net; MALUNet further reduces the model size by reducing the model channel number and introducing an attention mechanism; EGE-UNet proposes a group multi-axis Hadamard product attention module and a group aggregation bridging module, controlling the parameters at about 50KB; LB-UNet enhances the model's perception ability of lesion boundaries through a boundary assistance module.

[0005] However, lightweight segmentation models still face two key challenges: (1) how to balance boundary accuracy and overall region consistency; (2) how to make a reasonable trade-off between model computational efficiency and segmentation accuracy. These problems are particularly prominent in skin lesion segmentation because lesions often have irregular edges, blurred contours, and diverse texture features. In clinical applications, accurately identifying the boundaries of skin lesions is crucial for subsequent diagnosis and decision-making; at the same time, real-time segmentation on resource-constrained platforms such as mobile devices requires a model with extremely high computational efficiency.

[0006] The prior art generally has the following deficiencies: First, traditional segmentation methods adopt the same processing strategy for boundaries and internal regions, resulting in insufficient boundary accuracy; Second, existing lightweight strategies often sacrifice segmentation performance in exchange for improved computational efficiency; Third, the quadratic computational complexity of the standard self-attention mechanism limits its application in high-resolution medical images. These technical problems severely restrict the application effect of skin lesion segmentation algorithms in clinical practice. Summary of the Invention

[0007] In order to improve the accuracy of skin lesion segmentation and ensure computational efficiency at the same time, the present invention provides a skin lesion image segmentation method and device based on boundary enhancement dual optimization, and the technical solution is as follows:

[0008] The first object of the present invention is to provide a skin lesion image segmentation method, which realizes precise segmentation of skin lesion regions by constructing a boundary-enhanced sparse attention network. The boundary-enhanced sparse attention network includes: an encoder module, a decoder module, and a boundary-segmentation bidirectional enhancement framework;

[0009] The encoder module includes encoders in 6 stages, where the first 3 stages adopt standard convolutional blocks, the 4th and 5th stages adopt boundary-aware dynamic sparse attention modules, and the 6th stage adopts the boundary-aware dynamic sparse attention module and a max pooling layer;

[0010] The boundary-segmentation bidirectional enhancement framework is connected to the output end of the encoder in each stage, and is also connected to the corresponding level of the decoder;

[0011] The decoder module includes decoders in 6 stages, and adopts a symmetric structure to be connected to the corresponding layer of the encoder module through the boundary-segmentation bidirectional enhancement framework;

[0012] The boundary-segmentation bidirectional enhancement framework includes a representation space adaptive transformation and boundary tolerance perception loss architecture, a multi-scale depth supervision architecture based on an adaptive training process, and a boundary-region bidirectional feature interaction and decoupling architecture; the representation space adaptive transformation and boundary tolerance perception loss architecture transforms boundary labels and predictions through a dynamic soft binary function and constructs a boundary tolerance perception loss function; the multi-scale depth supervision architecture based on an adaptive training process dynamically adjusts the supervision weight according to the training stage; the boundary-region bidirectional feature interaction and decoupling architecture realizes the collaborative optimization of boundary and region features through feature decoupling and an attention mechanism to obtain comprehensively optimized features;

[0013] The boundary-aware dynamic sparse attention module includes: an adaptive dynamic sparsification strategy, a boundary-aware regularization and feature enhancement module, and a multi-scale shared memory space processing module; the adaptive dynamic sparsification strategy generates a dynamic sparsity rate according to the feature complexity, the boundary-aware regularization and feature enhancement module maintains the boundary information priority through the boundary detection branch, and the multi-scale shared memory space processing module captures multi-resolution context information through multi-scale memory units.

[0014] Optionally, the boundary tolerance-aware loss function is:

[0015]

[0016] Where, represents the predicted boundary map, represents the ground truth boundary label, is the Dice loss of the boundary region, is the balance coefficient, represents the gradient weight field, expressed as:

[0017]

[0018] Where, is the boundary core region, is the boundary peripheral region, and are the weight coefficients of the boundary core region and the boundary peripheral region respectively.

[0019] Optionally, the multi-scale deep supervision architecture based on the adaptive training process dynamically adjusts the supervision weights of each layer according to the current training step t :

[0020]

[0021] Where, t is the current training step, 、 and represent the step nodes in the early, middle, and late stages of training respectively, is the basic weight;

[0022] Combined with the dynamic weight adjustment strategy, the adaptive weighted composite loss is finally obtained:

[0023]

[0024] Where, is the hierarchical weight coefficient, N is the number of supervision layers, is the boundary tolerance-aware loss of the i th layer, Represents the hierarchical segmentation loss.

[0025] Optionally, the calculation process of the boundary-region bidirectional feature interaction and decoupling architecture includes:

[0026] First, separate the boundary and internal region information through feature decoupling operations:

[0027]

[0028] Among them, is the segmentation prediction map of the current layer, is the boundary prediction map, Erode and Dilate are morphological erosion and dilation operations respectively, k is the kernel size, is based on P b Adaptive threshold of statistical characteristics;

[0029] Generate the boundary attention map:

[0030]

[0031] Among them, is the boundary feature extracted from the segmentation prediction map by the Sobel operator, Concat represents concatenation in the channel dimension, is a 3×3 convolution, is the Sigmoid activation function;

[0032] Calculate the comprehensive optimization feature:

[0033]

[0034] Among them, F seg is the segmentation feature map of the current layer, is the boundary enhancement coefficient, is the regional consistency preservation coefficient.

[0035] Optionally, the calculation process of the adaptive dynamic sparsification strategy includes:

[0036] Dynamically generate the optimal sparsity rate according to the characteristics of the image content:

[0037]

[0038] Among them is the global feature summary, generated by the lightweight predictor, and are the minimum and maximum sparsity rates respectively;

[0039] Generate a statistically aware dynamic threshold:

[0040]

[0041] Among them, and are the mean and standard deviation of the importance map respectively, and are learnable parameters;

[0042] Construct a boundary-aware importance estimator that processes the input features in parallel through the main feature branch and the boundary detection branch:

[0043] Among them, the main feature branch uses depthwise separable convolutions to extract feature representations:

[0044]

[0045] The boundary detection branch uses the Sobel operator and lightweight convolutions to detect boundaries:

[0046]

[0047] Fuse the features of the two branches to generate an importance map:

[0048]

[0049] Generate the final sparse mask:

[0050]

[0051] Among them is the boundary detection map, is the boundary importance threshold, m is the activation steepness parameter.

[0052] Optionally, the calculation process of the multi-scale shared memory space processing module includes:

[0053] Group the input feature F and evenly divide it into 3 groups according to the channel dimension:

[0054]

[0055] Input the first group and the third group into the multi-scale storage space. The multi-scale storage space is divided into two scales of 4×4 and 8×8, each focusing on the context information of a specific scale, and jointly constituting a multi-resolution feature representation system:

[0056]

[0057] Among them represents the j th shared memory space, Represents a feature transformation function:

[0058]

[0059] Q 、 K 、 V respectively represent query, key-value, and value transformation functions;

[0060] Process the second set of features using depth convolution:

[0061]

[0062] Apply a sparse mask to each feature:

[0063]

[0064] Retain boundary information through residual connections:

[0065]

[0066] Feature recombination and channel shuffling recombine the three processed sets of features:

[0067]

[0068] Finally, process through depthwise separable convolution:

[0069]

[0070] Among them, represents the features finally output by the multi-scale shared memory space processing module.

[0071] Optionally, the boundary-aware regularization and feature enhancement module constructs a boundary importance regularization loss:

[0072]

[0073] Among them is the entropy of the attention mask, is the boundary region importance metric, and are weight coefficients;

[0074] Boundary region importance metric is expressed as:

[0075]

[0076] is the average importance of the boundary center region, is the boundary detection intensity, is the boundary strength threshold.

[0077] Optionally, the overall loss function for training the boundary-enhanced sparse attention network is:

[0078]

[0079] where represents the segmentation loss, represents the boundary loss, represents the region consistency loss, represents the importance regularization loss, represents the auxiliary supervision loss, , , respectively represent the weights of the boundary loss, region consistency loss, and importance regularization loss.

[0080] The second object of the present invention is to provide a skin lesion image segmentation device, including a memory and a processor;

[0081] The memory is used to store a computer program;

[0082] The processor is used to implement the skin lesion image segmentation method as described in any one of the above when executing the computer program.

[0083] The third object of the present invention is to provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the skin lesion image segmentation method as described in any one of the above is implemented.

[0084] The beneficial effects of the present invention are:

[0085] The present invention realizes the accurate segmentation of skin lesion regions by constructing a boundary-enhanced sparse attention network (BES-UNet). The advantages of the BES-UNet network are as follows:

[0086] First, a boundary-segmentation bidirectional enhancement framework (BF) is proposed. This framework constructs a dynamic adaptive soft binarization function, breaks through the limitations of traditional fixed-threshold binarization, and realizes the accurate conversion of boundary representation through adaptively calculated thresholds and temperature parameters; constructs a boundary loss mechanism with a gradient weight field, fundamentally solving the problem of boundary uncertainty caused by differences in expert annotations in medical images; proposes a theoretical framework for decoupling and recombining boundary and region features, and establishes a two-way promotion mechanism between boundary features and region features through the ingenious combination of boundary attention maps and internal region masks, effectively improving the accuracy of medical image boundary segmentation.

[0087] Secondly, the present invention proposes a Boundary-Aware Dynamic Sparse Attention Module (BDAM), constructs a dynamic computing resource allocation theory, and realizes the collaborative optimization of boundary enhancement and computing efficiency. This module generates an adaptive sparsity rate based on the complexity of image content, breaking through the limitations of traditional fixed sparsity rates; designs a dynamic threshold algorithm based on the statistical characteristics of importance distribution, solving the problem of inconsistent sparsification effects of fixed thresholds on different images; and proposes a sparsification strategy that preferentially protects the boundary region, ensuring the complete retention of boundary information during the sparse computing process. The proposed BDAM module not only guarantees the high precision of the attention mechanism at the algorithm level but also balances the computing efficiency, effectively improving the segmentation accuracy while reducing the computational complexity.

[0088] In one embodiment of the present invention, a multi-component loss function and an adaptive supervision mechanism can be constructed to achieve multi-objective collaborative optimization of boundary accuracy, regional consistency, and computing efficiency. By organically combining segmentation loss, boundary loss, regional consistency loss, importance regularization loss, and hierarchical auxiliary supervision loss, a complete model optimization theory system is established, ensuring excellent performance of the model in all dimensions.

[0089] Experimental results prove that the present invention achieves extreme compression of model parameters and computational complexity while maintaining high segmentation accuracy. Compared with LB-UNet, the number of parameters of BES-UNet is reduced from 0.038M to 0.032M (a decrease of approximately 15.8%), the computational volume is reduced from 0.098 GFLOPs to 0.089 GFLOPs (a reduction of approximately 9.2%), and the training time is shortened from 4 hours to approximately 2.5 hours (an acceleration of approximately 60%). This extremely lightweight design makes BES-UNet one of the most efficient skin lesion segmentation models currently, providing a technological innovation with great theoretical significance and application value for the field of medical image segmentation and being particularly suitable for deployment in resource-constrained environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0091] Figure 1 It is the overall structure diagram of the Boundary Enhancement Sparse Attention Network (BES-UNet) framework of the present invention.

[0092] Figure 2 It is the structure diagram of the Boundary-Segmentation Bidirectional Enhancement Framework (BF) of the present invention.

[0093] Figure 3 It is the architecture diagram of the boundary-aware dynamic sparse attention module (BDAM) of the present invention.

[0094] Figure 4 It is the flowchart for constructing the framework of the boundary-enhanced sparse attention network (BES-UNet) of the present invention.

[0095] Figure 5 It is the visualization effect diagram of the segmentation results of different models on the ISIC2018 dataset.

[0096] Figure 6 It is the performance comparison result of different methods on the ISIC2017 and ISIC2018 datasets.

[0097] Figure 7 It is the analysis result of the segmentation accuracy of the boundary regions of different methods.

[0098] Figure 8 It is the ablation experiment result of the contributions of different components of the present invention.

[0099] Figure 9 It is the ablation experiment result of the boundary-aware dynamic sparse attention mechanism. Detailed implementation manners

[0100] To make the objectives, technical solutions and advantages of the present invention clearer, the following will further describe the implementation manners of the present invention in detail with reference to the accompanying drawings.

[0101] Example 1:

[0102] This example provides a precise skin lesion segmentation method based on boundary-enhanced dual optimization, which realizes the precise segmentation of skin lesion regions by constructing the framework of the boundary-enhanced sparse attention network (BES-UNet), and designs a boundary-segmentation bidirectional enhancement adaptive learning framework to realize the collaborative optimization of boundary precise positioning and region segmentation through the boundary tolerance perception loss and the bidirectional feature interaction strategy; designs a boundary-aware dynamic sparse attention mechanism, which significantly reduces the model calculation complexity through the content-adaptive dynamic sparsification strategy and the multi-scale shared memory space, and at the same time enhances the recognition ability of the lesion edges.

[0103] The construction and training of the BES-UNet framework in this example include the following processes:

[0104] Step 1: Dataset preparation and preprocessing.

[0105] First, collect and organize the publicly available ISIC2017 and ISIC2018 skin lesion datasets required for the experiment. ISIC2017 contains 2,150 dermoscopic images, and ISIC2018 contains 2,694 dermoscopic images. Each image has a segmentation mask annotated by experts. Randomly divide the training set and test set in a 7:3 ratio to ensure the balance of data distribution. Standardize all images, including pixel value normalization (subtract the mean and divide by the standard deviation) and size adjustment (unify to a 256×256 pixel resolution). To enhance the generalization ability of the model, design a data augmentation strategy, including random horizontal and vertical flipping (both with a probability of 0.5), random rotation (±15 degrees), random brightness and contrast adjustment (±0.2), etc., to enrich the diversity of training samples.

[0106] Step 2: Construct the overall architecture of BES-UNet and initialize it.

[0107] As Figure 1 shown, BES-UNet includes: an encoder, a decoder, a boundary-segmentation bidirectional enhancement framework, and a boundary-aware dynamic sparse attention module.

[0108] The encoder consists of 6 stages, and the number of channels is {8, 16, 24, 32, 48, 64} in sequence. The first three stages use standard convolutional block chains (CBCs), and the last 3 stages use boundary-aware dynamic sparse attention modules (BDAMs). Among them, the 4th and 5th stages use the BDAM module + max pooling layer (Boundary-aware Dynamic Attention Pooling, BDAP).

[0109] As Figure 1As shown in the figure, the input image is first processed by the CBC module in Stage 1 to generate the feature image F1, and the boundary feature B1 and the region feature R1 are extracted through the Boundary-segmentation Bidirectional Enhanced Framework (BF); then the feature image F1 is downsampled by the ResidualBoundary Pooling (RBP) and input into the CBC module in Stage 2 to generate the feature image F2, while the BF module extracts the boundary feature B2 and the region feature R2; similarly, the feature image F2 is downsampled by RBP and enters the CBC module in Stage 3 to generate the feature image F3, and the BF framework extracts the boundary feature B3 and the region feature R3; the feature image F3 is then input into the BDAP module in Stage 4 through RBP to generate the feature image F4, and the BF module extracts the boundary feature B4 and the region feature R4; similarly, the feature image F4 passes through the BDAP module in Stage 5 to generate the feature image F5, and the feature image F5 finally passes through the BDAM module in Stage 6 to generate the feature image F6.

[0110] The BF module is connected to the output end of the encoder in each stage and is also connected to the decoder at the corresponding level. The output of the BF module includes the boundary features (B1 - B6) and the region features (R1 - R6), and these features are passed to the corresponding levels in the decoder for feature fusion.

[0111] The decoder adopts a symmetric structure and fuses with the corresponding layer of the encoder through skip connections. The top layer of the decoder upsamples the feature image F6 through the Upsampling Module (UP) and fuses it with the feature image F5, and at the same time combines the boundary feature B5 and the region feature R5 for boundary enhancement; each subsequent layer adopts a similar upsampling fusion structure to generate the D5, D4, D3, D2, and D1 features in turn. Finally, D1 generates the prediction mask P through a 1×1 convolution and the Sigmoid function.

[0112] When initializing the model parameters, the He initialization is used for the convolutional layer, and the Xavier initialization is used for the attention module and the linear layer to ensure a good starting point for training.

[0113] Step 3: Construct the Boundary-segmentation Bidirectional Enhanced Framework (BF module).

[0114] The structure of the Boundary-segmentation Bidirectional Enhanced Framework is as Figure 2 shown, mainly including the architecture representing spatial adaptive transformation and boundary tolerance perception loss, the multi-scale depth supervision architecture based on the adaptive training process, and the boundary-region bidirectional feature interaction and decoupling architecture.

[0115] Among them, it means that the spatio - adaptive transformation and boundary tolerance - aware loss architecture obtains the multi - scale features extracted by the BES - UNet encoder , and first converts them into boundary features through a 1×1 convolutional layer . It means that the spatio - adaptive transformation mechanism seamlessly transforms boundary labels and predictions through a dynamically adaptive soft - binarization function:

[0116]

[0117] Among them, represents the boundary probability value (0~1), τ represents the adaptive threshold, represents the temperature parameter;

[0118] The adaptive threshold and the temperature parameter are represented by and respectively, and are dynamically adjusted according to the statistical characteristics of the input image :

[0119]

[0120] Among them, and represent the mean and standard deviation of the input image features respectively, and are learnable parameters, and this design ensures the complete transmission of the gradient flow and training stability.

[0121] The boundary filtering mechanism selectively retains significant boundary features through an adaptive gating unit, and its calculation process is as follows:

[0122]

[0123] Among them, and are the learning parameters of the gating unit, is the Sigmoid activation function. This mechanism can effectively filter out noise and weak boundaries, and only retain the strong boundary information meaningful for the segmentation task.

[0124] The multi - scale boundary fusion module captures boundary information through dilated convolutions of different scales:

[0125]

[0126] Among them, , and are 3×3 dilated convolutions with different dilation rates (1, 2, 4) respectively, and Concat represents the concatenation operation in the channel dimension, It is a 1×1 convolution. This design can expand the receptive field without increasing the computational complexity and capture boundary features at different scales.

[0127] To solve the problem of boundary uncertainty caused by differences in expert annotations in medical images, the present invention introduces a boundary tolerance-aware loss function:

[0128]

[0129] Where, represents the predicted boundary map, represents the true boundary label, is the Dice loss of the boundary region, is the balance coefficient.

[0130] This loss uses a gradient weight field with a fine transition from the boundary core to the periphery :

[0131]

[0132] Where, is the boundary core region, is the boundary periphery region, and are the weight coefficients of the boundary core region and the periphery region respectively.

[0133] Based on the multi-scale deep supervision architecture of the adaptive training process, the feature maps output by each layer of the encoder and the features of each layer of the decoder are obtained. First, multi-scale segmentation predictions are generated through the deep supervision layer on the decoder side:

[0134]

[0135] Where, is a 1×1 convolution operation, is the Sigmoid activation function. These multi-scale predictions are compared with the true segmentation annotation to calculate the hierarchical segmentation loss:

[0136]

[0137] Where BCE is the binary cross-entropy loss and Dice is the Dice coefficient loss, is the balance coefficient.

[0138] The training process adaptive weight regulator dynamically adjusts the supervision weights of each layer according to the current training step t :

[0139]

[0140] Among them, t is the current training step number, , and respectively represent the time nodes in the early, middle, and late stages of training (set to 20%, 50%, and 80% of the total training step number), is the basic weight (set to 0.8). This mechanism enables the model to pay more attention to the learning of overall regional features in the early stage of training and gradually increase the attention to boundary details as the training progresses.

[0141] Combined with the multi-level weight assignment strategy, the adaptive weighted composite loss is finally obtained:

[0142]

[0143] Among them, is the hierarchical weight coefficient, which increases from 0.1 to 0.5 as the feature level is refined (for example = 0.1, = 0.2,..., = 0.5), N is the number of supervision layers (usually 5), is the application of the boundary tolerance perception loss in the i th layer.

[0144] The boundary-region bidirectional feature interaction and decoupling architecture obtains the boundary-aware adaptive loss and the adaptive weighted composite loss , as well as the current feature and . First, the boundary and internal region information are separated through the feature decoupling operation:

[0145]

[0146] Among them, is the segmentation prediction map of the current layer, is the boundary prediction map, Erode and Dilate are the morphological erosion and dilation operations respectively, k is the kernel size, is the adaptive threshold based on the statistical characteristics of Pb.

[0147] Then, the boundary attention map is generated:

[0148]

[0149] Among them, is the boundary feature extracted from the segmentation prediction map through the Sobel operator, Concat represents the concatenation in the channel dimension, is a 3×3 convolution, is the Sigmoid activation function.

[0150] Next, perform boundary-region bidirectional feature interaction calculation:

[0151]

[0152] Among them, F seg is the segmentation feature map of the current layer, is the boundary enhancement coefficient, is the region consistency preservation coefficient. This mechanism allows boundary features and region features to promote each other rather than compete with each other.

[0153] To further ensure the consistency of the internal characteristics of the region, an additional region statistical consistency constraint is introduced:

[0154]

[0155] Among them, and respectively represent the mean and variance of the region, and respectively represent the internal regions of the predicted region and the ground-truth region, and are the weight coefficients for mean and variance consistency.

[0156] This architecture finally outputs the enhanced comprehensive optimized features , and at the same time extracts the boundary features B and the region features R :

[0157]

[0158] Step 4: Construct a boundary-aware dynamic sparse attention module (BDAM).

[0159] The structure of the boundary-aware dynamic sparse attention module is as shown in Figure 3 , including an adaptive dynamic sparsification strategy, a boundary-aware regularization and feature enhancement module, and a multi-scale shared memory space processing module. The adaptive dynamic sparsification strategy is used to dynamically allocate computing resources according to the content complexity, the boundary-aware regularization and feature enhancement module is used to maintain the priority of boundary information, and the multi-scale shared memory space processing module is used to efficiently process context information at different scales.

[0160] First, the adaptive dynamic sparsification strategy dynamically generates the optimal sparsity rate according to the characteristics of the image content:

[0161]

[0162] Among them It is a global feature summary, generated by a lightweight predictor. and are the minimum and maximum sparsity rates respectively. This mechanism can automatically adjust the processing density according to the image complexity, using a high sparsity for simple regions and reserving more computing resources for complex regions.

[0163] Secondly, a statistical perception dynamic threshold is generated:

[0164]

[0165] where and are the mean and standard deviation of the importance map respectively, and are learnable parameters. This design enables the threshold to adapt to both the central tendency and the dispersion of importance, significantly improving the sparsification accuracy.

[0166] The boundary-aware importance estimator processes the input features in parallel through the main feature branch and the boundary detection branch. The main feature branch uses depthwise separable convolutions to extract feature representations:

[0167]

[0168] The boundary detection branch uses the Sobel operator and lightweight convolutions to detect boundaries:

[0169]

[0170] The feature fusion layer fuses the features of the two branches to generate an importance map:

[0171]

[0172] The boundary enhancement mask process obtains the dynamic threshold and the importance map output by the boundary-aware importance estimator and the boundary detection map to generate the final sparse mask:

[0173]

[0174] where is the boundary detection map, is the boundary importance threshold, m is the activation steepness parameter.

[0175] The multi-scale shared memory space processing module first groups the input features, evenly dividing them into 3 groups along the channel dimension:

[0176]

[0177] The first group and the third group are input into the multi-scale storage space, which is divided into two scales of 4×4 and 8×8, each focusing on the context information of a specific scale, and jointly constituting a multi-resolution feature representation system:

[0178]

[0179] where represents the j th shared memory space, represents the feature transformation function:

[0180]

[0181] Q , K , V represent the query, key-value, and value transformation functions respectively. This design reduces the number of parameters by about 60% and captures the structural information at different semantic levels.

[0182] The second group of features is processed by depth convolution:

[0183]

[0184] In the mask application stage, a sparse mask is applied to each feature:

[0185]

[0186] The boundary enhancement process preserves the boundary information through residual connections:

[0187]

[0188] Feature recombination and channel shuffling recombine the three processed groups of features:

[0189]

[0190] Finally, it is processed by depthwise separable convolution:

[0191]

[0192] The boundary-aware regularization and feature enhancement module includes the boundary importance regularization loss:

[0193]

[0194] where is the entropy of the attention mask, is the boundary region importance metric:

[0195]

[0196] is the average importance of the boundary center region, is the boundary detection intensity, is the boundary strength threshold. This regularization mechanism ensures that the model does not ignore key boundary information while maintaining computational efficiency.

[0197] Step 5: Design an adaptive training strategy.

[0198] Design a gradient flow control strategy with an adaptive training process to dynamically adjust the optimization weights of boundary detection and region segmentation according to the training stages (early, middle, and late). Implement a multi-scale deep supervision mechanism, introduce boundary-aware auxiliary supervision signals at different decoding levels, and form a progressive optimization framework from coarse to fine. Construct a hierarchical weight allocation strategy to gradually increase the supervision weight (from 0.1 to 0.5) as the feature hierarchy is refined. Design a comprehensive loss function, including segmentation loss, boundary loss, region consistency loss, and importance regularization loss, as well as auxiliary supervision loss, to achieve multi-objective collaborative optimization. The overall loss function is:

[0199]

[0200] Step 6: Model training configuration and optimization.

[0201] Configure the training environment, including hardware settings (AMD EPYC 9654 96-core processor, NVIDIA GeForce RTX3090 GPU) and software environment (PyTorch framework, Ubuntu 20.04 system). Set training hyperparameters, including optimizer (AdamW, initial learning rate 0.001, weight decay 0.01), learning rate policy (CosineAnnealingLR, maximum number of iterations 50, minimum learning rate 0.00001), batch size (8), and number of training epochs (300 epochs). Implement a training process monitoring mechanism, including real-time tracking of the loss function curve, validation set performance metrics, and resource utilization, to ensure the stability and effectiveness of the training process.

[0202] Step 7: Model training.

[0203] Load the preprocessed training data and apply data augmentation strategies to generate diverse training samples. Initialize the model parameters, including the parameter settings of the encoder, decoder, boundary-segmentation bidirectional enhancement framework, and boundary-aware dynamic sparse attention module. Execute the iterative training process, including forward propagation, loss calculation, backpropagation, and parameter update. Apply the adaptive training strategy to dynamically adjust the weights of each loss component according to the training process to ensure the stability and effectiveness of model optimization. Regularly evaluate the model performance on the validation set, monitor the changing trends of boundary accuracy and overall segmentation accuracy, and adjust the training strategy in a timely manner.

[0204] Step 8: Performance Evaluation and Comparison.

[0205] Comprehensively evaluate the performance of BES-UNet on the test set using multiple evaluation metrics: mean intersection over union (mIoU), Dice similarity coefficient (DSC), number of parameters (Params), and floating-point operations (GFLOPs). Pay particular attention to the segmentation accuracy in the boundary region, and calculate the F1 score, precision, and recall rate of the boundary region. Conduct a comprehensive comparison with existing lightweight skin lesion segmentation models, including methods such as UNeXt, MALUNet, EGE-UNet, and LB-UNet. Generate visualization graphs of the segmentation results to intuitively demonstrate the advantages of BES-UNet in terms of boundary accuracy and overall region consistency.

[0206] Step 9: Ablation Experiments and Mechanism Verification.

[0207] Design and execute detailed ablation experiments to verify the effectiveness of each key component: separately remove the boundary-aware dynamic sparse attention mechanism (BDAM) and the boundary-segmentation bidirectional enhancement framework (BEBIF), and evaluate their impact on the model performance. Conduct an in-depth analysis of the boundary-aware dynamic sparse attention mechanism, and separately evaluate the contributions of three sub-components: dynamic sparse rate, boundary enhancement, and shared memory space. Explore the impact of different hyperparameter configurations on the model performance, including the sparse rate range, gradient weight field parameters, and loss function weights. Through detailed experimental data and analysis, verify the scientificity and effectiveness of the proposed technical solution.

[0208] Step 10: Model Optimization and Deployment.

[0209] Based on the evaluation results and ablation experiments, further optimize the BES-UNet model, including structural fine-tuning, parameter configuration adjustment, and training strategy improvement. Optimize the inference efficiency of the model, including techniques such as model quantization, knowledge distillation, and computational graph optimization, to further reduce the computational complexity and storage overhead of the model. Convert the optimized model into a format suitable for mobile deployment (such as ONNX, TFLite, or Core ML), and conduct actual tests on mobile devices to verify its real-time performance in resource-constrained environments. Summarize the technical characteristics and application value of the model to provide technical support for subsequent practical applications in intelligent dermoscopy and mobile medical devices.

[0210] To illustrate the beneficial effects of the skin lesion precise segmentation method based on BES-UNet of the present invention, the following experimental data are provided, and the following evaluation metrics are used:

[0211] Mean intersection over union (mIoU): Used to evaluate the segmentation accuracy, defined as the ratio of the intersection to the union of the predicted region and the ground truth region, and a higher value indicates better segmentation results.

[0212] Dice Similarity Coefficient (DSC): Measures the similarity between two sets, calculated as 2×|X∩Y| / (|X|+|Y|), where X is the predicted region and Y is the ground truth region, with a value range of [0,1]. A higher value indicates better segmentation performance.

[0213] Number of parameters (Params): The total number of parameters in the model, used to evaluate the model size, with the unit of million (M).

[0214] GigaFLOating-point Operations Per Second (GFLOPs): The number of floating-point operations per second, used to evaluate the computational complexity of the model.

[0215] Boundary Region F1 Score: Specifically calculates the F1 score for the boundary region, used to evaluate the segmentation accuracy of the model in the boundary region.

[0216] The experimental environment is shown in Table 1:

[0217] Table 1 Detailed Parameters for Network Model Training

[0218]

[0219] To comprehensively evaluate the performance of BES-UNet, multiple representative methods were selected as comparison algorithms, mainly including UNet, MobileViTv2, MobileNetv3, UNeXt-S, MALUNet, EGE-UNet, and LB-UNet. All experiments were conducted under the same hardware environment and training settings to ensure the fairness of the evaluation. The experimental results are as Figure 6 shown.

[0220] From Figure 6 it can be seen that BES-UNet achieved 81.98% mIoU and 89.65% DSC on the ISIC2018 dataset. Compared with the current best method LB-UNet, the mIoU increased by 0.61 percentage points and the DSC increased by 0.11 percentage points, while the number of parameters decreased by 15.8%. Compared with UNeXt-S, the mIoU increased by 2.89 percentage points, the DSC increased by 1.32 percentage points, and the number of parameters decreased by 90%.

[0221] The visualization results of the segmentation results of different models on the ISIC2018 dataset are as Figure 5 shown. From Figure 5 it can be clearly seen that the BES-UNet model shows significantly better segmentation performance in the boundary region of skin lesions than other comparison methods.

[0222] In the first row of examples, BES-UNet can accurately capture the bifurcated structure of the lesion, while LB-UNet, EGE-UNet, and MALUNet show boundary missing in the areas marked by the red circles.

[0223] In the second row of examples, BES-UNet successfully retains the concave features of the lesion, while the other models show inaccurate predictions at the upper left corner edge (the area marked by the red circle), and MALUNet even produces significant segmentation errors in the right area (marked by the oval red circle).

[0224] In the third row of examples, BES-UNet can completely segment the irregular edges of the lesion, while LB-UNet and EGE-UNet have inaccurate boundaries at the bottom edge (the area marked by the red circle), and MALUNet has serious boundary missing on both the left and right sides (the areas marked by the red circles).

[0225] In the fourth row of examples, BES-UNet accurately segments the complex boundaries at the bottom of the lesion, while the other models have obvious boundary missed detection or misdetection problems in the bottom and right areas (the areas marked by the red circles).

[0226] These visualization results fully demonstrate that the boundary-segmentation bidirectional enhancement framework and the boundary-aware dynamic sparse attention mechanism proposed in the present invention can significantly improve the model's perception ability of lesion boundaries, especially having obvious advantages in dealing with fuzzy, irregular, and complex boundary regions. Through the boundary tolerance perception loss and the adaptive feature interaction strategy, BES-UNet has successfully achieved the balance between boundary accuracy and regional consistency, providing reliable technical support for the precise segmentation of skin lesions.

[0227] The accuracy analysis of the boundary region segmentation of different models is as Figure 7 shown. In the accuracy analysis of the boundary region segmentation, the BES-UNet of the present invention reaches 60.17% in the F1 score of the boundary region, 1.05 percentage points higher than LB-UNet and 2.22 percentage points higher than EGE-UNet. This proves that the boundary-segmentation bidirectional enhancement framework and the boundary-aware dynamic sparse attention mechanism proposed in the present invention can significantly improve the segmentation accuracy of the model in the lesion boundary region.

[0228] The ablation experiment results of the contributions of different components are as Figure 8As shown in the figure, the results show that after adding the Boundary-Aware Dynamic Sparse Attention Mechanism (BDAM), the mIoU increased by 0.41 percentage points, the DSC increased by 0.07 percentage points, and the number of parameters decreased by 7.89%; after adding the Boundary-Segmentation Bidirectional Enhancement Framework (BF), the mIoU and DSC increased by 0.20 and 0.04 percentage points respectively; compared with the baseline model, the complete BES-UNet model had an mIoU increase of 0.61 percentage points, a DSC increase of 0.11 percentage points, while the number of parameters decreased by 15.79% and the computational cost decreased by 9.18%.

[0229] The ablation experiment results of the boundary-aware dynamic sparse attention mechanism are as Figure 9 shown. The dynamic sparsity rate increased the mIoU by 0.22 percentage points, verifying the effectiveness of the adaptive sparsity strategy; the boundary enhancement mechanism further increased the mIoU by 0.20 percentage points; the multi-scale shared memory space not only reduced the number of parameters from 0.044M to 0.032M and the computational cost from 0.105 to 0.089 GFLOPs, but also increased the mIoU by 0.19 percentage points, proving that each component makes significant contributions to performance improvement and model lightweighting.

[0230] In summary, the experimental results comprehensively verify the superior performance of BES-UNet. Through the innovative design of the boundary-segmentation bidirectional enhancement adaptive learning framework and the boundary-aware dynamic sparse attention mechanism, the present invention realizes the dual optimization of boundary accuracy and computational efficiency, providing a new technical solution for the field of skin lesion segmentation.

[0231] Some steps in the embodiments of the present invention can be implemented using software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.

[0232] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A skin lesion image segmentation method based on boundary enhancement and dual optimization, characterized in that, The method realizes the accurate segmentation of skin lesion areas by constructing a boundary-enhanced sparse attention network, and the boundary-enhanced sparse attention network includes: an encoder module, a decoder module, and a boundary-segmentation bidirectional enhancement framework; The encoder module includes encoders in 6 stages, where the first 3 stages adopt standard convolutional blocks, the 4th and 5th stages adopt boundary-aware dynamic sparse attention modules, and the 6th stage adopts the boundary-aware dynamic sparse attention module and a max pooling layer; The boundary-segmentation bidirectional enhancement framework is connected to the output end of the encoder in each stage, and is also connected to the decoder at the corresponding level; The decoder module includes decoders in 6 stages, and adopts a symmetric structure to be connected to the corresponding layer of the encoder module through the boundary-segmentation bidirectional enhancement framework; The boundary-segmentation bidirectional enhancement framework includes a representation space adaptive transformation and boundary tolerance perception loss architecture, a multi-scale deep supervision architecture based on an adaptive training process, and a boundary-region bidirectional feature interaction and decoupling architecture; the representation space adaptive transformation and boundary tolerance perception loss architecture transforms the boundary label and prediction through a dynamic soft binary function, and constructs a boundary tolerance perception loss function; the multi-scale deep supervision architecture based on an adaptive training process dynamically adjusts the supervision weight according to the training stage; the boundary-region bidirectional feature interaction and decoupling architecture realizes the collaborative optimization of boundary and region features through feature decoupling and an attention mechanism to obtain comprehensively optimized features; The boundary-aware dynamic sparse attention module includes: an adaptive dynamic sparsification strategy, a boundary-aware regularization and feature enhancement module, and a multi-scale shared memory space processing module; The adaptive dynamic sparsification strategy generates a dynamic sparsity rate according to the feature complexity, the boundary-aware regularization and feature enhancement module maintains the boundary information priority through a boundary detection branch, and the multi-scale shared memory space processing module captures multi-resolution context information through multi-scale memory units; The boundary tolerance perception loss function is: Among them, represents the predicted boundary map, represents the true boundary label, is the Dice loss of the boundary region, is the balance coefficient, represents the gradient weight field, expressed as: Among them, is the boundary core region, is the boundary peripheral region, and are the weight coefficients of the boundary core region and the boundary peripheral region, respectively.

2. The skin lesion image segmentation method based on boundary enhancement double optimization according to claim 1, wherein, The multi-scale deep supervision architecture based on the adaptive training process adjusts the supervision weights of each layer dynamically according to the current training step count. t Specifically, Among them, t is the current training step number, , and represent the step nodes in the early, middle, and late stages of training respectively, is the basic weight; Combined with a dynamic weight adjustment strategy, an adaptive weighted composite loss is finally obtained: Among them, is the hierarchical weight coefficient, N is the number of supervision layers, is the i boundary tolerance perception loss of the th layer, and represents the hierarchical segmentation loss.

3. The skin lesion image segmentation method based on boundary enhancement double optimization according to claim 1, characterized in that The calculation process of the boundary-region bidirectional feature interaction and decoupling architecture includes: First, separate the boundary and internal region information through a feature decoupling operation: Among them, is the segmentation prediction map of the current layer, is the boundary prediction map, and Erode and Dilate are morphological erosion and dilation operations respectively, k is the kernel size, is based on P b the adaptive threshold of statistical characteristics; Generate a boundary attention map: Among them, is the boundary feature extracted from the segmentation prediction map through the Sobel operator, Concat represents the concatenation in the channel dimension, is a 3×3 convolution, is the Sigmoid activation function; Calculate the comprehensively optimized feature: Among them, F seg is the segmentation feature map of the current layer, is the boundary enhancement coefficient, is the regional consistency preservation coefficient.

4. The skin lesion image segmentation method based on boundary enhancement and dual optimization according to claim 1, wherein, The calculation process of the adaptive dynamic sparsification strategy includes: Dynamically generate an optimal sparsity rate according to the image content characteristics: wherein is a global feature summary, generated by a lightweight predictor, and are the minimum and maximum sparsity rates, respectively; Generate a statistics-aware dynamic threshold: Among them, and are the mean and standard deviation of the importance map respectively, and are learnable parameters; Construct a boundary-aware importance estimator, and process the input features in parallel through a main feature branch and a boundary detection branch: Among them, the main feature branch uses depthwise separable convolution to extract feature representations: The boundary detection branch uses a Sobel operator and lightweight convolution to detect boundaries: Fuse the features of the two branches to generate an importance map: Generate a final sparse mask: wherein is the boundary detection graph, is the boundary importance threshold, m is the activation steepness parameter.

5. The skin lesion image segmentation method based on boundary enhancement and dual optimization according to claim 4, wherein The calculation process of the multi-scale shared memory space processing module includes: Group the input features F into 3 groups evenly according to the channel dimension: Input the first group and the third group into the multi-scale storage space. The multi-scale storage space is divided into two scales of 4×4 and 8×8, each focusing on the context information of a specific scale, and jointly constitutes a multi-resolution feature representation system: Among them represents the j th shared memory space, represents the feature conversion function: Q , K , V respectively represent query, key value, and value transformation function; Process the second set of features using depth convolution: Apply a sparse mask to each feature: Retain boundary information through residual connections: Feature recombination and channel shuffling recombine the three processed sets of features: Finally, process through depthwise separable convolution: Among them, represents the feature finally output by the multi-scale shared memory space processing module.

6. The skin lesion image segmentation method based on boundary enhancement double optimization according to claim 5, wherein The boundary-aware regularization and feature enhancement module constructs a boundary importance regularization loss: where is the entropy of the attention mask, is the importance measure of the boundary region, and are weight coefficients; Boundary Region Importance Measure Expressed as: is the average importance of the boundary center region, is the boundary detection intensity, is the boundary strength threshold.

7. The skin lesion image segmentation method based on boundary enhancement double optimization according to claim 1, wherein The overall loss function for training the boundary-enhanced sparse attention network is: Among them, represents the segmentation loss, represents the boundary loss, represents the region consistency loss, represents the importance regularization loss, represents the auxiliary supervision loss, and and represent the weights of the boundary loss, the region consistency loss, and the importance regularization loss, respectively.

8. An apparatus for segmenting skin lesion images based on boundary enhancement and dual optimization, characterized in that, Comprising a memory and a processor; The memory is used to store a computer program; The processor is configured to, when executing the computer program, implement the skin lesion image segmentation method based on boundary enhancement double optimization according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Medical image segmentation method based on boundary perception and attention mechanism

    CN117078930A

  • Medical image segmentation method based on boundary optimization

    CN118657956A