Dermoscope image segmentation method based on attention mechanism and UNet
By introducing attention mechanism and UNet's dermatoscope image segmentation method, combining depth separation convolution and efficient hollow fusion attention module, the problem of insufficient accuracy of lesion area segmentation in dermatology image segmentation is solved, and more efficient lesion area segmentation is achieved.
Patent Information
- Application Number
- CN202510418496.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
AI Technical Summary
The existing skin disease image segmentation methods have insufficient accuracy in segmentation in the lesion area, limited superficial information learning and recovery ability, and insufficient global information modeling ability, resulting in the loss of boundary information in the lesion area or inaccurate prediction.
The dermatoscope image segmentation method based on attention mechanism and UNet is adopted, and the depth-separable convolution and efficient hollow fusion attention module are combined with dense connection strategies to enhance feature extraction and global context modeling capabilities, and optimize the network structure to improve the segmentation accuracy of the lesion area.
It significantly improves the accuracy and robustness of skin disease image segmentation, can position the lesion area more accurately, and enhances the segmentation effect of the lesion area, especially in complex scenarios.
Smart Images

Figure CN120339300A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing methods, and in particular, to a dermoscopic image segmentation method based on an attention mechanism and UNet. Background Art
[0002] The main objective of dermoscopic image segmentation is to accurately locate and segment the lesion areas in the image. However, dermopathy images pose significant challenges, mainly because the differences between lesion types are small, while the feature differences within the same lesion type are large. This characteristic increases the complexity and uncertainty of the segmentation task, making it difficult for the model to accurately identify the lesion areas in practical applications. Therefore, how to design an efficient and robust segmentation network has become the focus of research.
[0003] In the dermoscopic image segmentation task, the learning and utilization of local information are crucial. Local information usually includes boundaries, textures, and subtle lesion features, which are the key to distinguishing the lesion area from normal skin. Especially the shallow features, which carry the low-level information of the image (such as edges and colors), play a fundamental role in the accurate segmentation of the lesion area. Therefore, in the training process of deep convolutional neural networks, the learning and recovery of shallow information are particularly important. However, traditional deep segmentation networks usually focus more on extracting deep features and easily ignore the recovery of shallow information, resulting in the loss of boundary information of the lesion area or inaccurate prediction.
[0004] Currently, high-performance image segmentation networks usually require a large amount of computing resources and hardware support, such as deep convolutional networks or multi-scale feature fusion networks. Although these networks perform excellently in terms of performance, in practical application scenarios with limited resources, their high hardware requirements make deployment difficult. Against this background, lightweight segmentation networks have become a research hotspot. Among them, Unet, as a classic medical image segmentation model, with its simple structure, excellent performance, and the design of skip connections, has become a representative of lightweight segmentation networks. Unet directly transfers the shallow features extracted by the encoder to the decoder through skip connections, thereby effectively recovering the shallow information and improving the segmentation accuracy to a certain extent. The simplified Unet not only has a relatively low number of parameters but also can still achieve good segmentation results under the condition of limited computing resources.
[0005] However, simply using Unet still has some limitations. For example, Unet is insufficient in extracting long-term dependent features and global information, and it is difficult to capture the global context relationship of the lesion area only relying on local convolution operations. This lack of global information extraction ability will limit the learning ability of the network in complex scenarios, thereby affecting the segmentation accuracy.
[0006] Through the above analysis, introducing the attention mechanism into the Unet structure is an effective improvement strategy. This method can not only make up for the deficiencies of Unet in extracting global information and modeling long-term dependent features, but also further improve the segmentation performance by focusing on key regions and suppressing interference from irrelevant backgrounds. Therefore, designing a Unet image segmentation method based on the attention mechanism, which can not only give full play to the shallow information learning ability of Unet, but also strengthen the global context modeling ability, has become an effective solution worthy of in-depth exploration. Summary of the Invention
[0007] The technical problem to be solved by the present invention is aimed at the three major difficulties in skin disease image segmentation: insufficient accuracy in lesion area segmentation, limited ability to learn and recover shallow information, and insufficient global information modeling ability. A dermoscopic image segmentation method based on the attention mechanism and UNet is proposed. By optimizing the network structure and feature extraction method, it aims to achieve more efficient and accurate lesion area segmentation.
[0008] In order to achieve the above invention purpose, the technical solution adopted by the present invention is specifically as follows: A dermoscopic image segmentation method based on the attention mechanism and UNet, comprising the following steps:
[0009] Step 1: Data preprocessing, performing image denoising, image enhancement, and partitioning on the data set;
[0010] Step 2: Construct a dermoscopic image segmentation network based on the attention mechanism and UNet. The network consists of an encoder and a decoder. The first three layers of the encoder and decoder use depthwise separable convolutions to effectively reduce the number of model parameters, and introduce a dilation rate in the depth convolution to more accurately locate the lesion area. At the same time, an Efficient Dilated Fusion Attention (EDFA) module is introduced in the last two layers of the encoder and decoder to enhance the interaction ability of position information and improve the expression ability of channel and spatial features. In addition, dense connections are used in the decoder, and the shallow feature outputs are transmitted to the deep network through skip connections;
[0011] Step 3: Put the processed skin disease image training set into the dermoscopic image segmentation network model based on the attention mechanism and UNet for training to obtain the optimal model;
[0012] Step 4: After training, input the test set into the optimal model to detect the lesion segmentation results in the skin disease images.
[0013] The specific process in the above Step 1 is as follows:
[0014] Step 1.1: Uniformly adjust the size of each image in the original skin disease image data set to 256×256.
[0015] Step 1.2: For the dermatological image dataset whose size has been adjusted in Step 1.1, a series of preprocessing techniques such as grayscale conversion, black hat operation, threshold segmentation, image inpainting, bilateral filtering, and histogram equalization are adopted to effectively remove hair interference and enhance the distinguishability of the lesion area. First, convert the RGB image to a grayscale image to reduce the computational complexity; then perform a black hat operation using a 5x5 elliptical structuring element to highlight areas with darker backgrounds and prominent hair, and generate a binary image through global threshold segmentation to enhance the contrast; next, use a fast matching method to repair the missing areas, and calculate the weighted combination of the direction factor, geometric distance factor, and gray-scale similarity factor to ensure the integrity of the edge features; subsequently, adopt bilateral filtering for denoising, taking into account both noise reduction and detail preservation; finally, convert the image to the YCbCr color space and perform histogram equalization on the luminance channel to enhance the contrast and details.
[0016] The formulas used in Step 1.2 are as follows:
[0017] (1) Grayscale image conversion formula
[0018] G(i) = 0.299R(i) + 0.587G(i) + 0.114B(i)
[0019] where G(i) is the grayscale value at image index i, and R(i), G(i), B(i) are the pixel values of the image in the red, green, and blue channels respectively.
[0020] (2) Black hat operation
[0021] Perform a morphological black hat operation using a 5x5 elliptical structuring element. The formula is as follows:
[0022]
[0023] where src is the original input image, is the morphological closing operation, which means dilation followed by erosion. element is the elliptical structuring element used to adapt to the shape characteristics of the hair, and dst is the result image after the black hat operation;
[0024] (3) Threshold segmentation
[0025] Perform global threshold segmentation on the result of the black hat operation to generate a binary image. The formula is as follows:
[0026]
[0027] Among them, I′(i) represents the new value of this pixel after threshold processing. Thresh is the threshold. When the pixel value is higher than the threshold, it is set to maxVal, otherwise it is 0. maxVal is the maximum pixel value after binarization.
[0028] (4) Image inpainting
[0029] Use the fast matching method FMM to inpaint the missing area of the binarized image. The inpainted pixel value is calculated by weighting the surrounding known pixels:
[0030] w(p,q) = f d (p,q)·f g (p,q)·f c (p,q)
[0031] Among them, w(p, q) is the weight of pixel q for the inpainting point p, f d (p, d) is the direction factor, which controls the inpainting to proceed preferentially along the normal direction, f g (p, q) is the geometric distance factor, f c (p, q) is the gray-scale similarity factor, σ g and σ c are smoothing parameters;
[0032]
[0033] (5) Bilateral filtering for denoising
[0034] Apply bilateral filtering to the inpainted image, and combine the spatial position and gray-scale information to achieve noise reduction and edge preservation:
[0035]
[0036] Among them, I′(p) is the output value of pixel p, Ω is the filtering window, f g (p, q) is the spatial distance weight, which measures the geometric distance between pixel p and q, f c (p, d) is the gray-scale similarity weight, which measures the similarity degree of gray-scale values, and W(p) is the normalization factor;
[0037]
[0038] The specific process in step 2 is as follows:
[0039] Step 2.1: In the first three layers of the encoder and decoder of the model, replace the conventional convolution with depthwise separable convolution. First, process the input feature map channel by channel through depthwise convolution, and introduce a dilation rate r = 2 to expand the receptive field. Subsequently, use pointwise convolution (1×1 convolution) to linearly combine the output of the depthwise convolution to generate the final feature map, thereby reducing the number of parameters and retaining the feature expression ability. At the same time, add a 2×2 max pooling layer to each layer of the encoder, and use a fixed window and stride to select the maximum value in the local area of the input feature map to generate the downsampled feature map, further reducing the size of the feature map while retaining significant information, providing more effective feature support for subsequent lesion area localization and segmentation.
[0040] The formulas used in Step 2.1 are as follows:
[0041] (1) Depthwise convolution
[0042] Depthwise convolution performs convolution operations independently for each channel. After introducing the dilation rate r = 2, the calculation formula is:
[0043]
[0044] where x c (u, v) is the pixel value of the input feature map at position (u, v) and channel c, and w c (i, j) is the weight of the depthwise convolution kernel in the channel, k is the size of the convolution kernel, and r is the dilation rate, representing the interval between the elements of the convolution kernel, which is used to expand the receptive field.
[0045] (2) Pointwise convolution
[0046] Linearly combine the output of the depthwise convolution through a 1×1 convolution kernel, and the formula is:
[0047]
[0048] where z m (u, v) is the pixel value of the output feature map of the pointwise convolution at position (u, v) and channel m, y c (u, v) is the pixel value of the output of the depthwise convolution, and w c,m is the weight of the pointwise convolution kernel, connecting the input channel c and the output channel m, and b m is the bias term of the pointwise convolution, and C is the number of input channels.
[0049] (3) Max pooling downsampling
[0050] In the pooling operation, add a 2×2 max pooling layer to each layer of the encoder, and use the maximum value in the local area as the downsampled feature value. The calculation formula is:
[0051]
[0052] where z c (h′, w′) is the pixel value of the pooled output feature map at channel c and position (h′, w′), and x c (h, w) is the pixel value of the input feature map at channel c and position (h, w). is the pixel region corresponding to the pooling window.
[0053] Step 2.2: An efficient dilated fusion attention module is introduced in the fourth and fifth layers of the encoder and decoder. This module consists of a dilated convolution fusion unit and a spatial and channel efficient attention module. First, the input features are divided into four branches in the dilated convolution fusion unit, and features are extracted through convolutions with different dilation rates. Subsequently, the four-way features are concatenated in the channel dimension to enhance the feature representation ability. Then, after 1x1 convolution processing on the concatenated features, the results are input into the channel efficient attention module.
[0054] In the channel efficient attention module, the features are divided into upper and lower branches for separate processing. The upper half is further divided into two paths through feature separation operations, and depthwise separable convolutions with kernel sizes of 3x3 and 5x5 are used to extract features respectively. Subsequently, the two-way features are concatenated in the channel dimension to achieve multi-scale feature interaction, and then processed through global average pooling and 1D convolution (kernel size is 3), similar to the "squeeze and excitation" mechanism of the SE module, to complete the interaction optimization of channel information. After Sigmoid and Softmax activations, cross-channel information weights are generated for selecting important channel features.
[0055] The lower half performs convolution operations on the input features to compress the channel information of the feature map into a single dimension. Subsequently, spatial weights are generated through activation functions and multiplied element-wise with the original features to enhance the feature representation at key positions. Finally, the features extracted from the upper and lower branches are added element-wise to generate the fused output features.
[0056] Construction of the dilated convolution fusion unit:
[0057] (1) Input feature segmentation
[0058] The input feature X (C x H x W) is split into 4 parts along the channel dimension, and the number of channels for each part is C / 4:
[0059] X = [X1, X2, X3, X4]
[0060] where [] represents the splitting operation along the channel dimension.
[0061] (2) Dilated convolution extraction
[0062] For each sub-feature X iApply convolution operations with different dilation rates d and output X i1 :
[0063] X i1 = Conv(X i , d), d ∈ {7, 7, 3, 1}, i = 1, 2, 3, 4
[0064] where Conv(X, d) represents the convolution operation with dilation rate d.
[0065] (3) Channel concatenation
[0066] Concatenate the convolved sub - features along the channel dimension to obtain the output feature X concat :
[0067] X concat = [X 11 , X 21 , X 31 , X 41
[0068] where [] represents the concatenation operation along the channel dimension.
[0069] Construction of channel - efficient attention module:
[0070] Upper branch: Channel attention mechanism
[0071] (1) Local feature extraction: Perform 3×3 depth convolution on the input feature X input to obtain the local feature X local :
[0072] X local = DepthConv 3×3 (X input )
[0073] (2) Global feature extraction: Perform 5×5 depth convolution on the input feature X input to obtain the global feature X global :
[0074] X global = DepthConv 5×5 (X input )
[0075] (3) Feature fusion: Fuse the local feature X local and the global feature X global by pixel - by - pixel addition to obtain X fuscd :
[0076] X fused = X local + X global
[0077] (4) Channel weight generation: Perform global average pooling on X fuscd to compress it to the channel dimension Z global , and then extract the weight Z through a one-dimensional convolution Conv1D3 conv :
[0078] Z global = GlobalAvgPool(X fused )
[0079] Z conv = Conv1D3(Z global )
[0080] (5) Activation and channel selection: Perform Sigmoid and Softmax activation operations on Z conv in sequence to generate the channel weight W:
[0081] Z sig = Sigmoid(Z conv )
[0082] W = Softmaxx(Z sig )
[0083] (6) Channel weighted output: Apply the channel weight W to the original input feature X input to obtain the channel enhanced feature A global :
[0084] A global = X input ·W
[0085] Lower branch: Spatial attention mechanism
[0086] (1) Spatial weight generation: Perform a 1×1×1 convolution on the input feature X input to obtain the spatial feature Z spatial , and generate the spatial weight A through Sigmoid activation spatial :
[0087] Z spatial = Conv 1×1×1 (X input )
[0088] A spatial = Sigmoid(Z spatial )
[0089] (2) Spatial weighted output: Apply the spatial weight A spatial to the input feature X input , to obtain the spatial enhanced feature:
[0090] X spatial= X input · A spatial
[0091] Final Feature Fusion
[0092] Feature Weighted Fusion: Add the channel-enhanced feature A global and the spatial-enhanced feature X spatial element-wise to obtain the final output feature:
[0093] X output = A global + X spatial
[0094] Step 2.2: Adopt dense connections in the decoder and pass the shallow feature output to the deep network through skip connections. Introduce the dense connection strategy in the decoder part and pass the features generated by the shallow network to the deep network through skip connections. Each layer of the decoder is regarded as part of the dense connection. Since there are differences in the spatial dimensions of the shallow and deep feature maps, it is necessary to first perform transposed convolutional upsampling on the shallow feature map and then splice it with the deep feature map.
[0095] The specific process in Step 3 is as follows:
[0096] Step 3.1: Initialize the parameters of the skin disease image segmentation network;
[0097] Step 3.2: Set the training parameters. The number of epochs for training is 120, the sample batch size (Batch Size) is 4, the initial learning rate (lr) of the network is set to 1e-4, and the Adam optimizer is used;
[0098] Step 3.3: Use three evaluation metrics, Accuracy, Jaccard Index, and F1-Score, to measure the segmentation effect during training;
[0099] Accuracy, that is, the proportion of pixels correctly predicted by the model among all pixels, is shown by the following formula:
[0100]
[0101] Precision, that is, the proportion of pixels correctly classified as lesions among all pixels predicted as lesions, is shown by the following formula:
[0102]
[0103] Recall, that is, the proportion of pixels correctly classified as lesions among all pixels that are lesions, is shown by the following formula:
[0104]
[0105] The F1-Score is the harmonic mean of Precision and Recall. It is a comprehensive evaluation metric, and the formula is as follows:
[0106]
[0107] The Jaccard coefficient is a metric used to measure the similarity and difference of sample data. Here, X represents the lesion area in the segmentation result, Y represents the lesion area in the mask image, and ‖ represents calculating the number of pixels in the image. The closer the Jaccard coefficient is to 1, the better the segmentation result. The formula is as follows:
[0108]
[0109] Among them, True Positive (TP) means the model predicts a lesion and it is actually a lesion; True Negative (TN) means the model predicts normal skin and it is actually normal skin; False Positive (FP) means the model predicts a lesion but it is actually normal skin; False Negative (FN) means the model predicts normal skin but it is actually a lesion.
[0110] Step 3.4: Use a dermoscopic image segmentation model based on the attention mechanism and UNet for iterative training, extract the feature information of skin disease images, and obtain the segmentation result to output the optimal model.
[0111] Compared with the existing technologies, the beneficial effects of the technical solution of the present invention are:
[0112] (1) A comprehensive preprocessing method is proposed for the problems of hair interference and lesion area enhancement in dermoscopic image segmentation. First, the calculation complexity is simplified through grayscale conversion, and the contrast of hair and lesion areas is highlighted using black hat operation and global threshold segmentation; then, a fast matching repair technique is adopted, combined with direction, geometric distance, and grayscale similarity factors, to achieve high-fidelity repair of missing areas; next, bilateral filtering is used to effectively reduce noise and retain edge details; finally, histogram equalization is performed on the luminance channel based on the YCbCr color space to enhance image contrast and detail distinguishability. This multi-step preprocessing scheme effectively enhances the segmentation effect of the lesion area and provides higher-quality data input for subsequent image analysis.
[0113] (2) By combining depthwise separable convolution and an efficient dilated fusion attention module, the performance and efficiency of the dermatological image segmentation network are improved. First, depthwise separable convolution is introduced in the first three layers of the encoder and decoder, and the receptive field is extended by combining dilation rates. At the same time, max pooling downsampling is used to reduce the number of parameters while retaining significant features. Second, an efficient dilated fusion attention module is designed in the last two layers of the encoder and decoder, which fuses multi-dilation rate feature extraction with channel and spatial attention mechanisms to enhance the multi-scale feature expression ability. Channel attention generates channel weights through multi-path depth convolution, global average pooling, and activation functions to highlight key channel features; spatial attention strengthens the expression of important regions through convolution and element-wise weighting. The above method not only effectively balances the lightweight and expression ability of the network, but also significantly improves the segmentation accuracy and robustness of the lesion area.
[0114] (3) A dense connection strategy is introduced in the decoder. Through skip connections, shallow features are effectively transmitted to the deep network, enhancing the utilization rate of feature information and the segmentation effect. To solve the problem of inconsistent spatial dimensions between shallow and deep feature maps, transposed convolution is used to upsample the shallow feature maps and concatenate them with the deep feature maps in the channel dimension, fully integrating the edge and detail information of the shallow layer with the global semantic features of the deep layer. This method not only strengthens the information flow between features, but also improves the network's ability to capture multi-scale features, thus significantly enhancing the accuracy and robustness of dermatological image segmentation. Description of the Drawings
[0115] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.
[0116] Figure 1 It is the overall flowchart of a dermoscopic image segmentation method based on an attention mechanism and UNet provided by the present invention.
[0117] Figure 2 It is the network framework diagram of a dermoscopic image segmentation method based on an attention mechanism and UNet provided by the present invention.
[0118] Figure 3 It is the structural schematic diagram of the efficient dilated fusion attention module EDFA in the present invention.
[0119] Figure 4 It is the structural schematic diagram of the spatial and channel efficient attention module scECA in the present invention. Detailed Embodiments
[0120] In order to make the objectives, technical solutions, and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0121] Embodiment 1
[0122] Figure 1 It is a schematic diagram of the overall process of an embodiment of the present invention. First, denoising, enhancement, and dataset division are performed on the input image to ensure data quality and diversity. In network construction, the first three layers of the encoder use depthwise separable convolutions and introduce dilated convolutions to extract multi-scale features. At the same time, an efficient spatial fusion attention module is added to the last two layers of the encoder and decoder to enhance the ability to focus on important regions. In the decoder, high-resolution features are retained through dense connection and skip connection strategies, thereby improving the segmentation accuracy. Then, through the optimization training of the segmentation network, the model learns the mapping relationship between the input image and the segmentation result. Finally, the trained model is applied to the test data to generate high-quality segmentation results.
[0123] The specific implementation steps of the above method are as follows:
[0124] Step 1: Data preprocessing, perform image denoising, image enhancement, and division on the dataset;
[0125] Step 2: Construct a dermoscopic image segmentation network based on the attention mechanism and UNet. The network consists of an encoder and a decoder. The first three layers of the encoder and decoder use depthwise separable convolutions, effectively reducing the number of model parameters. At the same time, a dilation rate is introduced in the depth convolution to more accurately locate the lesion area. In addition, an efficient dilated fusion attention module (EDFA) is introduced into the last two layers of the encoder and decoder to enhance the interaction ability of position information and improve the expression ability of channel and spatial features. In the decoder, dense connection is adopted, and the shallow feature output is transmitted to the deep network through skip connection;
[0126] Step 3: Put the processed dermopathy image training set into the dermoscopic image segmentation network model based on the attention mechanism and UNet for training to obtain the optimal model;
[0127] Step 4: After training is completed, input the test set into the optimal model to detect the lesion segmentation result in the dermopathy image.
[0128] The specific process in Step 1 is as follows:
[0129] Step 1.1: Uniformly adjust the size of each image in the original dermopathy image dataset to 256×256;
[0130] Step 1.2: For the dermatological image dataset whose size has been adjusted in Step 1.1, a series of preprocessing techniques such as grayscale conversion, black-hat operation, threshold segmentation, image inpainting, bilateral filtering, and histogram equalization are adopted to effectively remove hair interference and enhance the distinguishability of the lesion area. First, convert the RGB image to a grayscale image to reduce computational complexity; then use a 5x5 elliptical structuring element for the black-hat operation to highlight areas with darker backgrounds and prominent hair, and generate a binary image through global threshold segmentation to enhance contrast; next, use a fast matching method to repair the missing areas, combining weighted calculations of the direction factor, geometric distance factor, and grayscale similarity factor to ensure the integrity of edge features; subsequently, adopt bilateral filtering for denoising, taking into account both noise reduction and detail preservation; finally, convert the image to the YCbCr color space and perform histogram equalization on the luminance channel to enhance contrast and details.
[0131] The formulas used in Step 1.2 are as follows:
[0132] (1) Grayscale image conversion formula
[0133] G(i) = 0.299R(i) + 0.587G(i) + 0.114B(i)
[0134] where G(i) is the grayscale value at image index i, and R(i), G(i), B(i) are the pixel values of the image in the red, green, and blue channels respectively.
[0135] (2) Black-hat operation
[0136] Perform a morphological black-hat operation using a 5x5 elliptical structuring element. The formula is as follows:
[0137]
[0138] where src is the original input image, is the morphological closing operation, which means dilation followed by erosion, element is the elliptical structuring element used to adapt to the shape characteristics of hair, and dst is the result image after the black-hat operation;
[0139] (3) Threshold segmentation
[0140] Perform global threshold segmentation on the result of the black-hat operation to generate a binary image. The formula is as follows:
[0141]
[0142] where I′(i) represents the new value of this pixel after threshold processing, thresh is the threshold, when the pixel value is higher than the threshold, it is set to maxVal, otherwise it is 0, and maxVal is the maximum pixel value after binarization.
[0143] (4) Image inpainting
[0144] The missing area of the binary image is inpainted using the Fast Matching Method (FMM), and the inpainted pixel value is calculated by weighting the surrounding known pixels:
[0145] w(p,q) = f d (p,q)·f g (p,q)·f c (p,q)
[0146] where w(p, q) is the weight of pixel q for the inpainting point p, f d (p, q) is the direction factor, which controls the inpainting to proceed preferentially along the normal direction, f g (p, q) is the geometric distance factor, f c (p, q) is the gray-scale similarity factor, σ g and σ c are smoothing parameters;
[0147]
[0148] (5) Bilateral filtering for denoising
[0149] Bilateral filtering is applied to the inpainted image to achieve denoising and edge preservation by combining spatial position and gray-scale information:
[0150]
[0151] where I′(p) is the output value of pixel p, Ω is the filtering window, f g (p, q) is the spatial distance weight, which measures the geometric distance between pixels p and q, f c (p, q) is the gray-scale similarity weight, which measures the similarity degree of gray-scale values, and W(p) is the normalization factor;
[0152]
[0153] The specific process in step 2 is as follows:
[0154] Step 2.1: In the first three layers of the encoder and decoder of the model, replace the conventional convolution with depthwise separable convolution. First, process the input feature map channel by channel through depthwise convolution, and introduce a dilation rate r = 2 to expand the receptive field. Subsequently, use pointwise convolution (1×1 convolution) to linearly combine the output of the depthwise convolution to generate the final feature map, thereby reducing the number of parameters while retaining the feature expression ability. At the same time, add a 2×2 max pooling layer to each layer of the encoder. Use a fixed window and stride to select the maximum value in the local area of the input feature map to generate the downsampled feature map, further reducing the size of the feature map while retaining significant information, providing more effective feature support for subsequent lesion area localization and segmentation.
[0155] The formulas used in Step 2.1 are as follows:
[0156] (1) Depthwise convolution
[0157] Depthwise convolution performs convolution operations independently for each channel. After introducing the dilation rate r = 2, the calculation formula is:
[0158]
[0159] where x c (u, v) is the pixel value of the input feature map at position (u, v) and channel c, w c (i, j) is the weight of the depthwise convolution kernel on channel c, k is the size of the convolution kernel, and r is the dilation rate, representing the interval between the elements of the convolution kernel, used to expand the receptive field.
[0160] (2) Pointwise convolution
[0161] Linearly combine the output of the depthwise convolution through a 1×1 convolution kernel. The formula is:
[0162]
[0163] where z m (u, v) is the pixel value of the output feature map of the pointwise convolution at position (u, v) and channel m, y c (u, v) is the pixel value of the output of the depthwise convolution, w c,m is the weight of the pointwise convolution kernel, connecting the input channel c and the output channel m, b m is the bias term of the pointwise convolution, and C is the number of input channels.
[0164] (3) Max pooling downsampling
[0165] In the pooling operation, add a 2×2 max pooling layer to each layer of the encoder. Use the maximum value in the local area as the downsampled feature value. The calculation formula is:
[0166]
[0167] where z c (h′, w′) is the pixel value of the pooled output feature map at channel c and position (h′, w′), and x c (h, w) is the pixel value of the input feature map at channel c and position (h, w), is the pixel region corresponding to the pooling window.
[0168] Step 2.2: An efficient dilated fusion attention module is introduced in the fourth and fifth layers of the encoder and decoder. This module consists of a dilated convolution fusion unit and a spatial and channel efficient attention module. First, the input features are divided into four branches in the dilated convolution fusion unit, and features are extracted through convolutions with different dilation rates. Subsequently, the four-way features are concatenated in the channel dimension to enhance the feature representation ability. Then, after 1x1 convolution processing on the concatenated features, the result is input into the channel efficient attention module.
[0169] In the channel efficient attention module, the features are divided into upper and lower branches for separate processing. The upper half is further divided into two paths through feature separation operations, and depthwise separable convolutions with 3x3 and 5x5 are used respectively to extract features. Subsequently, the two-way features are concatenated in the channel dimension to achieve multi-scale feature interaction, and then processed through global average pooling and one-dimensional convolution (kernel size is 3), similar to the "squeeze and excitation" mechanism of the SE module, to complete the interaction optimization of channel information. After Sigmoid and Softmax activations, cross-channel information weights are generated for selecting important channel features.
[0170] The lower half performs a convolution operation on the input features, compressing the channel information of the feature map into a single dimension. Subsequently, spatial weights are generated through an activation function and multiplied element-wise with the original features, thereby enhancing the feature representation at key positions. Finally, the features extracted from the upper and lower branches are added element-wise to generate the fused output features.
[0171] Construction of the dilated convolution fusion unit:
[0172] (1) Input feature segmentation
[0173] The input feature X (CxHxW) is split into 4 parts along the channel dimension, and the number of channels for each part is C / 4:
[0174] X = [X1, X2, X3, X4]
[0175] where [] represents the splitting operation along the channel dimension.
[0176] (2) Dilated convolution extraction
[0177] For each sub-feature X iApply convolution operations with different dilation rates d and output X i1 :
[0178] X i1 = Conv(X i , d), d ∈ {7, 7, 3, 1}, i = 1, 2, 3, 4
[0179] where Conv(X, d) represents the convolution operation with dilation rate d.
[0180] (3) Channel concatenation
[0181] Concatenate the convolved sub - features along the channel dimension to obtain the output feature X concat :
[0182] X concat = [X 11 , X 21 , X 31 , X 41
[0183] where [] represents the concatenation operation on the channel dimension.
[0184] Construction of channel - efficient attention module:
[0185] Upper branch: Channel attention mechanism
[0186] (1) Local feature extraction: Apply 3×3 depthwise convolution to the input feature X input to obtain the local feature X local :
[0187] X local = DepthConv 3×3 (X input )
[0188] (2) Global feature extraction: Apply 5×5 depthwise convolution to the input feature X input to obtain the global feature X global :
[0189] X global = DepthConv 5×5 (X input )
[0190] (3) Feature fusion: Fuse the local feature X local and the global feature X global by pixel - wise addition to obtain X fuscd :
[0191] X fused = X local + X global
[0192] (4) Channel weight generation: Perform global average pooling on X fuscd to compress it to the channel dimension Z global , and then extract the weight Z through one-dimensional convolution Conv1D3 conv :
[0193] Z global = GlobalAvgPool(X fused )
[0194] Z conv = Conv1D3(Z global )
[0195] (5) Activation and channel selection: Perform Sigmoid and Softmax activation operations on Z conv in sequence to generate the channel weight W:
[0196] Z sig = Sigmoid(Z conv )
[0197] W = Softmax(Z sig )
[0198] (6) Channel weighted output: Apply the channel weight W to the original input feature X input to obtain the channel enhanced feature A global :
[0199] A global = X input ·W
[0200] Lower branch: Spatial attention mechanism
[0201] (1) Spatial weight generation: Perform 1×1×1 convolution on the input feature X input to obtain the spatial feature Z spatial , and generate the spatial weight A through Sigmoid activation spatial :
[0202] Z spatial = Conv 1×1×1 (X input )
[0203] A spatial = Sigmoid(Z spatial )
[0204] (2) Spatial weighted output: Apply the spatial weight A spatial to the input feature X input , to obtain the spatial enhanced feature:
[0205] X spatial= X input ·A spatial
[0206] Final feature fusion
[0207] Feature weighted fusion: Add the channel-enhanced feature A global and the spatial-enhanced feature X spatial element-wise to obtain the final output feature:
[0208] X output = A global + X spatial
[0209] Step 2.2: Adopt dense connections in the decoder and pass the shallow feature output to the deep network through skip connections. Introduce the dense connection strategy in the decoder part and pass the features generated by the shallow network to the deep network through skip connections. Each layer of the decoder is regarded as a part of the dense connection. Since there are differences in the spatial dimensions of the shallow and deep feature maps, it is necessary to first perform transposed convolution upsampling on the shallow feature map and then splice it with the deep feature map.
[0210] The specific process in Step 3 is as follows:
[0211] Step 3.1: Initialize the parameters of the skin disease image segmentation network;
[0212] Step 3.2: Set the training parameters. The number of epochs for training is 120, the sample batch size (Batch Size) is 4, the initial learning rate (lr) of the network is set to 1e-4, and the Adam optimizer is used;
[0213] Step 3.3: Use three evaluation metrics, namely Accuracy, Jaccard Index, and F1-Score, to measure the segmentation effect during training;
[0214] Accuracy, that is, the proportion of pixels correctly predicted by the model among all pixels, is shown by the following formula:
[0215]
[0216] Precision, that is, the proportion of pixels correctly classified as lesions among all pixels predicted as lesions, is shown by the following formula:
[0217]
[0218] Recall, that is, the proportion of pixels correctly classified as lesions among all pixels that are lesions, is shown by the following formula:
[0219]
[0220] The F1-Score is the harmonic mean of Precision and Recall. It is a comprehensive evaluation metric, and the formula is as follows:
[0221]
[0222] The Jaccard coefficient is an index used to measure the similarity and difference of sample data. Here, X represents the lesion area in the segmentation result, Y represents the lesion area in the mask image, and ‖ represents calculating the number of pixel points in the image. The closer the Jaccard coefficient is to 1, the better the segmentation result. The formula is as follows:
[0223]
[0224] Among them, True Positive (TP) means the model predicts a lesion and it is actually a lesion. True Negative (TN) means the model predicts normal skin and it is actually normal skin. False Positive (FP) means the model predicts a lesion but it is actually normal skin. False Negative (FN) means the model predicts normal skin but it is actually a lesion.
[0225] Step 3.4: Use the dermoscopic image segmentation model based on the attention mechanism and UNet for iterative training, extract the feature information of skin disease images, and obtain the segmentation result to output the optimal model.
[0226] Comparative experiment:
[0227] The ISIC 2017 dataset is one of the commonly used datasets for skin lesion analysis, mainly used for lesion segmentation and classification tasks in dermoscopic images. This dataset contains approximately 2,000 high-quality RGB images and corresponding lesion segmentation masks and classification labels, covering two major lesion types: melanoma and benign nevi. The annotation is completed by professional dermatologists and has high precision and medical authority. The ISIC 2017 dataset provides a standard benchmark for skin lesion detection and segmentation tasks and is widely used in the development and performance evaluation of deep learning models.
[0228] Compare the present invention with the current mainstream dermoscopic image segmentation models on the ISIC 2017 dataset, and the results are shown in Table 1.
[0229] Table 1 Performance comparison of segmentation models on the ISIC 2017 dataset
[0230] Model Accuracy F1-Score Jaccard Index UNet 0.9624 0.9201 0.8536 U-Net++ 0.9485 0.8831 0.7977 ResUNet++ 0.9600 0.9136 0.8278 The method of the present invention 0.9643 0.9306 0.8702
[0231] As can be seen from Table 1, the segmentation performance of the method of the present invention on the ISIC 2017 dataset is superior to that of UNet, U-Net++, and ResUNet++. It achieves the highest values in terms of Accuracy, F1-Score, and Jaccard Index, and particularly shows significant performance in the Jaccard Index, reaching 0.8702, indicating its obvious advantage in the precise segmentation of lesion areas. In contrast, the performance of U-Net++ is the lowest, and ResUNet++ and UNet perform similarly, but they are still slightly inferior to the method of the present invention in terms of shallow feature learning and global information capture, fully demonstrating the effectiveness of the new method in feature fusion and attention mechanism design.
[0232] Example 2
[0233] Based on Example 1, the same tests were conducted on the ISIC 2018 dataset. Compared with ISIC 2017, the ISIC 2018 dataset has a significant improvement in terms of scale, complexity, and diversity of lesion types. It contains more high-quality RGB images and precise segmentation masks, and introduces richer classification labels and higher-precision expert annotations. As an important extension of the ISIC series, ISIC 2018 provides a more comprehensive and authoritative standardized platform for artificial intelligence research in dermatological diagnosis.
[0234] The present invention was compared with the current mainstream dermatological image segmentation models on the ISIC 2018 dataset, and the results are shown in Table 2.
[0235] Table 2 Performance comparison of segmentation models on the ISIC 2018 dataset
[0236]
[0237]
[0238] As can be seen from Table 2, the segmentation performance of the method of the present invention on the ISIC 2018 dataset is significantly superior to other models. Among the three indicators of Accuracy, F1-Score, and Jaccard Index, the method of the present invention reaches 0.9643, 0.9288, and 0.8605 respectively. Especially in the Jaccard Index, it has an improvement of about 0.42% compared with UNet. Compared with U-Net++ and Attention-UNet, the method of the present invention significantly enhances the segmentation accuracy and coverage rate of the lesion area by optimizing the feature extraction and fusion strategy, demonstrating stronger model robustness and superior segmentation ability.
[0239] Example 3
[0240] Based on Example 2, the method of the present invention is further compared with other methods. The present invention is compared with the current mainstream dermatological image segmentation models on the ISIC 2018 dataset, and the results are shown in Table 3 below.
[0241] Table 3 Performance comparison of segmentation models on the ISIC 2018 dataset
[0242] Model Accuracy F1-Score Jaccard Index TransFuse 0.9456 0.9136 0.8597 UTNetV2 0.9427 0.8861 0.8126 SANet 0.9433 0.9213 0.8265 The method of the present invention 0.9643 0.9288 0.8605
[0243] As can be seen from Table 3, the segmentation performance of the method of the present invention on the ISIC 2018 dataset is better than that of TransFuse, UTNetV2, and SANet. Among the three indicators of Accuracy, F1-Score, and Jaccard Index, the method of the present invention reaches 0.9643, 0.9288, and 0.8605 respectively. Especially, the improvement in F1-Score and Jaccard Index is significant, indicating its obvious advantages in the segmentation accuracy and integrity of the lesion area. In contrast, although TransFuse performs well in Accuracy and Jaccard Index, it is still slightly inferior to the method of the present invention; while the overall performance of UTNetV2 and SANet is weak, especially in capturing the details and boundaries of the lesions. The method of the present invention demonstrates stronger segmentation ability and robustness with a better feature optimization and fusion strategy.
[0244] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A dermoscopic image segmentation method based on the attention mechanism and UNet, characterized in that, It includes the following steps: Step 1: Data preprocessing, including image denoising, image enhancement and division of the dataset; Step 2: Construct a dermoscopic image segmentation network integrating an attention mechanism and UNet. This network consists of an encoder and a decoder. The first three layers of the encoder and decoder use depthwise separable convolutions. An Efficient Dilated Fusion Attention module (EDFA) is introduced in the last two layers of the encoder and decoder. Dense connections are adopted in the decoder, and shallow feature outputs are passed to the deep network through skip connections; Step 3: Put the processed dermopathy image training set into the dermoscopic image segmentation network model based on the attention mechanism and UNet for training to obtain the optimal model; Step 4: After training, input the test set into the optimal model to detect the lesion segmentation results in dermopathy images.
2. The dermoscopic image segmentation method based on the attention mechanism and UNet according to claim 1, characterized in that The said Step 1 includes the following steps: Step 1.1: Uniformly adjust the size of each image in the original dermopathy image dataset to 256×256; Step 1.2: Apply grayscale conversion, black hat operation, threshold segmentation, image inpainting, bilateral filtering and histogram equalization to the dermopathy image dataset whose size has been adjusted in Step 1.1; The formulas used in Step 1.2 are as follows: (1) Grayscale image conversion formula G(i) = 0.299R(i) + 0.587G(i) + 0.114B(i) where G(i) is the grayscale value at image index i, and R(i), G(i), B(i) are the pixel values of the image in the red, green and blue channels respectively. (2) Black hat operation Perform morphological black hat operation using a 5x5 elliptical structuring element. The formula is as follows: Among them, src is the original input image, is a morphological closing operation, which means dilation followed by erosion. element is an elliptical structuring element used to adapt to the shape characteristics of hair, and dst is the result image after the black hat operation; (3) Threshold segmentation Perform global threshold segmentation on the result of the black hat operation to generate a binary image. The formula is as follows: where I′(i) represents the new value of this pixel after threshold processing, thresh is the threshold, when the pixel value is higher than the threshold, it is set to maxVal, otherwise it is 0, and maxVal is the maximum pixel value after binarization. (4) Image inpainting Use the Fast Marching Method (FMM) to repair the missing areas of the binary image. The inpainted pixel values are calculated by weighting the surrounding known pixels: w(p,q) = f d (p,q)·f g (p,q)·f c (p,q) Among them, w(p, q) is the weight of pixel q for the inpainting point p, f d (p, q) is the direction factor, which controls the inpainting to be preferentially performed along the normal direction, f g (p, q) is the geometric distance factor, f c (p, q) is the gray-scale similarity factor, σ g and σ c are smoothing parameters; (5) Bilateral filtering for denoising Apply bilateral filtering to the repaired image to achieve noise reduction and edge preservation by combining spatial position and grayscale information: where I′(p) is the output value of pixel p, Ω is the filtering window, f g (p, q) is the spatial distance weight that measures the geometric distance between pixels p and q, f c (p, q) is the gray-scale similarity weight that measures the degree of gray-scale value similarity, and W(p) is the normalization factor; 3. A dermoscopic image segmentation method based on an attention mechanism and UNet according to claim 1, characterized in that, The said Step 2 includes the following steps: Step 2.1: In the first three layers of the encoder and decoder of the model, replace the conventional convolution with depthwise separable convolution. First, process the input feature map channel by channel through depth convolution, and introduce a dilation rate r = 2 to expand the receptive field. Subsequently, use pointwise convolution to linearly combine the output of the depth convolution to generate the final feature map. Add a 2×2 max pooling layer to each layer of the encoder to select the maximum value in the local area of the input feature map using a fixed window and stride to generate the downsampled feature map; The formulas used in Step 2.1 are as follows: (1) Depth convolution Depth convolution performs convolution operations independently for each channel. After introducing a dilation rate r = 2, the calculation formula is: Among them, y c (u, v) represents the convolution output value at position (u, v) after deep convolution calculation, and x c (u, v) is the pixel value of the input feature map at position (u, v) and channel c, and w c (i, j) is the weight of the deep convolution kernel on channel c, k is the size of the convolution kernel, and r is the dilation rate, indicating the interval between convolution kernel elements, which is used to expand the receptive field; (2) Pointwise convolution Linearly combine the output of the depth convolution through a 1×1 convolution kernel. The formula is: Among them, z m (u, v) is the pixel value of the pointwise convolution output feature map at position (u, v) and channel m, y c (u, v) is the pixel value of the depth convolution output, w c,m is the weight of the pointwise convolution kernel, connecting the input channel c and the output channel m, b m is the bias term of the pointwise convolution, and C is the number of input channels; (3) Max pooling for downsampling Pooling operations add a 2×2 max pooling layer in each layer of the encoder, using the maximum value in the local area as the feature value after downsampling. The calculation formula is as follows: where z c (h′, w′) is the pixel value of the pooled output feature map at channel c and position (h′, w′), and x c (h, w) is the pixel value of the input feature map at channel c and position (h, w), is the pixel region corresponding to the pooling window; Step 2.2: An efficient dilated fusion attention module is introduced in the fourth and fifth layers of the encoder and decoder. This module consists of a dilated convolution fusion unit and a spatial and channel efficient attention module; First, the input features are divided into four branches in the dilated convolution fusion unit, and features are extracted through convolutions with different dilation rates. Subsequently, the four-way features are concatenated in the channel dimension; Next, after performing 1x1 convolution on the concatenated features, the result is input into the channel efficient attention module; In the channel efficient attention module, the features are divided into upper and lower branches for separate processing. The upper half is further divided into two paths through feature separation operations, and depthwise separable convolutions with 3x3 and 5x5 are used to extract features respectively. Subsequently, the two-way features are concatenated in the channel dimension to achieve multi-scale feature interaction, and then through global average pooling and one-dimensional convolution processing, the channel information interaction is optimized. After activation by Sigmoid and Softmax, cross-channel information weights are generated for selecting important channel features; The lower half performs a convolution operation on the input features, compressing the channel information of the feature map into a single dimension. Subsequently, spatial weights are generated through an activation function and multiplied element-wise with the original features to enhance the feature expression at key positions. Finally, the features extracted from the upper and lower branches are added element-wise to generate the fused output features; Construction of the dilated convolution fusion unit: (1) Input feature segmentation The input feature X (CxHxW) is split into 4 parts along the channel dimension, and the number of channels for each part is C / 4: X = [X1, X2, X3, X4] Among them, [] represents the splitting operation on the channel dimension; (2) Dilated convolution extraction For each sub-feature X i Apply a convolution operation with different dilation rates d and output X i1 : X i1 = Conv(X i , d), d ∈ {7, 7, 3, 1}, i = 1, 2, 3, 4 Among them, Conv(X, d) represents a convolution operation with a dilation rate of d; (3) Channel concatenation Concatenate the convolved sub - features along the channel dimension to obtain the output feature X concat : X concat = [X 11 , X 21 , X 31 , X 41 Among them, [] represents the concatenation operation on the channel dimension; Construction of the channel efficient attention module: Upper branch: Channel attention mechanism (1) Local feature extraction: For the input feature X input Perform 3×3 depth convolution to obtain the local feature X local : X local = DepthConv 3×3 (X input ) (2) Global feature extraction: For the input feature X input Perform a 5×5 depth convolution to obtain the global feature X global : X global = DepthConv 5×5 (X input ) (3) Feature fusion: Fuse the local feature X local and the global feature X global by pixel-by-pixel addition to obtain X fuscd : X fused = X local + X global (4) Channel weight generation: Perform global average pooling on X fuscd to compress it to the channel dimension Z global , and then extract the weight Z through one-dimensional convolution Conv1D3 conv : Z global = GlobalAvgPool(X fused ) Z conv = Conv1D3(Z global ) (5) Activation and Channel Selection: For Z conv Perform Sigmoid and Softmax activation operations in sequence to generate channel weights W: Z sig = Sigmoid(Z conv ) W = Softmax(Z sig ) (6) Channel weighted output: Apply the channel weight W to the original input feature X input to obtain the channel enhanced feature A global : A global = X input ·W Lower branch: Spatial attention mechanism (1) Spatial weight generation: Convolve the input feature X input with a 1×1×1 convolution to obtain the spatial feature Z spatial , and generate the spatial weight A spatial through Sigmoid activation: Z spatial = Conv 1×1×1 (X input ) A spatial = Sigmoid(Z spatial ) (2) Spatial weighted output: Apply the spatial weight A spatial to the input feature X input , and obtain the spatially enhanced feature: X spatial = X input · A spatial Final feature fusion Feature weighted fusion: Add the channel-enhanced feature A global and the space-enhanced feature X spatial element by element to obtain the final output feature: X output = A global + X spatial Step 2.2: Dense connections are adopted in the decoder. The shallow feature outputs are passed to the deep network through skip connections. A dense connection strategy is introduced in the decoder part, and the features generated by the shallow network are passed to the deep network through skip connections. Each layer of the decoder is regarded as part of the dense connection. Due to the difference in the spatial sizes of the shallow and deep feature maps, it is necessary to first perform transposed convolution upsampling on the shallow feature map and then concatenate it with the deep feature map.
4. A dermoscopic image segmentation method based on an attention mechanism and UNet according to claim 1, characterized in that Step 3 includes the following steps: Step 3.1: Initialize the parameters of the skin disease image segmentation network; Step 3.2: Set the training parameters. The number of training epochs is 120, the sample batch size (Batch Size) is 4, the initial learning rate (lr) of the network is set to 1e-4, and the Adam optimizer is used; Step 3.3: The training uses three evaluation metrics, namely Accuracy, Jaccard Index, and F1-Score, to measure the segmentation effect; Accuracy, that is, the proportion of pixels correctly predicted by the model among all pixels, is shown by the following formula: Precision is the proportion of pixels correctly classified as lesions among all pixels predicted as lesions, and the formula is shown as follows: Recall is the proportion of pixels correctly classified as lesions among all pixels that are lesions, and the formula is shown as follows: F1-Score is the harmonic mean of Precision and Recall, and the formula is shown as follows: X represents the lesion area in the segmentation result, Y represents the lesion area in the mask image, and ‖ represents calculating the number of pixels in the image. The formula is shown as follows: Among them, True Positive (TP) means the model predicts as a lesion and it is actually a lesion; True Negative (TN) means the model predicts as normal skin and it is actually normal skin; False Positive (FP) means the model predicts as a lesion but it is actually normal skin; False Negative (FN) means the model predicts as normal skin but it is actually a lesion; Step 3.4: Use a dermoscopic image segmentation model based on the attention mechanism and UNet for iterative training, extract the feature information of skin disease images, and obtain the segmentation result to output the optimal model.
Citation Information
Cited By
Automatic bacterial colony identifying and counting method based on machine learning segmentation
CN121482784A
Key structure image segmentation method and system for obstetrical ultrasound
CN122049379A