Dermoscope image segmentation method based on attention mechanism and multi-scale information interaction
Through the improved hollow space pyramid pooling module, axial gating attention mechanism and reverse enhancing attention module, combined with data preprocessing, the problems of multi-scale characteristics and computational complexity in skin lesion area segmentation are solved, and high-precision and efficient skin lesion segmentation are achieved.
Patent Information
- Application Number
- CN202510357385.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
AI Technical Summary
The traditional skin lesion area segmentation method is difficult to take into account both global and local information when the multi-scale characteristics are strong and the size and shape of the lesion area are large. The edges are blurred and the contrast is low, the calculation complexity is high, and the reverse enhancement strategy for low-response areas is lacking, resulting in poor segmentation accuracy and boundary recognition.
The dermatoscope image segmentation method based on attention mechanism and multi-scale information interaction is adopted. Through the improved hollow space pyramid pooling module, axial gating attention mechanism and reverse enhancing attention module, combined with data preprocessing strategies, multi-scale features are extracted and computational complexity is reduced, and boundary recognition capabilities are enhanced.
It improves the segmentation accuracy and boundary recognition ability of skin lesion areas, adapts to lesion targets of different sizes and morphology, reduces the computational complexity, and enhances the generalization ability of the model and the meticulousness and accuracy of the segmentation results.
Smart Images

Figure CN120339607A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning image segmentation, and particularly to a dermoscopic image segmentation method based on an attention mechanism and multi-scale information interaction. Background Art
[0002] Skin cancer is one of the most common malignant tumors worldwide. Early detection and accurate segmentation of the lesion area are crucial for improving the diagnostic accuracy and treatment effect. Dermoscopy is a commonly used non-invasive detection method that can provide high-resolution skin images and is helpful for observing skin lesions. However, due to the complex color, morphology, and edge changes of skin lesions, and the large differences in skin types among different patients, traditional image segmentation methods based on thresholds, region growing, and manual feature extraction are difficult to adapt to complex lesion areas and are easily affected by interference such as illumination, noise, and skin texture, resulting in low segmentation accuracy.
[0003] In recent years, deep learning technologies, especially convolutional neural networks (CNNs), have made significant progress in the field of medical image segmentation. Classic semantic segmentation models such as U-Net and DeepLabV3+ can effectively extract deep features of skin lesion areas. However, the existing methods still have the following problems: (1) Dermoscopic images have strong multi-scale characteristics, and the size and shape of lesion areas vary greatly. Single-scale feature extraction methods are difficult to balance global and local information; (2) The edges of lesion areas are blurred and the contrast with normal skin tissue is low, which easily leads to the problem of unclear boundaries; (3) Although the attention mechanism can improve the feature expression ability, most existing methods use standard self-attention mechanisms, which have a large computational complexity and are difficult to balance computational efficiency and feature extraction ability; (4) Most existing feature enhancement strategies are based on explicit attention enhancement, and lack a reverse enhancement strategy for low-response areas, making it difficult to fully mine the detailed features of lesion areas. How to solve the above technical problems is the subject faced by the present invention. Summary of the Invention
[0004] The purpose of the present invention is to provide a dermoscopic image segmentation method based on an attention mechanism and multi-scale information interaction, which can effectively combine global and local feature information, improve the segmentation accuracy of skin lesion areas, and enhance the recognition ability of boundary details. Compared with traditional segmentation methods, the present invention can not only maintain a high segmentation accuracy under complex backgrounds and low-contrast conditions, but also reduce the computational complexity through an improved attention mechanism, making it more practical.
[0005] To achieve the above invention purpose, the technical solution adopted by the present invention is specifically: a dermoscopic image segmentation method based on an attention mechanism and multi-scale information interaction, including the following steps:
[0006] Step 1: Construct a high-precision skin disease image dataset;
[0007] Step 2: Preprocess the images in the dataset, including enhancing and normalizing skin melanosis images, and dividing them into a training set and a test set according to a certain proportion to ensure the optimization of data quality and model performance;
[0008] Step 3: Construct a dermoscopic image segmentation model based on the attention mechanism and multi-scale information interaction;
[0009] Step 4: Input the preprocessed skin disease image training set into the dermoscopic image segmentation network based on the attention mechanism and multi-scale information interaction for training to optimize the model performance to obtain the optimal segmentation effect;
[0010] Step 5: After the training is completed, input the test set into the optimal model to segment the lesions of the skin disease images and evaluate the segmentation effect.
[0011] The specific process in Step 1 is as follows:
[0012] Step 1.1: Read and prepare the data of the relevant skin disease image dataset;
[0013] Step 1.2: Remove low-quality images, calculate the sharpness of the images using the Laplace transform, and a low variance indicates blurriness;
[0014] (1) Calculate the Laplace transform of the image:
[0015]
[0016] where I ij is the pixel value, and ΔI ij is the gradient value after the Laplace transform. Calculate the variance V of the whole image as the sharpness index.
[0017] (2) Set the sharpness threshold:
[0018] If V ≤ 100, the image is determined to be blurry and excluded.
[0019] Step 1.3: Detect duplicate data, calculate the structural similarity (SSIM), and remove similar or duplicate images in the dataset to improve data diversity;
[0020] (1) Calculate the mean μ x , μ y , the standard deviation σ x , σ y and the covariance σ xy
[0021]
[0022] (2) Set the similarity threshold
[0023] If SSIM(x, y) ≥ 0.95, it is considered that the two images are duplicates, and one of them is deleted.
[0024] Step 1.4: Output a high-quality and non-duplicate skin disease image dataset.
[0025] The specific process in Step 2 is as follows:
[0026] Step 2.1: Data augmentation, geometrically flipping, translating, and scaling dermoscopic images;
[0027] (1) Random rotation, randomly rotate the image to simulate the scenario of observing skin lesions from multiple angles, which helps improve the model's adaptability to rotational invariance.
[0028] Rotation matrix formula
[0029]
[0030] where θ is the rotation angle, and the range is ±20°.
[0031] Each pixel coordinate (x, y) in the image calculates the new coordinate (x′, y′) through the rotation matrix:
[0032]
[0033] (2) Horizontal / vertical offset, perform random horizontal and vertical translation operations on the image to enable the model to better adapt to the offset of the target position.
[0034] Translation matrix formula:
[0035]
[0036] where Δx is the horizontal offset, and the value range is ±5% of the image width, Δy is the vertical offset, and the value range is ±5% of the image height.
[0037] After translation, the new coordinates of each pixel point:
[0038]
[0039] (3) Random scaling, simulate lesion areas at different distances by randomly scaling the image to improve the model's recognition ability for multi-scale lesions.
[0040] Scaling matrix formula:
[0041]
[0042] where s x ,sy is the scaling ratio, ranging from 0.95 to 1.05.
[0043] New coordinates of the scaled pixel points:
[0044]
[0045] (4) Random flipping, simulating a mirror scene by horizontally flipping the image to enhance the model's adaptability to symmetric lesions.
[0046] Horizontal flipping matrix formula:
[0047]
[0048] where W is the image width.
[0049] New coordinates of the flipped pixel points:
[0050]
[0051] Step 2.2: Standardize the data using Min - Max normalization;
[0052]
[0053] where I min is the minimum pixel value in the dataset, and I max is the maximum pixel value in the dataset.
[0054] Step 2.3: Divide the dataset into an 80% training set, a 10% validation set, and a 10% test set.
[0055] The specific process in Step 3 is as follows:
[0056] Step 3.1: Use the improved atrous spatial pyramid pooling module to perform spatial pyramid pooling operations by adjusting the dilation rate, effectively capturing multi - scale information under filters with different receptive fields, thereby enhancing the ability to understand global and local features;
[0057] Branch 1: 1x1 convolution branch
[0058] Use a 1×1 convolution operation alone to extract global semantic information.
[0059] Branch 2: Global average pooling branch
[0060] The input feature map first undergoes a global average pooling (GAP) operation to map the features of each channel to a global average value, thereby extracting global context information. Subsequently, the output of GAP passes through a 1×1 convolution to adjust the number of channels and generate a feature representation with semantic information. Finally, through an upsampling operation, the features are restored to the same spatial size as the original input, providing global semantic support for subsequent feature fusion.
[0061] Branch 3: Dilated Convolution Branch
[0062] By using multiple 3×3 dilated convolution operations with different dilation rates, multi-scale context information can be extracted at different receptive fields. The size of the dilation rate r directly affects the receptive field range: a smaller dilation rate (e.g., r = 2) can capture local details, while a larger dilation rate (e.g., r = 16) is used to extract global context information. The dilation rates adopted in the dilated convolution branch design are r = 2, 4, 8, 16, which can effectively combine local and global features and provide support for the extraction of multi-scale information.
[0063] Step 3.2: To capture multi-scale information, an axial gating mechanism is adopted to decompose self-attention into longitudinal and lateral attention calculations, combined with learnable gating parameters, reducing the computational complexity and enhancing the feature modeling ability in the height and width dimensions;
[0064] (1) Input Feature Processing
[0065] Input feature map X in First, a 1×1 convolutional layer is used to reduce the dimensionality of the channels, reducing the computational complexity. The formula is:
[0066] X conv = Conv 1×1 (X in )
[0067] Then, it passes through a batch normalization layer to normalize the feature distribution. The formula is:
[0068] X bn = BN(X conv )(2) Axial Attention Calculation
[0069] The features are respectively input into the longitudinal attention and lateral attention modules for axial decomposition calculation:
[0070] 1) Longitudinal Attention Modeling
[0071] For the features at each position (i, j), calculate the longitudinal attention output based on the height dimension H
[0072]
[0073] Among them, W Q and W K and W V are learnable projection matrices for queries, keys, and values. G Q and G K and G V 1 and G V2 are gating parameters that dynamically adjust the attention distribution. is the position bias term for vertical attention.
[0074] 2) Horizontal attention modeling
[0075] For the features at each position (i, j), calculate the horizontal attention output based on the width dimension w
[0076]
[0077] Among them, W Q and W K and W V are learnable projection matrices for queries, keys, and values. G Q and G K and G V1 and G V2 are gating parameters that dynamically adjust the attention distribution. is the position bias term for horizontal attention.
[0078] (3) Vertical and horizontal feature fusion
[0079] Perform weighted fusion on the outputs of vertical attention and horizontal attention :
[0080] X attn = Y H + Y W
[0081] (4) Output feature processing
[0082] The fused feature X attn is passed through a 1×1 convolution again to adjust the channel dimension and is normalized by BN:
[0083] X out = BN(Conv 1×1 (X attn ))
[0084] Finally, add X out to the input X in through a residual connection to obtain the final output:
[0085] X final = X in + Xout
[0086] Step 3.3: By using the reverse enhanced attention module, deeply mine the tumor region features through the reverse deletion strategy, combine complementary regions and detailed information, effectively obtain accurate tumor region segmentation results, and improve the fineness and accuracy of segmentation;
[0087] (1) High-level feature input
[0088] After decoding and upsampling, the obtained high-resolution feature map is input into the Sigmoid activation function to perform pixel-level classification. The decoder restores the deep semantic information extracted in the encoding stage and combines upsampling to improve the spatial resolution, making the feature expression more refined. Finally, these features are mapped to probability values between 0 and 1 through Sigmoid, and the value of each pixel represents the possibility of its belonging to the target class, thereby achieving an accurate pixel-level classification task.
[0089] (2) Generate reverse attention weights
[0090] 1) Sigmoid activation
[0091] Input the upsampled feature S up into the Sigmoid function to generate a probability distribution:
[0092] W = 1 - Sigmoid(S up )
[0093] where S up is the upsampled feature and W represents the reverse attention weight.
[0094] 2) Reverse operation
[0095] Use the 1 - Sigmoid method to highlight the low-response regions and mine the detailed information of the tumor boundary.
[0096] (3) Feature fusion
[0097] Multiply the reverse attention weight W and the high-level feature H element-wise (Multiplication) to generate a weighted feature map:
[0098] F weighted = W ⊙ H
[0099] where H represents the high-level feature and ⊙ represents element-wise multiplication.
[0100] (4) Output features
[0101] The weighted feature F weighted represents the enhanced convolutional feature, which can be used for further optimization of the segmentation result.
[0102] The specific process in Step 4 is as follows:
[0103] Step 4.1: Initialize the parameters of the skin disease image segmentation network;
[0104] Step 4.2: Set the training parameters, the number of epochs for training is 200, the sample batch size is 4, the initial learning rate (lr) of the network is set to 0.0001, and the Adam optimizer is used;
[0105] Step 4.3: The training is quantitatively evaluated using two metrics (i.e., Dice and IoU). The calculation processes of the Dice metric and the IoU metric are as follows:
[0106]
[0107] Among them, |x| and |y| represent the label value and the predicted value respectively. |x∩y| represents the intersection of |x| and |y|. |x∪y| represents the union of |x| and |y|. Use 0.5 times the binary cross-entropy (BCE) loss and the Dice loss as the model loss function. The calculation process is as follows:
[0108] Loss = 0.5L BCE (y i , x i ) + L Dice (y i , x i )
[0109] Among them, x i represents the true value, and y i represents the predicted value. L Dice is 1 minus the Dice value.
[0110] Step 4.4: Use the dermoscopic image segmentation network model based on the attention mechanism and multi-scale information interaction for iterative training, extract the feature information of the skin disease image, and obtain the segmentation result to output the optimal model.
[0111] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0112] 1. The present invention adopts an improved atrous spatial pyramid pooling module, combines convolutional operations with different receptive fields, and effectively extracts multi-scale feature information. The global semantic information is extracted through the 1×1 convolution branch, the global average pooling branch enhances the global context expression ability, and the atrous convolution branch extracts local and global feature information at different dilation rates (r = 2, 4, 8, 16), thereby improving the segmentation accuracy of the skin lesion area and adapting to skin lesion targets of different sizes and shapes.
[0113] 2. Through the axial gated attention mechanism, the present invention decomposes self-attention calculation into vertical attention and horizontal attention, avoiding the high computational overhead caused by directly calculating global attention. By reducing the dimension through convolution to reduce the amount of calculation and combining batch normalization to standardize the feature distribution, the calculation of the attention mechanism becomes more efficient, ensuring both the global feature extraction ability and improving the inference speed and computational efficiency.
[0114] 3. Traditional attention mechanisms mainly focus on high-response regions, while the present invention adopts a reverse enhancement attention mechanism, which highlights low-response regions through a 1 - Sigmoid reverse operation, effectively mining the detailed information of the tumor edge and combining pointwise multiplication to enhance the feature expression ability of the lesion boundary. This strategy can more accurately segment lesion regions with blurred contours or low contrast, improving the fineness and accuracy of the segmentation results.
[0115] 4. In the data preprocessing stage, the present invention calculates the image sharpness using Laplace transform, eliminates low-quality images, and combines structural similarity to remove similar or duplicate images in the dataset, ensuring data diversity. In addition, data augmentation strategies such as rotation, translation, scaling, and flipping are used to enhance the model's adaptability to different lighting conditions, shooting angles, and lesion morphological changes, improving the model's generalization ability on different datasets and enabling it to be applicable to various dermoscopic image segmentation tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0116] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0117] Figure 1 It is the overall flowchart of a dermoscopic image segmentation method based on attention mechanism and multi-scale information interaction provided by the present invention.
[0118] Figure 2 It is a schematic diagram of an improved atrous spatial pyramid pooling module in a dermoscopic image segmentation method based on attention mechanism and multi-scale information interaction provided by the present invention.
[0119] Figure 3 It is a schematic diagram of three branches in an improved atrous spatial pyramid pooling module in a dermoscopic image segmentation method based on attention mechanism and multi-scale information interaction provided by the present invention.
[0120] Figure 4 It is a schematic diagram of an axial gated attention mechanism in a dermoscopic image segmentation method based on attention mechanism and multi-scale information interaction provided by the present invention.
[0121] Figure 5 It is a schematic diagram of a reverse enhancement attention module in a dermoscopic image segmentation method based on attention mechanism and multi-scale information interaction provided by the present invention.
[0122] Figure 6 Schematic diagram of the network model constructed in a dermoscopic image segmentation method based on attention mechanism and multi-scale information interaction provided by the present invention. Specific embodiments
[0123] The following will further explain the present invention in detail with reference to the accompanying drawings, so that those skilled in the art can understand the present invention more deeply and be able to implement it. However, the following is only used to explain the present invention by reference to examples and does not limit the present invention.
[0124] Embodiment 1
[0125] See Figures 1 to 6 , this embodiment proposes a dermoscopic image segmentation method based on attention mechanism and multi-scale information interaction. In this embodiment, the experiments of the present invention are carried out on Pytorch, and network training and testing are carried out on NVIDIA 4060 GPU. During the experiment, the model is trained and evaluated on the ISIC public dataset. And this dataset is used for comparison with other methods.
[0126] A dermoscopic image segmentation method based on attention mechanism and multi-scale information interaction includes the following steps:
[0127] Step 1: Construct a high-precision skin disease image dataset;
[0128] Step 2: Preprocess the images in the dataset, including enhancement and normalization of skin melanosis images, and divide them into training set and test set according to a certain proportion to ensure data quality and model performance optimization;
[0129] Step 3: Construct a dermoscopic image segmentation model based on attention mechanism and multi-scale information interaction;
[0130] Step 4: Input the preprocessed skin disease image training set into the dermoscopic image segmentation network based on attention mechanism and multi-scale information interaction for training, and optimize the model performance to obtain the optimal segmentation effect;
[0131] Step 5: After the training is completed, input the test set into the optimal model to segment the lesions of the skin disease images and evaluate the segmentation effect.
[0132] The specific process in Step 1 is as follows:
[0133] Step 1.1: Read and prepare the data of relevant skin disease image datasets;
[0134] Step 1.2: Remove low-quality images, calculate the sharpness of the images using Laplace transform, and low variance indicates blurriness;
[0135] (1) Calculate the Laplace transform of the image:
[0136]
[0137] where I ij is the pixel value, and ΔI ij is the gradient value after the Laplace transform. Calculate the variance V of the entire image as the clarity index.
[0138] (2) Set the clarity threshold:
[0139] If V ≤ 100, the image is determined to be blurred and excluded.
[0140] Step 1.3: Detect duplicate data, calculate the structural similarity (SSIM), and remove similar or duplicate images in the dataset to improve data diversity;
[0141] (1) Calculate the mean μ x , μ y , the standard deviation σ x , σ y and the covariance σ xy
[0142]
[0143] (2) Set the similarity threshold
[0144] If SSIM(x, y) ≥ 0.95, it is considered that the two images are duplicates, and one of them is deleted.
[0145] Step 1.4: Output a high-quality and non-duplicate skin disease image dataset.
[0146] The specific process in Step 2 is as follows:
[0147] Step 2.1: Data augmentation, geometrically flipping, translating, and scaling the dermoscopic images;
[0148] (1) Random rotation, randomly rotate the image to simulate the scenario of observing skin lesions from multiple angles, which helps to improve the model's adaptability to rotational invariance.
[0149] The rotation matrix formula
[0150]
[0151] where θ is the rotation angle, and the range is ±20°.
[0152] Each pixel coordinate (x, y) in the image calculates the new coordinate (x′, y′) through the rotation matrix:
[0153]
[0154] (2) Horizontal / vertical offset: Perform random horizontal and vertical translation operations on the image to enable the model to better adapt to the offset of the target position.
[0155] Translation matrix formula:
[0156]
[0157] Among them, Δx is the horizontal offset, and its value range is ±5% of the image width. Δy is the vertical offset, and its value range is ±5% of the image height.
[0158] After translation, the new coordinates of each pixel point:
[0159]
[0160] (3) Random scaling: Simulate lesion areas at different distances by randomly scaling the image to improve the model's recognition ability for multi-scale lesions.
[0161] Scaling matrix formula:
[0162]
[0163] Among them, s x , s y is the scaling ratio, and the range is 0.95, 1.05.
[0164] After scaling, the new coordinates of the pixel points:
[0165]
[0166] (4) Random flipping: Simulate a mirror image scenario by horizontally flipping the image to enhance the model's adaptability to symmetric lesions.
[0167] Horizontal flipping matrix formula:
[0168]
[0169] Among them, W is the image width.
[0170] After flipping, the new coordinates of the pixel points:
[0171]
[0172] Step 2.2: Standardize the data using Min-Max normalization;
[0173]
[0174] Among them, I min is the minimum value of the pixel values in the dataset, I maxis the maximum value of the pixel values in the dataset.
[0175] Step 2.3: Divide the dataset into an 80% training set, a 10% validation set, and a 10% test set.
[0176] The specific process in Step 3 is as follows:
[0177] Step 3.1: Use the improved atrous spatial pyramid pooling module to perform spatial pyramid pooling operations by adjusting the dilation rate, effectively capturing multi-scale information under filters with different receptive fields, thereby enhancing the ability to understand global and local features;
[0178] Branch 1: 1x1 convolution branch
[0179] Use the 1×1 convolution operation alone to extract global semantic information.
[0180] Branch 2: Global average pooling branch
[0181] The input feature map first undergoes a global average pooling (GAP) operation to map the features of each channel to a global average value, thereby extracting global context information. Subsequently, the output of GAP passes through a 1×1 convolution to adjust the number of channels and generate a feature representation with semantic information. Finally, the features are restored to the same spatial size as the original input through an upsampling operation, providing global semantic support for subsequent feature fusion.
[0182] Branch 3: Atrous convolution branch
[0183] By using 3×3 atrous convolution operations with multiple different dilation rates, multi-scale context information can be extracted under different receptive fields. The size of the dilation rate r directly affects the receptive field range: a smaller dilation rate (e.g., r = 2) can capture local details, while a larger dilation rate (e.g., r = 16) is used to extract global context information. The dilation rates adopted in the atrous convolution branch design are r = 2, 4, 8, 16, which can effectively combine local and global features and provide support for the extraction of multi-scale information.
[0184] Step 3.2: For the captured multi-scale information, adopt an axial gating mechanism to decompose self-attention into longitudinal and lateral attention calculations, combined with learnable gating parameters, to reduce the computational complexity and enhance the feature modeling ability in the height and width dimensions;
[0185] (1) Input feature processing
[0186] Input feature map X in First, reduce the number of channels through a 1×1 convolutional layer to reduce the computational complexity. The formula is:
[0187] X conv = Conv1×1 (X in )
[0188] Then, the feature distribution is normalized through the batch normalization layer, and the formula is:
[0189] X bn = BN(X conv )
[0190] (2) Axial attention calculation
[0191] The features are respectively input into the vertical attention and horizontal attention modules for axial decomposition calculation:
[0192] 1) Vertical attention modeling
[0193] For the features at each position (i, j), the vertical attention output is calculated based on the height dimension H
[0194]
[0195] where W Q , W K , W V are learnable projection matrices for queries, keys, and values, G Q , G K , G V 1, G V2 are gating parameters that dynamically adjust the attention distribution, is the position bias term of the vertical attention.
[0196] 2) Horizontal attention modeling
[0197] For the features at each position (i, j), the horizontal attention output is calculated based on the width dimension w
[0198]
[0199] where W Q , W k , W V are learnable projection matrices for queries, keys, and values, G Q , G K , G V1 , G V2 are gating parameters that dynamically adjust the attention distribution, is the position bias term of the horizontal attention.
[0200] (3) Vertical and horizontal feature fusion
[0201] The outputs of the vertical attention and the horizontal attention are weighted and fused:
[0202] X attn = Y H + Y W
[0203] (4) Output feature processing
[0204] The fused feature X attn Is adjusted in channel dimension again through a 1×1 convolution and normalized by BN:
[0205] X out = BN(Conv 1×1 (X attn ))
[0206] Finally, X out Is added to the input X in To obtain the final output:
[0207] X final = X in + X out
[0208] Step 3.3: By using the reverse enhancement attention module, deeply mine the tumor region features through the reverse deletion strategy, combine the complementary regions and detailed information, and effectively obtain accurate tumor region segmentation results, improving the fineness and accuracy of the segmentation;
[0209] (1) High-level feature input
[0210] After decoding and upsampling, the obtained high-resolution feature map is input into the Sigmoid activation function to perform pixel-level classification. The decoder restores the deep semantic information extracted in the encoding stage and combines upsampling to improve the spatial resolution, making the feature expression more refined. Finally, these features are mapped to probability values between 0 and 1 through Sigmoid, and the value of each pixel point represents the possibility of its belonging to the target category, thus achieving an accurate pixel-level classification task.
[0211] (2) Generate reverse attention weights
[0212] 1) Sigmoid activation
[0213] The upsampled feature S up Is input into the Sigmoid function to generate a probability distribution:
[0214] W = 1 - Sigmoid(S up )
[0215] Where S up Is the upsampled feature, and W represents the reverse attention weight.
[0216] 2) Reverse operation
[0217] Use the 1 - Sigmoid method to highlight the low - response regions and mine the detailed information of the tumor boundary.
[0218] (3) Feature fusion
[0219] Multiply the reverse attention weight W and the high - level feature H point - by - point (Multiplication) to generate a weighted feature map:
[0220] F weighted = W ⊙ H
[0221] where H represents the high - level feature and ⊙ represents point - by - point multiplication.
[0222] (4) Output feature
[0223] The weighted feature F weighted represents the enhanced convolutional feature and can be used for further optimization of the segmentation result.
[0224] The specific process in step 4 is as follows:
[0225] Step 4.1: Initialize the parameters of the skin disease image segmentation network;
[0226] Step 4.2: Set the training parameters, the number of epochs for training is 200, the sample batch size is 4, the initial learning rate (lr) of the network is set to 0.0001, and use the Adam optimizer;
[0227] Step 4.3: Use two metrics (i.e., Dice, IoU) for quantitative evaluation during training. The calculation processes of the Dice metric and the IoU metric are as follows:
[0228]
[0229] where |x| and |y| represent the label value and the predicted value respectively. |x ∩ y| represents the intersection of |x| and |y|. |x ∪ y| represents the union of |x| and |y|. Use 0.5 times the binary cross - entropy (BCE) loss and Dice loss as the model loss function. The calculation process is as follows:
[0230] Loss = 0.5L BCE (y i , x i ) + L Dice (y i , x i )
[0231] where, x i represents the true value, y i represents the predicted value. LDice It is 1 minus the Dice value.
[0232] Step 4.4: Use the dermoscopic image segmentation network model based on the attention mechanism and multi-scale information interaction for iterative training, extract the feature information of skin disease images, and obtain the segmentation result to output the optimal model.
[0233] Comparative experiment:
[0234] Compare the present invention with the current mainstream skin disease image segmentation models on the ISIC 2018 dataset, and the results are shown in Table 1.
[0235] Table 1 Performance comparison of segmentation models on the ISIC 2018 dataset
[0236] Segmentation model Dice IoU UNet 0.7677 0.6367 SegNet 0.7088 0.5643 UNet++ 0.7611 0.6413 ResUNet++ 0.8143 0.7032 The method of the present invention 0.8210 0.7121
[0237] As can be seen from Table 1, the method of the present invention achieved the best results of 0.8210 (Dice) and 0.7121 (IoU) on the ISIC 2018 dataset, significantly better than other methods, indicating that it has higher accuracy and stronger robustness in the skin lesion segmentation task.
[0238] Example 2
[0239] On the basis of Example 1, perform the same operation steps to test the ISIC 2017 dataset of Example 2.
[0240] Compare the present invention with the current mainstream skin disease image segmentation models on the ISIC 2017 dataset, and the results are shown in Table 2.
[0241] Table 2 Performance comparison of segmentation models on the ISIC 2017 dataset
[0242] Segmentation model Dice IoU UNet 0.7256 0.5956 SegNet 0.6509 0.5365 UNet++ 0.6988 0.5571 ResUNet++ 0.7913 0.6736 The method of the present invention 0.8013 0.6873
[0243] As can be seen from Table 2, the method of the present invention achieved the best results of 0.8013 (Dice) and 0.6873 (IoU) on the ISIC 2018 dataset, significantly better than other methods, indicating that it has higher accuracy and stronger robustness in the skin lesion segmentation task.
[0244] Example 3
[0245] On the basis of Example 1, perform the same operation steps to test the ISIC 2016 dataset of Example 3.
[0246] Compare the present invention with the current mainstream skin disease image segmentation models on the ISIC 2016 dataset, and the results are shown in Table 3.
[0247] Table 3 Performance Comparison of Segmentation Models on the ISIC 2016 Dataset
[0248] Segmentation model Dice IoU UNet 0.8611 0.7663 SegNet 0.8365 0.7253 UNet++ 0.8664 0.7586 ResUNet++ 0.8887 0.8015 The method of the present invention 0.8906 0.8151
[0249] As can be seen from Table 3, the method of the present invention achieved the best results of 0.8906 (Dice) and 0.8151 (IoU) on the ISIC 2016 dataset, significantly outperforming other methods, indicating its higher accuracy and stronger robustness in the skin lesion segmentation task.
[0250] The present invention is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A dermoscopic image segmentation method based on an attention mechanism and multi-scale information interaction, characterized in that It includes the following steps: Step 1: Construct a skin disease image dataset; Step 2: Preprocess the images in the dataset, including enhancing and normalizing skin melanin lesion images, and dividing them into a training set and a test set according to a certain proportion to ensure the optimization of data quality and model performance; Step 3: Construct a dermoscopic image segmentation model based on the attention mechanism and multi-scale information interaction; Step 4: Input the preprocessed skin disease image training set into the dermoscopic image segmentation network based on the attention mechanism and multi-scale information interaction for training to optimize the model performance to obtain the optimal segmentation effect; Step 5: After the training is completed, input the test set into the optimal model to segment the lesions of the skin disease images and evaluate the segmentation effect.
2. The dermoscopic image segmentation method based on the attention mechanism and multi-scale information interaction according to claim 1, characterized in that In the above-mentioned Step 1, constructing a skin disease image dataset includes the following steps: Step 1.1: Read and prepare the data of the relevant skin disease image dataset; Step 1.2: Remove low-quality images, calculate the sharpness of the images using the Laplace transform, and a low variance indicates blurriness; (1) Calculate the Laplace transform of the image: Among them, I ij is the pixel value, and ΔI ij is the gradient value after Laplace transform. The variance V of the whole image is calculated as the clarity index; (2) Set the sharpness threshold: If V≤100, the image is determined to be blurry and excluded; Step 1.3: Detect duplicate data, calculate the structural similarity SSIM, and remove similar or duplicate images from the dataset; (1) Calculate the mean μ x , μ y , the standard deviation σ x , σ y and the covariance σ x y (2) Set the similarity threshold If SSIM(x, y)≥0.95, it is considered that the two images are duplicates, and one of them is deleted; Step 1.4: Output a high-quality and non-duplicate skin disease image dataset.
3. The dermoscopic image segmentation method based on attention mechanism and multi-scale information interaction according to claim 1, characterized in that In the above-mentioned Step 2, preprocessing the images in the dataset, including enhancing and normalizing skin melanin lesion images, and dividing them into a training set and a test set according to a certain proportion to ensure the optimization of data quality and model performance, specifically: Step 2.1: Data augmentation, geometrically flipping, translating, and scaling the dermoscopic images; (1) Random rotation, randomly rotate the image to simulate the scenario of observing skin lesions from multiple angles; Rotation matrix formula where θ is the rotation angle, with a range of ±20°; Each pixel coordinate (x, y) in the image calculates the new coordinate (x′, y′) through the rotation matrix (2) Horizontal / vertical offset, perform random horizontal and vertical translation operations on the image to make the model adapt to the offset of the target position; Translation matrix formula: where Δx is the horizontal offset, with a value range of ±5% of the image width, and Δy is the vertical offset, with a value range of ±5% of the image height; After translation, the new coordinates of each pixel point: (3) Random scaling, simulate lesion areas at different distances by randomly scaling the image; Scaling matrix formula: where s x and s y is the scaling ratio, ranging from 0.95 to 1.05; The new coordinates of the pixel points after scaling: (4) Random flipping, simulate the mirror image scenario by horizontally flipping the image; Horizontal flipping matrix formula: where W is the image width; The new coordinates of the pixel points after flipping: Step 2.2: Standardize the data using Min-Max normalization; where, I min is the minimum value of the pixel values in the dataset, and I max is the maximum value of the pixel values in the dataset; Step 2.3: Divide the dataset according to the ratio of 80% training set, 10% validation set, and 10% test set.
4. The dermoscopic image segmentation method based on the attention mechanism and multi-scale information interaction according to claim 1, wherein In the above-mentioned Step 3, constructing a dermoscopic image segmentation model 1 based on the attention mechanism and multi-scale information interaction, specifically: Step 3.1: Use the improved atrous spatial pyramid pooling module to perform spatial pyramid pooling by adjusting the dilation rate to effectively capture multi-scale information under filters with different receptive fields; Branch 1: 1x1 convolution branch Use 1×1 convolution operation alone to extract global semantic information; Branch 2: Global Average Pooling Branch The input feature map first passes through the global average pooling (GAP) operation to map the features of each channel to a global average, thereby extracting global context information. Subsequently, the output of GAP passes through a 1×1 convolution to adjust the number of channels and generate feature representations with semantic information. Finally, the features are restored to the same spatial size as the original input through an upsampling operation to provide global semantic support for subsequent feature fusion. Branch 3: Dilated Convolution Branch By using multiple 3×3 dilated convolution operations with different dilation rates, multi-scale context information is extracted under different receptive fields. The size of the dilation rate r directly affects the range of the receptive field. Step 3.2: To capture multi-scale information, an axial gating mechanism is used to decompose self-attention into vertical and horizontal attention calculations, combined with learnable gating parameters to reduce computational complexity and enhance feature modeling capabilities in height and width dimensions; (1) Input feature processing Input feature map X in First, reduce the number of channels through a 1×1 convolutional layer. The formula is as follows: X conv = Conv 1×1 (X in ) Then, the feature distribution is normalized through the batch normalization layer, the formula is: X bn = BN(X conv ) (2) Axial attention calculation The features are input into the vertical attention and horizontal attention modules respectively, and the axial decomposition calculation is performed: 1) Vertical Attention Modeling For the features at each position (i, j), calculate the vertical attention output based on the height dimension H Among them, W Q , W K , W V are learnable projection matrices for queries, keys, and values, G Q , G K , G V 1, G V2 are gating parameters that dynamically adjust the attention distribution, is the position bias term for vertical attention; 2) Horizontal Attention Modeling For the features at each position (i, j), calculate the horizontal attention output based on the width dimension w Among them, W Q , W K , W V are learnable projection matrices for queries, keys, and values, G Q , G K , G V1 , G V2 are gating parameters that dynamically adjust the attention distribution, the position bias term of the horizontal attention; (3) Vertical and horizontal feature fusion Fuse the outputs of vertical attention and horizontal attention by weighted fusion: X attn = Y H + Y W (4) Output feature processing Fused feature X attn The channel dimension is adjusted again through 1×1 convolution, and normalization is performed through BN: X out = BN(Conv 1×1 (X attn )) Finally, add X out to the input X in to obtain the final output: X final = X in + X out Step 3.3: By using the reverse enhanced attention module, the tumor region features are deeply mined through the reverse deletion strategy, combining complementary regions and detail information; (1) High-level feature input After decoding and upsampling, the high-resolution feature map is input into the Sigmoid activation function to perform pixel-level classification. The decoder recovers the deep semantic information extracted in the encoding stage and combines upsampling to improve the spatial resolution. These features are mapped to probability values between 0 and 1 by Sigmoid. (2) Generate reverse attention weights 1) Sigmoid activation Input the upsampled feature S up into the Sigmoid function to generate a probability distribution: W = 1 - Sigmoid(S up ) Among them, S up is the feature after upsampling, and W represents the reverse attention weight; 2) Reverse operation Use 1-sigmoid to highlight low-response areas and mine detailed information about tumor boundaries; (3) Feature Fusion Multiply the reverse attention weight W and the high-level feature H element by element to generate a weighted feature map: F weighted = W☉H Where H represents high-level features and ⊙ represents point-by-point multiplication; (4) Output features Weighted feature F weighted Represents the enhanced convolutional feature for further optimization of the segmentation result.
5. The dermoscopic image segmentation method based on attention mechanism and multi-scale information interaction according to claim 1, characterized in that In the step, the preprocessed skin disease image training set is input into the dermatoscope image segmentation network based on the attention mechanism and multi-scale information interaction for training, and the model performance is optimized to obtain the optimal segmentation effect: Step 4.1: Initialize the parameters of the skin disease image segmentation network; Step 4.2: Set the training parameters, the training epoch is 200, the sample batch is 4, the initial learning rate lr of the network is set to 0.0001 and the Adam optimizer is used; Step 4.3: The training is quantitatively evaluated using two metrics, Dice and IoU. The calculation processes of the Dice metric and the IoU metric are as follows: Among them, |x| and |y| represent the label value and the predicted value respectively, |x∩y| represents the intersection of |x| and |y|, |x∪y| represents the union of |x| and |y|. 0.5 times the binary cross-entropy BCE loss and the Dice loss are used as the model loss function, and the calculation process is as follows: Loss=0.5L BCE (y i ,x i )+L Dice (y i ,x i ) Among them, x i represents the true value, y i represents the predicted value, and L Dice is 1 minus the Dice value; Step 4.4: Use a dermoscopic image segmentation network model based on the attention mechanism and multi-scale information interaction for iterative training, extract the skin disease image feature information, and obtain the segmentation result to output the optimal model.
Citation Information
Cited By
Coprocessing method and system for medical image lesion segmentation and classification
CN121258940A
Bird identification and monitoring method and system based on cross-modal feature learning
CN121281104A
A rice field weed semantic segmentation method based on multispectral unmanned aerial vehicle image
CN122530598A