A Skin Lesion Segmentation Method Based on Error Localization and Adaptive Optimization
By introducing mislocalization and adaptive optimization techniques into the skin lesion segmentation method, dynamically adjusting the region of focus of the model and suppressing potential errors, the problem of unsatisfactory segmentation effect in the existing technology is solved, and the robustness and generalization ability of the model are significantly improved.
Patent Information
- Application Number
- CN202510188569.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The existing skin lesion segmentation methods are not ideal in dealing with complex morphology, fuzzy boundaries and limited data, and are prone to ignore subtle changes in local characteristics and potential error information.
Using a skin lesion segmentation method based on error positioning and adaptive optimization, a segmentation network containing encoder and decoder is constructed, and error optimization units are used to perform error positioning and feature optimization, dynamically adjust the region of focus of the model and suppress the impact of potential errors.
It significantly improves the robustness and generalization ability of the model in complex scenarios, can more accurately segment the lesion area, reduce the interference of error accumulation on decision-making, and improve the details of segmentation results.
Smart Images

Figure CN119672041B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of skin lesion diagnosis and relates to a skin lesion segmentation method based on misalignment and adaptive optimization. Background Art
[0002] Skin lesion diagnosis generally includes image preprocessing, lesion segmentation, feature extraction, classification and other steps, among which lesion segmentation is the most fundamental and crucial step. Accurately segmenting the lesion area not only helps to clarify the boundary, shape and distribution of the lesion, but also lays the foundation for subsequent feature analysis and classification tasks.
[0003] Using computer-aided diagnosis can actively assist dermatologists in clinical diagnosis, has broad prospects in helping doctors diagnose skin cancer earlier and more accurately, and can also reduce the workload of doctors. In recent years, with the development of deep learning technology, more and more studies have adopted deep learning methods such as convolutional neural networks to improve the performance of skin cancer segmentation. Compared with traditional medical image segmentation algorithms, deep learning image segmentation methods are more effective and stable.
[0004] Most of the skin lesion segmentation methods in the past were mainly based on the improvement of the encoder-decoder structure. For example, introducing the Transformer structure into the UNet framework, designing a multi-scale feature extraction architecture, or adding spatial and channel attention mechanisms, etc. However, the following difficulties in skin lesion samples lead to the unsatisfactory overall segmentation effect of these models. First, the morphology of skin lesions is complex and diverse, with highly heterogeneous shapes, textures and distributions, and the lesion characteristics may vary significantly due to lesion types or individual differences. Second, the color and contrast within the lesion are often uneven, and there may be color gradients or low-contrast regions, interfering with visual identification. At the same time, the boundary between the lesion and normal skin is usually blurred and difficult to clearly separate. In addition, since skin lesion image annotation requires professional medical knowledge and is usually carried out by dermatology experts, it is time-consuming and laborious, which results in a relatively small scale of existing skin lesion image datasets and is difficult to support high-quality model training and the improvement of generalization performance.
[0005] To address these challenges, some methods attempt to introduce prior knowledge such as lesion contours and boundary key points to enhance the model's perception of key information. This prior knowledge helps the model segment more accurately under complex shapes and ambiguous boundary conditions by providing additional supervision signals. For example, in the literature (Boundary-aware transformers for skin lesion segmentation[C] / / Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part I 24. Springer International Publishing, 2021: 206-216.), boundary key points extracted before training are used to supervise the network to focus on key lesion boundary regions, thereby improving the model's segmentation performance in the boundary region. However, such methods usually rely on pre-computed fixed attention regions and lack the flexibility to dynamically adjust according to the actual performance of the model during training, which may overlook subtle changes in local features. In addition, existing models often fail to effectively detect and exclude potential error information when dealing with difficult-to-segment lesion regions. This neglect of error accumulation may lead to deviations in the segmentation results in details, further limiting the performance and robustness of the model.
[0006] Therefore, it is of great significance to study a skin lesion segmentation method based on error localization and adaptive optimization to effectively learn the subtle differences between lesions and normal skin under limited data, dynamically adjust the regions of interest of the model, and suppress the impact of potential errors on the segmentation results. Summary of the Invention
[0007] The objective of the present invention is to solve the problems existing in the prior art and provide a skin lesion segmentation method based on error localization and adaptive optimization.
[0008] To achieve the above objective, the technical solution adopted by the present invention is as follows:
[0009] A skin lesion segmentation method based on error localization and adaptive optimization, comprising the following steps:
[0010] (1) Divide the skin lesion images captured by dermoscopy and the corresponding segmentation labels into a training set, a validation set, and a test set;
[0011] (2) Construct a segmentation network;
[0012] The segmentation network consists of an encoder and a decoder; the encoder contains 5 encoding stages. The first encoding stage consists of a CNN (Convolutional Neural Network) module. The second to fourth encoding stages consist of a CNN module and a Transformer module. The fifth encoding stage consists of a Transformer module. The first encoding stage only performs the CNN module and downsampling to perform preliminary feature extraction and dimensionality reduction on the original image. The feature maps of the second to fourth encoding stages are respectively input into the CNN module and the Transformer module to complete local and global information extraction. Since the resolution of the fifth encoding stage is relatively small, only the Transformer module is used for processing. Finally, the feature map encoded by the encoder is input into the decoder; the decoder contains 3 decoding stages. The feature map input to each decoding stage is processed by an error optimization unit. The error optimization unit includes the functions of error localization and feature optimization; the feature map output by the fifth encoding stage is input into the first decoding stage. The feature map output by the first decoding stage is then input into the second decoding stage. The feature map output by the second decoding stage is then output to the third decoding stage. After performing a skip connection between the feature map output by the third decoding stage and the output feature map of the second encoding stage, the predicted segmentation map is output through the segmentation head (i.e., the final segmentation result);
[0013] (3) Use the training set to train the constructed segmentation network, and continuously monitor the generalization ability of the segmentation network through the validation set during the training process, and finally obtain the trained segmentation network;
[0014] (4) Input the test set into the trained segmentation network to output the segmentation result of skin lesions.
[0015] Different from traditional methods that simply rely on initial features or fixed feature extraction, the present invention transforms the conventional segmentation task into an adaptive error correction process, and fully utilizes the error region information in the feedback prediction to dynamically adjust the focus and weight of feature extraction, and uses accurate class feature information to help the model enhance feature representation. This design can not only effectively exclude potential error regions in the existing features, but also avoid the interference of error accumulation on subsequent decisions through iterative optimization, significantly improving the robustness and generalization ability of the model in complex scenarios.
[0016] As a preferred technical solution:
[0017] For a skin lesion segmentation method based on error localization and adaptive optimization as above, in step (1), the data in the divided training set is preprocessed and enhanced, and the data in the divided validation set and test set is preprocessed; the preprocessing includes resizing the image to Pixels; The enhancement includes randomly rotating 90 degrees with a 50% probability, horizontal or vertical flipping, and random scaling operations with a scaling ratio between 0.9 and 1.1.
[0018] A skin lesion segmentation method based on error localization and adaptive optimization as described above. After constructing the segmentation network in step (2), first, for the input original image X, the encoder uses the CNN module and downsampling operations in the first encoding stage to extract image features and reduce the resolution to obtain a feature map , and then is input into the CNN module of the second encoding stage; in the second encoding stage, first use the CNN module to extract the local information map , and then use the Transformer module to extract the global information map , and then and are fused to obtain the output feature map of the second encoding stage ; After downsampling, it is input into the CNN module of the third encoding stage. In the third encoding stage, first use the CNN module to extract the local information map , and then and the feature map formed by fusing the downsampled features are input into the Transformer module to obtain the global information map , and then and are fused to obtain the output feature map of the third encoding stage ; After downsampling, it is input into the CNN module of the fourth encoding stage. In the fourth encoding stage, first use the CNN module to extract the local information map , and then and the feature map formed by fusing the downsampled features are input into the Transformer module to obtain the global information map , and then and are fused to obtain the output feature map of the fourth encoding stage ; In the fifth encoding stage, the downsampled feature map and the feature map formed by fusing the downsampled features are input into the Transformer module to obtain the global information map , and then and are fused to obtain the output feature map of the fifth encoding stage ;
[0019] The calculation method of feature extraction in the CNN module is as follows:
[0020] ;
[0021] where conv represents a convolution with a convolution kernel size of , BN represents batch normalization, ReLU is an activation function, and down represents downsampling;
[0022] In the Transformer module of the second encoding stage, the multi-head self-attention mechanism is first used to capture the global dependency relationship to obtain , and then the feed-forward network is used to enhance the non-linear expression ability to obtain the final output of the Transformer module. The calculation method of the Transformer module is as follows:
[0023] ;
[0024] ;
[0025] where LN represents layer normalization, FFN represents a feed-forward network containing two linear transformations and a ReLU activation function, and MHSA represents multi-head self-attention;
[0026] The calculation methods of the Transformer modules in the 3rd to 5th encoding stages are as follows:
[0027] ;
[0028] ;
[0029] where concat represents channel connection, and conv1 represents a convolution with a convolution kernel size of ;
[0030] In the 2nd to 5th encoding stages, the fusion of and to output the feature map is calculated as follows:
[0031] ;
[0032] The inputs of the 1st to 3rd decoding stages in the decoder are the feature maps , and ; is equal to the output feature map of the 5th encoding stage; the feature map is composed of the feature map The feature map after being processed by the error optimization unit and the feature map are obtained through skip connection; the feature map is composed of the feature map The feature map after being processed by the error optimization unit and the feature map are obtained through skip connection; the feature map is composed of the feature map The feature map after being processed by the error optimization unit and the feature map are obtained through skip connection;
[0033] The process of inputting the feature map input in each decoding stage into the error optimization unit for processing is as follows:
[0034] Input the feature map into the error optimization unit. In the error optimization unit, first perform error localization. Input the rough segmentation map predicted from the feature map and the original image X connected by channel dimension into the UNet network to predict the error of the rough segmentation map, and output the predicted error region map ; then, use the predicted error region map , the rough segmentation map and the feature map to perform feature optimization, respectively enhance the category features corresponding to the foreground lesions and the background normal skin in the feature map , and output the feature map after being processed by the error optimization unit; among them, .
[0035] For a skin lesion segmentation method based on error localization and adaptive optimization as above, the error localization in the error optimization unit is to predict the missegmented regions in the segmentation result through a UNet network synchronously trained with the overall segmentation network. The label used in the training of the UNet network is the sum of the regions that actually belong to the target category but are not predicted in the rough segmentation map , and the non-target regions that are mispredicted as the target category;
[0036] The feature optimization in the error optimization unit is to first refer to the predicted error region map to extract the foreground category map and the background category map ; then multiply the category maps , , respectively, with the feature map pixel by pixel, retain the features of the matching categories, and then take the average to obtain the category prototype ; then, use convolution to compress the channel dimension of the category prototype and the feature map to obtain and ; By performing linear transformations with two learnable weight matrices and to obtain the key vector and the value vector ; By multiplying element-wise with to weight the error regions, and then performing a linear transformation on the weighted features with the learnable weight matrix to obtain the query vector ; Calculate the attention matrix according to the calculation formula of the attention mechanism, and dot-multiply and to obtain the class-enhanced feature map ; Finally, concatenate with two class-enhanced feature maps and using channel concatenation and fuse them with convolution to output the feature map processed by the error optimization unit;
[0037] The calculation process of the class map is as follows: First, extract the regions in the rough segmentation map where the value is greater than 0.5. The value of each pixel represents the confidence of predicting the foreground class; then extract the regions in where the value is less than or equal to 0.5. The value of each pixel represents the confidence of predicting the background class; then, according to the predicted error region map , reduce the prediction confidence of these regions by a weight of 0.1 to obtain the class map , .
[0038] A skin lesion segmentation method based on error localization and adaptive optimization as described above, and The specific calculation formulas are:
[0039] ;
[0040] ;
[0041] where, , represents element-wise multiplication;
[0042] The specific calculation formula is:
[0043] ;
[0044] where, is an indicator function, and M represents the number of pixels in the class map . represents the m-th pixel of the class map . represents the m-th pixel of the feature map ;
[0045] , , , , , and The specific calculation formulas are as follows:
[0046] ;
[0047] ;
[0048] ;
[0049] ;
[0050] ;
[0051] ;
[0052] ;
[0053] where e is a matrix of all 1s with the same size as , and is the dimension of the vector .
[0054] For a skin lesion segmentation method based on error localization and adaptive optimization as above, the process of outputting the predicted segmentation map from the feature map output in the third decoding stage through the segmentation head is as follows: The segmentation head first maps the number of channels of the feature map to a single-channel output through a convolutional layer, then upsamples the single-channel feature map to the original image size using a transposed convolution, and finally converts the value of each pixel to the probability of belonging to the foreground class through a sigmoid activation function to obtain the final segmentation map .
[0055] For a skin lesion segmentation method based on error localization and adaptive optimization as above, the training process of the segmentation network in step (3) is specifically as follows:
[0056] First, randomly initialize the parameters of each layer of the segmentation network, select the Adam optimizer, set the initial learning rate to 0.0001, and train for 100 epochs with mini-batch data of batch size 8;
[0057] In each training epoch, first perform forward propagation, input the images in the training set into the segmentation network, and generate a predicted coarse segmentation map in the decoding stage The features obtained by decoding are predicted by the segmentation head to obtain the final segmentation map ;
[0058] For the coarse segmentation map predicted in each decoding stage , calculate the error region label of the UNet network in the training error localization part :
[0059] ;
[0060] Then calculate the total loss , including the total loss of the predicted segmentation map and the total loss of the predicted error region map ; Since the constructed segmentation network uses deep supervision, the total loss is the weighted sum of the losses of the predicted segmentation maps at 4 different scales according to the size of the coarse segmentation map; The calculation formula of
[0061] is as follows:
[0062] Among them, the Focal Loss is used to calculate the loss ; The calculation formula of
[0063] is as follows:
[0064] The binary cross-entropy loss and the intersection over union loss are used to calculate the segmentation loss of each predicted segmentation map, and their loss values are weighted according to the size of the predicted segmentation maps at different scales, and then summed to obtain ; The calculation formula of
[0065] is as follows:
[0066] ;
[0067] ;
[0068] Among them, is the label of the i-th error region of the k-th sample, is the predicted error region map of the i-th of the k-th sample, N is the batch size, N = 8; α is set to 0.25 to balance the number of classes, and γ is set to 2 to balance the degree of error; represents the i-th segmentation label of the k-th sample, represents the i-th predicted segmentation map of the k-th sample, is set to 1, 0.5, 0.4, and 0.3, representing the weights added when calculating the loss for predictions at different resolutions. ϵ is set to 1e-6 to prevent the denominator from being zero;
[0069] Finally, according to the calculation result, use the chain rule to calculate the gradients of each layer in the network, and pass the gradients to each layer of the network through backpropagation; then, use the Adam optimizer to update the parameters in the network based on the gradients obtained from backpropagation;
[0070] Every 10 training epochs, use the validation set to check the generalization ability of the segmentation network, and calculate its performance metrics on the validation set, including DSC, IoU, and Acc; if these metrics do not improve or show a downward trend on the validation set, then halve the learning rate to further optimize the segmentation network.
[0071] Beneficial effects:
[0072] (1) A skin lesion segmentation method based on error localization and adaptive optimization of the present invention proposes an encoder that extracts local and global feature information at different resolution levels, providing richer feature expressions for the decoder.
[0073] (2) A skin lesion segmentation method based on error localization and adaptive optimization of the present invention proposes an error localization process and integrates it into the network decoder, which can identify the regions where the model makes mistakes to reveal the current performance limitations.
[0074] (3) A skin lesion segmentation method based on error localization and adaptive optimization of the present invention proposes a feature optimization process and integrates it into the network decoder; combining the segmentation error map generated by the error localization process, analyzes reliable class regions, and provides support for subsequent predictions by aggregating more accurate class feature prototypes, thereby reducing the impact of errors and refining the segmentation results.
[0075] (4)A skin lesion segmentation method based on error localization and adaptive optimization of the present invention can identify error regions at multiple scales and use the error map as feedback on the learning effect to further optimize the category understanding by performing error localization and feature optimization at each decoding stage. This mechanism effectively focuses on the key regions that need to be optimized in actual training, helps the model better understand and distinguish the features of lesions and normal skin, and thus improves its segmentation ability when dealing with complex lesions such as fuzzy boundaries, complex shapes, and background noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 is a structural diagram of a skin lesion segmentation network based on error localization and adaptive optimization;
[0077] Figure 2 is a structural diagram of the Transformer module and the CNN module;
[0078] Figure 3 is a structural diagram of the error optimization unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0079] The present invention will be further described below in conjunction with the specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0080] A skin lesion segmentation method based on error localization and adaptive optimization includes the following steps:
[0081] (1)Divide the skin lesion images captured by dermoscopy and the corresponding segmentation labels into a training set, a validation set, and a test set, and then preprocess and enhance the data in the divided training set, and preprocess the data in the validation set and the test set;
[0082] (2)Construct a segmentation network;
[0083] The segmentation network consists of an encoder and a decoder; the encoder contains 5 encoding stages, the first encoding stage consists of CNN modules, the 2nd to 4th encoding stages consist of CNN modules and Transformer modules, and the 5th encoding stage consists of Transformer modules; the decoder contains 3 decoding stages, and the feature maps input to each decoding stage are processed by an error optimization unit, which includes the functions of error localization and feature optimization; the feature map output by the 5th encoding stage is input to the first decoding stage, the feature map output by the first decoding stage is then input to the second decoding stage, and the feature map output by the second decoding stage is then output to the third decoding stage. After performing a skip connection between the feature map output by the third decoding stage and the output feature map of the second encoding stage, the predicted segmentation map is output through the segmentation head ;
[0084] The process of outputting the predicted segmentation map through the segmentation head for the feature map output by the third decoding stage is as follows: First, the segmentation head maps the number of channels of the feature map to a single-channel output through a convolutional layer, then uses a transposed convolution to upsample the single-channel feature map to the original image size, and finally converts the value of each pixel to the probability of belonging to the foreground class through a sigmoid activation function, thereby obtaining the final segmentation map ;
[0085] After constructing the segmentation network, as Figure 1 shown, first, for the input original image X, the encoder uses the CNN 1 module and downsampling operation in the first encoding stage to extract image features and reduce the resolution to obtain the feature map , and then is input to the CNN 2 module in the second encoding stage; in the second encoding stage, first use the CNN 2 module to extract the local information map , then use the Transformer 1 module to extract the global information map , and then and are fused to obtain the output feature map of the second encoding stage; After downsampling, it is input to the CNN 3 module in the third encoding stage. In the third encoding stage, first use the CNN 3 module to extract the local information map , and then and the feature map formed by fusing the downsampled features are input to the Transformer 2 module to obtain the global information map , and then and Fuse to obtain the output feature map of the third encoding stage ; After downsampling, it is input into the CNN 4 module of the fourth encoding stage. In the fourth encoding stage, first use the CNN 4 module to extract the local information map , then and The feature map fused with the downsampled features is input into the Transformer 3 module to obtain the global information map , then and Fuse to obtain the output feature map of the fourth encoding stage ; In the fifth encoding stage, The downsampled feature map and The feature map fused with the downsampled features is input into the Transformer 4 module to obtain the global information map , then and Are fused to obtain the output feature map of the fifth encoding stage ;
[0086] As Figure 2 Shown, the calculation method of feature extraction by the CNN module is as follows:
[0087] ;
[0088] Among them, conv represents convolution with a convolution kernel size of , BN represents batch normalization, ReLU is the activation function, and down represents downsampling;
[0089] The Transformer module in the second encoding stage first uses the multi-head self-attention mechanism to capture Global dependencies to obtain , and then use the feed-forward network to enhance the non-linear expression ability to obtain the final output of the Transformer module , the calculation method of the Transformer module is as follows:
[0090] ;
[0091] ;
[0092] Among them, LN (Layer Norm) represents layer normalization, FFN represents a feed-forward network containing two linear transformations and one ReLU activation function, and MHSA represents multi-head self-attention;
[0093] The calculation method of the Transformer module in the 3rd to 5th encoding stages is as follows:
[0094] ;
[0095] ;
[0096] Among them, concat represents channel concatenation, and conv1 represents a convolution with a convolution kernel size of ;
[0097] The fusion of the 2nd to 5th encoding stages and output feature maps The calculation method is as follows:
[0098] ;
[0099] The inputs of the 1st to 3rd decoding stages in the decoder are the feature maps , and ; is equal to the output feature map of the 5th encoding stage ; The feature map is obtained by performing a skip connection between the feature map after being processed by the error optimization unit of the feature map and the feature map ; The feature map is obtained by performing a skip connection between the feature map after being processed by the error optimization unit of the feature map and the feature map ; The feature map is obtained by performing a skip connection between the feature map after being processed by the error optimization unit of the feature map and the feature map ;
[0100] The process of inputting the feature map input in each decoding stage into the error optimization unit for processing is as follows:
[0101] (2.1) As Figure 3 shown, input the feature map into the error optimization unit. In the error optimization unit, first use the feature map to predict the rough segmentation map , and perform error localization on ;
[0102] Error localization is to predict the mis-segmented area in the segmentation result through a UNet network that is synchronously trained with the overall segmentation network. The label used during the training of the UNet network is the rough segmentation map The sum of the regions that actually belong to the target category but are not predicted, and the non-target regions that are mispredicted as the target category;
[0103] (2.2)Connect the rough segmentation map and the original image X along the channel dimension and input them into the UNet network to predict the errors in the rough segmentation map, and output the predicted error region map ;
[0104] (2.3)Utilize the predicted error region map , the rough segmentation map and the feature map to perform feature optimization, enhance the category features corresponding to the foreground lesions and the background normal skin in the feature map respectively, and output the feature map after being processed by the error optimization unit; where, ;
[0105] The feature optimization in the error optimization unit first refers to the predicted error region map to extract the foreground category map and the background category map , f represents the foreground, b represents the background, and then multiply the category maps , with the feature map pixel by pixel, retain the features of the matching categories, and then average them to obtain the category prototype ; then, use to convolve and compress the channel dimension of the category prototype and the feature map to obtain and ; Through linear transformation with two learnable weight matrices and , obtain the key vector and the value vector , K represents the key vector, V represents the value vector; Through multiplication with pixel by pixel, weight the error regions, and then linearly transform the weighted features with the learnable weight matrix to obtain the query vector ; Calculate the attention matrix according to the calculation formula of the attention mechanism, and dot multiply and to obtain the category-enhanced feature map ; Finally, combine and the two category-enhanced feature maps and using channel concatenation and After convolution fusion, the feature map processed by the output error optimization unit is output;
[0106] and The specific calculation formula is:
[0107] ;
[0108] ;
[0109] Among them, , represents element-wise multiplication;
[0110] The specific calculation formula is:
[0111] ;
[0112] Among them, is the indicator function, M represents the number of pixels of the category map , represents the m-th pixel of the category map , represents the m-th pixel of the feature map ;
[0113] , , , , , and The specific calculation formulas are respectively:
[0114] ;
[0115] ;
[0116] ;
[0117] ;
[0118] ;
[0119] ;
[0120] ;
[0121] Among them, e is a matrix of all 1s with the same size as , is the dimension of the vector ;
[0122] Category map The calculation process is as follows: First, extract the regions in the rough segmentation map where the median value is greater than 0.5. The value of each pixel represents the confidence of being predicted as the foreground class; then extract the regions where the median value is less than or equal to 0.5. The value of each pixel represents the confidence of being predicted as the background class; then, respectively, according to the predicted error region map , reduce the prediction confidence of the mispredicted regions by a weight of 0.1 to obtain the class map , ;
[0123] (3) Use the training set to train the constructed segmentation network, and continuously monitor the generalization ability of the segmentation network through the validation set during the training process, and finally obtain the trained segmentation network;
[0124] The training process of the segmentation network is specifically as follows:
[0125] First, randomly initialize the parameters of each layer of the segmentation network, select the Adam optimizer, set the initial learning rate to 0.0001, and perform 100 rounds of training with mini-batch data of batch size 8;
[0126] In each training round, first perform forward propagation, input the images in the training set into the segmentation network, and generate the predicted rough segmentation map in the decoding stage. The features obtained by decoding are predicted by the segmentation head to obtain the final segmentation map ;
[0127] For each predicted rough segmentation map in the decoding stage, calculate the error region labels of the UNet network in the training error localization part :
[0128] ;
[0129] Then calculate the total loss , including the total loss of the predicted segmentation map and the total loss of the predicted error region map
[0130] ;
[0131] Among them, use Focal Loss to calculate the loss ; The calculation formula of
[0132] ;
[0133] Using binary cross - entropy loss and intersection - over - union loss Calculate the segmentation loss for each predicted segmentation map, and weight its loss value according to the size of the predicted segmentation map at different scales, and then sum to obtain ; The calculation formula of
[0134] ;
[0135] ;
[0136] ;
[0137] where, is the label of the i - th error region of the k - th sample, is the predicted error region map of the i - th of the k - th sample, N is the batch size, N = 8; α is set to 0.25 to balance the number of classes, and γ is set to 2 to balance the degree of error; represents the i - th segmentation label of the k - th sample, represents the i - th predicted segmentation map of the k - th sample, is set to 1, 0.5, 0.4, and 0.3, indicating the weights added when calculating the loss for predictions at different resolutions, and ϵ is set to 1e - 6 to prevent the denominator from being zero;
[0138] Finally, according to the calculation result of calculate the gradients of each layer in the network using the chain rule, and transmit the gradients to each layer of the network through backpropagation; then, use the Adam optimizer to update the parameters in the network based on the gradients obtained from backpropagation;
[0139] Every 10 training epochs, use the validation set to check the generalization ability of the segmentation network, and calculate its performance metrics on the validation set, including DSC (Dice similarity coefficient), Acc (accuracy), IoU (intersection - over - union); if these metrics do not improve or show a downward trend on the validation set, then halve the learning rate to further optimize the segmentation network;
[0140] (4)Input the test set into the trained segmentation network and output the segmentation result of skin lesions.
[0141] Specifically, the present invention conducts experiments on two datasets, ISIC 2017 and ISIC 2018. Among them, for the ISIC 2017 dataset, it is randomly divided into 1250 training samples, 150 validation samples, and 600 test samples. For the ISIC 2018 dataset, it is randomly divided into 1815 training samples, 259 validation samples, and 520 test samples. The pre - processing includes resizing the image to Pixels; The augmentation includes randomly rotating 90 degrees with a 50% probability, horizontally or vertically flipping, and randomly scaling with a scaling ratio between 0.9 and 1.1. The comparison of the effects of the present invention with other skin lesion segmentation methods in the prior art is shown in Table 1, and statistical analysis is specifically carried out on the three indicators of DSC, Acc, and IoU. Assume and are the foreground region and background region predicted by the model respectively, and are the segmentation labels of the true foreground region and background region respectively. The calculation formulas of these three indicators are as follows:
[0142] ;
[0143] ;
[0144] ;
[0145] Among them, DSC is more focused on measuring the segmentation performance of the model for lesion categories. Acc considers the prediction accuracy of all pixels, which reflects the overall segmentation accuracy. IoU focuses on the segmentation of the boundary by evaluating the spatial overlap quality of the predicted region and the true region.
[0146] It can be seen from Table 1 that the present invention has achieved the best comprehensive results on these two datasets.
[0147] Table 1
[0148] .
[0149] The specific literature information of the prior art [1] - [5] is as follows:
[0150] [1] U-net: Convolutional networks for biomedical image segmentation[C] / / Medical image computing and computer-assisted intervention–MICCAI 2015:18th international conference, Munich, Germany, October 5-9, 2015,proceedings, part III 18. Springer International Publishing, 2015: 234-241.
[0151] [2] Attention gated networks: Learning to leverage salient regions in medical images[J]. Medical image analysis, 2019, 53: 197-207.
[0152] [3] Transunet: Transformers make strong encoders for medical image segmentation[J]. arXiv preprint arXiv:2102.04306, 2021.
[0153] [4] Medical transformer: Gated axial-attention for medical image segmentation[C] / / Medical image computing and computer assisted intervention–MICCAI 2021: 24th international conference, Strasbourg, France, September 27–October 1, 2021, proceedings, part I 24. Springer International Publishing, 2021: 36-46.
[0154] [5] Hiformer: Hierarchical multi-scale representations using transformers for medical image segmentation[C] / / Proceedings of the IEEE / CVF winter conference on applications of computer vision. 2023: 6202-6212.
Claims
1. A skin lesion segmentation method based on error localization and adaptive optimization, characterized in that The steps include: (1) The skin lesion images taken by the dermatoscope and the corresponding segmentation labels are divided into a training set, a validation set, and a test set; (2) Construct a segmentation network; The segmentation network consists of an encoder and a decoder; the encoder contains 5 encoding stages, the first encoding stage consists of a CNN module, the second to fourth encoding stages consist of a CNN module and a Transformer module, and the fifth encoding stage consists of a Transformer module; the decoder contains 3 decoding stages, and the feature map input to each decoding stage is processed by an error optimization unit, which includes the functions of error location and feature optimization; the feature map output by the fifth encoding stage is input to the first decoding stage, the feature map output by the first decoding stage is input to the second decoding stage, and the feature map output by the second decoding stage is output to the third decoding stage. After the feature map output by the third decoding stage and the output feature map of the second encoding stage are jump-connected, the predicted segmentation map is output through the segmentation head. (3) Use the training set to train the constructed segmentation network, and continuously monitor the generalization ability of the segmentation network through the validation set during the training process, and finally obtain a trained segmentation network; (4) Input the test set into the trained segmentation network and output the segmentation result of skin lesions; The process of inputting the feature map of each decoding stage into the error optimization unit for processing is as follows: The feature map D i Input error optimization unit, in which error optimization unit, firstly, error location is performed, and feature map D i Predicted coarse segmentation map After being connected to the original image X according to the channel dimension, it is input into the UNet network to predict the errors of the coarse segmentation map and output the predicted error area map Then, using the predicted error area map Coarse segmentation map and feature map D i Perform feature optimization to enhance the feature map D i The category features corresponding to the foreground lesions and the background normal skin are output, and the feature map processed by the error optimization unit is output; where i∈{2,3,4}; Error localization in the error optimization unit is done by predicting the segmentation error areas in the segmentation results through a UNet network trained synchronously with the overall segmentation network. The labels used in UNet training are the coarse segmentation maps. The sum of the areas that actually belong to the target category but are not predicted, and the non-target areas that are incorrectly predicted as the target category; The feature optimization in the error optimization unit first refers to the predicted error area map Extracting foreground category maps and background category map Then the category map Respectively with the feature map D i Multiply pixel by pixel, retain the features of the matching category, and then average to get the category prototype P i j ; Then, use 1×1 convolution to compress the category prototype P i j and feature map D i The channel dimension is obtained and By combining two learnable weight matrices and Perform a linear transformation to obtain the key vector Sum value vector V i j ; Through Multiply pixel by pixel to weight the error area, and then combine the weighted features with the learnable weight matrix Perform linear transformation to obtain the query vector Calculate the attention matrix according to the calculation formula of the attention mechanism and will and V i j Point product to get the feature map of category enhancement Finally, And the feature maps of two categories enhancement and After fusion with channel connection and 1×1 convolution, the feature map processed by the error optimization unit is output.
2. The skin lesion segmentation method based on error localization and adaptive optimization according to claim 1, characterized in that: In step (1), the data in the training set is preprocessed and enhanced, and the data in the validation set and the test set are preprocessed.
3. The skin lesion segmentation method based on error localization and adaptive optimization according to claim 1, characterized in that: Step (2) After the segmentation network is constructed, first, for the input original image X, the encoder uses the CNN module and downsampling operation in the first encoding stage to extract image features and reduce the resolution to obtain the feature map L0, and then inputs L0 into the CNN module of the second encoding stage; in the second encoding stage, the CNN module is first used to extract the local information map L1, and then the Transformer module is used to extract the global information map G1 of L1, and then L1 and G1 are fused to obtain the output feature map E1 of the second encoding stage; L1 is downsampled and input into the CNN module of the third encoding stage. In the third encoding stage, the CNN module is first used to extract the local information map L2, and then the feature map formed by the fusion of the downsampled features of L2 and G1 is input into the Transformer module. The r module obtains the global information graph G2, and then L2 and G2 are fused to obtain the output feature graph E2 of the third encoding stage; L2 is input into the CNN module of the fourth encoding stage after downsampling. In the fourth encoding stage, the local information graph L3 is first extracted by the CNN module, and then the feature graph formed by fusion of the downsampled features of L3 and G2 is input into the Transformer module to obtain the global information graph G3, and then L3 and G3 are fused to obtain the output feature graph E3 of the fourth encoding stage; in the fifth encoding stage, the feature graph L4 after L3 downsampling and the feature graph formed by fusion of the downsampled features of G3 are input into the Transformer module to obtain the global information graph G4, and then G4 and L4 are fused to obtain the output feature graph E4 of the fifth encoding stage; The calculation method of CNN module feature extraction is as follows: Where conv represents a convolution with a kernel size of 3×3, BN represents batch normalization, and ReLU is the activation function. down means downsampling; The Transformer module in the second encoding stage first uses the multi-head self-attention mechanism to capture the global dependency of L1 to obtain G1 ′ , and then use the feedforward network to enhance the nonlinear expression ability to obtain the final output G1 of the Transformer module. The calculation method of the Transformer module is as follows: G′1=L1+MHSA(LN(L1)); G1=G′1+FFN(LN(G′1)); Among them, LN represents layer normalization, FFN represents a feedforward network containing two linear transformations and a ReLU activation function, and MHSA represents multi-head self-attention; The Transformer module calculation method for the 3rd to 5th encoding stages is as follows: G′ i =conv1(concat([L i ,down(G i-1 )]))+MHSA(LN(conv1(concat([L i ,down(G i-1 )])))),i∈{2,3,4}; G i =G′ i +FFN(LN(G′ i )),i∈{2,3,4}; Among them, concat represents channel connection, conv1 represents convolution with a convolution kernel size of 1×1; The 2nd to 5th encoding stages integrate L i and G i Output feature map E i The calculation method is as follows: E i =conv1(concat([L i ,G i ])),i∈{1,2,3,4}; The inputs of the 1st to 3rd decoding stages in the decoder are feature maps D4, D3 and D2 respectively; D4 is equal to the output feature map E4 of the 5th encoding stage; feature map D3 is obtained by skipping the feature map D4 after being processed by the error optimization unit and feature map E3; feature map D2 is obtained by skipping the feature map D3 after being processed by the error optimization unit and feature map E2; feature map D1 is obtained by skipping the feature map D2 after being processed by the error optimization unit and feature map E1.
4. The skin lesion segmentation method based on error localization and adaptive optimization according to claim 3, characterized in that: Category diagram The calculation process is as follows: First, extract the coarse segmentation map In the area where the median value is greater than 0.5, the value of each pixel represents the confidence of the prediction as the foreground category; then extract In the area where the median value is less than or equal to 0.5, the value of each pixel represents the confidence of the prediction as the background category; Then, according to the predicted error area map Reduce the prediction confidence of these areas by 0.1 to get the category map 5. The skin lesion segmentation method based on error localization and adaptive optimization according to claim 4, characterized in that: and The specific calculation formula is: Among them, μ1=0.1, ⊙ represents pixel-by-pixel multiplication; P i j The specific calculation formula is: in, is the indicator function, M represents the feature map D i The number of pixels, Representation category diagram The mth pixel, D i,m Represents the feature map D i The mth pixel of V i j , and The specific calculation formulas are: Among them, e is and All-1 matrices of the same size, d k is a vector Dimension.
6. The skin lesion segmentation method based on error localization and adaptive optimization according to claim 5, characterized in that: The feature map output by the third decoding stage is used to output the predicted segmentation map through the segmentation head The process is as follows: the segmentation head first maps the number of channels of the feature map D1 to a single-channel output through a convolution layer, then uses a transposed convolution to upsample the single-channel feature map to the original image size, and finally uses a sigmoid activation function to convert the value of each pixel into the probability of belonging to the foreground category, thereby obtaining the final segmentation map.
7. The skin lesion segmentation method based on error localization and adaptive optimization according to claim 6, characterized in that: The training process of the segmentation network in step (3) is as follows: First, the parameters of each layer of the segmentation network are randomly initialized, the Adam optimizer is selected, the initial learning rate is set to 0.0001, and 100 rounds of training are performed with a small batch size of 8; In each training round, forward propagation is first performed to input the images in the training set into the segmentation network, and the predicted coarse segmentation map is generated in the decoding stage. The decoded features are used to predict the final segmentation map through the segmentation head. Coarse segmentation map predicted at each decoding stage Calculate the error region label ε of the UNet network in the training error localization part i : Then calculate the total loss The total loss including the predicted segmentation map and the total loss of the predicted error area map The calculation formula is as follows: Among them, ε i,k is the i-th wrong region label of the k-th sample, is the error area map of the i-th prediction of the k-th sample, N is the batch size, N = 8; α is set to 0.25 to balance the number of categories, and γ is set to 2 to balance the degree of error; y i,k represents the i-th segmentation label of the k-th sample, represents the i-th predicted segmentation map of the k-th sample, w i Set to 1, 0.5, 0.4, and 0.3, indicating the weight added when calculating the loss for predictions of different resolutions, and ϵ is set to 1e-6 to prevent the denominator from being zero; Finally, according to The chain rule is used to calculate the gradients of each layer in the network, and the gradients are passed to each layer of the network through back propagation. Then, the Adam optimizer is used to update the parameters in the network based on the gradients obtained through back propagation. After every 10 rounds of training, the generalization ability of the segmentation network is checked using the validation set, and its performance indicators on the validation set are calculated; if these indicators do not improve or show a downward trend on the validation set, the learning rate is halved to further optimize the segmentation network.
Citation Information
Patent Citations
Human face detection-based method for testing incorrect reading posture
CN101539989A
Detection label for skin surface shape change and real-time detection method
CN109247915A