Skin dermoscopy image segmentation method, segmentation network and segmentation network construction method
The dermoscopy image segmentation network enhances precision by modeling global context and reducing uncertainty through a multi-head self-attention module and uncertainty reduction techniques, addressing the challenges of varying scales and shapes in dermoscopy images.
Patent Information
- Application Number
- CN202111051917.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-09-08
AI Technical Summary
In the existing automatic dermatoscope image analysis system, the fuzzy boundaries and variable scales of the lesion area lead to inaccurate segmentation, which becomes a bottleneck for automatic analysis.
Build a multi-head self-attention context perception module and an uncertainty reduction module, combine it with the hollow space pyramid module to capture global context relationships, reduce uncertainty, and improve segmentation accuracy.
By modeling global contextual relationships and reducing uncertainty, the problems of different sizes and morphology of the areas to be segmented are alleviated, and the accuracy of dermatoscope image segmentation is improved.
Smart Images

Figure CN113744255B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital image processing, and particularly to a method for segmenting dermoscopic images, a segmentation network, and a method for constructing the segmentation network. Background Art
[0002] A dermoscope is a non-invasive microscopic image analysis technique for observing the fine structures beneath the surface of living skin. The segmentation of regions of interest is one of the core tasks in medical image analysis and an indispensable step in quantitative analysis. For example, the segmentation of lesion regions in dermoscopic images, the shape, appearance, and location of the segmented regions are of great significance for the early diagnosis of diseases. Currently, there have been reports on automatic analysis and diagnosis systems for dermoscopic images internationally. However, due to the often blurred boundaries, variable scales, and shapes of lesion regions in dermoscopic images, the accurate segmentation of lesion regions has become a bottleneck in current automatic analysis systems for dermoscopic images. Summary of the Invention
[0003] In order to overcome the deficiencies in the prior art, embodiments of the present invention provide a method for segmenting dermoscopic images, a segmentation network, and a method for constructing the segmentation network. Using this segmentation network to segment dermoscopic images can capture and model the global context relationship, alleviate the problems of different scales and variable morphologies of regions to be segmented, perform secondary prediction on points with high uncertainty, reduce uncertainty, and improve the segmentation accuracy.
[0004] To achieve the above object, the technical solution adopted by the present invention is: A method for segmenting dermoscopic images, comprising the following steps:
[0005] Generate at least a second feature map and a third feature map based on the acquired image;
[0006] Extract and process the features in the second feature map to form a second corrected map;
[0007] Fuse the third feature map and the second corrected map to generate a first fused map;
[0008] Process the uncertain points in the first fused map, and process the certain points in the first fused map and the uncertain points in the processed first fused map to generate a first fused corrected map.
[0009] In the above technical solution, the step "Generate at least a second feature map and a third feature map based on the acquired image" includes:
[0010] Perform downsampling processing on the second feature map to generate a third preset feature map, and generate a third feature map according to the atrous spatial pyramid module and the third preset feature map.
[0011] In the above technical solution, the step of "extracting and processing the features in the second feature map to form a second corrected map" includes:
[0012] Extracting and processing the features in the second feature map based on the multi-head self-attention context awareness module to form a second corrected map.
[0013] In the above technical solution, the step of "extracting and processing the features in the second feature map to form a second corrected map" includes:
[0014] The encoding block extracts the feature F from the second feature map;
[0015] The feature F extracted by the encoding block is mapped to the embedding space using three different convolutional layers Q, K, and V to obtain the feature maps F Q , F K , F V ;
[0016] The channel dimensions of F Q , F K , F V are mapped to C to reduce the number of parameters, and the channel dimensions of the feature maps F Q , F K , F V in the embedding space are all H×W×C;
[0017] Adjust the shapes of the feature maps F Q and F k to HW×C, and transpose F Q so that the shape of F Q becomes C×HW;
[0018] Use matrix multiplication for F Q and F k to obtain a feature map F T of size HW×HW;
[0019] Use the position encoding map E to encode the relative positions in the feature map. The size of the position encoding map E is HW×HW, where the value E(i,j) of each pixel represents the relative relationship between the i-th position and the j-th position;
[0020] According to Att = softmax(F T +E), obtain the attention weight map Att, where the size of the attention weight map Att is HW×HW, and Att(i,j) represents the dependence relationship between the i-th pixel value and the j-th pixel value. F T is a feature map of size HW×HW, and E represents the position encoding;
[0021] Transpose Att and multiply it with the feature map F VPerform matrix multiplication, add the feature F in the form of residual connection to obtain the filtered feature F';
[0022] In the decoding block, concatenate the feature F' to restore the image, and form the second corrected image;
[0023] In the above technical solution, the step "process the uncertain points in the first fused image, and process the certain points in the first fused image and the uncertain points in the processed first fused image to generate a first fused corrected image" includes:
[0024] Based on the uncertainty reduction module, process the uncertain points in the first fused image, and process the certain points in the first fused image and the uncertain points in the processed first fused image to generate a first fused corrected image.
[0025] In the above technical solution, the step "process the uncertain points in the first fused image based on the uncertainty reduction module" includes:
[0026] After each upsampling in the encoding block, use a 1×1 convolution to reduce the number of channels of the feature map to 1;
[0027] Activate using the Sigmoid function, normalize the value of each pixel point to 0-1 as the first prediction value;
[0028] For all pixel points activated by the Sigmoid function, select N pixel points with the first prediction value of 0.5, and extract the point-level features at their positions in the channel dimension;
[0029] Use a two-layer fully connected layer to construct a perceptron to re-predict the second prediction value of N pixel points, and update the corresponding first prediction value using the re-predicted second prediction value.
[0030] In the above technical solution, in the uncertainty reduction module, use the cross-entropy loss function as the auxiliary loss for deep supervision; the calculation of the cross-entropy loss is as follows: L BCE =-G·logP-(1-G)log(1-P), where P represents the output probability map of the segmentation network, G represents the segmentation gold standard, and L BCE represents the cross-entropy loss between the output probability map and the segmentation gold standard.
[0031] In the above technical solution, it further includes:
[0032] Generate an intermediate feature map according to the obtained image, extract and process the features in the intermediate feature map to form an intermediate corrected image;
[0033] Fuse the intermediate corrected image and the fused corrected image of the previous step to generate a second fused image;
[0034] Process the uncertain points in the second fusion graph, and process the certain points in the second fusion graph and the processed uncertain points in the second fusion graph to generate a second fusion correction graph.
[0035] A segmentation network for dermoscopic images, the architecture of the segmentation network includes:
[0036] A baseline network, the baseline network is based on the encoding-decoding structure of a fully convolutional network, includes a plurality of encoding blocks and decoding blocks corresponding to the encoding blocks, the plurality of encoding blocks are connected by downsampling operations, the plurality of decoding blocks are connected by upsampling operations, the encoding blocks are used to extract features of the image, and the decoding blocks restore the image according to the features to generate prediction values;
[0037] A dilated spatial pyramid module is provided between the top encoding block and the top decoding block for extracting multi-scale features; a multi-head self-attention context awareness module is provided in other encoding blocks for capturing the dependencies of the features in the encoding stage and the corresponding decoding stage; an uncertainty reduction module is provided in other decoding blocks for performing secondary prediction on the uncertain points corresponding to the prediction values.
[0038] A method for constructing a segmentation network for dermoscopic images, including the following steps:
[0039] Preprocess the data set to obtain an image sample set;
[0040] Construct a segmentation network;
[0041] Use the image sample set to train the segmentation network; and the training process is as follows:
[0042] Use the nearest neighbor interpolation method to downsample the segmentation gold standard, thereby obtaining segmentation gold standards of sizes 1 / 2, 1 / 4, and 1 / 8 for depth supervision to calculate the auxiliary loss;
[0043] Use the Jaccard loss to alleviate the foreground-background imbalance; the Jaccard loss is calculated as follows: where k represents each training sample, N is the batch size, P k represents the output probability map of the segmentation network, and G k represents the gold standard of this sample;
[0044] The loss function of the entire network is calculated as follows: L total = L BCE + λL JA , where, L total represents the total loss of the entire network, and λ is a weight parameter used to balance the weights of the two losses in the parameter update of the entire network.
[0045] Due to the application of the above technical solution, the present invention has the following advantages compared with the prior art:
[0046] 1. In the present invention, by constructing a multi-head self-attention context-aware module, it is possible to capture and model the global context relationship, and alleviate the problems of different scales and various shapes of the regions to be segmented.
[0047] 2. By constructing an uncertainty reduction module, the points with high uncertainty are re-predicted using a multi-layer perceptron to reduce uncertainty and improve the segmentation accuracy.
[0048] To make the above and other objects, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a schematic diagram of the segmentation network architecture in the embodiment of the present invention;
[0051] Figure 2 It is a schematic diagram of the flowchart of the segmentation method in the embodiment of the present invention;
[0052] Figure 3 It is the segmentation result on the dermoscopic image in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0054] Embodiment 1: Refer to Figures 1 to 3 As shown, a method for segmenting dermoscopic images includes the following steps:
[0055] Generate at least a second feature map and a third feature map according to the acquired image;
[0056] Extract and process the features in the second feature map to form a second corrected map;
[0057] Fuse the third feature map and the second corrected map to generate a first fused map;
[0058] Process the uncertain points in the first fused map, and process the certain points in the first fused map and the uncertain points in the processed first fused map to generate a first fused corrected map.
[0059] By extracting and processing features in the encoding stage, connecting them to the decoding stage, modeling the global context relationship, alleviating the problems of different scales and variable shapes of the regions to be segmented. Process the uncertain points to reduce uncertainty and improve the segmentation accuracy.
[0060] Among them, the step of "generating at least a second feature map and a third feature map according to the acquired image" includes: performing downsampling processing on the second feature map to generate a third preset feature map, and generating a third feature map according to the atrous spatial pyramid module and the third preset feature map.
[0061] The step of "extracting and processing the features in the second feature map to form a second corrected map" includes: extracting and processing the features in the second feature map based on the multi-head self-attention context-aware module to form a second corrected map.
[0062] The step of "processing the uncertain points in the first fused map, and processing the certain points in the first fused map and the uncertain points in the processed first fused map to generate a first fused corrected map" includes: processing the uncertain points in the first fused map based on the uncertainty reduction module, and processing the certain points in the first fused map and the uncertain points in the processed first fused map to generate a first fused corrected map. Among them, the uncertain points are the points with a predicted value of 0.5.
[0063] The method for segmenting a dermoscopic image further includes:
[0064] Generate an intermediate feature map according to the acquired image, extract and process the features in the intermediate feature map to form an intermediate corrected map;
[0065] Fuse and process the intermediate corrected map and the fused corrected map of the previous step to generate a second fused map;
[0066] Process the uncertain points in the second fused map, and process the certain points in the second fused map and the uncertain points in the processed second fused map to generate a second fused corrected map.
[0067] A dermoscopic image segmentation network, which uses the above segmentation method to segment a dermoscopic image, and the architecture of the segmentation network includes:
[0068] A baseline network, which is based on the encoding-decoding structure of a fully convolutional network, includes multiple encoding blocks and corresponding decoding blocks. The multiple encoding blocks are connected through downsampling operations, and the multiple decoding blocks are connected through upsampling operations. The encoding blocks are used to extract features of an image, and the decoding blocks restore the image according to the features to generate prediction values.
[0069] A dilated spatial pyramid module is provided between the top encoding block and the top decoding block for extracting multi-scale features; a multi-head self-attention context awareness module is provided in other encoding blocks for capturing the dependencies of the features in the encoding stage and the corresponding decoding stage; an uncertainty reduction module is provided in other decoding blocks for performing secondary prediction on the uncertainty points corresponding to the prediction values.
[0070] Specifically, for example, in the encoder, a second feature map is generated in the second stage through downsampling, its features are extracted, and after being screened by the multi-head self-attention context awareness module to remove redundant features, it restores the image in the second stage of the decoder to form a second corrected map. In the encoder, a third preset feature map is generated in the first stage through downsampling, its features are extracted, and it restores the image in the first stage of the decoder to generate a third feature map. The third feature map and the second corrected map are fused to generate a first fused map. The uncertainty reduction module in the decoder performs secondary prediction on the uncertainty points in the first fused map to improve the segmentation accuracy.
[0071] See Figure 1 As shown, the segmentation network includes four such encoding blocks and corresponding four decoding blocks. Among them, the multi-head self-attention context awareness module is respectively provided in three sequentially connected encoding blocks, and the uncertainty reduction module is respectively provided in the three corresponding decoding blocks.
[0072] By constructing a multi-head self-attention context awareness module to capture non-local information and global context dependencies, redundant feature expressions are screened, and the problems of inconsistent scale sizes and variable shapes of the regions to be segmented are alleviated. Specifically, the features in the encoding stage are mainly high-resolution shallow features, and the features in the decoding stage are mainly low-resolution deep semantic features. Skip connections fuse the shallow features and the deep features, but this simple and inefficient fusion often brings many interferences. On the one hand, it is because the shallow features contain a large amount of redundant information, and on the other hand, it is because the context modeling ability of the encoder is weak and cannot effectively model global dependencies. Therefore, by constructing a multi-head self-attention context awareness module in the segmentation network, the context modeling ability of the encoder is enhanced, and redundant feature expressions are screened.
[0073] At the top layer of the baseline network, there is a dilated spatial pyramid module for extracting multi-scale features, such as feature H, feature W, and feature C. The features of different scales are used as query values in the convolutional layer, which can provide richer and more accurate context information. The deep features and shallow features of multiple stages are fused, which can not only represent semantic information but also contain sufficient detailed information.
[0074] Build an uncertainty reduction module to reduce uncertainty and improve segmentation accuracy by re-predicting points with high uncertainty using a multi-layer perceptron.
[0075] See Figure 2 As shown, the construction method of the segmentation network includes the following steps:
[0076] S1. Dataset preprocessing;
[0077] S2. Establish an encoder-decoder structure baseline network;
[0078] S3. Build a multi-head self-attention context-aware module;
[0079] S4. Build an uncertainty reduction module;
[0080] S5. Establish a multi-head self-attention uncertainty reduction segmentation network;
[0081] S6. Train the multi-head self-attention uncertainty reduction segmentation network;
[0082] S7. The segmentation network automatically segments the lesions.
[0083] By combining the baseline network with the multi-head self-attention context-aware module, the uncertainty reduction module, and the dilated spatial pyramid module through the above steps, a multi-head self-attention uncertainty reduction segmentation network is constructed. The multi-head self-attention context-aware module is used to model global context dependencies, the dilated spatial pyramid module is used to extract multi-scale feature expressions, and the uncertainty reduction module re-predicts points with high uncertainty to obtain a more accurate segmentation result. The following provides more detailed steps.
[0084] S1. Dataset preprocessing
[0085] The dataset is a publicly available dermoscopic image dataset. The images in the publicly available dermoscopic image dataset vary in size from 540×576 to 6688×6780. To accelerate the network reading speed and calculation efficiency, the image sizes in the dataset are uniformly adjusted to 192×256 while maintaining the original aspect ratio of the images. The pixel values of the three RGB channels of the images are normalized from 0-255 to 0-1. To further increase the number of image samples in the dataset for constructing, training, and testing the segmentation network, and to improve the accuracy of the segmentation network in segmenting the lesion area by reducing the uncertainty of the multi-head self-attention, data augmentation is performed on the samples in the dataset. The specific augmentation methods can include horizontal flipping, vertical flipping, and rotation between -90° and 90°. Random transformations can be performed under each augmentation method to achieve online data augmentation of the sample size.
[0086] S2. Establish an encoder-decoder structure baseline network
[0087] The segmentation network is based on the encoder-decoder structure of the fully convolutional network, including an encoder and a decoder corresponding to the encoder. The encoder includes multiple encoding blocks, and the multiple encoding blocks are connected through downsampling operations. The multiple decoding blocks are connected through upsampling operations. The encoder uses the ResNet34 residual network with the last fully connected layer removed as the backbone network of the encoder to reduce the spatial resolution, extract more discriminative features, and promote convergence. The decoder includes multiple decoding blocks, and each decoding block includes a bilinear interpolation function and two 3×3 convolutions. The decoder is used to gradually restore the original size of the image and generate the final pixel-by-pixel classification prediction.
[0088] S3. Construct a multi-head self-attention context-aware module
[0089] The process of capturing non-local information and global context dependencies through the multi-head self-attention context-aware module includes:
[0090] Using three different convolutional layers Q, K, and V to map the features F extracted from the encoding blocks in the encoder to the embedding space to obtain the feature maps F Q 、F K 、F V 。
[0091] Mapping the channel dimensions of F Q 、F K 、F V to C to reduce the number of parameters, and obtaining the feature maps F Q 、F K 、F V in the embedding space with channel dimensions of H×W×C.
[0092] Adjust the feature map F Q and F k have the shape of HW×C, and transpose F Q so that the shape of F Q becomes C×HW.
[0093] Perform matrix multiplication on F Q and F k to obtain a feature map F T of size HW×HW.
[0094] To utilize the position information in the feature map, use the position encoding map E to encode the relative positions in the feature map. The position encoding map E has a size of HW×HW and is a learnable parameter. Among them, the value E(i,j) of each pixel represents the relative relationship between the ith position and the jth position.
[0095] According to Att = softmax(F T +E), obtain the attention weight map Att. Among them, the attention weight map Att has a size of HW×HW, and Att(i,j) represents the dependence relationship between the ith pixel value and the jth pixel value. F T is a feature map of size HW×HW, and E represents the position encoding.
[0096] Transpose Att and perform matrix multiplication with the feature map F V and add the feature F in the form of residual connection to obtain the filtered feature F’. This feature F’ effectively models the global dependence relationship. In the decoding block, concatenate the feature F’ to restore the image. In addition, when generating the attention weight map Att, multiple attention heads are used to more effectively model the non-local global dependence and remove redundant features.
[0097] S4. Construct an uncertainty reduction module
[0098] For a pixel point in the image, if the prediction value given by the segmentation network for this point approaches 1, it indicates that this point is probably part of the foreground; if the prediction value given by the segmentation network for this point approaches 0, it indicates that this point is probably part of the background. However, if the prediction value of this point is close to 0.5, it indicates that the segmentation network cannot make an effective prediction for this point, that is, this point is a point with high uncertainty. By constructing an uncertainty reduction module, re-predict the points with high uncertainty generated during the upsampling process in the decoding stage to reduce the uncertainty and improve the segmentation accuracy.
[0099] To determine which points are points with high uncertainty, the present invention uses 1×1 convolution after each upsampling in the encoder to reduce the number of channels of the feature map to 1;
[0100] Activated by the Sigmoid function, the value of each pixel is normalized to 0-1 as the first prediction value.
[0101] For all pixels activated by the Sigmoid function, select N pixels with the first prediction value of 0.5, and extract the point-level features at their positions in the channel dimension. For example, if the size of the feature map is H×W×C, then N point-level features of size 1×1×C can be obtained.
[0102] Next, use a two-layer fully connected layer to construct a perceptron to re-predict the prediction values of the N pixels, and use the second prediction value obtained by re-prediction to update the corresponding first prediction value, reducing uncertainty and improving segmentation accuracy.
[0103] Among them, the cross-entropy loss function is used as the auxiliary loss for deep supervision in the uncertainty reduction module. By introducing the deep supervision mechanism, the auxiliary loss is used to promote the convergence of the network. The calculation of the cross-entropy loss is as follows: L BCE =-G·logP-(1-G)log(1-P), where P represents the output probability map of the segmentation network, G represents the segmentation gold standard, and L BCE represents the cross-entropy loss between the output probability map and the segmentation gold standard.
[0104] S5. Establish a multi-head self-attention uncertainty reduction segmentation network
[0105] Combine the encoder-decoder structure baseline network with the multi-head self-attention context perception module, the uncertainty reduction module, and the atrous spatial pyramid module to construct a multi-head self-attention uncertainty reduction segmentation network. The atrous spatial pyramid module is used at the top of the encoder-decoder structure to extract multi-scale feature expressions. The atrous spatial pyramid module consists of multiple atrous convolution operations with different dilation rates and can effectively extract multi-scale features. In all the remaining skip connections of the encoder-decoder structure, a multi-head self-attention context perception module is added. This module can capture global dependencies, effectively model global context information, screen the features in the encoding stage, and leave more representative feature expressions. After each upsampling in the decoding stage, an uncertainty reduction module is used to re-predict the positions with high uncertainty (points where the prediction probability value approaches 0.5) to obtain a higher confidence, reduce uncertainty, and thus obtain a more accurate segmentation result.
[0106] S6. Train the multi-head self-attention uncertainty reduction segmentation network
[0107] To train the segmentation network, segmentation gold standards of different scales are required. The specific training process is as follows:
[0108] The segmentation gold standard is downsampled using the nearest neighbor interpolation method to obtain segmentation gold standards of sizes 1 / 2, 1 / 4, and 1 / 8, which are used for calculating the auxiliary loss in depth supervision;
[0109] For the output layer of the segmentation network, the Jaccard loss is used to alleviate the foreground-background imbalance; the Jaccard loss is calculated as follows: where k represents each training sample, N is the batch size, P k represents the output probability map of the segmentation network, and G k represents the gold standard of this sample;
[0110] Thus, the loss function of the entire network is calculated as follows: L total = L BCE + λL JA , where L total represents the total loss of the entire network, λ is a weight parameter used to balance the weights of the two losses in the update of the entire network's parameters, and it is set empirically.
[0111] S7. The segmentation network automatically segments the lesion
[0112] The picture to be tested is input into the system, and the multi-head self-attention uncertainty reduction segmentation network automatically segments the lesion area according to the input picture to be tested.
[0113] The multi-head self-attention uncertainty reduction segmentation network that has completed the test can automatically segment the lesion area of the dermoscopic image to be segmented.
[0114] In the present invention, specific embodiments are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for segmenting dermoscopic images, characterized in that, Including the following steps: At least generate a second feature map and a third feature map based on the acquired image. Among them, perform downsampling processing on the second feature map to generate a third preset feature map, and generate a third feature map according to the atrous spatial pyramid module and the third preset feature map; Extract and process the features in the second feature map to form a second corrected map; Fuse the third feature map and the second corrected map to generate a first fused map; Based on the uncertainty reduction module, process the uncertain points in the first fused map, and process the certain points in the first fused map and the uncertain points in the processed first fused map to generate a first fused corrected map, where the uncertain points are points with a predicted value of 0.5; The "processing the uncertain points in the first fused map based on the uncertainty reduction module" includes: using a 1×1 convolution after each upsampling in the encoding block to reduce the number of channels of the feature map to 1; Activate using the Sigmoid function to normalize the value of each pixel point to 0-1 as the first prediction value; For all pixel points activated by the Sigmoid function, select N pixel points with a first prediction value of 0.5, and extract the point-level features at their positions in the channel dimension; Use a two-layer fully connected layer to construct a perceptron to re-predict the second prediction values of N pixel points, and update the corresponding first prediction values using the re-predicted second prediction values.
2. The segmentation method of dermoscopic images according to claim 1, wherein: The step "at least generate a second feature map and a third feature map based on the acquired image" includes: Perform downsampling processing on the second feature map to generate a third preset feature map, and generate a third feature map according to the atrous spatial pyramid module and the third preset feature map.
3. The segmentation method of the dermoscopic image according to claim 1, wherein: The step "extract and process the features in the second feature map to form a second corrected map" includes: Based on the multi-head self-attention context perception module, extract and process the features in the second feature map to form a second corrected map.
4. The segmentation method of the dermoscopic image according to claim 3, wherein: The step "extract and process the features in the second feature map to form a second corrected map" includes: The encoding block extracts feature F from the second feature map; The feature F extracted from the encoding block is mapped to the embedding space using three different convolutional layers Q, K, and V to obtain the feature map F Q 、F K 、F V ; Map the channel dimensions of F Q , F K , F V to C, reduce the number of parameters, and obtain the feature maps F Q , F K , F V in the embedding space, whose channel dimensions are all H×W×C; Adjust the feature map F Q and F k have a shape of HW×C, and transpose F Q so that the shape of F Q becomes C×HW; For F Q and F k Using matrix multiplication, a feature map F of size HW×HW is obtained T ; Use the position encoding map E to encode the relative positions in the feature map. The size of the position encoding map E is HW×HW, where the value E(i,j) of each pixel point represents the relative relationship between the i-th position and the j-th position; According to Att = softmax(F T + E), the attention weight map Att is obtained, where the size of the attention weight map Att is HW × HW, and Att(i, j) represents the dependence relationship between the i-th pixel value and the j-th pixel value. F T is a feature map of size HW × HW, and E represents the positional encoding; Transpose Att and perform matrix multiplication with the feature map F V Add the feature F in the form of residual connection to obtain the filtered feature F'. In the decoding block, concatenate the feature F' to restore the image to form the second corrected map.
5. The segmentation method of a dermoscopic image according to claim 1, wherein: In the uncertainty reduction module, the cross-entropy loss function is used as an auxiliary loss for deep supervision; the calculation of the cross-entropy loss is as follows: L BCE = -G·logP - (1 - G)log(1 - P), where P represents the output probability map of the segmentation network, G represents the segmentation ground truth, and L BCE represents the cross-entropy loss between the output probability map and the segmentation ground truth.
6. The segmentation method of a dermoscopic image according to claim 1, wherein It also includes: Generate an intermediate feature map based on the acquired image, extract and process the features in the intermediate feature map to form an intermediate corrected map; Fuse and process the intermediate corrected map and the fused corrected map of the previous step to generate a second fused map; Process the uncertain points in the second fused map, and process the certain points in the second fused map and the uncertain points in the processed second fused map to generate a second fused corrected map.
7. A method for constructing a segmentation network of dermoscopic images. The segmentation network applies the segmentation method of dermoscopic images described in claim 1. The architecture of the segmentation network includes: The baseline network, which is based on the encoder-decoder structure of a fully convolutional network, includes a plurality of encoding blocks and decoding blocks corresponding to the encoding blocks. The plurality of encoding blocks are connected through downsampling operations, and the plurality of decoding blocks are connected through upsampling operations. The encoding blocks are used to extract features of an image, and the decoding blocks restore the image according to the features to generate prediction values; A dilated spatial pyramid module is provided between the top encoding block and the top decoding block for extracting multi-scale features; a multi-head self-attention context awareness module is provided in other encoding blocks for capturing the dependencies of the features in the encoding stage and the corresponding decoding stage; an uncertainty reduction module is provided in other decoding blocks for performing secondary prediction on the uncertainty points corresponding to the prediction values; It is characterized in that the construction method includes the following steps: Preprocess the data set to obtain an image sample set; Construct a segmentation network; Train the segmentation network using the image sample set; and the training process is as follows: Use the nearest neighbor interpolation method to downsample the segmentation gold standard, thereby obtaining segmentation gold standards of sizes 1 / 2, 1 / 4, and 1 / 8 for calculating the auxiliary loss by deep supervision; Use the Jaccard loss to alleviate foreground and background imbalance; the Jaccard loss is calculated as follows: where k represents each training sample, N is the batch size, P k represents the output probability map of the segmentation network, and G k represents the ground truth of this sample; The loss function of the entire network is calculated as follows: L total = L BCE + λL JA , where L total represents the total loss of the entire network, and λ is a weight parameter used to balance the weights of the two losses in the update of the entire network parameters.
Citation Information
Patent Citations
Medical image segmentation method and system based on boundary and neighborhood guidance
CN112489062A
Software defined network performance intelligent prediction method
CN113158543A