Melanoma segmentation method based on dilated convolution and multi-scale fusion
By adopting hollow convolution and multi-scale fusion methods in dermatoscope image segmentation, the problem of low melanoma segmentation accuracy in the prior art is solved, and higher segmentation accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202011094831.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-10-14
AI Technical Summary
The prior art is difficult to accurately segment melanomas in dermatoscope images, especially in the presence of blurred boundaries, different shapes and interferers, resulting in low segmentation accuracy.
The segmentation method based on hollow convolution and multi-scale fusion is adopted. The channel attention void convolution module expands the receptive field, suppresses information redundant noise, and obtains multi-scale information through the aggregation interaction module to make up for the semantic gap between the encoding layer and the decoding layer.
It improves the adaptability to melanoma areas with large size differences, improves the accuracy of melanoma segmentation, and can more accurately assist dermatologists in segmenting dermatoscope images.
Smart Images

Figure CN112446890B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for segmenting skin cancer melanoma Background Art
[0002] Melanoma is one of the most dangerous skin diseases. Early studies have found that the 5-year survival rate for the most aggressive melanomas can be as high as 99%; however, delayed diagnosis can cause the survival rate to drop to 23%. Dermoscopy, the examination of skin lesions with a dermatoscope, is commonly used to diagnose melanoma. However, manual review of dermoscopic images is an error-prone and time-consuming task, even for professional dermatologists.
[0003] Therefore, it is necessary to develop a computational support system to assist dermatologists in accurately segmenting melanomas. This task remains challenging because melanomas exist in different sizes, shapes, and textures. Moreover, some dermoscopic images may contain distractors such as hair, ruler markings, and color calibration. Convolutional neural networks are widely used to solve semantic segmentation tasks. Among them, U-Net is widely used in medical image segmentation. It adopts an encoder (downsampling) and decoder (upsampling) structure, and uses jump connections to integrate low-level texture features and corresponding high-level semantic features. Since shallow features are not processed and directly integrated with deep features, information redundancy will occur, which will affect the segmentation accuracy results.
[0004] Since melanoma regions have fuzzy boundaries and various shapes, it is difficult for general segmentation networks to accurately segment melanoma. There may be strong connections between large-scale pixels in medical images, but general segmentation networks usually use fixed-size convolution kernels to downsample images, which results in the network only capturing local context information. The proposed Atrous Spatial Convolution Pooling Pyramid (ASPP) can only extract partial context information after downsampling and cannot produce compact multi-scale features. Summary of the invention
[0005] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provide a melanoma segmentation method based on dilated convolution and multi-scale fusion.
[0006] In order to cope with the interference in dermoscopic images that causes the neural network to be unable to obtain good segmentation accuracy, the present invention makes corresponding changes to the downsampling structure and simple skip connection in the traditional U-Net, effectively expands the receptive field and suppresses the noise caused by the direct fusion of shallow features and deep features, and provides a melanoma segmentation method based on void convolution and multi-scale fusion.
[0007] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention is further described below. A melanoma segmentation method based on dilated convolution and multi-scale fusion comprises the following steps:
[0008] Step 1) preprocessing medical images;
[0009] The collected dermoscopic image data were divided into training set, validation set and test set according to the ratio of 7:1:2, and data augmentation was performed on the training set images used for network training;
[0010] Step 2) construct a multi-scale aggregation network model with flexible receptive field;
[0011] 2.1 Construct a channel-attention dilated convolution module for feature extraction;
[0012] The encoding layer in U-Net is replaced by a channel-attention dilated convolutional layer, and the dermoscopic image in step 1) is used as input. The extracted feature map featuremap is output to provide input for the subsequent network; each layer uses three parallel dilated convolutions to extract features, and the dilation rate of each dilated convolution is different; then the extracted feature map is subjected to a global average pooling operation, and the cross-channel interaction information is captured by considering each channel and its k neighboring channels, and the weight is reallocated to each channel of the feature map, and then the three feature maps are added together to obtain the output feature map of this layer; the shallow layer usually learns simple texture information, and as the number of layers increases, complex abstract information will be captured;
[0013] 2.2 Building the Aggregate Interaction Module, or AIM;
[0014] The aggregation interaction module is proposed to make up for the semantic gap between the feature maps of the encoding layer and the corresponding decoding layer and suppress the noise that may be caused by jump connections. In U-Net, the two are directly aggregated. Since the semantic information between the two is quite different, redundant information noise will be generated, which will affect the final segmentation result. AIM receives feature maps from adjacent encoding layers, reduces the number of feature channels through a 3*3 convolutional layer to reduce the amount of calculation, and then uses the convolutional layer to obtain multi-scale information aggregation into the final feature map.
[0015] 2.3 Construct the decoding layer;
[0016] The decoding layer upsamples the feature map obtained by the encoding layer, and then adds the coefficients to the feature map output by the aggregation interaction module. After two 3*3 convolution layer operations, the output feature map is obtained; after the last decoding layer, a 1*1 convolution layer and Sigmoid function processing are connected to obtain the final segmentation result;
[0017] Step 3) Input the training set data into the model for training;
[0018] Input the processed training set in step 1) into the network model constructed in step 2), use random initialization and Adam optimization method; set the initial learning rate, momentum, and number of iterations, and train according to the set training strategy. First, perform data augmentation on the input training set and then train, then use the validation set on the trained network model to obtain the validation result, then update once according to the gradient, and repeat this step until the number of iterations is reached;
[0019] Step 4) segmentation of lesion area in dermoscopic image;
[0020] The test set data is input into the prediction model trained in step 3) to obtain the segmentation result. According to the evaluation index, it can be shown that the present invention can assist in the segmentation of dermoscopic images.
[0021] The present invention proposes a melanoma segmentation method based on dilated convolution and multi-scale fusion. By adding three parallel dilated convolutions and assigning different weights to the feature maps generated by each dilated convolution, the effect of flexibly expanding the receptive field is achieved; by fusing the feature information obtained from adjacent downsampling layers, the noise interference caused by information redundancy in the jump connection is suppressed. The adaptability to melanoma areas with large size differences is improved, and the accuracy of melanoma segmentation is improved.
[0022] The present invention has the following advantages:
[0023] The present invention proposes a multi-scale aggregation network with a flexible receptive field for segmenting melanoma, wherein the channel attention hole convolution module can adaptively expand the receptive field according to image features, obtain tighter contextual information, and alleviate the problem of insufficient features caused by a fixed receptive field; the aggregation interaction module can aggregate the features output by the encoding layer with the features of the adjacent encoding layer to obtain multi-scale information, and reduce the semantic gap between the encoding layer and the corresponding decoding layer, suppressing the information redundancy noise caused by the jump connection structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic diagram of the overall network framework structure of the system for implementing the method of the present invention.
[0025] Figure 2 It is a schematic diagram of a channel attention hole convolution module in a network of a system implementing the method of the present invention.
[0026] Figure 3 It is a schematic diagram of a channel attention module in a network of a system for implementing the method of the present invention.
[0027] Figure 4 It is a schematic diagram of the aggregation interaction module in the network of the system for implementing the method of the present invention. DETAILED DESCRIPTION
[0028] The present invention will be further described below in conjunction with the accompanying drawings:
[0029] The melanoma segmentation method based on dilated convolution and multi-scale fusion of the present invention comprises the following steps:
[0030] Step 1) preprocessing medical images;
[0031] The collected dermoscopic images were divided into training set, validation set and test set according to the ratio of 7:1:2, and the image pixel size was set to 128*128; the training set images used for network training were augmented, randomly rotated in the range of -30° to 30°, randomly flipped horizontally and randomly scaled to between 0.8 and 1.2 times of the original image;
[0032] Step 2) construct a multi-scale aggregation network model with flexible receptive field;
[0033] 2.1 Construct a channel-attention dilated convolution module for feature extraction;
[0034] The encoding layer in U-Net is replaced by a channel-attention dilated convolutional layer, and the dermoscopic image in step 1) is used as input. The extracted feature map featuremap is output to provide input for the subsequent network; each layer uses three parallel dilated convolutions to extract features, and the dilation rate of each dilated convolution is different; then the extracted feature map is subjected to a global average pooling operation, and the cross-channel interaction information is captured by considering each channel and its k neighboring channels, and the weight is reallocated to each channel of the feature map, and then the three feature maps are added together to obtain the output feature map of this layer; the shallow layer usually learns simple texture information, and as the number of layers increases, complex abstract information will be captured;
[0035] Downsampling uses a total of 5 channel-attention dilated convolution layers. Each layer uses three parallel dilated convolutions with a kernel size of 3*3. The dilation rates are set to 1, 2, and 3, respectively. The stride is 1, and the padding is the same as the dilation rate. The pooling operation uses 2*2 maximum pooling. The 128*128*3 image is input, and three feature maps with 64 channels are obtained through three dilated convolutions with 64 convolution kernels. After that, a vector of 1*1*C is obtained through the global average pooling operation of the channel attention module. Then, a 1*1 convolution with a kernel size of 3 is used to obtain cross-channel information. After that, the Sigmoid function is used for activation, and each channel is assigned its own weight by multiplying the original feature map coefficient. Finally, a 128*128*64 feature map is obtained by adding them together, and a pooling operation is used as the input of the next encoding layer. Repeat this operation five times, and the encoding layer obtains feature maps with 64, 128, 256, 512, and 1024 channels respectively. The module can be described as:
[0036]
[0037] Where D represents the dilated convolution, k represents the dilation rate, C represents the channel attention module, and f represents the input feature;
[0038] 2.2 Building the Aggregate Interaction Module, or AIM;
[0039] The aggregation interaction module is proposed to make up for the semantic gap between the feature maps of the encoding layer and the corresponding decoding layer and suppress the noise that may be caused by the jump connection. In U-Net, the two are directly aggregated. Since the semantic information between the two is quite different, redundant information noise will be generated, which will affect the final segmentation result. AIM receives feature maps f from adjacent encoding layers. i-1 、f i 、f i+1 , a 3*3 convolution layer is used to reduce the number of feature channels to reduce the amount of calculation; then each branch B is scaled to the size of the adjacent branch feature map using pooling or interpolation operations, and each branch is fused by coefficient addition; finally, all branches are integrated into a convolution layer, and a residual module is added to the output to make training easier to optimize; the entire module process can be written as:
[0040]
[0041]
[0042] Where I and M represent residual mapping and branch merging, respectively, i represents the operation of branch i, and f represents the input feature;
[0043] 2.3 Construct the decoding layer;
[0044] The decoding layer upsamples the feature map obtained by the encoding layer, and then adds the coefficients to the feature map output by the aggregation interaction module. After two 3*3 convolutional layer operations, the output feature map is obtained; after the last decoding layer, a 1*1 convolutional layer and Sigmoid function are used to obtain the final segmentation result; the Sigmoid function is defined as follows:
[0045]
[0046] Step 3) Input the training set data into the model for training;
[0047] Input the processed training set in step 1) into the network model constructed in step 2), and use random initialization and Adam optimization method; set the initial learning rate, momentum, and number of iterations, and train according to the set training strategy; first perform data augmentation on the input training set and then train, then obtain the verification result on the trained network model, and then update it once according to the gradient, and repeat this step until the number of iterations is reached;
[0048] The batch size is 12, the epoch is 80, the initial learning rate is 0.0001, and the momentum is 0.9. The prediction and groundtruth obtained by training are trained using tverskyloss+consistency-enhancedloss. The loss function can be written as:
[0049]
[0050]
[0051] L total =L tver (p, g, α, β)+L cel (p, g) (7)
[0052] Among them, α is set to 0.3, β is set to 0.7, p and g represent the predicted image and the calibrated standard image respectively;
[0053] Step 4) segmentation of lesion area in dermoscopic image;
[0054] Input the test set data into the prediction model trained in step 3) to obtain the segmentation result. The evaluation indicators are used to evaluate the segmentation result, including accuracy (AC), dice coefficient (DI), Jacquard index (JA), and sensitivity (SE). The calculation method of the evaluation indicators is as follows:
[0055]
[0056]
[0057] Among them, TP is true positive, TN is true negative, FP is false positive, and FN is false negative; according to the evaluation index, it can be shown that the present invention can assist in segmenting dermoscopic images.
[0058] The channel attention hole convolution module proposed in the present invention can adaptively expand the receptive field according to the image features, obtain more compact context information, and alleviate the problem of insufficient features caused by the fixed receptive field; the aggregation interaction module can aggregate the features output by the encoding layer with the features of the adjacent encoding layer to obtain multi-scale information, and reduce the semantic gap between the encoding layer and the corresponding decoding layer, suppressing the noise caused by direct aggregation. The present invention can segment accurate dermoscopic images and play an auxiliary role.
Claims
1. A melanoma segmentation method based on dilated convolution and multi-scale fusion, comprising the following steps: Step 1) preprocessing medical images; The collected dermoscopic images were divided into training set, validation set and test set according to the ratio of 7:1:2, and the image pixel size was set to 128*128; the training set images used for network training were augmented, randomly rotated in the range of -30° to 30°, randomly flipped horizontally and randomly scaled to between 0.8 and 1.2 times of the original image; Step 2) construct a multi-scale aggregation network model with flexible receptive field; 2.1 Construct a channel-attention dilated convolution module for feature extraction; The encoding layer in U-Net is replaced by a channel-attention dilated convolutional layer, and the dermoscopic image in step 1) is used as input. The extracted feature map is output as input for the subsequent network. Each layer uses three parallel dilated convolutions to extract features, and the dilation rate of each dilated convolution is different. Then, a global average pooling operation is performed on the extracted feature map. By considering each channel and its k neighboring channels, cross-channel interaction information is captured, and the weights are reallocated to each channel of the feature map. Then, the three feature maps are added together to obtain the output feature map of this layer. The shallow layer usually learns simple texture information, and as the number of layers increases, complex abstract information will be captured. Downsampling uses a total of 5 channel-attention dilated convolution layers. Each layer uses three parallel dilated convolutions with a kernel size of 3*3. The dilation rates are set to 1, 2, and 3 respectively, with stride=1. The padding is the same as the dilation rate, and the pooling operation uses 2*2 maximum pooling. Input a 128*128*3 image, and at the same time, three dilated convolutions with 64 convolution kernels are used to obtain three feature maps with 64 channels. After that, a global average pooling operation is performed on the channel attention module to obtain a vector of size 1*1*C. Then, a 1*1 convolution with a kernel size of 3 is used to obtain cross-channel information. After that, the Sigmoid function is used for activation, and each channel is assigned its own weight by multiplying the original feature map coefficient. Finally, a 128*128*64 feature map is added together, and a pooling operation is used as the input of the next encoding layer. Repeat this operation five times, and the encoding layer obtains feature maps with 64, 128, 256, 512, and 1024 channels respectively. The module can be described as: Where D represents the dilated convolution, u represents the dilation rate, C represents the channel attention module, and f represents the input feature; 2.2 Building the Aggregate Interaction Module, or AIM; The aggregation interaction module is proposed to make up for the semantic gap between the feature maps of the encoding layer and the corresponding decoding layer and suppress the noise that may be caused by the jump connection. In U-Net, the two are directly aggregated. Since the semantic information between the two is quite different, redundant information noise will be generated, which will affect the final segmentation result. AIM receives feature maps f from adjacent encoding layers. i-1 、f i 、f i+1 , a 3*3 convolution layer is used to reduce the number of feature channels to reduce the amount of calculation; then each branch B is scaled to the size of the adjacent branch feature map using pooling or interpolation operations, and each branch is fused by coefficient addition; finally, all branches are integrated into a convolution layer, and a residual module is added to the output to make training easier to optimize; the entire module process can be written as: Where I and M represent residual mapping and branch merging, respectively, i represents the operation of branch i, and f represents the input feature; 2.3 Construct the decoding layer; The decoding layer upsamples the feature map obtained by the encoding layer, and then adds the coefficients to the feature map output by the aggregation interaction module. After two 3*3 convolutional layer operations, the output feature map is obtained; after the last decoding layer, a 1*1 convolutional layer and Sigmoid function are used to obtain the final segmentation result; the Sigmoid function is defined as follows: Step 3) Input the training set data into the model for training; Input the processed training set in step 1) into the network model constructed in step 2), using random initialization and Adam optimization method; Set the initial learning rate, momentum, and number of iterations, and train according to the set training strategy; first perform data augmentation on the input training set and then train, then use the validation set to obtain the validation results on the trained network model, then update once according to the gradient, and repeat this step until the number of iterations is reached; The batch size is 12, the epoch is 80, the initial learning rate is 0.0001, and the momentum is 0.
9. The prediction and groundtruth obtained by training are trained using tversky loss + consistency-enhanced loss. The loss function can be written as: L total =L tver ( p i ,g i ,α,β)+L cel ( p i ,g i ) (7) Among them, α is set to 0.3, β is set to 0.7, p and g represent the predicted image and the calibrated standard image respectively; Step 4) segmentation of lesion area in dermoscopic image; Input the test set data into the prediction model trained in step 3) to obtain the segmentation result. The evaluation indicators are used to evaluate the segmentation result, including accuracy (AC), dice coefficient (DI), Jacquard index (JA), and sensitivity (SE). The calculation method of the evaluation indicators is as follows: Among them, TP is true positive, TN is true negative, FP is false positive, and FN is false negative.
Citation Information
Patent Citations
Method and system for melanoma image tissue segmentation based on deep neural network
CN108510502A
Medical image automatic segmentation method based on multi-path attention fusion
CN111681252A