Intelligent melanoma segmentation method and system based on DDPM

By introducing the multi-scale feature fusion and dual attention module of the DDPM segmentation model, combined with distance map supervision, the problems of multi-scale feature extraction, boundary segmentation and noise processing in melanoma segmentation are solved, and high-precision melanoma image segmentation is achieved.

CN120388036APending Publication Date: 2025-07-29SHANGHAI UNIV

Patent Information

Application Number
CN202510530124.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing melanoma segmentation algorithm has shortcomings in multi-scale feature extraction, boundary segmentation accuracy and noise processing, resulting in inaccurate segmentation results, affecting diagnosis and treatment decisions.

Method used

The segmentation method based on DDPM is adopted, and the multi-scale feature fusion module and dual attention module are combined with the distance graph to supervise the signal, optimize boundary segmentation, eliminate noise interference, and improve feature alignment accuracy.

Benefits of technology

High-precision segmentation of melanoma images is achieved, the accuracy of segmentation results and boundary recognition capabilities are improved, and the impact of noise interference is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388036A_ABST
    Figure CN120388036A_ABST
Patent Text Reader

Abstract

The invention relates to a DDPM-based melanoma intelligent segmentation method and system. The method comprises the following steps of obtaining a dermatoscope image of melanoma and performing image preprocessing; inputting the dermatoscope image of the melanoma after image preprocessing into a DDPM segmentation model, gradually adding noise through a forward denoising process, gradually removing noise through a backward denoising process, further learning probability distribution of the dermatoscope image of the melanoma, and obtaining a segmentation result of the melanoma; the DDPM segmentation model comprises an encoder and a decoder, the encoder comprises a multi-scale feature fusion module and a double-attention module, and the encoder extracts multi-scale features of a dermatoscope image of the melanoma through a multi-scale convolutional layer; and the decoder gradually recovers the segmentation result of the dermatoscope image of the melanoma according to the multi-scale features of the dermatoscope image of the melanoma through up-sampling and feature fusion operations. Compared with the prior art, the method has the advantage that accurate segmentation of the melanoma image is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image analysis, and particularly to an intelligent melanoma segmentation method and system based on DDPM. Background Art

[0002] In the field of medical image analysis, early diagnosis of melanoma is crucial for improving patient survival rate. In recent years, with the rapid development of deep learning technology, intelligent segmentation algorithms based on convolutional neural networks (CNNs) and Transformers have made remarkable progress in melanoma diagnosis. These methods assist doctors in making more accurate diagnoses by automatically segmenting the lesion areas in dermoscopic images.

[0003] In the development of the prior art, U-Net, as a classic medical image segmentation network in CNNs, can effectively capture local and global features of images through an encoder-decoder structure and skip connections. U-Net has performed excellently in melanoma segmentation tasks and has become the basis for many subsequent studies. Subsequently, scholars have also improved segmentation accuracy and model robustness by introducing dense connections, nested skip connections, dilated convolutions, and multi-scale feature fusion. With the rise of Transformers as a method for capturing global features of images, Swin-Unet introduced it into medical image segmentation tasks and improved the feature capture ability through self-attention mechanisms. TransUNet combines the advantages of CNNs and Transformers, captures global features through Transformers, and captures local features through CNNs, achieving more accurate segmentation. The rapid development of image generation networks has been widely applied in the fields of image segmentation, classification, and detection. Generative adversarial networks (GANs) can generate high-quality segmentation results through adversarial training between a generator and a discriminator. GAN-based methods have performed excellently in melanoma segmentation tasks, especially in dealing with complex boundaries and noise.

[0004] However, existing melanoma segmentation algorithms still face some key challenges. First, there is insufficient multi-scale feature extraction. When existing segmentation frameworks based on CNNs and Transformers process melanoma images, it is often difficult to effectively capture multi-scale features of the lesion areas. The morphology and size of melanoma vary greatly among different patients, and single-scale feature extraction is difficult to meet the requirements of accurate segmentation.

[0005] Secondly, the boundary segmentation accuracy is insufficient. The boundaries of melanoma are usually irregular and blurred, and existing methods are prone to errors when segmenting the boundaries, resulting in inaccurate segmentation results. For example, patent application CN115689993A discloses a skin cancer image segmentation method based on attention and multi-feature fusion. This method uses the attention mechanism to extract context information and edge features. However, the self-attention mechanism focuses more on local features and ignores global information. This method uses the CGE module to combine ordinary convolution and dilated convolution to fuse image information. Although it can fuse image features at multiple scales, it lacks feature weight matching and the fusion of high and low frequency information. This inaccuracy in boundary segmentation may affect subsequent diagnostic and treatment decisions. The problem of noise interference also exists in many models. There are often noises and artifacts in dermoscopic images, and these interference factors will affect the performance of the segmentation algorithm. Existing methods often have poor effects when dealing with noise, resulting in unnecessary noise points and artifacts in the segmentation results. There may also be an impact from the added noise in the DDPM. Thirdly, there is the problem of feature alignment. During the multi-scale feature fusion process, the alignment problem between different scale features has not been effectively solved. Inaccurate feature alignment will lead to information loss and affect the segmentation accuracy. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide an intelligent melanoma segmentation method and system based on DDPM, which realizes the accurate segmentation of melanoma images.

[0007] The purpose of the present invention can be achieved through the following technical solutions:

[0008] An intelligent melanoma segmentation method based on DDPM, comprising the following steps:

[0009] Obtain the dermoscopic image of melanoma and perform image preprocessing;

[0010] Input the dermoscopic image of melanoma after image preprocessing into the DDPM segmentation model, gradually add noise through the forward denoising process, and gradually remove noise through the backward denoising process, thereby learning the probability distribution of the dermoscopic image of melanoma and obtaining the segmentation result of melanoma;

[0011] The DDPM segmentation model includes an encoder and a decoder. The encoder includes a multi-scale feature fusion module and a dual attention module. The encoder extracts multi-scale features of the dermoscopic image of melanoma through multi-scale convolutional layers, and the decoder gradually restores the segmentation result of the dermoscopic image of melanoma according to the multi-scale features of the dermoscopic image of melanoma through upsampling and feature fusion operations.

[0012] Further, the image preprocessing includes image normalization, data augmentation, and noise removal.

[0013] Further, the multi-scale feature fusion module fuses the diffusion noise embedding and the conditional semantic embedding through a cross-attention mechanism, encodes the diffusion noise embedding into the conditional semantic embedding, and encodes the conditional semantic embedding into the diffusion noise embedding. The diffusion noise embedding is generated by the multi-scale convolutional layer and is gradually added to the dermoscopic image of melanoma. The conditional semantic embedding is extracted from the dermoscopic image of melanoma through a multi-scale convolutional layer.

[0014] Further, the specific steps for the multi-scale feature fusion module to perform multi-scale feature fusion are as follows:

[0015] Use an encoder to extract multi-scale features of the dermoscopic image of melanoma through a multi-scale convolutional layer;

[0016] Transfer the multi-scale features to the frequency domain through Fourier transform for feature fusion to obtain the fused multi-scale features;

[0017] Obtain the final result of multi-scale feature fusion by performing inverse Fourier transform on the fused multi-scale features.

[0018] Further, the fused multi-scale features are:

[0019] F fused = A · V

[0020]

[0021] Q = W q E ∈

[0022] K = W k E c

[0023] V = W v E c

[0024] where F fused is the fused multi-scale feature, A is the attention weight, softm is the activation function, Q is the feature after noise embedding transformation, W q is the weight matrix, E ∈ is the diffusion noise embedding, K is the key feature after conditional semantic feature transformation, W k is the weight matrix, E c is the conditional semantic embedding, V is the value feature after conditional semantic feature transformation, W v is the weight matrix, d k is the feature dimension.

[0025] Further, the dual attention module includes a channel attention unit and a spatial attention unit;

[0026] The channel attention unit generates a channel attention map by calculating average pooling and max pooling in the channel dimension. The channel attention map is:

[0027] F c = A c ·F

[0028] where F c is the channel attention map, A c is the channel attention weight, and F is the multi-scale feature of the dermoscopic image of melanoma;

[0029] The spatial attention unit generates a spatial attention map by calculating average pooling and max pooling in the spatial dimension. The spatial attention map is:

[0030] F s = A s ·F

[0031] where F s is the spatial attention map, A s is the spatial attention weight, and F is the multi-scale feature of the dermoscopic image of melanoma.

[0032] Further, the global average pooling result and the global max pooling result of the dermoscopic image of melanoma are non-linearly transformed through a multi-layer perceptron to obtain the channel attention weight:

[0033] A c = σ(MLP(GAP(F)) + MLP(GMP(F)))

[0034]

[0035] where A c is the channel attention weight, σ is the Sigmoid activation function, GAP(F) is the feature map obtained by global average pooling, GMP(F) is the feature map obtained by global max pooling, H is the image height, W is the image width, F(i, j, c) is the value of a specific point in the feature map, c is the channel index, i is the height index, and j is the width index;

[0036] The spatial attention weight is generated by concatenating the average pooling result and the max pooling result of the multi-scale feature of the dermoscopic image of melanoma along the channel dimension and passing through a convolutional layer:

[0037] A s= σ(Conv(Concat(Avgpool(F),Maxpool(F))))

[0038]

[0039] Wherein, A s is the spatial attention weight, Conv is the convolution operation, Concat is the depth linking operation, Avgpool(F) is the average pooling result of the multi-scale features of the dermoscopic image of melanoma, and Maxpool(F) is the maximum pooling result of the multi-scale features of the dermoscopic image of melanoma.

[0040] Furthermore, a distance map is used as an additional supervision signal to train the DDPM segmentation model, guiding the DDPM segmentation model to focus on the segmentation accuracy of the boundary region.

[0041] Furthermore, the loss function of the DDPM segmentation model is:

[0042] L total = L seg + L diffusion + L attention + L distance

[0043] Wherein, L total is the loss function of the DDPM segmentation model, L seg is the segmentation loss function, L diffusion is the diffusion loss function, L attention is the attention loss function, L distance is the boundary loss function;

[0044] The segmentation loss function is:

[0045]

[0046] Wherein, L seg is the segmentation loss function, M(i,j) is the pixel value at a specific position of the predicted segmentation mask, is the pixel value at a specific position of the true segmentation mask, and (i,j) is the specific position coordinate;

[0047] The diffusion loss function is:

[0048]

[0049] Wherein, L diffusion is the diffusion loss function, E is the mathematical expectation operator, ∈ is the true Gaussian noise added to the original data at time step t, is the predicted noise of the DDPM segmentation model, x tis the noisy data at time step t, where t is the time index of the diffusion process;

[0050] The attention loss function is as follows:

[0051]

[0052] In the formula, L attention is the attention loss function, C is the number of channels of the feature map, is the channel attention weight predicted by the DDPM segmentation model, c is the channel index, is the true channel attention weight, H is the image height, W is the image width, is the spatial attention weight predicted by the DDPM segmentation model, is the true spatial attention weight;

[0053] The boundary loss function is as follows:

[0054]

[0055] In the formula, L distance is the boundary loss function, is the distance transform map, is the segmentation mask predicted by the DDPM segmentation model, and M(i, j) is the true segmentation mask.

[0056] According to another aspect of the present invention, there is provided a melanoma intelligent segmentation system based on DDPM, including:

[0057] An image preprocessing module for acquiring a dermoscopic image of melanoma and performing image preprocessing;

[0058] An image segmentation module for inputting the dermoscopic image of melanoma after image preprocessing into the DDPM segmentation model, gradually adding noise through the forward denoising process and gradually removing noise through the backward denoising process, thereby learning the probability distribution of the dermoscopic image of melanoma and obtaining the segmentation result of melanoma; the DDPM segmentation model includes an encoder and a decoder, the encoder includes a multi-scale feature fusion module and a dual attention module, the encoder extracts multi-scale features of the dermoscopic image of melanoma through a multi-scale convolutional layer, and the decoder gradually restores the segmentation result of the dermoscopic image of melanoma according to the multi-scale features of the dermoscopic image of melanoma through upsampling and feature fusion operations.

[0059] Compared with the prior art, the present invention has the following beneficial effects:

[0060] 1. The dermoscopic images of melanoma after image preprocessing are input into the DDPM segmentation model in the present invention. By introducing a multi-scale feature fusion module and a dual attention module into the encoder of the DDPM segmentation model, noise is gradually added and removed through the multi-scale feature fusion module, thereby learning the probability distribution of the dermoscopic images of melanoma and generating high-quality segmentation results. The dual attention module enhances the features of important channels and spatial positions, improving the segmentation accuracy of melanoma.

[0061] 2. The present invention adopts a training strategy based on the distance map. By optimizing the boundary segmentation, the influence of the blurred boundary of the melanoma region on the segmentation accuracy is reduced. The distance map is used as an additional supervision signal to train the DDPM segmentation model, guiding the DDPM segmentation model to focus on the segmentation accuracy of the boundary region and improving the accuracy of the segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a schematic flowchart of a melanoma intelligent segmentation method based on DDPM proposed by the present invention;

[0063] Figure 2 is a schematic structural diagram of the DDPM segmentation model;

[0064] Figure 3 is a schematic diagram of the multi-scale feature fusion module;

[0065] Figure 4 is a flowchart of the dual attention module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manner and specific operation process are given, but the protection scope of the present invention is not limited to the following embodiments.

[0067] English abbreviations involved:

[0068] Denoising Diffusion Probabilistic Models: DDPM

[0069] Multilayer Perceptron: MLP

[0070] Global Average Pooling: GAP

[0071] Global Max Pooling: GMP

[0072] Embodiment 1

[0073] This embodiment provides a melanoma intelligent segmentation method based on DDPM, as Figure 1 shown, which includes the following steps:

[0074] S1. Obtain the dermoscopic image of melanoma and perform image preprocessing.

[0075] To enhance the performance of the DDPM segmentation model, first perform image preprocessing on the input dermoscopic image. Image preprocessing includes image normalization, data augmentation, and noise removal. Image normalization scales the pixel values of the image to a unified range to reduce the impact of illumination changes on the segmentation result. Data augmentation increases the diversity of training data through operations such as random rotation, scaling, and flipping, and improves the generalization ability of the model. Noise removal uses a frequency domain transformation method to eliminate high-frequency noise in the image, enabling better fusion of conditional features and noise features.

[0076] S2. Input the dermoscopic image of melanoma after image preprocessing into the DDPM segmentation model. Gradually add noise through the forward denoising process and gradually remove noise through the backward denoising process, thereby learning the probability distribution of the dermoscopic image of melanoma and obtaining the segmentation result of melanoma.

[0077] The DDPM segmentation model includes an encoder and a decoder. The encoder includes a multi-scale feature fusion module and a dual attention module. The encoder extracts multi-scale features of the dermoscopic image of melanoma through multi-scale convolutional layers. The decoder gradually restores the segmentation result of the dermoscopic image of melanoma according to the multi-scale features of the dermoscopic image of melanoma through upsampling and feature fusion operations. The multi-scale feature fusion module fuses the diffusion noise embedding and the conditional semantic embedding through a cross-attention mechanism, encodes the diffusion noise embedding into the conditional semantic embedding, and encodes the conditional semantic embedding into the diffusion noise embedding. The diffusion noise embedding is generated through multi-scale convolutional layers and is gradually added to the dermoscopic image of melanoma. The conditional semantic embedding is extracted from the dermoscopic image of melanoma through multi-scale convolutional layers.

[0078] Encoder EM i has a structure as Figure 2 shown. Use two paths to respectively input the noisy segmentation mask at the current time step T and the original dermoscopic image g o-1 of melanoma, and embed the time step T. The noisy mask image is sent into the multi-scale feature fusion module FF-Parser through two residual blocks RB1 and RB2. The original dermoscopic image g i-1After passing through a residual block RB1, it is concatenated with the masked image of the same-layer residual output and fed into the dual attention module CAB-SAB to extract features and output through another residual block RB2. This output is then fed into the multi-scale feature fusion module FF-Parser for feature extraction and fusion. The fused feature images are each passed through the linear attention mechanism L-Att and sent to the next encoder for repeated operations. Each encoder has the same structure, responsible for extracting multi-scale features, gradually compressing the spatial resolution, and increasing the number of channels at the same time. The structure of the multi-scale feature fusion module FF-Parser is as Figure 3 shown. The principle of the multi-scale feature fusion module is based on the cross-attention mechanism. This cross-attention mechanism constructs a frequency-domain-spatial-domain joint modeling framework by combining frequency-domain feature extraction of the fast Fourier transform (FFT), dynamic temporal control of the temporal embedding vector, and cross-scale interaction of multi-level features (c0-c4). Its core uses the FFT to perform global correlation analysis on the query (q), key (k), and noise embedding (e), generates attention weights through the maximum inner product, and introduces cropping and rotation operations to enhance the robustness of local features. After the temporal embedding is fused with the spatial feature W, the noise prediction strategy at different diffusion stages is dynamically adjusted, and at the same time, multi-scale encoder features are integrated to balance high-frequency details and low-frequency semantic information. Finally, the optimized feature map output by the fully connected layer MLP significantly improves the boundary accuracy of melanoma image segmentation and effectively suppresses dynamic noise interference.

[0079] Decoder DM i has a structure as Figure 2 shown. The input of the decoder is the downsampled output after passing through the BM layer. BM is composed of residual blocks, and its main function is to match the number of channels to ensure that the channels and sizes of upsampling and downsampling are consistent. Decoder DM i also receives the feature extraction output sent from the downsampling layer at the same scale. The output after passing through BM is first concatenated with the feature output of multi-scale fusion and fed into the first residual block then concatenated with the feature extraction output of the dual attention module and fed into the second residual block and then output through the linear attention mechanism L-Att. The residual block contains an interpolation operation to help the feature map gradually recover its size, and then this operation is repeated. The structure of each decoder is also the same. Finally, a noise map of the same size as the original image is output through the activation function and the 1×1 convolutional layer. Subsequently, this noise map ∈ θ is fed into the denoising iteration step for denoising to generate a mask.

[0080] The specific steps for the multi-scale feature fusion module to perform multi-scale feature fusion are as follows:

[0081] The encoder is used to extract multi-scale features of the dermoscopic image of melanoma through multi-scale convolutional layers. The multi-scale features are:

[0082] F = {f1, f2,..., f n}

[0083] where F is the multi-scale feature of the dermoscopic image of melanoma, and f n is the nth scale feature.

[0084] The multi-scale features are transformed to the frequency domain through Fourier transform for feature fusion to obtain the fused multi-scale features, and the fused multi-scale features are:

[0085] F fused = A · V

[0086]

[0087] Q = W q E ∈

[0088] K = W k E c

[0089] V = W v E c

[0090] where F fused is the fused multi-scale feature, A is the attention weight, softm is the activation function, Q is the feature after noise embedding transformation, W q is the weight matrix, E ∈ is the diffusion noise embedding, K is the key feature after conditional semantic feature transformation, W k is the weight matrix, E c is the conditional semantic embedding, V is the value feature after conditional semantic feature transformation, W v is the weight matrix, and d k is the feature dimension.

[0091] The final result of multi-scale feature fusion is obtained by inverse Fourier transform of the fused multi-scale features:

[0092] F' fused = F -1 (F fused )

[0093] where F' fused is the final result of multi-scale feature fusion, F -1 is the inverse Fourier transform, and F fused is the fused multi-scale feature.

[0094] Through frequency domain feature fusion, the DDPM segmentation model can effectively eliminate high-frequency noise, fuse global and local features at the same time, and further improve the segmentation accuracy.

[0095] Improve the feature alignment problem in different feature spaces through a dual attention module, which includes a channel attention unit and a spatial attention unit.

[0096] As Figure 4 shown, the channel attention unit generates a channel attention map by calculating average pooling and max pooling in the channel dimension. The channel attention map is:

[0097] F c = A c · F

[0098] In the formula, F c is the channel attention map, F is the multi-scale feature of the dermoscopic image of melanoma, and A c is the channel attention weight. The channel attention weight is obtained by performing a non-linear transformation on the global average pooling result and the global max pooling result of the dermoscopic image of melanoma through a multi-layer perceptron:

[0099] A c = σ(MLP(GAP(F)) + MLP(GMP(F)))

[0100]

[0101] In the formula, A c is the channel attention weight, σ is the Sigmoid activation function, GAP(F) is the feature map obtained by global average pooling, GMP(F) is the feature map obtained by global max pooling, H is the image height, W is the image width, F(i, j, c) is the value of a specific point in the feature map, c is the channel index, i is the height index, and j is the width index.

[0102] As Figure 4 shown, the spatial attention unit generates a spatial attention map by calculating average pooling and max pooling in the spatial dimension. The spatial attention map is:

[0103] F s = A s · F

[0104] In the formula, F s is the spatial attention map, F is the multi-scale feature of the dermoscopic image of melanoma, and A s is the spatial attention weight. The spatial attention weight is generated by concatenating the average pooling result and the max pooling result of the multi-scale feature of the dermoscopic image of melanoma along the channel dimension and passing it through a convolutional layer:

[0105] A s = σ(Conv(Concat(Avgpool(F), Maxpool(F))))

[0106]

[0107] Wherein, A s is the spatial attention weight, Conv is the convolution operation, Concat is the depth linking operation, Avgpool(F) is the average pooling result of the multi-scale features of the dermoscopic image of melanoma, and Maxpool(F) is the maximum pooling result of the multi-scale features of the dermoscopic image of melanoma.

[0108] The dual attention module extracts rough semantic features from the original image path through Gaussian spatial attention and channel attention and integrates them into the diffusion model. After obtaining the channel attention map and the spatial attention map through feature extraction, they are weighted and summed to obtain the final dual attention enhanced feature:

[0109] F dual = αF C + βF S

[0110] Wherein, F dual is the dual attention enhanced feature, and α and β are learnable parameters. Finally, the dual attention enhanced feature is integrated into the DDPM segmentation model to provide the correct noise prediction range for the segmentation model and enhance the effect of feature alignment.

[0111] To reduce the influence of the blurred regional boundary of melanoma on the segmentation accuracy, a distance map is used as an additional supervision signal to train the DDPM segmentation model, guiding the DDPM segmentation model to focus on the segmentation accuracy of the boundary region. The distance map is generated by searching for the distance between the foreground points in the image and their closest background points through the Euclidean distance transform algorithm. The distance map is used to define the training task, and the greater the added noise, the weaker the guiding effect on the boundary. During the training process, the distance map is used as an additional supervision signal to guide the model to focus on the segmentation accuracy of the boundary region. In this way, the model can better handle the problem of boundary blur and improve the accuracy of the segmentation result. The generation of the distance map requires the prior definition of the foreground and the background. Let the binary mask of the input image be M, where the foreground pixels (melanoma region) are M(i,j)=1 and the background pixels are M(i,j)=0. For each foreground pixel point (i,j), calculate its Euclidean distance to the closest background pixel point. The Euclidean distance is:

[0112]

[0113] Wherein, (k,l) is the position of the background pixel point.

[0114] Then, the distances of all foreground pixel points are combined into a distance map D, and each pixel value in the distance map D represents the distance of that pixel point to the closest background pixel point.

[0115] The DDPM segmentation model is trained using the distance map as an additional supervision signal. First, the distance map is normalized to obtain the distance transformation map. During the training process, noise is added according to the normalized distance transformation map. The greater the added noise, the weaker the guiding effect on the boundary. Specifically, the intensity σ of noise addition is inversely proportional to the normalized distance transformation map:

[0116]

[0117] where α max is the maximum noise intensity, is the distance transformation map.

[0118] The loss function of the DDPM segmentation model is:

[0119] L total = L seg + L diffusion + L attention + L distance

[0120] where L total is the loss function of the DDPM segmentation model, L seg is the segmentation loss function, L diffusion is the diffusion loss function, L attention is the attention loss function, L distance is the boundary loss function;

[0121] The segmentation loss function is:

[0122]

[0123] where L seg is the segmentation loss function, M(i,j) is the pixel value at a specific position of the predicted segmentation mask, is the pixel value at a specific position of the true segmentation mask, and (i,j) is the coordinate of the specific position;

[0124] The diffusion loss function is:

[0125]

[0126] where L diffusion is the diffusion loss function, E is the mathematical expectation operator, ∈ is the true Gaussian noise added to the original data at time step t, is the predicted noise of the DDPM segmentation model, x t is the noise-added data at time step t, and t is the time index of the diffusion process;

[0127] The attention loss function is:

[0128]

[0129] In the formula, L attention is the attention loss function, C is the number of channels of the feature map, is the channel attention weight predicted by the DDPM segmentation model, c is the channel index, is the true channel attention weight, H is the image height, and W is the image width, is the spatial attention weight predicted by the DDPM segmentation model, is the true spatial attention weight;

[0130] The boundary loss function is:

[0131]

[0132] In the formula, L distance is the boundary loss function, is the distance transformation map, is the segmentation mask predicted by the DDPM segmentation model, and M(i, j) is the true segmentation mask.

[0133] To better quantify and highlight the performance of the model of the present invention, comparative experiments were carried out on two important data sets (ISIC2018 and HAM10000) in this embodiment, and the experimental results are shown in the following table.

[0134] Table 1 Experimental data on ISIC2018 and HAM10000 data sets

[0135] Dataset Dice coefficient IoU Accuracy ISIC2018 89.5 81.2 92.3% HAM10000 88.7 80.5 91.8%

[0136] As can be seen from Table 1, the melanoma intelligent segmentation method based on DDPM proposed in this embodiment shows remarkable results on the ISIC2018 data set, specifically as follows: the Dice coefficient is 89.5%, the IoU is 81.2%, and the accuracy rate is 92.3%. These indicators show that the melanoma intelligent segmentation method based on DDPM proposed in this embodiment performs excellently in the melanoma segmentation task and demonstrates balanced and efficient performance on multiple key indicators. This is mainly attributed to the fact that this method improves the segmentation accuracy by improving the interaction of multi-scale features and optimizing boundary segmentation.

[0137] The experimental results of the DDPM segmentation model proposed by this method on the ISIC2018 data set are shown in the following table.

[0138] Table 2 Experimental data of different modules on the ISIC2018 data set

[0139] Model name Dice coefficient IoU Accuracy Basic segmentation model 85.3 75.6 89.2% Basic segmentation model + DDPM framework 87.8 78.9 90.7% Basic segmentation model + Dual attention module 86.5 77.2 90.1% Basic model + Distance map training 88.1 79.5 91.2% Complete model (DDPM + Dual attention + Distance map) 89.5 81.2 92.3%

[0140] The experimental results show that the present invention is significantly superior to the existing methods in the melanoma segmentation task on the ISIC2018 dataset, demonstrating its effectiveness.

[0141] Example 2

[0142] This embodiment provides a melanoma intelligent segmentation system based on DDPM, including:

[0143] An image preprocessing module, configured to obtain a dermoscopic image of melanoma and perform image preprocessing;

[0144] An image segmentation module, configured to input the dermoscopic image of melanoma after image preprocessing into the DDPM segmentation model, gradually add noise through the forward denoising process, and gradually remove noise through the backward denoising process, thereby learning the probability distribution of the dermoscopic image of melanoma to obtain the segmentation result of melanoma; the DDPM segmentation model includes an encoder and a decoder, the encoder includes a multi-scale feature fusion module and a dual attention module, the encoder extracts multi-scale features of the dermoscopic image of melanoma through multi-scale convolutional layers, and the decoder gradually restores the segmentation result of the dermoscopic image of melanoma according to the multi-scale features of the dermoscopic image of melanoma through upsampling and feature fusion operations.

[0145] The rest is the same as in Example 1.

[0146] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. An intelligent melanoma segmentation method based on DDPM, characterized in that, It includes the following steps: Obtain a dermoscopic image of melanoma and perform image preprocessing; Input the dermoscopic image of melanoma after image preprocessing into the DDPM segmentation model, gradually add noise through the forward denoising process, and gradually remove noise through the backward denoising process, thereby learning the probability distribution of the dermoscopic image of melanoma and obtaining the segmentation result of melanoma; The DDPM segmentation model includes an encoder and a decoder. The encoder includes a multi-scale feature fusion module and a dual attention module. The encoder extracts multi-scale features of the dermoscopic image of melanoma through multi-scale convolutional layers, and the decoder gradually restores the segmentation result of the dermoscopic image of melanoma according to the multi-scale features of the dermoscopic image of melanoma through upsampling and feature fusion operations.

2. The melanoma intelligent segmentation method based on DDPM according to claim 1, wherein The image preprocessing includes image normalization, data augmentation, and noise removal.

3. The melanoma intelligent segmentation method based on DDPM according to claim 1, wherein, The multi-scale feature fusion module fuses the diffusion noise embedding and the conditional semantic embedding through a cross-attention mechanism, encodes the diffusion noise embedding into the conditional semantic embedding, and encodes the conditional semantic embedding into the diffusion noise embedding. The diffusion noise embedding is generated through the multi-scale convolutional layer and is gradually added to the dermoscopic image of melanoma. The conditional semantic embedding is extracted from the dermoscopic image of melanoma through the multi-scale convolutional layer.

4. The melanoma intelligent segmentation method based on DDPM according to claim 1, wherein, The specific steps for the multi-scale feature fusion module to perform multi-scale feature fusion are as follows: Use the encoder to extract multi-scale features of the dermoscopic image of melanoma through multi-scale convolutional layers; Transfer the multi-scale features to the frequency domain through Fourier transform for feature fusion to obtain the fused multi-scale features; Obtain the final result of multi-scale feature fusion through inverse Fourier transform of the fused multi-scale features.

5. The melanoma intelligent segmentation method based on DDPM according to claim 4, wherein, The fused multi-scale features are: F fused = A·V Q = W q E ∈ K = W k E c V = W v E c Where F fused is the fused multi-scale feature, A is the attention weight, softm is the activation function, Q is the feature after noise embedding transformation, W q is the weight matrix, E ∈ is the diffusion noise embedding, K is the key feature after conditional semantic feature transformation, W k is the weight matrix, E c is the conditional semantic embedding, V is the value feature after conditional semantic feature transformation, W v is the weight matrix, d k is the feature dimension.

6. The melanoma intelligent segmentation method based on DDPM according to claim 1, characterized in that The dual attention module includes a channel attention unit and a spatial attention unit; The channel attention unit generates a channel attention map by calculating average pooling and max pooling in the channel dimension. The channel attention map is: F c = A c ·F where F c is the channel attention map, A c is the channel attention weight, and F is the multi-scale feature of the dermoscopic image of melanoma; The spatial attention unit generates a spatial attention map by calculating average pooling and max pooling in the spatial dimension. The spatial attention map is: F s = A s ·F In the formula, F s is the spatial attention map, A s is the spatial attention weight, and F is the multi-scale feature of the dermoscopic image of melanoma.

7. The melanoma intelligent segmentation method based on DDPM according to claim 6, characterized in that, Perform a non-linear transformation on the global average pooling result and the global max pooling result of the dermoscopic image of melanoma through a multi-layer perceptron to obtain the channel attention weight: A c = σ(MLP(GAP(F)) + MLP(GMP(F))) Where, A c is the channel attention weight, σ is the Sigmoid activation function, GAP(F) is the feature map obtained by global average pooling, GMP(F) is the feature map obtained by global maximum pooling, H is the image height, W is the image width, F(i, j, c) is the value of a specific point in the feature map, c is the channel index, i is the height index, and j is the width index; Generate the spatial attention weight by concatenating the average pooling result and the max pooling result of the multi-scale features of the dermoscopic image of melanoma along the channel dimension and passing through a convolutional layer: A s = σ(Conv(Concat(Avgpool(F),Maxpool(F)))) Where, A s is the spatial attention weight, Conv is the convolution operation, Concat is the depth linking operation, Avgpool(F) is the average pooling result of the multi-scale features of the dermoscopic image of melanoma, and Maxpool(F) is the maximum pooling result of the multi-scale features of the dermoscopic image of melanoma.

8. The melanoma intelligent segmentation method based on DDPM according to claim 1, wherein Use the distance map as an additional supervision signal to train the DDPM segmentation model to guide the DDPM segmentation model to focus on the segmentation accuracy of the boundary region.

9. The melanoma intelligent segmentation method based on DDPM according to claim 1, wherein The loss function of the DDPM segmentation model is: L total = L seg + L diffusion + L attention + L distance where L total is the loss function of the DDPM segmentation model, L seg is the segmentation loss function, L diffusion is the diffusion loss function, L attention is the attention loss function, L distance is the boundary loss function; The segmentation loss function is: Where L seg is the segmentation loss function, M(i,j) is the pixel value at a specific position of the predicted segmentation mask, is the pixel value at a specific position of the ground truth segmentation mask, and (i,j) is the coordinate of the specific position; The diffusion loss function is: Where, L diffusion is the diffusion loss function, E is the mathematical expectation operator, ∈ is the true Gaussian noise added to the original data at time step t, is the predicted noise of the DDPM segmentation model, x t is the noise-added data at time step t, and t is the time index of the diffusion process; The attention loss function is: Where L attention is the attention loss function, C is the number of channels of the feature map, is the channel attention weight predicted by the DDPM segmentation model, c is the channel index, is the true channel attention weight, H is the image height, W is the image width, is the spatial attention weight predicted by the DDPM segmentation model, is the true spatial attention weight; The boundary loss function is: Where L distance is the boundary loss function, is the distance transform map, is the segmentation mask predicted by the DDPM segmentation model, and M(i,j) is the ground truth segmentation mask.

10. An intelligent melanoma segmentation system based on DDPM, characterized in that, It includes: An image preprocessing module for obtaining a dermoscopic image of melanoma and performing image preprocessing; An image segmentation module, which is used to input the dermoscopic image of melanoma after image preprocessing into the DDPM segmentation model, gradually add noise through the forward denoising process, and gradually remove noise through the backward denoising process, so as to learn the probability distribution of the dermoscopic image of melanoma and obtain the segmentation result of melanoma; the DDPM segmentation model includes an encoder and a decoder, the encoder includes a multi-scale feature fusion module and a dual attention module, the encoder extracts multi-scale features of the dermoscopic image of melanoma through a multi-scale convolutional layer, and the decoder gradually restores the segmentation result of the dermoscopic image of melanoma according to the multi-scale features of the dermoscopic image of melanoma through upsampling and feature fusion operations.

Citation Information

Patent Citations

  • Skin cancer image segmentation method and system based on attention and multi-feature fusion

    CN115689993A

Cited By

  • Photoacoustic image segmentation method and device, equipment and storage medium

    CN122453856A