Medical image segmentation method and system based on tree structure denoising diffusion probability model
Through the medical image segmentation method based on the tree structure denoising diffusion probability model, the information interaction of binary tree structure and prompt image is solved, and the problems of low segmentation accuracy and noise sensitivity in the existing methods are achieved, achieving higher segmentation accuracy and robustness.
Patent Information
- Application Number
- CN202510332704.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When existing medical image segmentation methods process large-area or unevenly distributed medical images, there are problems such as low segmentation accuracy, noise sensitivity and lack of comprehensive understanding of the global and local features of the image.
The tree structure denoising diffusion probability model is adopted, and the encoder and decoder are constructed through the binary tree structure, and the prompt image of the same lesions is used for information interaction and feature fusion, and the model parameters are optimized by combining the forward diffusion process and the cross entropy loss function to achieve multi-scale information capture and noise suppression.
It improves the accuracy and robustness of medical image segmentation, enhances the model's processing ability of noise and outliers, and improves the generalization performance and segmentation efficiency of the model.
Smart Images

Figure CN120298424A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a medical image segmentation method and system, and in particular to a medical image segmentation method and system based on a tree-structured denoising diffusion probability model, belonging to the technical field of medical image segmentation. Background Art
[0002] Medical image segmentation is one of the key tasks in medical image analysis, which involves separating the region of interest (ROI) in a medical image from the background or other tissues. This task is crucial for disease diagnosis, treatment, and efficacy evaluation. However, medical image segmentation faces many challenges, such as image noise, tissue complexity, and individual differences. Traditional methods such as manual segmentation, threshold segmentation, and region growing can achieve the segmentation goal to a certain extent, but they still have disadvantages such as low segmentation accuracy and sensitivity to noise.
[0003] With the rapid development of computer vision and deep learning technologies, deep learning-based medical image segmentation methods have gradually emerged. These methods use convolutional neural networks (CNNs) or Transformer technologies to achieve high-precision segmentation by automatically learning image features. However, the CNN-based segmentation model cannot fully consider the relationships and dependencies between different regions in a medical image in the medical image segmentation task, lacking a comprehensive insight into the entire image, and facing serious challenges when segmenting large-area or unevenly distributed medical images; the Transformer-based model overemphasizes global features and lacks the ability to recognize local details, resulting in a decline in the segmentation performance of the model for fuzzy boundaries.
[0004] In recent years, diffusion models have received extensive attention in the image segmentation task. It has excellent feature learning ability in the reverse denoising process and has good segmentation effects, but its generalization performance is poor, and further research is needed to improve the segmentation accuracy and robustness. Summary of the Invention
[0005] Object of the Invention: The object of the present invention is to provide a medical image segmentation method and system based on a tree-structured denoising diffusion probability model that can improve the segmentation performance of the model.
[0006] Technical Solution: A medical image segmentation method based on a tree-structured denoising diffusion probability model according to the present invention includes:
[0007] Step 1: Obtain a medical image dataset containing segmentation masks, and randomly select three segmentation masks belonging to the same lesion from it as prompt images;
[0008] Step 2: Preprocess the medical image dataset, divide the preprocessed dataset into a training set and a test set according to a preset ratio, construct a denoising diffusion probabilistic model, and initialize the model parameters;
[0009] Step 3: Based on the training set, for a segmentation mask X0 in the training set, add Gaussian noise to X0 through the forward diffusion process to corrupt it, generating a series of noisy images X t , t = 1, 2, …, T, X t represents the noisy image at the t-th time step;
[0010] Step 4: Construct an encoder and a decoder based on a binary tree structure. Receive the medical image I corresponding to X0 and the noisy image X t , as well as the prompt image, enable information interaction between the medical image I and the prompt image and X t , output a feature representation. Receive the feature representation output by the encoder through the decoder, randomly sample the feature representation to generate different segmentation masks, and perform feature fusion output;
[0011] Step 5: Repeat Step 4 until the predicted segmentation mask map X0' of X0 is output, calculate the loss between X0' and X0, and update the model parameters;
[0012] Step 6: Use the test set to test the model with updated parameters.
[0013] Furthermore, in Step 2, the preprocessing includes normalization and data augmentation. The denoising diffusion probabilistic model uses UNet as the model architecture. The initialization parameters specifically include:
[0014] Initialize the noise schedule parameter betas to control the amount of noise added at each step;
[0015] Assign initial weights to each layer of the model, which is completed through random initialization;
[0016] Define the training parameters, and the training parameters include: learning rate, batch size, and number of training epochs;
[0017] Configure the optimizer and set the learning rate and momentum.
[0018] Furthermore, the forward diffusion process in Step 3 is as follows:
[0019]
[0020] where, β t is a hyperparameter representing the amplitude of the noise added at time step t, N represents the normal distribution, and E is the identity matrix. Simplify the above formula to:
[0021]
[0022] Among them, ε t is the noise sampled from the standard normal distribution and is used to be added to the data at each step. The value of β t increases as the time step t increases;
[0023] The noise distribution at any time step t is as follows:
[0024]
[0025] α t = 1 - β t ,
[0026]
[0027] Among them, α t is the decay coefficient, represents the product of all α t during the T-step diffusion process, and ε is the noise sampled from the standard Gaussian distribution.
[0028] Furthermore, the specific process of the encoder described in step 4 includes:
[0029] Taking X t , I, and the prompt images S1, S2, and S3 as leaf nodes, input X t and I into the first-layer network, fuse them in the dimension of channels, and then perform a convolution operation and input it into the second-layer network;
[0030] Select S1 as a leaf node and input it into the second-layer network, fuse it with the output of the first-layer network and then perform convolution, and input the result of the convolution into the third-layer network;
[0031] Select S2 as a leaf node and input it into the third-layer network, fuse it with the output of the second-layer network and then perform convolution, and input the result of the convolution into the fourth-layer network;
[0032] Select S3 as a leaf node and input it into the fourth-layer network, fuse it with the output of the third-layer network and then perform convolution, and output the encoded feature representation.
[0033] Furthermore, the specific process of the decoder described in step 4 includes:
[0034] Randomly sample the feature representation generated by the encoder twice to generate two branch features, randomly sample each of the two branch features twice to generate four segmentation masks, and merge the segmentation masks into one output through accumulation and averaging.
[0035] Furthermore, the said step 5 includes:
[0036] For X t Repeat step 4 for T times, output the predicted segmentation mask map X0', use the cross - entropy loss function and Dice loss function to evaluate the difference between the segmentation result and the ground truth label, and optimize the segmentation network composed of the encoder and decoder, which is expressed as:
[0037]
[0038] Furthermore, the said step 6 includes:
[0039] Step 6.1: For a medical image Y0 in the test set, add Gaussian noise to the medical image Y0 to generate a noisy image Y t , t = 1, 2, …, T, Y represents the noisy image at the t - th time step;
[0040] Step 6.2: Input Y t and Y0 into the encoder at the same time. First, fuse and then perform convolution operations. Input the prompt images S1, S2, and S3 into the encoder in sequence to guide the segmentation. The decoder generates multiple segmentation masks through random sampling, and combines the multiple segmentation masks into one output through accumulation and averaging;
[0041] Step 6.3: After T - step iteration of step 6.2, obtain the predicted segmentation mask map Y0′ of Y0;
[0042] Step 6.4: Traverse the test set and use the evaluation metrics Dice coefficient and IoU index to analyze the performance of the network model.
[0043] Furthermore, in the said step 6.4, use the Dice coefficient to measure the performance of the model by calculating the overlap degree between the predicted segmentation region and the ground truth segmentation region, ranging from 0 to 1. The higher the value, the better the segmentation effect. Its formula is expressed as:
[0044]
[0045] where X and Y represent the pixel sets of the predicted segmentation region and the ground truth segmentation region respectively, |X∩Y| represents the number of pixels in their intersection, and |X| and |Y| represent the number of pixels in their respective sets;
[0046] Use the IoU index to measure the ratio of the intersection to the union of the predicted segmentation region and the ground truth segmentation region, ranging from 0 to 1. The higher the value, the more accurate the segmentation result. Its formula is expressed as:
[0047]
[0048] Based on the same inventive concept, the present invention also provides a medical image segmentation system based on a tree - structured denoising diffusion probability model, including:
[0049] A data acquisition module, configured to obtain a medical image dataset containing segmentation masks, and randomly select three segmentation masks belonging to the same lesion from the dataset as prompt images;
[0050] An initialization module, configured to preprocess the medical image dataset, divide the preprocessed dataset into a training set and a test set according to a preset ratio, construct a denoising diffusion probabilistic model and initialize the model parameters;
[0051] A noise addition module, configured to, based on the training set, for a segmentation mask X0 in the training set, add Gaussian noise to X0 through a forward diffusion process to corrupt it, and generate a series of noisy images X t , where t = 1, 2, …, T, and X t represents the noisy image at the t-th time step;
[0052] A fusion and output module, configured to construct an encoder and a decoder based on a binary tree structure, receive the medical image I corresponding to X0 and the noisy image X t , as well as the prompt images, to enable information interaction between the medical image I and the prompt images and X t , output a feature representation, receive the feature representation output by the encoder through the decoder, randomly sample the feature representation to generate different segmentation masks, and perform feature fusion and output;
[0053] An update module, configured to repeat the fusion and output module until a predicted segmentation mask map X0' of X0 is output, calculate the loss between X0' and X0, and update the model parameters;
[0054] A test module, configured to test the model with updated parameters using the test set.
[0055] Based on the same inventive concept, the present invention also provides a computing device, including: one or more processors, one or more memories, and one or more programs, where the programs are stored in the memories and are configured to be executed by the processors, and when the programs are loaded into the processors, the steps of the medical image segmentation method based on a tree-structured denoising diffusion probabilistic model according to any one of the above are implemented.
[0056] Beneficial effects: Compared with the prior art, the tree structure of the present invention allows the model to fuse features at different levels, which helps to capture multi-scale information in medical images, thereby improving the accuracy of segmentation; each node in the encoding stage of the present invention can be regarded as a feature selection point, and the model can learn to select the most relevant features for segmentation at these points, thereby improving the segmentation quality; the present invention provides multiple denoising paths through branches, increasing the robustness of the model to noise and outliers, and improving the quality of the final segmentation result through accumulation and averaging; the present invention uses the hint images of the same lesion to assist denoising layer by layer, promotes the effective transmission and fusion of information, realizes the suppression of image noise, helps the model to better understand and process specific regions in the image. When segmenting new medical images, only the hint images of the same lesion need to be replaced, improving the generalization performance of the model; by constructing the denoising diffusion probability model into a tree structure, the present invention makes full use of the advantages of the tree, improving the performance and efficiency of medical image segmentation. Brief Description of the Drawings
[0057] Figure 1 It is a flowchart of the method according to an embodiment of the present invention;
[0058] Figure 2 It is a schematic diagram of the denoising diffusion probability model according to an embodiment of the present invention. Detailed Embodiments
[0059] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0060] As shown in the Figure 1 accompanying drawings, the medical image segmentation method based on a tree structure denoising diffusion probability model in this embodiment includes:
[0061] Step 1: Obtain a medical image dataset containing segmentation masks, and randomly select three segmentation masks belonging to the same lesion from it as hint images;
[0062] Step 2: Preprocess the medical image dataset, divide the preprocessed dataset into a training set and a test set according to a preset ratio, and construct a denoising diffusion probability model and initialize the model parameters;
[0063] Step 3: Based on the training set, for a segmentation mask X0 in the training set, add Gaussian noise to X0 through the forward diffusion process to corrupt it, generating a series of noisy images X t , where t = 1, 2, …, T, and X t represents the noisy image at the t-th time step;
[0064] Step 4: Build an encoder and a decoder based on a binary tree structure. Receive the medical image I corresponding to X0 and the noisy image X t , as well as the prompt image, through the encoder, enabling information interaction between the medical image I and the prompt image with X t , output a feature representation. Receive the feature representation output by the encoder through the decoder, randomly sample the feature representation to generate different segmentation masks, and perform feature fusion output;
[0065] Step 5: Repeat Step 4 until the predicted segmentation mask map X0' of X0 is output, calculate the loss between X0' and X0, and update the parameters of the model;
[0066] Step 6: Use the test set to test the model with updated parameters.
[0067] Specifically, in the embodiment of the present application, the building of the model and the initialization of parameters in Step 2 specifically include:
[0068] Select UNet as the model architecture. As Figure 2 shown, the initialized parameters include:
[0069] The noise scheduling parameter betas, used to control the amount of noise added at each step;
[0070] Assign initial weights to each layer of the model, completed through random initialization;
[0071] Define the parameters during the training process, including the learning rate, batch size, and number of training epochs;
[0072] Configure the optimizer, set the learning rate and momentum;
[0073] Step 2 also includes dataset preparation, loading and preprocessing the medical image dataset, including image normalization and data augmentation, and dividing the dataset into a training set and a test set.
[0074] Furthermore, in this embodiment, the forward diffusion process in Step 3 is as follows:
[0075]
[0076] where β tis a hyperparameter representing the magnitude of the noise added at time step t. N represents the normal distribution, and E is the identity matrix. This formula can be further simplified as:
[0077]
[0078] where ∈ t is the noise sampled from the standard normal distribution and is used to be added to the data at each step. The value of β t will increase as the time step t increases. This formula defines the amount of noise added at each time step. Through iterative derivation, the noise distribution at any time t can be obtained:
[0079]
[0080] α t = 1 - β t ,
[0081]
[0082] where α t is the decay coefficient related to β t , represents the product of all α t in the T-step diffusion process.
[0083] Furthermore, in the embodiments of the present application, the specific processes of the encoder and decoder in step 4 include:
[0084] Construct an encoder similar to a tree structure. The noisy image X t and the medical image I are input into the first-layer network and fused in the dimension of channels. Cross convolution is introduced to interact the feature maps of the noisy image with the feature maps of the medical image I, and then input into the second-layer network;
[0085] Select a prompt image S1 as a leaf node and input it into the second-layer network. First, fuse it with the output of the first-layer network in the dimension of channels, then perform cross convolution, and input it into the third-layer network;
[0086] The third layer and the fourth layer respectively input the prompt images S2 and S3. The operations are the same as above. Finally, the encoded feature representation is output. The prompt images and medical images in this process interact with the noisy image to guide the model to denoise and improve the final segmentation result;
[0087] The decoder randomly repeats the sampling of the feature representation generated by the encoder twice to generate two branch features, and then randomly repeats the sampling of the two branch features twice respectively to generate four segmentation masks, and combines them into one output through accumulation and averaging.
[0088] Further, in the embodiments of the present application, the specific operations in step 5 are as follows:
[0089] Noisy image X T Through the denoising operation in step 4 for T steps, the predicted segmentation mask map X0' is output. The cross-entropy loss function and the Dice loss function are used to evaluate the difference between the segmentation result and the ground truth label, and the segmentation network is optimized, including two parts. The first part is the cross-entropy loss between the ground truth label and the predicted result, and the second part is the Dice loss function between the ground truth label and the predicted result. The function is expressed as:
[0090]
[0091] Further, in the embodiments of the present application, the model test in step 6 uses the trained network for image segmentation. The specific operation steps are as follows:
[0092] Step 6-1, for a medical image in the test set, in the forward diffusion process, Gaussian noise is gradually added to the medical image to generate a noisy image;
[0093] Step 6-2, the noisy image and the medical image are simultaneously input into the encoder network. First, they are fused and then convolutional operations are performed. The prompt images S1, S2, and S3 are sequentially input into the encoder to guide the segmentation. The decoder generates multiple segmentation masks through random sampling and merges them into one output by accumulation and averaging.
[0094] Step 6-3, after T-step iteration of step 6-2, the predicted segmentation mask map is obtained;
[0095] Step 6-4, traverse the test data set, and use the evaluation metrics Dice coefficient and IoU index to analyze the performance of the network model.
[0096] Based on the same inventive concept, this embodiment also provides a medical image segmentation system based on a tree-structured denoising diffusion probability model, including:
[0097] An acquisition module, configured to obtain a medical image data set containing segmentation masks, and randomly select three segmentation masks belonging to the same lesion from it as prompt images;
[0098] An initialization module, configured to preprocess the medical image data set, divide the preprocessed data set into a training set and a test set according to a preset ratio, and construct a denoising diffusion probability model and initialize the model parameters;
[0099] A noise addition module, configured to, based on the training set, for a segmentation mask X0 in the training set, add Gaussian noise to X0 through the forward diffusion process to damage it and generate a series of noisy images X t , t = 1, 2,..., T, Xt Denote the noisy image at the \(t\)-th time step;
[0100] A fusion output module, which is used to construct an encoder and a decoder based on a binary tree structure, and receive the medical image \(I\) corresponding to \(X_0\) and the noisy image \(X\) through the encoder t , as well as a prompt image, so that the medical image \(I\) and the prompt image interact with \(X\) t for information interaction, output a feature representation, receive the feature representation output by the encoder through the decoder, randomly sample the feature representation to generate different segmentation masks, and perform feature fusion output;
[0101] An update module, which is used to repeat the fusion output module until the predicted segmentation mask map \(X_0'\) of \(X_0\) is output, calculate the loss between \(X_0'\) and \(X_0\), and update the parameters of the model;
[0102] A test module, which is used to test the model with updated parameters using a test set.
[0103] Based on the same inventive concept, this embodiment also provides a computing device, including: one or more processors, one or more memories, and one or more programs, where the programs are stored in the memory and configured to be executed by the processor, and when the programs are loaded into the processor, the steps of the medical image segmentation method based on the tree structure denoising diffusion probability model according to any one of the above are implemented.
[0104] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product, and can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code, and can include the processes of the embodiments of the above method. The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features, and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A medical image segmentation method based on a tree-structured denoising diffusion probabilistic model, characterized in that, including: Step 1: Obtain a medical image dataset containing segmentation masks, and randomly select three segmentation masks belonging to the same lesion from it as prompt images; Step 2: Preprocess the medical image dataset, divide the preprocessed dataset into a training set and a test set according to a preset ratio, construct a denoising diffusion probability model and initialize the model parameters; Step 3: Based on the training set, for a segmentation mask X0 in the training set, add Gaussian noise to X0 through the forward diffusion process to corrupt it, generating a series of noisy images X t , where t = 1, 2, …, T, and X t represents the noisy image at the t-th time step; Step 4: Construct an encoder and a decoder based on the binary tree structure. Receive the medical image I corresponding to X0 and the noise image X through the encoder t , as well as the prompt image, and enable information interaction between the medical image I and the prompt image with X t to output a feature representation. Receive the feature representation output by the encoder through the decoder, randomly sample the feature representation to generate different segmentation masks, and perform feature fusion for output; Step 5: Repeat Step 4 until the predicted segmentation mask map X0' of X0 is output, calculate the loss between X0' and X0, and update the model parameters; Step 6: Use the test set to test the model with updated parameters.
2. The medical image segmentation method based on the tree structure denoising diffusion probability model according to claim 1, wherein, In Step 2, the preprocessing includes normalization and data augmentation. The denoising diffusion probability model uses UNet as the model architecture. The initialization parameters specifically include: Initialize the noise scheduling parameter betas to control the amount of noise added at each step; Assign initial weights to each layer of the model, completed by random initialization; Define training parameters, which include: learning rate, batch size, and number of training epochs; Configure the optimizer and set the learning rate and momentum.
3. The medical image segmentation method based on the tree-structured denoising diffusion probability model according to claim 1, characterized in that The forward diffusion process in Step 3 is as follows: where β t is a hyperparameter representing the magnitude of the noise added at time step t, N represents the normal distribution, and E is the identity matrix. The above equation is simplified as follows: where ε t is the noise sampled from the standard normal distribution and is added to the data at each step, and the value of β t increases as the time step t increases; The noise distribution at any time step t is as follows: α t =1-β t , Among them, α t is the attenuation coefficient, represents the product of all α t in the T-step diffusion process, and ε is the noise sampled from the standard Gaussian distribution.
4. The medical image segmentation method based on the tree-structured denoising diffusion probability model according to claim 1, wherein The specific process of the encoder in Step 4 includes: With X t , I, and hint images S1, S2, and S3 as leaf nodes, input X t and I into the first-layer network, fuse them in terms of the channel dimension, and then perform a convolution operation and input it into the second-layer network; Select S1 as the leaf node and input it into the second-layer network, fuse it with the output of the first-layer network and then perform convolution, and input the result of the convolution into the third-layer network; Select S2 as the leaf node and input it into the third-layer network, fuse it with the output of the second-layer network and then perform convolution, and input the result of the convolution into the fourth-layer network; Select S3 as the leaf node and input it into the fourth-layer network, fuse it with the output of the third-layer network and then perform convolution, and output the encoded feature representation.
5. The medical image segmentation method based on the tree-structured denoising diffusion probability model according to claim 1, wherein, The specific process of the decoder in Step 4 includes: Randomly resample the feature representation generated by the encoder twice to generate two branch features, randomly resample each of the two branch features twice to generate four segmentation masks, and merge each segmentation mask into one output through accumulation and averaging.
6. The medical image segmentation method based on the tree structure denoising diffusion probability model according to claim 1, wherein Step 5 includes: For X t Repeat step 4 for T times, output the predicted segmentation mask graph X0', use the cross-entropy loss function and the Dice loss function to evaluate the difference between the segmentation result and the ground truth label, and optimize the segmentation network composed of the encoder and the decoder, which is expressed as:
7. The medical image segmentation method based on the tree-structured denoising diffusion probability model according to claim 1, wherein Step 6 includes: Step 6.1: For a medical image Y0 in the test set, Gaussian noise is added to the medical image Y0 to generate a noisy image Y t , where t = 1, 2, …, T, and Y represents the noisy image at the t-th time step; Step 6.2: Input Y t and Y0 into the encoder simultaneously. First, fuse them and then perform a convolution operation. Input the prompt images S1, S2, and S3 into the encoder in sequence to guide the segmentation. The decoder generates multiple segmentation masks through random sampling, and merges the multiple segmentation masks into one output by accumulation and averaging; Step 6.3: After T-step iteration of Step 6.2, obtain the predicted segmentation mask map Y0' of Y0; Step 6.4: Traverse the test set and use the evaluation metrics Dice coefficient and IoU index to analyze the performance of the network model.
8. The medical image segmentation method based on the tree-structured denoising diffusion probabilistic model according to claim 7, wherein, In Step 6.4, the Dice coefficient is used to measure the performance of the model by calculating the overlap degree between the predicted segmentation region and the true segmentation region, ranging from 0 to 1. The higher the value, the better the segmentation effect. Its formula is expressed as: where X and Y respectively represent the pixel sets of the predicted segmentation region and the true segmentation region, |X∩Y| represents the number of pixels in their intersection, and |X| and |Y| respectively represent the number of pixels in their respective sets; The IoU index is used to measure the ratio of the intersection to the union of the predicted segmentation region and the true segmentation region, ranging from 0 to 1. The higher the value, the more accurate the segmentation result. Its formula is expressed as:
9. A medical image segmentation system based on a tree structure denoising diffusion probability model, characterized in that, including: An acquisition module for obtaining a medical image dataset containing segmentation masks, and randomly selecting three segmentation masks belonging to the same lesion from it as prompt images; An initialization module for preprocessing a medical image dataset, dividing the preprocessed dataset into a training set and a test set according to a preset ratio, constructing a denoising diffusion probabilistic model and initializing model parameters; Noise addition module, which is used to add Gaussian noise to a segmentation mask X0 in the training set through a forward diffusion process for destruction based on the training set, generating a series of noisy images X t , where t = 1, 2, …, T, and X t represents the noisy image at the t-th time step; A fusion output module is used to construct an encoder and a decoder based on a binary tree structure. The encoder receives the medical image I corresponding to X0 and the noise image X t , as well as the prompt image, enabling information interaction between the medical image I and the prompt image with X t , outputting a feature representation. The decoder receives the feature representation output by the encoder, randomly samples the feature representation to generate different segmentation masks, and performs feature fusion output; An update module for repeating the fusion output module until a predicted segmentation mask map X0' of X0 is output, calculating the loss between X0' and X0, and updating the model parameters; A test module for testing the model with updated parameters using the test set.
10. A computing device, characterized in that, Including: One or more processors, one or more memories, and one or more programs, where the programs are stored in the memory and configured to be executed by the processors. When the programs are loaded into the processors, the steps of the medical image segmentation method based on the tree-structured denoising diffusion probabilistic model according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Medical image segmentation system based on multi-architecture fusion and diffusion model
CN121053147A