Method and system for image segmentation based om corrected diffusion model
The correction diffusion model iteratively corrects systematic errors in image segmentation by calculating noise differences and reconstructing errors, enhancing accuracy and reliability in image segmentation tasks.
Patent Information
- Application Number
- JP2024225468
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Conventional image segmentation models based on machine learning, such as convolutional neural networks and U-Net, fail to accurately correct systematic errors, leading to low segmentation accuracy due to factors like model characteristics, image noise, and artifacts.
An image segmentation method using a correction diffusion model that calculates the difference between a systematic error image with and without Gaussian noise, iteratively corrects and reconstructs the error using a loss function, and combines the primary segmentation result with the reconstructed error to improve accuracy.
The method enhances image segmentation accuracy by effectively addressing systematic errors, reducing training costs, and improving flexibility and reliability with vector quantization and cross-attention mechanisms.
Smart Images

Figure 2025100514000001_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image segmentation, and particularly relates to an image segmentation method and system based on a corrected diffusion model.
Background Art
[0002] The description of this part is only for providing information on the background art related to the present invention, and does not necessarily constitute the prior art.
[0003] Generally, image segmentation errors can be classified into two types: random errors and systematic errors. Random errors are caused by random effects such as image noise or unpredictable image artifacts, and can be reduced by fusing multiple information labels generated by independent segmentations. However, since random errors are unpredictable, it is very difficult to detect and correct random errors. Systematic errors are errors due to the system mode, that is, the difference between the standard segmentation definition and the manual segmentation protocol used by the system developer to train the segmentation model. Since systematic errors are related to a specific model, they always occur under specific model conditions. Therefore, the systematic errors of a specific model can be effectively corrected using machine learning techniques.
[0004] In recent years, methods based on machine learning have been widely used in image segmentation. Direct segmentation of images by convolutional neural networks or U-Net and its variant models is one of the most widely used techniques. In the prior art, image segmentation based on machine learning can be realized, but due to the following problems, the accuracy of the final segmentation is still low.
[0005] Conventional segmentation models based on machine learning do not consider the problem of segmentation error. That is, when performing image segmentation using a convolutional neural network or a U-Net model, it does not consider the problem that errors occur in the accuracy of the image segmentation result due to factors such as the characteristics of the model, image noise, and artifacts.
[0006] Conventional segmentation models based on machine learning do not consider the correction of systematic errors. That is, although deep learning models such as convolutional neural networks and U-Nets can improve the accuracy of image segmentation to a certain extent, many of them cannot automatically process systematic errors. Also, it does not consider how to correct systematic errors.
Summary of the Invention
[0007] To overcome the above deficiencies of the prior art, the present invention provides an image segmentation method and system based on a correction diffusion model. In the correction diffusion model, the difference between the systematic error image with Gaussian noise added and the systematic error image without noise added is calculated by a loss function, and the systematic error image is corrected and reconstructed to spread the systematic error, thereby realizing further improvement in image segmentation accuracy.
[0008] To achieve the above object, the first aspect of the present invention is performing preprocessing on the image to be segmented to obtain the preprocessed image to be segmented; performing segmentation on the preprocessed image to be segmented by a trained U-Net segmentation network to obtain a primary segmentation result image; performing a difference operation on the ground truth label of the segmentation result of the image to be segmented and the primary segmentation result image to obtain a systematic error image; Under the condition of a Markov chain, Gaussian noise is added to the systematic error image, and the difference between the systematic error image with Gaussian noise added and the systematic error image without noise added is calculated by a constructed loss function. By repeatedly correcting and reconstructing the systematic error image according to the calculation result of the loss function, a reconstructed systematic error image is obtained; Performing an addition operation on the primary segmentation result image and the obtained reconstructed systematic error image to obtain a final segmentation result image; Provided is an image segmentation method based on a correction diffusion model, including the above steps.
[0009] The second aspect of the present invention is An acquisition module that performs preprocessing on the image to be segmented to obtain a preprocessed image to be segmented; A primary segmentation module that performs segmentation on the preprocessed image to be segmented by a trained U-Net segmentation network to obtain a primary segmentation result image; A calculation module that performs a difference operation on the ground truth label of the segmentation result of the image to be segmented and the primary segmentation result image to obtain a systematic error image; A reconstruction module that, under the condition of a Markov chain, adds Gaussian noise to the systematic error image, calculates the difference between the systematic error image with Gaussian noise added and the systematic error image without noise added by a constructed loss function, and repeatedly corrects and reconstructs the systematic error image according to the calculation result of the loss function to obtain a reconstructed systematic error image; A segmentation module that performs an addition operation on the primary segmentation result image and the obtained reconstructed systematic error image to obtain a final segmentation result image; Provided is an image segmentation system based on a correction diffusion model, including the above modules.
[0010] A third aspect of the present invention provides a computer device comprising a processor, a memory, and a bus, wherein the memory stores machine-readable commands executable by the processor, and when the computer device operates, the processor and the memory communicate via the bus, and when the machine-readable commands are executed by the processor, steps of the image segmentation method based on the corrected diffusion model described above are executed.
[0011] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, steps of the image segmentation method based on the corrected diffusion model described above are executed.
[0012] The one or more technical means described above have the following beneficial effects.
[0013] In the present invention, the difference between the systematic error image with added Gaussian noise and the systematic error image without added noise is calculated by a loss function, and by correcting and reconstructing the systematic error image to spread the systematic error, a further improvement in image segmentation accuracy is realized.
[0014] In the present invention, by introducing a vector quantization variational autoencoder and applying the corrected diffusion model to the latent space of a powerful autoencoder, the training cost can be reduced, and the training of the corrected diffusion model can be realized with limited computing resources.
[0015] In the present invention, by introducing a cross-attention mechanism to improve the flexibility and reliability of the corrected diffusion model, the corrected diffusion model can obtain more flexible and reliable results based on the segmentation target image as a conditional input.
[0016] Advantages of additional aspects of the present invention are partially explained by the following description, and are also partially clarified by the following description, or are understood by the practice of the present invention.
[0017] The attached drawings of the specification constituting a part of the present invention are for providing a further understanding of the present invention, and the exemplary embodiments of the present invention and the description thereof are for explaining the present invention and do not constitute an undue limitation to the present invention.
Brief Description of the Drawings
[0018]
Figure 1
Figure 2
Figure 3
Modes for Carrying Out the Invention
[0019] It should be noted that all the following detailed descriptions are exemplary for further explaining the present invention. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0020] In addition, the terms used here are only for explaining specific embodiments and are not intended to limit the exemplary embodiments of the present invention.
[0021] Unless there is a contradiction, the features in the embodiments and the embodiments of the present invention can be combined with each other.
[0022] Example 1 In this example, performing preprocessing on the image to be segmented to obtain the preprocessed image to be segmented; Performing segmentation on the pre-processed image to be segmented using a trained U-Net segmentation network to obtain a primary segmentation result image; Performing a difference operation on the ground truth label of the segmentation result of the image to be segmented and the primary segmentation result image to obtain a systematic error image; Under the condition of a Markov chain, adding Gaussian noise to the systematic error image, calculating the difference between the systematic error image with added Gaussian noise and the systematic error image without added noise using a constructed loss function, and iteratively correcting and reconstructing the systematic error image based on the calculation result of the loss function to obtain a reconstructed systematic error image; Performing an addition operation on the primary segmentation result image and the obtained reconstructed systematic error image to obtain a final segmentation result image; Disclosed is an image segmentation method based on a correction diffusion model, including the above steps.
[0023] In this embodiment, the technical means will be described in detail using an MRI brain tumor image as an example of the image to be segmented.
[0024] In this embodiment, specifically for the MRI brain tumor image to be segmented, Performing a process of adjusting the size of the obtained MRI brain tumor image to be segmented to a unified size, including pre-processing that standardizes the size of all input images to 128×128.
[0025] By performing pre-processing on the original MRI brain tumor image, the calculation efficiency can be improved, the memory space can be saved, thereby improving the performance of the network model and making it suitable for the segmentation task.
[0026] Next, the pre-processed MRI brain tumor image to be segmented is input into the trained U-Net segmentation network for segmentation, and the brain tumor segmentation result image is output.
[0027] As the training process of the U-Net segmentation network, the labeled MRI brain tumor image is used as the input, the brain tumor segmentation result image is used as the output, and the U-Net segmentation network is trained.
[0028] Perform a difference operation on the actual brain tumor segmentation image and the primary segmentation result image obtained by the U-Net segmentation network, and output the systematic error image.
[0029] Input the systematic error image into the trained correction diffusion model, correct the generated systematic error image by the trained correction diffusion model, and use the trained encoder τ θ to encode the MRI brain tumor image and map it to the correction diffusion model to guide the systematic error image, and output the reconstructed systematic error image.
[0030] In this embodiment, a correction diffusion model segmentation network is constructed to segment the image. The correction diffusion model segmentation network includes a correction diffusion model and a U-Net segmentation network. Here, the correction diffusion model mainly consists of a diffusion model, an encoder, a decoder, a vector quantization variational autoencoder τ θ and cross-attention.
[0031] Specifically, in the correction diffusion model segmentation network, the U-Net segmentation network performs segmentation on the MRI brain tumor image to obtain the primary segmentation result image of the brain tumor. The difference module obtains a systematic error image based on the actual brain tumor segmentation image and the primary segmentation result image. The correction diffusion model module uses θ as the parameter, performs the forward diffusion process and the reverse diffusion process respectively, and obtains a reconstructed systematic error image. The vector quantization variational autoencoder composed of the encoder E and the decoder D maps the MRI image to a low-dimensional and highly efficient latent space. The cross-attention mechanism improves the flexibility and reliability of the correction diffusion model. The fusion module obtains the final brain tumor segmentation result image based on the primary segmentation result image and the reconstructed systematic error image.
[0032] In the training stage of the correction diffusion model segmentation network, the U-Net segmentation network performs segmentation on the MRI brain tumor image to obtain the primary segmentation result image of the brain tumor. The difference module obtains a systematic error image based on the actual brain tumor segmentation image and the primary segmentation result image. The correction diffusion model module obtains a reconstructed systematic error image. Specifically, for the systematic error image obtained by the difference module, first perform the forward diffusion process and encode it by the vector quantization variational autoencoder. Then, perform the forward noise addition process T times to obtain a Gaussian noise image. Next, perform the reverse diffusion process, condition on the MRI brain tumor image, perform the reverse noise removal process T times from the Gaussian noise image, and then decode it by the decoder to obtain a reconstructed systematic error image. The fusion module obtains the final brain tumor segmentation result image based on the primary segmentation result image and the reconstructed systematic error image.
[0033] In the actual application of the corrected diffusion model segmentation network, that is, in the inference stage, the U-Net segmentation network performs segmentation on the MRI brain tumor image to obtain the brain tumor segmentation result image. The corrected diffusion model module obtains the reconstructed systematic error image. Specifically, in the reverse diffusion process, taking the MRI brain tumor image as the input condition, performing the reverse noise removal process T times from the Gaussian noise image, and then decoding by the decoder to obtain the reconstructed systematic error image. The fusion module obtains the final brain tumor segmentation result image based on the primary segmentation result image and the reconstructed systematic error image.
[0034] In this embodiment, as the training process of the corrected diffusion model, in the forward diffusion process, the systematic error image is used as the input, and the Gaussian noise image is used as the output. In the reverse diffusion process, the difference between the predicted output and the actual Gaussian noise output image is calculated by the loss function, and the update of the connection weights is calculated by the gradient, and the reconstructed systematic error image is iteratively generated from the Gaussian noise image step by step.
[0035] Next, an addition operation is performed on the primary segmentation result image obtained by the U-Net segmentation network and the reconstructed systematic error image obtained by the corrected diffusion model, and the final brain tumor segmentation result image is output.
[0036] In the corrected diffusion model segmentation network, brain tumor segmentation is realized by the MRI brain tumor image X and the ground truth label M.
[0037] The processing process of the corrected diffusion model is
Equation
[0038] The vector quantization variational autoencoder in the above corrected diffusion model is specifically constructed as follows. Taking the U-Net segmentation result image as the primary mask S, and using the systematic error image obtained by the difference operation between the ground truth label M and the primary mask S as E. For a given systematic error image E, by encoding the systematic error image E with a vector quantization variational autoencoder, the latent space representation Z0 = E(E) is obtained. Then, the latent space representation Z0 is decoded by a decoder to obtain the reconstructed error image
Number
[0039] In this embodiment, the centers of the latent variables z are set to zero by two regularization methods. One is KL regularization that adds KL divergence to z. The other is VQ regularization that first learns codebooks for different z and then regularizes the latent space using a vector quantization layer. By compressing the model, the MRI brain tumor image can enter a latent representation space where high-frequency details and fine details are abstracted. Also, the likelihood-based generative model can focus on the important semantic bits of the data and can be trained in a lower-dimensional and computationally more efficient space, so it is more suitable for such a low-dimensional and highly efficient latent space.
[0040] The forward diffusion process in the above corrected diffusion model is specifically constructed as follows. In the diffusion process, under the condition of a Markov chain, Gaussian noise is continuously added to the initial data information, that is, Gaussian noise is continuously added to z0 to convert z0 to z T and in each step of noise addition,
Number
Number
Number
[0041] By the reparameterization technique,
Number
[0042]
Number
[0043] The reverse diffusion process in the above correction diffusion model is specifically constructed as follows. The initial data is obtained from Gaussian noise by the reverse diffusion process. Similarly
Number
Number
[0044] Using the neural network P θ , fit the reverse process q(Z t-1 |Z t , τ θ (X)). q(Z t-1 |Z t , τ θ (X), Z0) can be derived from P θ . Since q(Z t-1 |Z t , τ θ (X), Z0) is solved by Bayes' formula, finally an exponential function with an exponent that is a quadratic function is obtained. By Bayes' formula, the mean value and variance can be obtained. Since Z0 should not be known in advance in the reverse process, replace Z0 with Z t . Therefore,
Number
Number
[0045] In addition, a cross-attention mechanism is also introduced into the reverse diffusion process to enhance the underlying U-Net noise removal network. The reverse diffusion process is conditioned on the MRI brain tumor image encoded by the encoder. The MRI brain tumor image is mapped to the intermediate layer of the U-Net noise removal network in the diffusion model by the cross-attention mechanism. Next, the conditional noise predictor ε θ is used to perform T iterative noise removals, and the latent variable z T is converted to z0 and then input to the decoder for decoding to obtain the reconstruction error diagram
Number
[0046] Note that the principle of the diffusion model is T iterative noise removals. For each iterative noise removal step, one U-Net noise removal network is included. In each step, Z t is converted to Z t-1 by the U-Net noise removal network. For example, by inputting the sample at time step t into the U-Net noise removal network and performing implicit noise removal, the sample at time step t-1 can be obtained.
[0047] Therefore, the loss function of the corrected diffusion model is
Number
[0048] The cross-attention mechanism in the above corrected diffusion model is specifically constructed as follows. In principle, the diffusion model is a conditional noise removal autoencoder ε θ (z t , t, X) by
Number
[0049] In this embodiment, a CLIP image encoder is introduced, and τ θ is set. By doing so, preprocessing is performed on the MRI brain tumor image X. The CLIP image encoder projects the MRI brain tumor image X into the intermediate representation τ θ (X), and maps it to the intermediate layer of the U-Net noise removal network in the diffusion model through the cross-attention mechanism. Through the cross-attention mechanism, conditional encoding and the intermediate layer of the U-Net noise removal network are connected, and the input to the cross-attention mechanism is z t and τ θ (X), and the output of the cross-attention mechanism is input to the decoder and decoded.
[0050] The reverse diffusion process is used for learning systematic errors, and by adding information of the MRI brain tumor image X at each time step, the learning of the model is guided. By mapping the information of the MRI brain tumor image X to the intermediate layer of the U-Net noise removal network, it contributes to improving the performance and ability of the model and more accurately grasping the characteristics and distribution of the data.
[0051] In the reverse diffusion process, by adding information of the MRI brain tumor image X at each time step, the learning of systematic errors is guided.
[0052]
Number
[0053] To summarize the above, in this embodiment, the training steps of the above-mentioned corrected diffusion model segmentation network are as follows: Constructing a training sample set of MRI brain tumor images with tumor labels; Constructing a corrected diffusion model segmentation network; Training the corrected diffusion model segmentation network with the training sample set, and stopping the training when the loss function reaches the minimum value or the number of iterations meets the set requirements, to obtain a trained corrected diffusion model segmentation network; including.
[0054] In the training process, based on the sample set, the samples are randomly divided so that the ratio is 8:2, with 80% of them being training samples and the remaining 20% being test samples. To evaluate the model, the brain tumor datasets provided by BRATS2019, BRATS2020, and Jun Cheng are selected. Here, the sample values of BRATS2019 and BRATS2020 are brain tumor images labeled with the whole tumor WT, tumor core TC, and enhanced tumor ET, and the brain tumor dataset provided by Jun Cheng includes three types of brain tumor images, and the sample values are all brain tumor images labeled with tumor labels.
[0055] Here, the corrected diffusion model segmentation network to be trained mainly includes two parts, and these two parts respectively correspond to the two modules shown in Figure 1. The first part is the U-Net segmentation network, which takes the MRI brain tumor image as the input, performs segmentation on the MRI brain tumor image by the U-Net network, and obtains the primary mask s. The second part is the corrected diffusion model network, which uses the diffusion model principle to correct the segmentation systematic error and obtains the systematic error segmentation result image by the loss function.
[0056] In this embodiment, the loss function of the corrected diffusion model system is specifically as follows.
Number
[0057] In the training process, preprocessing is performed on the MRI brain tumor image, and the preprocessed MRI brain tumor image is input into the U-Net segmentation network to generate a primary mask s. The obtained error diagram E is input into the corrected diffusion model. By simultaneously optimizing the conditional noise predictor ε θ and the vector quantization variational autoencoder τ θ , the accuracy of the experiment is further improved by the loss function, and the reconstructed error diagram
Number
Number
[0058] In this embodiment, the inference process of the corrected diffusion model network is specifically as follows. As shown in Figure 2, the MRI brain tumor image is used as the input. Sample z T from a standard Gaussian distribution, encode the MRI brain tumor image X by the vector quantization variational autoencoder τ θ , map it to the U-Net noise removal network in the diffusion model by the cross-attention layer, and apply it in each iteration. Then, the conditional noise predictor ε θUsing z t , t and τ θ as inputs, calculate z t-1 . When t = 1, z0 is the prediction of the final systematic error. By decoding z0 with a decoder, a reconstruction error diagram
Number
Number
Number
[0059] In this example, brain tumor segmentation experiments were conducted on the brain tumor datasets provided by BRATS2019, BRATS2020, and Jun Cheng. The training set of BRATS2019 contains 335 images, which were randomly divided into 268 training images and 67 validation images. Similarly, the training set of BRATS2020 contains 369 images, which were randomly divided into 295 training images and 74 validation images. Also, the brain tumor dataset provided by Jun Cheng consists of 3064 contrast-enhanced images weighted by t1, which were similarly randomly divided into 2452 for training and 612 for validation.
[0060] Comparing the final segmentation results in the BRATS2020 dataset in this example with the conventional segmentation techniques, as shown in Figure 3 and Table 1 below, the technique described in this example can achieve a better segmentation effect.
[0061]
Table 1
[0062] When comparing the final segmentation results in the BRATS2019 dataset in this embodiment with the conventional segmentation techniques, as shown in FIG. 3 and Table 2 below, the technique described in this embodiment can achieve a better segmentation effect.
[0063]
Table 2
[0064] In this embodiment, the acquired MRI brain tumor image is input into the U-Net network after preprocessing, and segmentation is performed on the MRI brain tumor image to output a brain tumor segmentation result image and generate a primary mask s. Then, a difference operation is performed on the actual brain tumor segmentation image and the U-Net segmentation result image s, and the obtained error diagram E is input into the correction diffusion model network, and ε θ performs T times of iterative noise removal, and the reconstructed error diagram
Number
Number
[0065] The accurate MRI brain tumor image segmentation system based on the correction diffusion model provided by this embodiment diffusively corrects the systematic error generated by the segmentation algorithm by means of a diffusion model. By diffusing the systematic error, a further improvement in the accuracy of MRI brain tumor image segmentation is realized, assisting doctors in the diagnosis of brain tumor diseases.
[0066] The correction diffusion model algorithm provided by this embodiment applies the diffusion model to the correction of systematic errors, and introduces correction learning into the diffusion model to improve the segmentation accuracy of brain tumors.
[0067] This embodiment introduces a vector quantization variational autoencoder, and by applying the correction diffusion model to the latent space of a powerful autoencoder, the training cost can be reduced, and the training of the correction diffusion model can be realized with limited computing resources.
[0068] In this embodiment, a cross-attention mechanism is introduced to improve the flexibility and reliability of the correction diffusion model, so that the correction diffusion model can output more flexible and reliable results based on the conditional input of MRI brain tumor images.
[0069] Example 2 This embodiment includes an acquisition module that performs preprocessing on the segmentation target image to obtain a preprocessed segmentation target image, and a primary segmentation module that performs segmentation on the preprocessed segmentation target image by a trained U-Net segmentation network to obtain a primary segmentation result image, and a calculation module that performs a difference operation on the ground truth label of the segmentation result of the segmentation target image and the primary segmentation result image to obtain a systematic error image, and a reconstruction module that adds Gaussian noise to the systematic error image under the condition of a Markov chain, calculates the difference between the systematic error image with added Gaussian noise and the systematic error image without added noise by a constructed loss function, and repeatedly corrects and reconstructs the systematic error image according to the calculation result of the loss function to obtain a reconstructed systematic error image, and a segmentation module that performs an addition operation on the primary segmentation result image and the obtained reconstructed systematic error image to obtain a final segmentation result image, and Provided is an image segmentation system based on a corrected diffusion model, including
[0070] Example 3 This example aims to provide a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method described in Example 1 is realized when the program is executed by the processor.
[0071] Example 4 This example aims to provide a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and the method described in Example 1 is executed when the program is executed by a processor.
[0072] The steps and methods related to the devices in the above Examples 2, 3, and 4 correspond to those in Example 1. For specific embodiments, reference can be made to the relevant description parts in Example 1. The term "computer-readable storage medium" is understood to be a single medium or multiple media including one or more command sets, and can store, encode, or carry a command set for execution by a processor, and is understood to include any medium capable of causing a processor to execute any of the methods in the present invention.
[0073] As can be understood by those skilled in the art, each module or each step of the above-described present invention can be realized by a general-purpose computer device, or alternatively, by the code of a program executable by the computer device. Therefore, they can be stored in a storage and executed by the computer device, or they can be created into integrated circuit modules respectively, or a plurality of them can be created into a single integrated circuit module. The present invention is not limited to any specific combination of hardware and software.
[0074] As described above, the specific embodiments of the present invention have been described with reference to the accompanying drawings, but do not limit the protection scope of the present invention. As will be understood by those skilled in the art, various modifications or variations that can be made by those skilled in the art without making inventive efforts based on the technical means of the present invention are still included within the protection scope of the present invention.
Claims
1. Performing preprocessing on the image to be segmented to obtain the preprocessed image to be segmented; Performing segmentation on the preprocessed image to be segmented by a trained U-Net segmentation network to obtain a primary segmentation result image; Performing a difference operation on the ground truth label of the segmentation result of the image to be segmented and the primary segmentation result image to obtain a systematic error image; Adding Gaussian noise to the systematic error image under the condition of a Markov chain, calculating the difference between the systematic error image with added Gaussian noise and the systematic error image without added noise by a constructed loss function, and repeatedly correcting and reconstructing the systematic error image according to the calculation result of the loss function to obtain a reconstructed systematic error image; Performing an addition operation on the primary segmentation result image and the obtained reconstructed systematic error image to obtain a final segmentation result image; An image segmentation method based on a correction diffusion model, characterized by including the above steps.
2. The training of the U-Net segmentation network specifically includes using a labeled image as the input of the U-Net segmentation network and the segmentation result of the image as the output of the U-Net segmentation network, and training the U-Net segmentation network. The image segmentation method based on the correction diffusion model according to Claim 1, characterized by this.
3. Adding Gaussian noise to the systematic error image by the forward diffusion process of the correction diffusion model. Specifically, Encoding the systematic error image by an encoder; Continuously adding Gaussian noise to the encoded systematic error image; Obtaining a systematic error image with added Gaussian noise each time by the reparameterization technique. The image segmentation method based on the correction diffusion model according to Claim 1, characterized by this.
4. The segmentation target image is processed by a trained corrected diffusion model to obtain a reconstruction systematic error image. Specifically, in the reverse diffusion process of the trained corrected diffusion model, based on the segmentation target image encoded by the encoder, after performing the reverse noise removal process multiple times from a Gaussian noise image, it is decoded by the decoder to obtain a reconstruction systematic error image. The image segmentation system based on the corrected diffusion model according to claim 1, characterized in that.
5. Noise removal is performed by a U-Net noise removal network. In the U-Net noise removal network, the segmentation target image is encoded by a vector quantization variational autoencoder, the encoded segmentation target image is mapped to an intermediate layer of the U-Net noise removal network by a cross-attention mechanism, iterative noise removal is performed multiple times using a conditional noise predictor, the latent variable is transformed and then input to the decoder for decoding to obtain a reconstruction systematic error image. The image segmentation system based on the corrected diffusion model according to claim 4, characterized in that.
6. The training of the corrected diffusion model is specifically, in the forward diffusion process, the systematic error image is used as the input and the Gaussian noise image is used as the output. In the reverse diffusion process, the difference between the predicted output and the actual Gaussian noise output image is calculated by a loss function, and the update of the connection weights is calculated by the gradient, and gradually the reconstruction systematic error image is iteratively generated from the Gaussian noise image. The image segmentation system based on the corrected diffusion model according to claim 1, characterized in that.
7. In the reverse diffusion process, the segmentation target image is added at each time step to guide the learning of the systematic error. The image segmentation system based on the corrected diffusion model according to claim 4, characterized in that.
8. An acquisition module that performs preprocessing on the segmentation target image to obtain a preprocessed segmentation target image, A primary segmentation module that performs segmentation on the preprocessed segmentation target image by a trained U-Net segmentation network to obtain a primary segmentation result image, A calculation module that performs a difference operation on the ground truth label of the segmentation result of the image to be segmented and the primary segmentation result image to obtain a systematic error image, and Under the condition of a Markov chain, add Gaussian noise to the systematic error image, calculate the difference between the systematic error image with Gaussian noise added and the systematic error image without noise added by a constructed loss function, and iteratively correct and reconstruct the systematic error image according to the calculation result of the loss function to obtain a reconstructed systematic error image. A reconstruction module, A segmentation module that performs an addition operation on the primary segmentation result image and the obtained reconstructed systematic error image to obtain a final segmentation result image, and An image segmentation system based on a correction diffusion model, characterized by including the above.
9. A computer device comprising a processor, a memory, and a bus, wherein the memory stores machine-readable commands executable by the processor, and when the computer device operates, the processor and the memory communicate via the bus, and when the machine-readable commands are executed by the processor, the image segmentation method based on the correction diffusion model according to any one of claims 1 to 7 is executed.
10. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the image segmentation method based on the correction diffusion model according to any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Image processing device, image processing program, image recognition device, image recognition program, and image recognition system
JP2022078735A
Image processing method and computing device
JP2022103149A
Metallographic structure segmentation method
JP2022188504A