A method for synthesizing CT images from MR images based on mask-guided unsupervised adversarial diffusion model

By using a mask-guided unsupervised adversarial diffusion model in medical image translation, combined with large diffusion walk length and adversarial mapping, the problems of premature convergence and low computational efficiency in the prior art are solved, and efficient and accurate image translation effects are achieved, especially in the absence of paired training data.

CN118799432BActive Publication Date: 2025-05-13MEDMIND TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411158290.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2025-05-13
Estimated Expiration
2044-08-22

AI Technical Summary

Technical Problem

The prior art has problems in medical image translation with premature convergence of models, gradient disappearance and low computational efficiency, especially in the absence of paired training data, image translation effect is poor.

Method used

The unsupervised adversarial diffusion model based on mask guidance is adopted to directly capture the image distribution through a conditional diffusion process, combine large diffusion walk length and adversarial mapping to break through the calculation efficiency bottleneck of the traditional diffusion model, and introduce a cyclic consistency architecture based on non-diffusion modules and diffusion modules to achieve two-way translation between the source mode and the target mode.

Benefits of technology

It significantly improves the accuracy and efficiency of image translation, enables coherent and accurate image translation without paired training samples, and enhances the applicability of the model in scarce scenarios of paired data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118799432B_ABST
    Figure CN118799432B_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of MR image synthesis CT image, and discloses a method for synthesizing MR images into CT images based on a mask-guided unsupervised adversarial diffusion model. The adversarial diffusion model is introduced, and a conditional diffusion process is used to directly capture the image distribution. The image translation effect is improved by gradually mapping noise, a source image and a corresponding mask to a target image. In the reverse diffusion process, a large diffusion step size and adversarial mapping are used to break through the bottleneck of the calculation efficiency of the traditional diffusion model method. The mask is used to guide the learning of the diffusion model explicitly to achieve high-fidelity image translation. A cycle consistency architecture based on a non-diffusion module and a diffusion module is introduced to achieve two-way translation between the source modality and the target modality, ensuring that the image translation process is still accurate even in the absence of paired training samples, thereby significantly enhancing the applicability of the model in scenarios where paired data is scarce.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of synthesizing CT images from MR images, and specifically to a method for synthesizing CT images from MR images based on a mask-guided unsupervised adversarial diffusion model. Background Art

[0002] MR (magnetic resonance imaging) and CT (computed tomography) are two different medical imaging technologies, each with its own unique advantages and limitations. MR images have high resolution for soft tissues and can show the full picture of tumor tissues, but are not very sensitive to calcified lesions; while CT images are most sensitive to calcified lesions, but have low resolution for soft tissues. Therefore, combining MR images into CT images can take advantage of the advantages of both and improve the accuracy and comprehensiveness of diagnosis.

[0003] In the field of medical image synthesis, image translation technology can be used to synthesize missing CT images using MR images. Traditionally, image translation is often used in conjunction with generative adversarial networks (GANs). GAN models indirectly represent the target modality distribution through the interaction of the generator-discriminator. This implicit method can easily lead to premature convergence of the model or gradient vanishing. In addition, GAN models usually use a fast one-step sampling process without intermediate stages, which may limit the quality and diversity of synthesized images.

[0004] To address the above technical issues, a diffusion model can be introduced, which uses an explicit likelihood representation and a progressive sampling process to improve the sample fidelity in unconditional generative modeling tasks. However, the potential of diffusion methods in medical image translation remains unexplored, mainly due to the computational burden of image sampling and the lack of paired training in traditional diffusion models.

[0005] Therefore, we need to propose a method for synthesizing CT images from MR images based on a mask-guided unsupervised adversarial diffusion model, which can efficiently and highly fidelity achieve image translation between a given MR image (source modality) and a CT image (target modality); use the conditional diffusion process to directly capture the image distribution, and improve the image translation effect by gradually mapping the noise, source image and corresponding mask to the target image. Utility Model Content

[0006] The purpose of the invention is to provide a method for synthesizing CT images from MR images based on a mask-guided unsupervised adversarial diffusion model, introduce an adversarial diffusion model, use a conditional diffusion process to directly capture the image distribution, and improve the image translation effect by gradually mapping the noise, source image and corresponding mask to the target image; in the reverse diffusion process, a large diffusion step and adversarial mapping are used to break through the computational efficiency bottleneck of the traditional diffusion model method; mask guidance is used to clearly guide the learning of the diffusion model to achieve high-fidelity image translation; a cycle consistency architecture based on a non-diffusion module and a diffusion module is introduced to achieve bidirectional translation between the source modality and the target modality, ensuring that even in the absence of paired training samples, the image translation process is still coherent and accurate, thereby significantly enhancing the applicability of the model in scenarios where paired data is scarce, so as to solve the problems raised in the above background technology.

[0007] To achieve the above object, the invention provides the following technical solution: a method for synthesizing CT images from MR images based on a mask-guided unsupervised adversarial diffusion model, comprising the following steps:

[0008] S1. Introduce source modal condition mapper to capture complex transition probabilities under large step length in adversarial diffusion model , achieving fast and accurate reverse diffusion sampling, and through the conditional generator The image of the current t-th iteration of the diffusion model , source modality image and its corresponding mask As input, extract intermediate features and gradually denoise, synthesize ;in represents the denoised distribution of the diffusion model at the (tk)th iteration predicted by the conditional generator, Indicates that the above conditional generator is used to generate the image of the current tth iteration , source modality image and its corresponding mask As input, the conditional probability distribution of the denoised distribution of the (tk)th iteration of the diffusion model is predicted;

[0009] S2. Guidance of source modality image and boundary mask: In order to synthesize the image of the target modality, the parameterized inverse diffusion step needs the guidance of the source modality image. However, the training dataset contains unpaired modality images, which are MR images and CT images, denoted as x and y respectively. The MR image is set as the source modality and the CT image is set as the target modality. A cycle consistency architecture based on non-diffusion modules and diffusion modules is introduced to learn from the unpaired training dataset.

[0010] S3, sending the image and mask to the non-diffusion module, and using the non-diffusion method to synthesize the source modality image corresponding to each target image in the training data set;

[0011] S4, using the diffusion module to estimate the target modality image under the guidance of the source modality image generated by the non-diffusion module;

[0012] S5. Unsupervised learning: Compare the true target image and its reconstructed image using cycle consistency loss.

[0013] Preferably, in step S1, in order to improve the efficiency of image generation, a fast diffusion method of the forward process is proposed, as follows:

[0014] ; ;

[0015] In the above formula, t represents the tth iteration of the diffusion model, k represents the iteration step interval, and its value is greater than 1. represents the noise variance, For condition generator The generated estimated denoised distribution and actual denoised samples, represents the noise distribution of the forward process of the diffusion model, represents the normal distribution and I represents the identity matrix.

[0016] Preferably, in step S1, a time-based The computed learnable embeddings are incorporated into the extracted intermediate features to guide the model to generate accurate denoised predictions. Subsequently, the discriminator is introduced Distinguishing by generator Generated estimated denoised distribution and actual denoised samples , the temporal embedding is also added as a bias term to the discriminator in the feature map.

[0017] Preferably, in step S1, the condition generator Supervising a learnable discriminator using a non-saturating adversarial loss :

[0018] (1)

[0019] ;

[0020] (2)

[0021] Where E represents the expected value, x 0 Represents the input of the diffusion model at the 0th iteration, that is, the original image. The final denoising result is executed It is obtained after a reverse diffusion process, generating high-quality denoised output while maintaining robustness to adversarial noise.

[0022] Preferably, in step S2, the cycle consistency architecture performs bidirectional translation between the source modality and the target modality, effectively bridging the gap between unpaired data sets. The cycle consistency architecture ensures that the image translation process remains coherent and accurate even in the absence of paired training samples, thereby significantly enhancing the applicability of the model in scenarios where paired data is scarce.

[0023] Preferably, in step S3, the condition generator According to the target modality image and mask Generate source modality images. This process includes the following steps:

[0024] (3)

[0025] for , using a non-saturated adversarial loss function:

[0026] (4)

[0027] in, Represents the parameterized representation of the network under the source conditional distribution for a given target image, the discriminator The non-saturated adversarial loss function is used to distinguish the estimated source image from the actual source image. This process includes:

[0028] (5)

[0029] in, represents the true conditional distribution of target images given a source modality.

[0030] Preferably, in step S4, the diffusion module is used to estimate the target modality image, which consists of two adversarial diffusion processes, each of which is equipped with a specific discriminator , at each step in the reverse diffusion process ( ), the condition generator First generate an estimate of the target image :

[0031] (6)

[0032] Here, each step refers to the initial number of iterations t of the reverse diffusion process of the model.

[0033] Condition Builder Will Process it as a three-channel input and extract intermediate features ,in Represents the sub-block index in the encoding-decoding structure, for each time step Computational learnable temporal embeddings is added to the feature map in each sub-block as a channel-specific bias term, expressed as , then, the condition generator The target image is synthesized using the denoising distribution specific to each image modality.

[0034] Preferably, in step S4, the discriminator take over or and As input, distinguish the denoised distribution from the prediction and the true denoised distribution Samples of

[0035] Discriminator Will or Processed as a two-channel input, time-embedded is incorporated into the discriminator as a bias term In the feature map of When sampling, the condition generator The representation is as follows:

[0036] ( | , )= (7)

[0037] Among them, the condition generator Predicted distance have step, and then use the denoised distribution to synthesize the target image for each modality;

[0038] (8).

[0039] Preferably, in step S5, in the diffusion module, the reconstructed image is called the synthetic target image, and in the non-diffusion module, the estimated source image is passed through the conditional generator Mapped to the target modality, the diffusive and non-diffusive modules are jointly trained without any pre-training process;

[0040] Condition Builder The loss consists of the following:

[0041] (9)

[0042] in, Indicates loss The weight of Indicates loss The weight of represents a commonly used cyclic loss function;

[0043] Discriminator The overall loss consists of the following components:

[0044] (10)

[0045] in, Representation Discriminator overall loss;

[0046] During training, the non-diffuse module needs to generate source image estimates paired with a given target image, but during inference, the task changes to synthesizing an unacquired CT target image given an MR source image acquired from a specific medical image, so only the conditional generators specified in the diffuse module for the desired task need to be executed. That's it.

[0047] Compared with the prior art, the invention has the following beneficial effects:

[0048] 1. The present invention introduces an adversarial diffusion model, uses a conditional diffusion process to directly capture the image distribution, and improves the image translation effect by gradually mapping the noise, source image and corresponding mask to the target image;

[0049] 2. The present invention adopts a large diffusion step size and adversarial mapping in the reverse diffusion process, breaking through the bottleneck of computational efficiency of the traditional diffusion model method;

[0050] 3. The present invention uses mask guidance to explicitly guide the learning of the diffusion model and achieve high-fidelity image translation;

[0051] 4. The present invention introduces a cycle consistency architecture based on non-diffusion modules and diffusion modules, which realizes bidirectional translation between the source modality and the target modality, ensuring that the image translation process remains coherent and accurate even in the absence of paired training samples, thereby significantly enhancing the applicability of the model in scenarios where paired data is scarce. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 A schematic diagram of the invention to counter the diffusion model;

[0053] Figure 2 A schematic diagram of the invention of a non-diffusion module and a diffusion module;

[0054] Figure 3 A flowchart of the invention. DETAILED DESCRIPTION

[0055] The following will be combined with the drawings in the embodiments of the invention to clearly and completely describe the technical solutions in the embodiments of the invention. Obviously, the described embodiments are only part of the embodiments of the invention, not all of the embodiments. Based on the embodiments in the invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the invention.

[0056] See also Figure 1-3 The invention provides a technical solution: a method for synthesizing CT images from MR images based on a mask-guided unsupervised adversarial diffusion model, which uses unpaired medical MR images and CT image data to effectively train the image translation model of MR-to-CT synthesis, and achieves image translation from a given MR image (source modality) to a CT image (target modality). Different from the prior art solution of achieving image translation through generative adversarial networks (GANs), this patent adopts a mask-guided unsupervised adversarial diffusion model, combining the advantages of diffusion modeling and predictive modeling, and is significantly superior to the prior art in terms of accuracy and the ability to learn from unpaired data.

[0057] Two main modules are used in this method: an adversarial diffusion model for source modality conditional mapping, which can achieve fast and accurate inverse diffusion sampling; and a non-diffusion target to source modality paired image prediction module based on unsupervised learning. This method combines the advantages of diffusion modeling and predictive modeling to effectively learn the translation process of MR transformation to CT from unpaired MR and CT data.

[0058] The following steps are involved:

[0059] S1. Introduce source modal condition mapper to capture complex transition probabilities under large step length in adversarial diffusion model , achieving fast and accurate reverse diffusion sampling, and through the conditional generator The image of the current t-th iteration of the diffusion model , source modality image and its corresponding mask As input, extract intermediate features and gradually denoise, synthesize ;in represents the denoised distribution of the diffusion model at the (tk)th iteration predicted by the conditional generator, Indicates that the above conditional generator is used to generate the image of the current tth iteration , source modality image and its corresponding mask As input, the conditional probability distribution of the denoised distribution of the (tk)th iteration of the diffusion model is predicted;

[0060] like Figure 1As shown, in step S1, in order to improve the efficiency of image generation, a fast diffusion method of the forward process is proposed, as shown in the following formula:

[0061] ;

[0062] ; In the above formula, t represents the tth iteration of the diffusion model, k represents the iteration step interval, and its value is greater than 1. represents the noise variance, For condition generator The generated estimated denoised distribution and actual denoised samples, represents the noise distribution of the forward process of the diffusion model, represents the normal distribution and I represents the identity matrix.

[0063] In step S1, a time-based The computed learnable embeddings are incorporated into the extracted intermediate features to guide the model to generate accurate denoised predictions. Subsequently, the discriminator is introduced Distinguishing by generator Generated estimated denoised distribution and actual denoised samples , the temporal embedding is also added as a bias term to the discriminator in the feature map.

[0064] In step S1, the condition generator Supervising a learnable discriminator using a non-saturating adversarial loss :

[0065] (1)

[0066] ;

[0067] (2)

[0068] Where E represents the expected value, x 0 Represents the input of the diffusion model at the 0th iteration, that is, the original image. The final denoising result is executed It is obtained after a reverse diffusion process, generating high-quality denoised output while maintaining robustness to adversarial noise.

[0069] S2. Guidance of source modality image and boundary mask: In order to synthesize the image of the target modality, the parameterized inverse diffusion step needs the guidance of the source modality image. However, the training dataset contains unpaired modality images, which are MR images and CT images, denoted as x and y respectively. The MR image is set as the source modality and the CT image is set as the target modality. A cycle consistency architecture based on non-diffusion modules and diffusion modules is introduced to learn from the unpaired training dataset.

[0070] In step S2, the cycle consistency architecture performs bidirectional translation between the source modality and the target modality, effectively bridging the gap between unpaired datasets. The cycle consistency architecture ensures that the image translation process remains coherent and accurate even in the absence of paired training samples, thereby significantly enhancing the applicability of the model in scenarios where paired data is scarce.

[0071] like Figure 2 As shown, S3, the image and mask are sent to the non-diffusion module, and the source modality image corresponding to each target image in the training data set is synthesized using the non-diffusion method;

[0072] In step S3, the condition generator According to the target modality image and mask Generate source modality images. This process includes the following steps:

[0073] (3)

[0074] for , using a non-saturated adversarial loss function:

[0075] (4)

[0076] in, Represents the parameterized representation of the network under the source conditional distribution for a given target image, the discriminator The non-saturated adversarial loss function is used to distinguish the estimated source image from the actual source image. This process includes:

[0077] (5)

[0078] in, represents the true conditional distribution of target images given a source modality.

[0079] like Figure 2 As shown, S4, under the guidance of the source modality image generated by the non-diffusion module, the diffusion module is used to estimate the target modality image;

[0080] In step S4, the diffusion module is used to estimate the target modality image, which consists of two adversarial diffusion processes, each equipped with a specific discriminator , at each step in the reverse diffusion process ( ), the condition generator First generate an estimate of the target image :

[0081] (6)

[0082] Here, each step refers to the initial number of iterations t of the reverse diffusion process of the model.

[0083] Condition Builder Will Process it as a three-channel input and extract intermediate features ,in Represents the sub-block index in the encoding-decoding structure, for each time step Computational learnable temporal embeddings is added to the feature map in each sub-block as a channel-specific bias term, expressed as , then, the condition generator The target image is synthesized using the denoising distribution specific to each image modality.

[0084] In step S4, the discriminator take over or and As input, distinguish the denoised distribution from the prediction and the true denoised distribution Samples of

[0085] Discriminator Will or Processed as a two-channel input, time-embedded is incorporated into the discriminator as a bias term In the feature map of When sampling, the condition generator The representation is as follows:

[0086] ( | , )= (7)

[0087] Among them, the condition generator Predicted distance have step, and then use the denoised distribution to synthesize the target image for each modality;

[0088] (8).

[0089] S5. Unsupervised learning: Compare the true target image and its reconstructed image using cycle consistency loss.

[0090] In step S5, in the diffusion module, the reconstructed image is called the synthetic target image, and in the non-diffusion module, the estimated source image is passed through the conditional generator Mapped to the target modality, the diffusive and non-diffusive modules are jointly trained without any pre-training process;

[0091] Condition Builder The loss consists of the following:

[0092] (9)

[0093] in, Indicates loss The weight of Indicates loss The weight of represents a commonly used cyclic loss function;

[0094] Discriminator The overall loss consists of the following components:

[0095] (10)

[0096] in, Representation Discriminator overall loss;

[0097] During training, the non-diffuse module needs to generate source image estimates paired with a given target image, but during inference, the task changes to synthesizing an unacquired CT target image given an MR source image acquired from a specific medical image, so only the conditional generators specified in the diffuse module for the desired task need to be executed. That's it.

[0098] For example, to perform source-target mapping, use the generator ,in is a mask, is the MR image modality at time step CT target image sample, is the source modality MR image sample. Inference starts from time step Start with the initial sample is from The noisy target image sample generated at the end of each inverse diffusion step is used as the input of the next target image sample, using a total of A reverse diffusion step.

[0099] Although embodiments of the invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for synthesizing CT images from MR images based on a mask-guided unsupervised adversarial diffusion model, characterized in that: The following steps are involved: S1. Introduce source modal condition mapper to capture complex transition probabilities under large step length in adversarial diffusion model , achieving fast and accurate reverse diffusion sampling, and through the conditional generator The image of the current t-th iteration of the diffusion model , source modality image and its corresponding mask As input, extract intermediate features and gradually denoise, synthesize ;in represents the denoised distribution of the diffusion model at the (tk)th iteration predicted by the conditional generator, Indicates that the above conditional generator is used to generate the image of the current tth iteration , source modality image and its corresponding mask As input, the conditional probability distribution of the denoised distribution of the (tk)th iteration of the diffusion model is predicted; In step S1, in order to improve the efficiency of image generation, a fast diffusion method of the forward process is proposed, as follows: ; ; In the above formula, t represents the tth iteration of the diffusion model, k represents the iteration step interval, and its value is greater than 1. represents the noise variance, For condition generator The generated estimated denoised distribution and actual denoised samples, represents the noise distribution of the forward process of the diffusion model, represents normal distribution, I represents the identity matrix; In step S1, a time-based The computed learnable embeddings are incorporated into the extracted intermediate features to guide the model to generate accurate denoised predictions. Subsequently, the discriminator is introduced Distinguishing by generator Generated estimated denoised distribution and actual denoised samples , the temporal embedding is also added as a bias term to the discriminator In the feature map; In step S1, the condition generator Supervising a learnable discriminator using a non-saturating adversarial loss : (1) ; (2) Where E represents the expected value, x0 represents the input of the diffusion model at the 0th iteration, that is, the original image, and the final denoising result is executed After a reverse diffusion process, it generates high-quality denoised output while maintaining robustness to adversarial noise; S2. Guidance of source modality image and boundary mask: In order to synthesize the image of the target modality, the parameterized inverse diffusion step needs the guidance of the source modality image. However, the training dataset contains unpaired modality images, which are MR images and CT images, denoted as x and y respectively. The MR image is set as the source modality and the CT image is set as the target modality. A cycle consistency architecture based on non-diffusion modules and diffusion modules is introduced to learn from the unpaired training dataset. In step S2, the cycle consistency architecture performs bidirectional translation between the source and target modalities, effectively bridging the gap between unpaired datasets. The cycle consistency architecture ensures that the image translation process remains coherent and accurate even in the absence of paired training samples, significantly enhancing the applicability of the model in scenarios where paired data is scarce. S3, sending the image and mask to the non-diffusion module, and using the non-diffusion method to synthesize the source modality image corresponding to each target image in the training data set; In step S3, the condition generator According to the target modality image and mask Generate source modality images. This process includes the following steps: (3) for , using a non-saturated adversarial loss function: (4) in, Represents the parameterized representation of the network under the source conditional distribution for a given target image, the discriminator The non-saturated adversarial loss function is used to distinguish the estimated source image from the actual source image. This process includes: (5) in, represents the true conditional distribution of the target image given the source modality; S4, under the guidance of the source modality image generated by the non-diffusion module, the diffusion module is used to estimate the target modality image; In step S4, the diffusion module is used to estimate the target modality image, which consists of two adversarial diffusion processes, each equipped with a specific discriminator , at each step in the reverse diffusion process ( ), the condition generator First generate an estimate of the target image : (6) Here, each step refers to the initial iteration number t of the reverse diffusion process of the model; Condition Builder Will Process it as a three-channel input and extract intermediate features ,in Represents the sub-block index in the encoding-decoding structure, for each time step Computational learnable temporal embeddings is added to the feature map in each sub-block as a channel-specific bias term, expressed as , then, the condition generator Synthesize the target image using the denoising distribution specific to each image modality; In step S4, the discriminator take over or and As input, distinguish the denoised distribution from the prediction and the true denoised distribution Samples of Discriminator Will or Processed as a two-channel input, time-embedded is incorporated into the discriminator as a bias term In the feature map of When sampling, the condition generator The representation is as follows: ( | , )= (7) Among them, the condition generator Predicted distance have step, and then use the denoised distribution to synthesize the target image for each modality; (8); S5. Unsupervised learning: Compare the true target image and its reconstructed image using cycle consistency loss.

2. The method for synthesizing CT images from MR images based on a mask-guided unsupervised adversarial diffusion model according to claim 1, characterized in that: In step S5, in the diffusion module, the reconstructed image is called the synthetic target image, and in the non-diffusion module, the estimated source image is passed through the conditional generator Mapped to the target modality, the diffusive and non-diffusive modules are jointly trained without any pre-training process; Condition Builder The loss consists of the following: (9) in, Indicates loss The weight of Indicates loss The weight of represents a commonly used cyclic loss function; Discriminator The overall loss consists of the following components: (10) in, Representation Discriminator The overall loss of D represents the commonly used generative adversarial network discriminator loss function; During training, the non-diffuse module needs to generate source image estimates paired with a given target image, but during inference, the task changes to synthesizing an unacquired CT target image given an MR source image acquired from a specific medical image, so only the conditional generators specified in the diffuse module for the desired task need to be executed. That's it.

Citation Information

Patent Citations

  • MRI image and CT image conversion method and terminal based on deep learning

    CN114266929A

  • Method and system for synthesizing MRI image into CT image based on deep learning

    CN115311182A