Medical image generation method based on diffusion model
The diffusion model-based method generates high-quality PET images with controlled tumor characteristics, addressing data scarcity and variability issues, thereby enhancing AI model training and performance for tumor detection.
Patent Information
- Application Number
- CN202510463315.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-15
AI Technical Summary
The difficulty in obtaining and labeling high-quality PET image data leads to insufficient performance of AI models in tumor detection and segmentation tasks, and the existing synthesis algorithms cannot effectively simulate the location and morphology of tumors, resulting in data scarcity and labeling complexity problems.
Using a medical image generation method based on diffusion model, PET images with controllable tumor morphology are synthesized by inputting CT images and tumor masks, and high-quality PET images are generated using VQVAE and hidden space diffusion models, combined with a conditional encoder.
The generated PET images are more similar to the real image distribution, with rich details and correct anatomical structure, effectively expanding the data set and improving the performance of downstream tasks.
Smart Images

Figure CN120318359A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning image generation, and particularly relates to the research on high-quality generation methods for medical images. It is a method for generating medical images based on a diffusion model. Background Art
[0002] In the field of artificial intelligence (AI)-driven medical imaging, especially in clinical diagnosis, the demand for a large amount of accurately labeled high-quality training data is constantly rising. AI algorithms need this data to effectively learn and improve diagnostic accuracy. However, obtaining and labeling real medical images is a labor-intensive task that requires a large amount of manual labor and professional clinical knowledge, which poses a major challenge to the development and optimization of AI models.
[0003] Positron emission tomography (PET), as a key imaging tool widely used in oncology, cardiology, and neurology, faces multiple obstacles in data collection. First, PET requires the injection of radioactive tracers, which exposes patients to radiation for a long time, raises safety concerns, and limits large-scale data collection. Second, the high cost of PET equipment makes it difficult to popularize in many hospitals and medical centers, making PET data scarcer compared to imaging modalities such as CT and MRI. In addition, privacy protection issues also limit the access to public PET datasets. Moreover, the differences in tumor characteristics among different patients, such as location and morphology, lead to an imbalance in clinically sourced datasets, which further increases the difficulty of developing AI algorithms for tumor detection and segmentation.
[0004] Synthesizing PET images through tumor masks is a way to expand the dataset. The synthesis algorithm needs to be able to simulate the tumor distribution in a real clinical environment and provide the ability to precisely control the size, position, and shape of the tumor, thereby helping to overcome the challenges of data scarcity and annotation complexity, providing richer training data for the application of AI in the field of medical imaging, and ultimately improving the performance of AI diagnostic tools. Summary of the Invention
[0005] The present invention aims to solve the problem of synthesizing high-quality tumor PET images and proposes a method for generating medical images based on a diffusion model, which can synthesize PET images with highly controllable tumor morphology and satisfying CT anatomical structures by inputting artificially designed tumor masks and CT images.
[0006] The object of the present invention is achieved by the following technical solutions: A method for generating medical images based on a diffusion model, including the following steps:
[0007] The present invention discloses a method for generating medical images based on a diffusion model, including the following steps:
[0008] 1) Collect a batch of registered PET images, CT images with tumor lesions, and tumor masks as the dataset for training the model;
[0009] 2) Use all the PET images in the dataset to train a VQVAE model containing an encoder and a decoder, such that the PET image is consistent with the original image after being encoded, compressed, and then decoded and restored;
[0010] 3) Prepare a diffusion model, and connect the encoder and decoder of the VQVAE model trained in 2) before the noise addition of the diffusion model and after the denoising ends respectively, to obtain a latent space diffusion model based on the latent space;
[0011] 4) Introduce a conditional encoder into the latent space diffusion model in 3) to extract the features of the CT image and the tumor mask, and obtain an improved conditional diffusion model;
[0012] 5) Input the PET images, CT images, and tumor masks in the dataset into the improved conditional diffusion model for training, calculate the loss function, and use the gradient descent method to update the denoising network and the conditional encoder, and finally obtain a trained model;
[0013] 6) Input the tumor mask and CT image for testing into the model trained in 5) to output the synthesized PET image.
[0014] As a further improvement, the specific structure of the VQVAE model described in the present invention is:
[0015] The VQVAE model is divided into an encoder structure and a decoder structure. The encoder consists of 3 downsampling blocks. Each downsampling block contains a 3×3×3 convolutional layer, a batch normalization layer, and a ReLU activation function. The number of channels is 64, 128, and 256 in sequence. The 128×128×128 input image is compressed to 16×16×16 latent features through three downsamplings; the decoder symmetrically contains 3 upsampling blocks. Each block contains a 3×3×3 transposed convolutional layer, a batch normalization layer, and a ReLU activation function. The number of channels is 256, 128, and 64 in sequence; a vector quantization layer with a codebook size of 1024 and a vector dimension of 256 is connected to the output end of the encoder to further compress the image features.
[0016] As a further improvement, the specific process of training the PET image in step 5) of the present invention when input into the improved conditional diffusion model is:
[0017] The latent space diffusion model gradually adds noise to the PET image through the forward process until a pure noise sample is obtained. In the reverse denoising stage, a denoising network with a U-Net structure is used to gradually restore the original structure and details of the image. By learning the mapping of the noise addition and denoising processes, the denoising network can effectively remove noise and retain the detailed information of the image. Finally, the denoised result is input into the decoder of the VQVAE to be restored to the original image space, generating high-quality medical images.
[0018] As a further improvement, the specific process of inputting the CT image and the tumor mask into the improved conditional diffusion model for training in step 5) of the present invention is as follows:
[0019] The CT image and the tumor mask are introduced as conditions in the diffusion model. After splicing them by channels, they are input into the conditional encoder to obtain the conditional encoding. The conditional encoding is spliced with the output of the diffusion model denoising network by channels, and this process is repeated in each denoising step, so that the provided CT and tumor mask can be fused into the PET image. The structure of the conditional encoder is similar to the encoder in the VQVAE, and the input image is compressed from the image space of size 128 to the latent space of size 16 through three downsamplings.
[0020] As a further improvement, the specific mathematical formula of the loss function in step 5) of the present invention is:
[0021]
[0022] where ε t is the noise added at the t-th step in the noise addition stage of the diffusion model, ε θ is the denoising neural network in the diffusion model, ε is the standard Gaussian noise, E and E c are the encoder of the VQVAE and the conditional encoder respectively, x0 is the PET image in the dataset, c represents the conditional image obtained by splicing the CT image and the tumor mask in the dataset by channels, and the diffusion model is trained based on the above loss function.
[0023] The beneficial effects of the present invention are as follows:
[0024] The present invention uses a diffusion model to synthesize medical images. The diffusion model is the latest high-quality generative model proposed in recent years. Compared with earlier generative model paradigms such as variational autoencoders and generative adversarial networks, the diffusion model has more advantages in terms of training stability and generation quality, which is crucial for the medical image synthesis task.
[0025] From the perspective of data augmentation, the present invention can bring benefits to downstream tasks such as the training of segmentation models. Traditional data augmentation methods mainly perform simple geometric operations on the dataset, such as cropping, translation, and flipping. Such methods have limited improvement on model performance. The present invention adopts a generative model to massively expand the dataset from the level of image data distribution, effectively improving the performance of the segmentation model.
[0026] In addition, general methods for generating tumor PET images only consider using the tumor mask as the input during design. This only ensures that the position and shape of the tumor in the generated image are consistent with the input, while the rest of the structures need to be learned by the model, which increases the difficulty of model training. To further ensure that the generated PET images have reasonable anatomical structures, our diffusion model introduces both CT images and tumor masks as conditions into the training process, enabling the PET images output by the model to have accurate and reasonable anatomical structures and richer details. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is the flow structure diagram of the method;
[0028] Figure 2 is the model schematic diagram of the method;
[0029] Figure 3 is the output result diagram using the method;
[0030] (a) is the original PET image, and (b) is the PET image obtained using the method.
[0031] Figure 4 is the output result diagram using the existing method;
[0032] (a) is the PET image obtained using the conditional GAN network, and (b) is the PET image obtained using the diffusion model without CT constraint. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The present invention discloses a method for generating medical images based on a diffusion model, including the following steps:
[0034] 1) The flow structure of the present invention is as Figure 1 shown. First, the publicly available Hecktor dataset of head and neck tumors is collected. This dataset contains paired tumor PET images, tumor masks, and CT images, providing reliable training samples for the model. Specifically, the training set contains 200 pairs of three-dimensional image samples for model learning and parameter optimization; the validation set contains 24 pairs of image samples for evaluating model performance;
[0035] 2) Use all the PET images in the dataset to train a VQVAE model with an encoder and a decoder, such that the PET image is consistent with the original image after being encoded, compressed, and then decoded and restored. The training depends on the Python 3.11 and PyTorch 2.5.1 platforms and is completed based on the NVIDIA TITAN graphics card. The specific structure of the VQVAE model is as follows:
[0036] The VQVAE model is divided into an encoder structure and a decoder structure. The encoder consists of 3 downsampling blocks, each containing a 3×3×3 convolutional layer, a batch normalization layer, and a ReLU activation function. The number of channels is 64, 128, and 256 in sequence. Through three downsamplings, the 128×128×128 input image is compressed into a latent feature of 16×16×16. The decoder symmetrically contains 3 upsampling blocks, each containing a 3×3×3 transposed convolutional layer, a batch normalization layer, and a ReLU activation function. The number of channels is 256, 128, and 64 in sequence. The output end of the encoder is connected to a vector quantization layer with a codebook size of 1024 and a vector dimension of 256, which is used to further compress the image features.
[0037] 3) Prepare a diffusion model. Connect the encoder and decoder of the VQVAE model trained in 2) before the noise addition of the diffusion model and after the denoising ends respectively, to obtain a latent space diffusion model based on the latent space.
[0038] 4) Introduce a conditional encoder into the latent space diffusion model in 3) to extract the features of CT images and tumor masks, obtaining an improved conditional diffusion model.
[0039] 5) Input the PET images, CT images, and tumor masks in the dataset into the improved conditional diffusion model for training. Specifically, refer to Figure 2 the model schematic diagram shown, calculate the loss function, and use the gradient descent method to update the denoising network and the conditional encoder. Finally, obtain a trained model.
[0040] The specific process of inputting PET images into the improved conditional diffusion model for training is as follows:
[0041] The latent space diffusion model gradually adds noise to the PET image through the forward process until a pure noise sample is obtained. In the reverse denoising stage, a denoising network with a U-Net structure gradually restores the original structure and details of the image. By learning the mapping of the noise addition and denoising processes, the denoising network can effectively remove noise and retain the detailed information of the image. Finally, the denoised result will be input into the decoder of the VQVAE to restore to the original image space, generating high-quality medical images.
[0042] The specific process of inputting CT images and tumor masks into the improved conditional diffusion model for training is as follows:
[0043] In the diffusion model, CT images and tumor masks are introduced as conditions. After concatenating them by channel, they are input into the conditional encoder to obtain conditional encodings. The conditional encodings are concatenated with the output of the denoising network of the diffusion model by channel, and this process is repeated in each denoising step, thus enabling the provided CT and tumor masks to be fused into the PET image. The structure of the conditional encoder is similar to the encoder in VQVAE, and the input image is compressed from the image space of size 128 to the latent space of size 16 through three downsamplings. The specific mathematical formula of the loss function in step 5) is as follows:
[0044]
[0045] where ε t is the noise added at the t-th step in the noise addition stage of the diffusion model, ε θ is the denoising neural network in the diffusion model, ε is the standard Gaussian noise, E and E c are the encoder of VQVAE and the conditional encoder respectively, x0 is the PET image in the dataset, c represents the conditional image obtained by concatenating the CT image and the tumor mask in the dataset by channel, and the diffusion model is trained based on the above loss function.
[0046] 6) Input the tumor mask and CT image for testing into the model trained in 5), and output the synthesized PET image ( Figure 3 b). Compared with the result synthesized based on the GAN model ( Figure 4 a), the synthesized PET image ( Figure 3 b) by this method is closer to the distribution of the real PET image ( Figure 3 a). Compared with the synthesized image ( Figure 4 b) obtained without CT constraint, the synthesized image ( Figure 3 b) by this method has richer details and correct anatomical structures.
[0047] The above description of the embodiments is to facilitate those of ordinary skill in the art to understand and apply the present invention. Those familiar with the technology in the art can obviously make various modifications to the above embodiments easily, and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art according to the disclosure of the present invention should be within the protection scope of the present invention.
Claims
1. A medical image generation method based on a diffusion model, characterized in that, It includes the following steps: 1) Collect a batch of registered PET images with tumor lesions, CT images, and tumor masks as the dataset for training the model; 2) Use all the PET images in the dataset to train a VQVAE model containing an encoder and a decoder, so that the PET image is consistent with the original image after being encoded, compressed, and then decoded and restored; 3) Prepare a diffusion model, and connect the encoder and decoder of the VQVAE model trained in 2) before the noise addition of the diffusion model and after the denoising ends respectively, to obtain a latent space diffusion model based on the latent space; 4) Introduce a conditional encoder into the latent space diffusion model described in 3) to extract the features of the CT image and the tumor mask, and obtain an improved conditional diffusion model; 5) Input the PET images, CT images, and tumor masks in the dataset into the improved conditional diffusion model for training, calculate the loss function, and use the gradient descent method to update the denoising network and the conditional encoder, Finally, obtain a trained model; 6) Input the tumor mask and CT image used for testing into the model trained in 5) to output the synthesized PET image.
2. The medical image generation method based on a diffusion model according to claim 1, wherein The specific structure of the VQVAE model is as follows: The VQVAE model is divided into an encoder structure and a decoder structure. The encoder consists of 3 downsampling blocks, each of which contains a 3×3×3 convolutional layer, a batch normalization layer, and a ReLU activation function. The number of channels is 64, 128, and 256 in sequence. Through three downsamplings, the 128×128×128 input image is compressed into a latent feature of 16×16×16; The decoder symmetrically contains 3 upsampling blocks, each of which contains a 3×3×3 transposed convolutional layer, a batch normalization layer, and a ReLU activation function. The number of channels is 256, 128, and 64 in sequence. A vector quantization layer with a codebook size of 1024 and a vector dimension of 256 is connected to the output end of the encoder to further compress the image features.
3. The medical image generation method based on a diffusion model according to claim 1 or 2, characterized in that The specific process of training the PET image in step 5) by inputting it into the improved conditional diffusion model is as follows: The latent space diffusion model gradually adds noise to the PET image through the forward process until a pure noise sample is obtained. In the reverse denoising stage, a denoising network with a U-Net structure gradually restores the original structure and details of the image. By learning the mapping of the noise addition and denoising processes, the denoising network can effectively remove noise and retain the detail information of the image. Finally, the denoised result will be input into the decoder of the VQVAE to restore to the original image space and generate high-quality medical images.
4. The medical image generation method based on a diffusion model according to claim 3, wherein, The specific process of training the CT image and the tumor mask in step 5) by inputting them into the improved conditional diffusion model is as follows: Introduce CT images and tumor masks as conditions in the diffusion model. After concatenating them by channel, input them into the conditional encoder to obtain conditional encoding. The conditional encoding is concatenated with the output of the denoising network of the diffusion model by channel, and this process is repeated in each denoising step, then the provided CT and tumor masks can be fused into the PET image. The structure of the conditional encoder is similar to the encoder in VQVAE, and the input image is compressed from an image space of size 128 to a latent space of size 16 through three downsamplings.
5. The method for generating medical images based on a diffusion model according to claim 1 or 2 or 4, characterized in that The specific mathematical formula of the loss function in step 5) of claim 1 is: where ε t is the noise added at the t-th step of the denoising stage of the diffusion model, ε θ is the denoising neural network in the diffusion model, ε is the standard Gaussian noise, E and E c are the encoder and conditional encoder of the VQVAE respectively, x0 is the PET image in the dataset, c represents the conditional image obtained by concatenating the CT image and tumor mask in the dataset by channels, and the diffusion model is trained based on the above loss function.
Citation Information
Cited By
Condition-controllable image sample expansion method and system for surface defect detection
CN120525000A
Computed tomography (CT) image synthesis method and device for lung with tumor target region and storage medium
CN121304846A