Non-paired PET image enhancement method based on diffusion model
Through a two-stage learning framework, pseudo-paired data sets are generated and supervised training is carried out, cross-dose generalization problems of PET image quality improvement are solved, efficient PET image quality enhancement is achieved, and the generated pseudo-low-quality images are authentic and fidelity are good.
Patent Information
- Application Number
- CN202510533446.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-26
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art is difficult to effectively utilize unpaired PET image data for stable generalization across devices and doses, resulting in limited improvement in PET image quality and the direct use of diffusion model method to reason.
Using a two-stage learning framework, firstly, a pseudo-paired data set is generated through an unconditional diffusion model, followed by supervised training, using low-quality PET data sets to learn features, and gradually remove noise through the reverse diffusion process to generate pseudo-low-quality images paired with high-quality PET images.
The PET image quality improvement under different equipment and dose conditions is achieved, with strong generalization, reduced inference time, and the generated pseudo-low-quality images are authentic and fidelity.
Smart Images

Figure CN120471843A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning medical image processing, and in particular to an unpaired PET image enhancement method based on a diffusion model. Background Art
[0002] Positron emission tomography (PET) is a non-invasive functional imaging method that is widely used in oncology, neurology, and cardiology research. However, PET image quality is limited by the following factors:
[0003] Hardware performance limitations: Conventional scanners have poor spatial resolution and signal-to-noise ratio. Although new scanners (such as uEXPLORER) can obtain high-quality images, their high cost and maintenance expenses have led to low penetration.
[0004] Radiotracer dose and scan time limitations: In clinical practice, in order to reduce patient radiation exposure, the tracer dose is often reduced or the scan time is shortened, resulting in a significant increase in image noise and a decrease in quality.
[0005] Currently, the technical solutions for PET image enhancement mainly include the following two categories:
[0006] Supervised learning methods rely on paired low-quality and high-quality datasets to train models. However, obtaining paired data in clinical practice requires repeated scans, increasing the patient's radiation risk, and slight misalignments between paired images (due to patient / organ motion) can also adversely affect training.
[0007] Unsupervised / unpaired learning methods:
[0008] (1) Cycle consistency-based generative adversarial network method: Generate images through cycle consistency to achieve data mapping. However, the training stability of generative adversarial networks is poor and they are prone to mode collapse.
[0009] (2) Diffusion model-based methods: Use high-quality data to train an unconditional diffusion model, and use low-quality images as references when sampling to guide the model to generate corresponding high-quality results. The diffusion model has the characteristics of strong generation capabilities, stable training targets, and easy scalability, and the generation quality is better than the generative adversarial network. However, the diffusion model requires multiple iterative calculations when sampling, making it difficult to directly apply it in clinical practice. In addition, existing methods do not effectively use unpaired data for training, resulting in reduced generalization performance of the model in cross-device and cross-dose scenarios.
[0010] Therefore, how to utilize the unpaired low-quality and high-quality PET data that are widely available in clinical practice to achieve stable generalization across devices and doses is an urgent problem that needs to be solved in this field. Summary of the Invention
[0011] The purpose of the present invention is to address the problem that paired low-quality and high-quality PET images are difficult to obtain, and to provide an unpaired PET image enhancement method based on a diffusion model. The present invention relates to the research on the quality enhancement of PET images and is suitable for improving the quality of PET images in scenarios where there is a lack of real paired data.
[0012] The object of the present invention is achieved through the following technical solutions:
[0013] The present invention discloses an unpaired PET image enhancement method based on a diffusion model, comprising the following steps:
[0014] 1) Using a PET scanner to scan different patients injected with a standard dose of radiotracer to obtain high-quality PET datasets;
[0015] 2) Using PET scanners to scan different patients who were injected with substandard doses of radiotracer, resulting in low-quality PET datasets;
[0016] 3) Using the low-quality PET dataset obtained in 2) to train the unconditional diffusion model, so that it can learn the low-quality image features;
[0017] 4) Using the unconditional diffusion model obtained in 3) to perform conditional sampling, the high-quality PET images in the high-quality PET dataset are used as reference images. By gradually removing noise from the noisy images, pseudo low-quality PET images paired with the reference images are generated to construct a pseudo-paired PET dataset;
[0018] 5) Using the pseudo-paired PET dataset obtained in 4) for supervised training, the pseudo low-quality PET images are used as network input and the high-quality PET images are used as training labels. During training, the loss between the network output images and the training labels is calculated, the gradient is calculated through backpropagation, and the network parameters are updated using the gradient descent algorithm;
[0019] 6) Using the network obtained in 5), the low-quality PET image is used as the network input, and the PET image with improved quality is output.
[0020] As a further improvement, in step 3) of the present invention, the low-quality PET dataset obtained in step 2) is used to train the unconditional diffusion model as follows:
[0021] Set the noise variance coefficient β t , the coefficient increases with the time step 1~T, satisfying β1<β2<…<β T; For the input low-quality PET image x0, a time step t is randomly sampled from 1 to T, and random Gaussian noise is added to the image according to the pre-set noise variance coefficient to form a set of noise image-corresponding noise paired data; in the back diffusion process, the noise image and the corresponding time step t are used as the input of the noise prediction network, and the loss value between the output of the noise prediction network and the actual random noise added to the image is calculated. The gradient is calculated through back propagation, and the network parameters are updated using the gradient descent algorithm; after training, the noise prediction network can gradually remove the noise in the Gaussian noise image to generate a pseudo low-quality PET image.
[0022] As a further improvement, in step 4) of the present invention, conditional sampling is performed using the unconditional diffusion model obtained in step 3), and the high-quality PET image is gradually denoised to an intermediate state image, and reverse diffusion is performed starting from the intermediate state image instead of the pure Gaussian noise image:
[0023] By adding T0-step random Gaussian noise to the high-quality PET image (y0), an intermediate state image that only retains part of the structural information is obtained. Use it as the starting point for reverse diffusion
[0024]
[0025] in Indicates that given the initial high-quality PET image y0, the noisy image is obtained The conditional probability distribution of The variance is Gaussian distribution, where
[0026] As a further improvement, in step 4) of the present invention, the unconditional diffusion model obtained in step 3) is used for conditional sampling, and high-quality PET images are used for constraints in each step of the reverse diffusion process:
[0027] For the t-th step of back diffusion, the image and time step t is input into the noise prediction network ∈ θ , the noise of the output prediction Calculate the denoised image
[0028]
[0029] Where z~N(0,I) represents standard Gaussian noise, represents the noise intensity; The high-quality PET images are converted to the frequency domain by Fourier transform to obtain the corresponding frequency domain representation. The frequency domain representation of the high-quality PET image is subjected to high-pass filtering to extract the high-frequency component, the frequency domain representation of the high-quality PET image is subjected to low-pass filtering to extract the low-frequency component, the low-frequency component is added to the high-frequency component to form a composite frequency domain representation, and the composite frequency domain representation is converted back to the spatial domain by inverse Fourier transform to obtain the image
[0030] After the above back diffusion step is performed T0, a pseudo low-quality PET image paired with a high-quality PET image is obtained.
[0031] The beneficial effects of the present invention are:
[0032] This method utilizes a two-stage learning framework based on a diffusion model. In the first stage, the diffusion model is used to generate a pseudo-paired dataset, and in the second stage, supervised training is performed based on this pseudo-paired dataset. This framework overcomes the reliance on paired data, enabling training to be completed using unpaired data. This method can improve the quality of low-quality PET images from different patients, scanners, tracer doses, and scan times. It has strong generalization and wide applicability, addressing the difficulty of obtaining paired datasets, which hinders the use of supervised deep learning methods to improve PET image quality. By decoupling the two goals of data generation and quality enhancement, supervised deep learning methods can be flexibly adapted. Only the supervised deep learning network in the second stage is required for application, eliminating the time-consuming inference problem of directly using diffusion model methods.
[0033] This method uses an unconditional diffusion model to learn the characteristics of low-quality PET datasets. Leveraging the diffusion model's powerful generation capabilities and stable training, it achieves the goal of generating realistic pseudo-low-quality PET images. Backdiffusion sampling begins with an intermediate state image rather than a pure Gaussian noise image, reducing the number of backdiffusion steps and accelerating sampling. This overcomes the long sampling time associated with backdiffusion starting with a pure Gaussian noise image. By using high-quality PET images as constraints at each backdiffusion step, a balance is achieved between the realism and fidelity of the generated pseudo-low-quality PET images. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Schematic diagram of the process of the unpaired PET image quality enhancement method of the present invention;
[0035] Figure 2(a) shows a low-quality PET image;
[0036] Figure 2(b) shows a PET image processed using three-dimensional block matching denoising filtering;
[0037] FIG2( c ) is a PET image enhanced by the method of the present invention. DETAILED DESCRIPTION
[0038] In order to describe the present invention more specifically, the technical solution of the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0039] like Figure 1 As shown, the present invention is an unpaired PET image enhancement method based on a diffusion model, comprising the following steps:
[0040] 1) Using a PET scanner to scan different patients injected with a standard dose of radiotracer to obtain high-quality PET datasets;
[0041] 2) Using PET scanners to scan different patients who were injected with substandard doses of radiotracer, resulting in low-quality PET datasets;
[0042] 3) Using the low-quality PET dataset obtained in 2) to train the unconditional diffusion model, so that it can learn the low-quality image features;
[0043] 4) Using the unconditional diffusion model obtained in 3) to perform conditional sampling, the high-quality PET images in the high-quality PET dataset are used as reference images. By gradually removing noise from the noisy images, pseudo low-quality PET images paired with the reference images are generated to construct a pseudo-paired PET dataset;
[0044] 5) Using the pseudo-paired PET dataset obtained in 4) for supervised training, the pseudo low-quality PET images are used as network input and the high-quality PET images are used as training labels. During training, the loss between the network output images and the training labels is calculated, the gradient is calculated through backpropagation, and the network parameters are updated using the gradient descent algorithm;
[0045] 6) Using the network obtained in 5), the low-quality PET image is used as the network input, and the PET image with improved quality is output.
[0046] Furthermore, the step 3) includes the following sub-steps:
[0047] (3.1) Construct noise prediction network ∈ θ , using an improved UNet architecture, including a time-step embedding module and an attention mechanism;
[0048] (3.2) Set the noise variance coefficient β t , the coefficient increases with the time step 1~T, satisfying β1<β2<…<β T ;
[0049] (3.3) For the input low-quality PET image x0, a forward diffusion process is performed, a time step t is randomly sampled from 1 to T, and random Gaussian noise is added to the image according to the pre-set noise variance coefficient:
[0050]
[0051] Where: q(x t |x0) means that given the initial image x0, the noised image x is obtained. t The conditional probability distribution of The variance is Gaussian distribution, where α t =1-β t ;
[0052] (3.4) The noisy image x t and time step t are input into the noise prediction network, and the output prediction noise ∈ θ (x t ,t);
[0053] (3.5) Calculate the mean square error between the predicted noise and the true noise:
[0054]
[0055] (3.6) The gradient is calculated by back propagation and the noise prediction network parameters are updated using the gradient descent algorithm.
[0056] Furthermore, the step 4) includes the following sub-steps:
[0057] (4.1) Add T0 step noise to the high-quality PET image y0 to obtain an intermediate state image that only retains part of the structural information
[0058]
[0059] (4.2) The intermediate state image after adding noise As a starting point Perform back diffusion and use high-quality PET images for constraints in each step of back diffusion. For the t-th step of back diffusion:
[0060] (4.2.1) Image and time step t is input into the noise prediction network ∈ θ , the noise of the output prediction Calculate the denoised image
[0061]
[0062] Where: z~N(0,I), σ t represents the noise intensity,
[0063] (4.2.2) The high-quality PET images are converted to the frequency domain by Fourier transform to obtain the corresponding frequency domain.
[0064] Domain representation;
[0065] (4.2.3) Yes The frequency domain representation of the high-quality PET image is subjected to high-pass filtering to extract high-frequency components, and the frequency domain representation of the high-quality PET image is subjected to low-pass filtering to extract low-frequency components, wherein the cutoff frequencies of the high-pass filtering and the low-pass filtering are both D0;
[0066] (4.2.4) Add the low-frequency component and the high-frequency component described in (4.2.3) to form a composite frequency domain representation, and convert the composite frequency domain representation back to the spatial domain through inverse Fourier transform to obtain the image
[0067] (4.3) After the above back diffusion step is executed in step T0, a pseudo low-quality PET image paired with a high-quality PET image is obtained.
[0068] Implementation Examples
[0069] An embodiment of the present invention was implemented on a machine equipped with an Intel Core i9-10980XE CPU and an NVIDIA GeForce RTX 3090 GPU (24GB of video memory). High-quality PET datasets were obtained using PET images scanned with a total-body PET-CT uEXPLORER system (United Imaging Healthcare, Shanghai, China), with a scan time of 360 seconds and an image size of 3×256×256. Low-quality PET datasets were obtained using a Biograph 64Vision 600 PET / CT system (Siemens Healthineers, Erlangen, Germany), with a scan time of 60 seconds and an image size of 3×256×256. In the first phase, the diffusion model was trained for 75,000 steps, with the forward diffusion step number T0 and cutoff frequency D0 set to 150 and 10, respectively. In the second phase, the supervised network was trained for 30 rounds, resulting in the experimental results shown in Figure 2.
[0070] like Figure 2(a) to Figure 2(c) As shown in the figure, compared with the 3D block matching denoising filtering method, the proposed method can achieve quality enhancement of PET images, generate images with less noise, and better preserve tumor uptake.
[0071] The above description of the embodiments is intended to facilitate understanding and application of the present invention by those skilled in the art. It will be apparent that those skilled in the art can readily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without requiring inventive effort. Therefore, the present invention is not limited to the above embodiments, and improvements and modifications made by those skilled in the art based on the disclosure of the present invention should fall within the scope of protection of the present invention.
Claims
1. A method for enhancing an unpaired PET image based on a diffusion model, comprising the following steps: 1) Using a PET scanner to scan different patients injected with a standard dose of radiotracer to obtain high-quality PET datasets; 2) Using PET scanners to scan different patients who were injected with substandard doses of radiotracer, resulting in low-quality PET datasets; 3) Using the low-quality PET dataset obtained in 2) to train the unconditional diffusion model, so that it can learn the low-quality image features; 4) Using the unconditional diffusion model obtained in 3) to perform conditional sampling, the high-quality PET images in the high-quality PET dataset are used as reference images. By gradually removing noise from the noisy images, pseudo low-quality PET images paired with the reference images are generated to construct a pseudo-paired PET dataset; 5) Using the pseudo-paired PET dataset obtained in 4) for supervised training, the pseudo low-quality PET images are used as network input and the high-quality PET images are used as training labels. During training, the loss between the network output images and the training labels is calculated, the gradient is calculated through backpropagation, and the network parameters are updated using the gradient descent algorithm; 6) Using the network obtained in 5), the low-quality PET image is used as the network input, and the PET image with improved quality is output.
2. The unpaired PET image enhancement method based on a diffusion model according to claim 1, wherein: In step 3), the low-quality PET dataset obtained in step 2) is used to train the unconditional diffusion model as follows: Set the noise variance coefficient β t , the coefficient increases with the time step 1~T, satisfying β1<β2<…<β T For the input low-quality PET image x0, a time step t is randomly sampled from 1 to T, and random Gaussian noise is added to the image according to the pre-set noise variance coefficient to form a set of noise image-corresponding noise paired data; In the back-diffusion process, the noise image and the corresponding time step number t are used as the input of the noise prediction network. The loss value between the output of the noise prediction network and the actual random noise added to the image is calculated. The gradient is calculated through back-propagation, and the gradient descent algorithm is used to update the network parameters. After training, the noise prediction network can gradually remove the noise in the Gaussian noise image and generate a pseudo low-quality PET image.
3. The unpaired PET image enhancement method based on a diffusion model according to claim 1 or 2, characterized in that: In step 4), conditional sampling is performed using the unconditional diffusion model obtained in step 3), and the high-quality PET image is gradually denoised to an intermediate state image, and reverse diffusion is performed starting from the intermediate state image instead of the pure Gaussian noise image: By adding T0-step random Gaussian noise to the high-quality PET image (y0), an intermediate state image with only partial structural information is obtained. Use it as the starting point for reverse diffusion in Indicates that given the initial high-quality PET image y0, the noisy image is obtained The conditional probability distribution of The variance is Gaussian distribution, where 4. The unpaired PET image enhancement method based on a diffusion model according to claim 3, wherein: In step 4), the unconditional diffusion model obtained in step 3) is used for conditional sampling, and high-quality PET images are used for constraints in each step of the reverse diffusion process: For the t-th step of back diffusion, the image and time step t is input into the noise prediction network ∈ θ , the noise of the output prediction Calculate the denoised image Where z~N(0,I) represents standard Gaussian noise, represents the noise intensity; The high-quality PET images are converted to the frequency domain by Fourier transform to obtain the corresponding frequency domain representation. The frequency domain representation of the high-quality PET image is subjected to high-pass filtering to extract the high-frequency component, the frequency domain representation of the high-quality PET image is subjected to low-pass filtering to extract the low-frequency component, the low-frequency component is added to the high-frequency component to form a composite frequency domain representation, and the composite frequency domain representation is converted back to the spatial domain by inverse Fourier transform to obtain the image After the above back diffusion step is performed T0, a pseudo low-quality PET image paired with a high-quality PET image is obtained.
Citation Information
Cited By
Universal electron microscope image quality improvement method and device, and readable medium
CN122155977A