SPECT image reconstruction method and device based on VQGAN and generalized diffusion model
Through the unsupervised diffusion training model and CT image-guided method, combined with VQGAN and generalized diffusion model, the problem of noise and motion blur at low doses of SPECT images is solved, and the evaluation of high-quality cardiac functional parameters is achieved, which is suitable for a variety of scanning conditions and patients.
Patent Information
- Application Number
- CN202510552217.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-12
AI Technical Summary
Existing SPECT imaging reconstruction methods are difficult to effectively suppress noise and optimize cardiac motion blur after reducing radiation dose, which affects the accurate evaluation of cardiac functional parameters, and traditional methods rely on high-quality pairing data to obtain.
Unsupervised diffusion training model is used to combine VQGAN and generalized diffusion model, and hidden variables of cardiac-gated SPECT images are extracted through VQGAN, and reverse diffusion denoising is performed under the guidance of CT images. Unsupervised learning and CT image information assisted denoising to optimize image quality.
While reducing noise, reducing motion blur, improving the quality and diagnostic value of SPECT images, reducing dependence on high-quality paired data, adapting to different scanning conditions and patients, and meeting clinical real-time needs.
Smart Images

Figure CN120472027A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of medical image reconstruction and optimization, and in particular to a SPECT image reconstruction method and device based on VQGAN and a generalized diffusion model. Background Art
[0002] As a mature medical imaging technology, SPECT has been widely used in the early detection of coronary artery disease (CAD) and the assessment of myocardial infarction risk. Its working principle is to label specific physiological or pathological processes with radioactive tracers and capture the radioactive signals with a gamma camera, thereby generating functional images of organs and tissues in the body. Compared to other imaging technologies such as CT and MRI, SPECT can provide dynamic information on organ function and metabolic activity, and therefore has important clinical value in the diagnosis of cardiovascular disease, neurological diseases, and tumors.
[0003] Because SPECT imaging involves radioactive tracers, patients are exposed to a certain dose of radiation during the examination. Long-term or high-dose radiation exposure may increase the risk of cancer. Therefore, reducing the scan dose has become an important research direction to minimize radiation hazards while ensuring image quality.
[0004] The direct consequence of reducing the SPECT imaging dose is a reduction in the collected radioactive signal, resulting in a decrease in the image signal-to-noise ratio (SNR), manifested as an increase in artifacts, blurred edges, and loss of structural details. This quality degradation may affect doctors' identification and quantitative analysis of lesion areas, reducing the reliability of clinical diagnosis. Especially under low-dose conditions, traditional image reconstruction algorithms (such as filtered back projection (FBP) and maximum likelihood expectation maximization (MLEM)) may not be able to effectively suppress noise while maintaining key anatomical and functional information. Therefore, in order to maintain the high quality and diagnostic reliability of SPECT images at the lowest possible radiation dose, researchers are exploring advanced computer vision technologies and data-driven methods to optimize the quality of low-dose SPECT images. In recent years, deep learning technology has been introduced into SPECT image reconstruction and denoising tasks. Compared with traditional image filtering technology, deep learning methods can learn complex mapping relationships end-to-end, remove noise while maintaining key anatomical structures, and improve image quality.
[0005] However, cardiac gated SPECT images are not only affected by the high noise caused by low-dose acquisition, but also by the interference of cardiac motion, making it difficult for existing image reconstruction methods to optimize motion blur while suppressing noise, affecting the accurate assessment of cardiac function parameters. Summary of the Invention
[0006] Based on this, it is necessary to provide a SPECT image reconstruction method and device based on VQGAN and generalized diffusion model, which can improve image quality and enhance the accurate assessment of cardiac function parameters, in order to address the above technical problems.
[0007] In a first aspect, the present application provides a SPECT image reconstruction method based on VQGAN and a generalized diffusion model. The method comprises:
[0008] Train the unsupervised diffusion training model to obtain the target weight;
[0009] Obtain cardiac gated SPECT images and use VQGAN to extract features and obtain latent variables;
[0010] The latent variables are input into the back-diffusion inference model loaded with target weights, and the cardiac gated SPECT images are denoised under the guidance of CT images to obtain reconstructed images.
[0011] In one embodiment, denoising a cardiac gated SPECT image under the guidance of a CT image includes:
[0012] Construct a denoising network that includes CT image guidance information and combine it with the back-diffusion inference model for denoising.
[0013] The denoising probability of the denoising network is proportional to the product of the denoising probability of the back-diffusion inference model and the prior constraint of the CT image-guided information on the cardiac gated SPECT image.
[0014] In one embodiment, the back diffusion inference model includes a first-order neural network and a second-order neural network, the first-order neural network introduces a denoising network, and the second-order neural network introduces a denoising network and a time step adjustment module;
[0015] Among them, the time step adjustment module is used to adaptively adjust the time step of the second-order neural network according to the photon counting situation.
[0016] In one embodiment, the time step adjustment module includes a feature adjustment network and a feature extractor;
[0017] The feature adjustment network is used to estimate the optimization factor based on the Poisson noise of the current time step prediction and the initial input. The feature extraction is used to extract the time step features and adjust the time step in combination with the optimization factor.
[0018] In one embodiment, an attenuation coefficient degradation operator driven by patient information is obtained based on the photon attenuation mechanism of SPECT imaging; a restoration-degradation-restoration denoising process is completed using the first-order neural network, the degradation operator, and the second-order neural network; wherein the attenuation coefficient degradation operator is determined based on a set of arithmetic differences of photon intensities of cardiac-gated SPECT images, Poisson noise, noisy cardiac-gated SPECT images, and different training time steps, the set of arithmetic differences of photon intensities is determined based on the degree of photon attenuation, and the degree of photon attenuation is related to patient information including body fat percentage, lean body mass, and human cross-sectional thickness.
[0019] In one embodiment, the method further comprises:
[0020] Introducing a global attention module into the back-diffusion inference model.
[0021] In a second aspect, the present application also provides a SPECT image reconstruction device based on VQGAN and a generalized diffusion model. The device includes:
[0022] The weight acquisition module is used to train the unsupervised diffusion training model and obtain the target weight;
[0023] The feature extraction module is used to obtain cardiac gated SPECT images and use VQGAN to extract features and obtain latent variables;
[0024] The reconstruction module is used to input latent variables into the back diffusion inference model loaded with target weights, denoise the cardiac gated SPECT images under the guidance of CT images, and obtain reconstructed images.
[0025] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the above-mentioned SPECT image reconstruction method based on VQGAN and generalized diffusion model.
[0026] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned SPECT image reconstruction method based on VQGAN and the generalized diffusion model.
[0027] In a fifth aspect, the present application further provides a computer program product, comprising a computer program that, when executed by a processor, implements the steps in the above-mentioned SPECT image reconstruction method based on VQGAN and a generalized diffusion model.
[0028] The above-mentioned SPECT image reconstruction method and device based on VQGAN and generalized diffusion model trains an unsupervised diffusion training model to obtain target weights; obtains cardiac gated SPECT images, and uses VQGAN to extract features to obtain latent variables; inputs the latent variables into the reverse diffusion inference model loaded with target weights, and denoises the cardiac gated SPECT images under the guidance of CT images to obtain reconstructed images. This application breaks through the traditional supervised learning's dependence on high-quality matching image data, uses an unsupervised diffusion training model and a reverse diffusion inference model to perform reverse diffusion inference, reconstructs images in the latent space of VQGAN, and introduces CT image information to assist in denoising. Ultimately, this method can effectively reduce the noise interference of cardiac gated SPECT images, while reducing motion blur, and improving the quality and diagnostic value of static SPECT images. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 1 is a flow chart of a SPECT image reconstruction method based on VQGAN and a generalized diffusion model in one embodiment;
[0030] Figure 2 A schematic diagram of the structure of VQGAN implementation in one embodiment;
[0031] Figure 3 1 is a schematic diagram of the overall forward propagation and back propagation models in an unsupervised diffusion training model in one embodiment;
[0032] Figure 4 A simplified framework diagram of a back propagation model in an unsupervised diffusion training model in one embodiment;
[0033] Figure 5 Schematic diagram of the structure of a reverse diffusion inference model in one embodiment;
[0034] Figure 6 A schematic diagram of an implementation of a global attention module in one embodiment;
[0035] Figure 7 Schematic diagram of the time step adjustment module for photon counting adaptation;
[0036] Figure 8 FIG. 1 is a flow chart of a SPECT image reconstruction method using a three-dimensional stepwise mapping error-optimized diffusion model in one embodiment. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0038] In traditional supervised learning methods, model training usually relies on a large number of paired high-quality clean images and noisy images. However, in the actual 3D cardiac gated SPECT image acquisition process, due to the limitations of scanning conditions and factors such as radioactive tracer decay, it is often very difficult and costly to obtain high-quality static SPECT images as "clean" references. In addition, existing deep learning denoising methods (such as CNN, U-Net and GAN) and classical diffusion models still have limitations in the denoising task of cardiac gated SPECT images, such as the difficulty in modeling the complex noise distribution of SPECT images, and the denoising process may cause loss of image details or structural distortion. The present invention proposes a 3D cardiac gated SPECT image intelligent reconstruction method that combines VQGAN and CT image guidance to overcome the above challenges.
[0039] In one embodiment, as Figure 1 As shown in Figure 2, the SPECT image reconstruction method based on VQGAN and generalized diffusion model includes the following steps:
[0040] Step 102: Train an unsupervised diffusion training model to obtain target weights.
[0041] Unsupervised diffusion training models involve two steps: forward diffusion and backward denoising. The forward diffusion process, also known as forward propagation, gradually adds noise to data (such as an image), gradually converting the data into pure noise. The backward denoising process, also known as backpropagation, starts with pure Gaussian noise and gradually removes it, ultimately restoring the original data. This process is implemented using a neural network (typically a U-Net). The input of the neural network is the noisy data and the number of noise addition steps, and the output is the predicted noise.
[0042] In the forward diffusion process, only the original unlabeled data samples are needed for the model to learn the distribution structure of the data. In the reverse denoising process, by continuously optimizing the denoising network, the generated data can be close to the distribution of the real data without the need for additional label information. Therefore, the unsupervised diffusion training model does not require labeled data for training. Traditional supervised learning methods require high-quality full-dose images and low-dose images to be paired as training data, but it is difficult to obtain high-quality paired data for cardiac gated SPECT images due to factors such as scanning conditions and radioactive tracer decay. The present invention adopts an unsupervised strategy and only uses cardiac gated SPECT images for training, without the need for additional high-quality image pairing, which greatly reduces the difficulty of data preparation. In addition, since the unsupervised strategy does not need to strictly rely on paired data, the model can adaptively learn the characteristics of cardiac gated SPECT images under different imaging conditions, thereby having stronger generalization capabilities. At the same time, unsupervised learning can quickly complete model training when data accumulation is limited, is suitable for different patients and scanning conditions, and meets clinical needs for real-time image denoising and reconstruction.
[0043] In this embodiment, the training data for the unsupervised diffusion training model is collected from 3D cardiac gated SPECT images. After being encoded using VQGAN (Vector Quantized Generative Adversarial Network), the high-dimensional 3D image data is converted into a compact discrete latent representation. The VQGAN encoding process not only retains key structural and texture information, but also effectively reduces redundant features, improving the model's computational efficiency and reconstruction quality. The discrete latent representation extracted by VQGAN helps the diffusion model learn in a lower-dimensional latent space, reducing the computational complexity of directly modeling the original high-dimensional images.
[0044] Next, during the forward propagation of the unsupervised diffusion training model, the random perturbation mechanism of the diffusion process is utilized to gradually add noise to the latent representation of the cardiac gated SPECT images extracted by the VQGAN, causing them to gradually evolve into images with a random Poisson noise distribution. During this process, the model gradually transforms the input data from a known prior distribution (i.e., the distribution of cardiac gated SPECT images) to an unknown random distribution through multiple steps of random noise perturbation, thereby changing the statistical characteristics of the image data. Ultimately, after multiple forward diffusion steps, the characteristic information of the cardiac gated SPECT images is fully randomized, making it close to a pure noise distribution. The core goal of this process is to learn the reverse diffusion path from noise regression to the true image. This allows for intelligent reconstruction of cardiac gated SPECT images through unsupervised inversion during the denoising phase, ensuring the restoration of image details while suppressing artifacts and quality degradation caused by noise. Furthermore, the motion blur effect of the image is further optimized, so that the final reconstructed image can both reduce the noise level and maintain the consistency of the dynamic structure of the heart.
[0045] During the training process of an unsupervised diffusion training model, the model continuously adjusts its internal parameters (i.e., weights) to minimize a predefined loss function. For example, in diffusion models, the mean squared error loss function is often used to measure the difference between the noise predicted by the model and the actual noise added. In each round of training, the model updates the weights based on the gradient of the loss function using an optimization algorithm (such as the Adam optimizer). As training progresses, the performance of the model continues to change. At some point, the model's loss function on the validation set (or training set) reaches a minimum, and the corresponding model weights are considered to be the optimal weights, or target weights. These target weights represent the model's current best state in learning the data distribution and are able to most accurately predict the noise added to the data.
[0046] Step 104: Obtain cardiac gated SPECT images and use VQGAN to extract features to obtain latent variables.
[0047] The cardiac gated SPECT image in this step is the 3D cardiac gated SPECT image to be denoised, and VQGAN is also used for feature extraction to obtain latent variables.
[0048] Steps 102 and 104 are both feature extracted by VQGAN. VQGAN can encode high-dimensional cardiac gated SPECT images into compact discrete latent representations, retaining key structural information while reducing computational redundancy and making the diffusion process more efficient. Principle of improving computational efficiency: Assuming that the size of the original 3D SPECT image is H (height) × W (width) × D (depth), the latent variable z extracted by VQGAN is e The dimensions are: Where f is the downsampling factor of VQGAN. During the calculation process, the computational complexity will drop by about f 3 times, significantly reducing training time.
[0049] In addition, VQGAN improves the denoising ability of the diffusion model in low-dimensional space by learning more expressive latent representations, so that cardiac gated SPECT images can still retain rich anatomical details and texture information after denoising, enhancing the ability to restore image details.
[0050] In clinical applications, by performing diffusion modeling in the latent space of VQGAN, compared with directly operating in the original high-dimensional image space, it can effectively reduce computational overhead, improve the model's inference speed, and meet clinical application needs.
[0051] Step 106 : Input the latent variables into the back diffusion inference model loaded with target weights, and perform denoising on the cardiac gated SPECT image under the guidance of the CT image to obtain a reconstructed image.
[0052] Integrating the VQGAN backdiffusion inference model, the 3D cardiac gated SPECT image is used as the endpoint of the diffusion process. Leveraging the decoding capabilities of VQGAN, the latent representation learned during training is used as the initial input for reconstruction in a low-dimensional latent space, preserving key structural information and reducing computational complexity. Subsequently, the backdiffusion inference model is applied to the latent representation provided by VQGAN for progressive denoising. Using the backdiffusion path learned in step 102, cardiac gated SPECT images with high-quality texture and structural information are gradually recovered from the noisy state. Furthermore, a CT image-guided mechanism is introduced, enabling the model to simultaneously learn anatomical information about cardiac structure while denoising. This helps more accurately restore SPECT image details, improve signal consistency across different cardiac regions, avoid structural distortion or loss of key features during the denoising process, and predict the corresponding static SPECT image. By adaptively adjusting the balance between motion blur and noise, the resulting static SPECT image has a superior noise level to the original gated image, while exhibiting less motion blur than traditional static SPECT images, thereby improving overall image quality.
[0053] This embodiment uses random Poisson noise images and cardiac gated SPECT images as the endpoints of the diffusion process in the unsupervised diffusion training model and the back-diffusion inference model, respectively. This breaks away from the traditional diffusion model framework and expands to a generalized diffusion model. This significantly reduces training sampling time while enriching image feature information and improving the quality of image detail recovery.
[0054] The introduced CT images can provide high-resolution anatomical information, making up for the lack of structural details in SPECT images and helping the model to more accurately restore key features of the heart area. In addition, by fusing CT image information, the model can ensure the consistency of signal intensity in different areas of the heart during denoising, avoiding structural distortion or loss of key information during the denoising process. The CT guidance mechanism can also assist the model in further optimizing cardiac motion blur after denoising, so that the final static SPECT image has better noise level than the original gated image, and the degree of motion blur is lower than that of traditional static SPECT images, thereby improving the overall image quality.
[0055] The feature extraction encoder extracts multi-scale spatial structural features from the CT image, and then obtains high-level semantic features from the CT image by stacking convolutional layers, normalization layers (GroupNorm), and nonlinear activation functions (ReLu). This includes information from shallow textures to deep anatomical structures, which serves as guiding information and is represented by I CT I CT Constraining the denoising of cardiac gated SPECT images in the denoising network is specifically implemented with I CT As a modulation factor, it guides the self-attention calculation within the SPECT image features, so that key anatomical information can be retained during the feature interaction process to make up for the lack of structural details in SPECT images.
[0056] In this embodiment, the efficient latent variable extraction capability of VQGAN and the denoising characteristics of the generalized diffusion model are combined to achieve intelligent reconstruction of cardiac gated SPECT images under an unsupervised learning framework. VQGAN is used to extract three-dimensional latent feature representations to optimize the structural integrity and texture details of 3D cardiac gated SPECT images. At the same time, CT images are introduced as structural prior guidance to provide stable anatomical structure constraints in the back propagation process of the diffusion model to improve the anatomical consistency and noise suppression capabilities of the reconstructed images. In addition, the present invention uses static SPECT images as target images to denoise cardiac gated SPECT images to optimize the motion blur problem while reducing noise. Although traditional static SPECT images have low noise, they often have severe motion blur due to the influence of heart beating. Cardiac gated SPECT images can alleviate the motion blur problem, but the noise is relatively large under low-dose acquisition conditions. The method of the present invention uses the adaptive denoising mechanism of the diffusion model to suppress noise while adaptively controlling the degree of blur enhancement, so that the final reconstructed cardiac gated SPECT image is significantly better than the original gated image in terms of noise level, and at the same time has lower motion blur than the traditional static SPECT image, thereby improving image quality and enhancing the accurate assessment of cardiac function parameters.
[0057] In one embodiment, the backdiffusion inference model includes a first-order neural network and a second-order neural network. The first-order neural network incorporates a denoising network, and the second-order neural network incorporates a denoising network and a time step adjustment module. The time step adjustment module is configured to adaptively adjust the time step of the second-order neural network based on photon counting.
[0058] In one embodiment, a back-diffusion inference model is constructed that integrates an attenuation coefficient degradation operator driven by patient information and CT guidance information.
[0059] The degradation operator is represented by D(), and its construction is based on the photon attenuation mechanism of SPECT imaging. According to the exponential decay formula: I = I0e -μd , where μ is the attenuation coefficient, which is obtained from the patient's body fat percentage and lean body mass, d is the cross-sectional thickness of the human body, I0 is the initial intensity of the photon source, and I is the intensity after attenuation. The process of adding noise based on the patient's attenuation coefficient degradation operator can be expressed as: D(x q ,t)=(1-I t )x q +I t σ t =x t , where x q The cardiac gated SPECT image extracted by VQGAN. The feature extraction process of VQGAN is similar to the unsupervised forward propagation process. q =E(x0), x0 is the original cardiac gated SPECT image, I t is an arithmetic set of photon intensities at different training time steps, and gradually increases with the time step t, σ t is the noise sampled from the standard Poisson distribution, x t Poisson noise σ t After the image.
[0060] By rationally combining patient information (such as body fat percentage, lean body mass, and heart cross-sectional thickness), a personalized attenuation coefficient degradation operator is constructed to make the degradation process of SPECT images more consistent with the physical attenuation mechanism, while enhancing the model's adaptability to different patient data.
[0061] CT image-guided related content introduction: t-1 =τ θ (D(x q ,t),t,I CT ), where τ θ () is the denoising network guided by CT images. CT is the guidance information of CT image. θ (x t-1 |x t ,I CT )∝pθ (x t |x t-1 )P CT (x t ,I CT ), where p θ (x t |x t-1 ) is the denoising probability of the diffusion model itself, P CT (z t ,I CT ) is the prior constraint of CT images on cardiac gated SPECT images.
[0062] This embodiment upgrades the original first-order process of the generalized diffusion model to a second-order process, and the parameterized neural network algorithm is:
[0063]
[0064] Where θ and ω are the recovery network parameters and step-size adjustment network parameters under the guidance of CT images during the reverse diffusion process, x0 is the gold standard data, The image after denoising in the first-order process. The second-order process combines the annealing sampling strategy "restore-degrade-restore". The first-stage process adds I on the basis of the classic diffusion model. CT The imaging information guided, two-stage process combines the dose level dependent time step adjustment strategy and I CT , it can alleviate the problem of step-size cumulative error as much as possible and provide high-resolution anatomical structure information, thereby improving the intelligent reconstruction capability of the overall network model.
[0065] The time step adjustment module includes a feature adjustment network and a feature extractor. Designing the feature adjustment network It is predicted based on the current time step and add noise σ to the image t To estimate the optimization factor ρ t-1 ,γ t-1 The feature extractor is then used to extract the feature information of the time step, and the feature information is adjusted and mapped in combination with the optimization factor to optimize the dynamic adaptability of the time step.
[0066] This embodiment uses a feature adjustment network and combines photon counting information to optimize the time step adjustment in the diffusion process, so that the time step can adapt to the statistical characteristics of photons under different dose conditions, thereby improving the denoising stability and reconstruction accuracy of the model.
[0067] In one embodiment, the schematic diagram of VQGAN implementation is as follows Figure 2As shown in Figure 2, the VQGAN model is used to encode, quantize, and reconstruct cardiac gated SPECT images to extract compact low-dimensional latent representations. First, the input SPECT image is extracted through the encoder (E) and mapped to the codebook (Z) for vector quantization, thereby converting the continuous features into discrete representations z. q The decoder (D) passes z q The image is reconstructed and adversarial training is performed using the discriminator (G) to improve the authenticity of the generated image. In addition, the Transformer model further learns the sequential dependencies of the latent features to optimize the reconstruction quality. VQGAN constrains the encoding through reconstruction loss and commitment loss, enabling it to efficiently compress image information while maintaining key structures, ultimately providing high-quality latent representations for denoising and reconstruction of the diffusion model. The formula is expressed as: q =E(z); where E() is the encoder of VQGAN, z0=D(z q ); where D() is the decoder of VQGAN, the process needs to ensure The commitment loss is used to constrain the quantization error of discrete latent variables, z is the original SPECT image, and z q is the encoder output of VQGAN, z0 is the decoder output of VQGAN, and LVQ is the vector quantization loss VectorQuantizationLoss. The formula for the forward propagation process of the unsupervised training module is: t =α t z q +(1-α t )∈,∈~N(0,I), where α t is the cumulative noise attenuation coefficient, ∈~N(0,I) is the standard Gaussian noise (which can be extended to Poisson noise), z t Close to pure noise distribution.
[0068] In one embodiment, a simplified schematic diagram of the forward propagation and back propagation models of the unsupervised diffusion training model is as follows: Figure 3 In the forward diffusion stage, starting from the original cardiac gated SPECT image x0, as time step t advances, noise is gradually added to the image, gradually becoming a noisy image x T In the reverse diffusion stage, the noise image x T Initially, the model uses the learning parameter θ to follow the probability distribution p θ (x t-1 |x t ) gradually removes the noise and finally reconstructs a cardiac gated SPECT image x0 close to the original one.
[0069] In one embodiment, the simplified framework of the 3D unsupervised diffusion training model back propagation model is as follows: Figure 4As shown in the figure, the random Poisson noise image is gradually denoised and finally reconstructed into a static SPECT image intelligently. It includes the GAM module, the photon counting adaptive time step adjustment module and the I CT (CT image information guidance).
[0070] In one embodiment, the back diffusion inference model is as follows Figure 5 As shown. Taking x in a high noise state T Cardiac gated SPECT images are used as the starting point, replacing the end point of the traditional diffusion process, and introducing I CT The guidance information of the image is used to assist in processing. The reasoning process is divided into two stages. In Stage I, the input x T It is first combined with SinPE(t), and then processed by the multi-layer perceptron MLP. The global image features are captured by the global attention module GAM. After passing through the error-modulated module (EMM), the EMM module can use the latest prediction and the given cardiac gated SPECT image calibration time step to embed features. The model uses the ConvBlock convolution block to extract features, downsamples to obtain higher-level features, and upsamples to restore resolution. The SwitchModule switching module controls feature fusion, and finally obtains an intermediate image with preliminary denoising and feature restoration. Entering Stage II, Combined with SinPE(t-1), it goes through MLP, GAM and EMM modules again. In this stage, the photon counting adaptive time step adjustment module is specially introduced. The time step is adaptively adjusted according to the photon counting situation to optimize the back diffusion.
[0071] In one embodiment, the GAM module schematic diagram is as follows Figure 6 As shown. GAM is used to process the input feature F1. First, the channel attention mechanism M c , analyze the importance of different channel features, highlight key channels, and suppress unimportant channels. Then use the spatial attention mechanism M s , determining the importance of different spatial locations in the feature map and focusing on key areas. The features processed by channel attention are multiplied and fused with the original input features F1, and the fused features are convolved with the features processed by spatial attention to ultimately generate the output features F1. This output feature combines important information from both channel and spatial perspectives, better facilitating subsequent model tasks.
[0072] In one embodiment, the photon counting adaptive time step adjustment module is as follows: Figure 7 As shown. First, the optimization factor is extracted by feature adjustment network where φθ is a feature adjustment network, which is based on the prediction of the current time step and the Poisson noise of the initial input σ t To estimate the optimization factor ρ t-1 and γ t-1 The time step features are then extracted through the feature extractor, and the time step features are adjusted and mapped in combination with the optimization factor. The algorithm is expressed as:
[0073]
[0074] Where SinPE(t-1) represents the sinusoidal position encoding of time step t-1. FeatureExtractor() is the feature extractor, s t-1 is the time step feature, is the time step feature after adaptive optimization.
[0075] In one embodiment, the trained sum weights are placed in a reverse diffusion inference model, and the 3D cardiac gated SPECT image is used as the end point of the diffusion process. The reverse diffusion process is directly performed on it to predict the static SPECT image after noise reduction. The predicted image effectively reduces the motion blur caused by cardiac motion while reducing noise. This embodiment compares the predicted image with the full-dose static image to evaluate the prediction effect. The predicted effect is evaluated using the normalized mean square error (NMSE), structural similarity index (SSIM) and peak signal-to-noise ratio (PSNR) to evaluate the predicted cardiac gate SPECT image.
[0076] In one embodiment, the complete process of the 3D cardiac gated SPECT image intelligent reconstruction method is as follows: Figure 8 This method leverages the potential representation capabilities of VQGAN and the efficient denoising capabilities of the generalized diffusion model to achieve denoising reconstruction of cardiac gated SPECT images, ultimately generating high-quality static SPECT images that reduce noise and motion blur, improving image clarity and diagnostic value.
[0077] First, the original 3D TIF-format cardiac gated SPECT images were converted to the npy data format to facilitate training and processing of the deep learning model. During the data conversion process, a series of preprocessing operations were performed to ensure image data quality and consistency, including normalization, image intensity adjustment, and histogram matching. This ensured a more consistent grayscale distribution and reduced variations due to differences in imaging equipment or scanning conditions. Subsequently, the preprocessed npy-format data was fed into a VQGAN for feature extraction. VQGAN encodes high-dimensional cardiac gated SPECT images into a compact, low-dimensional discrete latent representation, reducing data redundancy, preserving key image features, and improving model stability. The image data were then fed into an unsupervised diffusion training model for training. Due to its unsupervised approach, the model does not rely on high-quality full-dose-low-dose paired data. Instead, it independently learns the noise distribution and structural characteristics of cardiac gated SPECT images, continuously optimizing its denoising capabilities. After training, the model saves the optimal weights for use in the back-diffusion inference phase. The 3D cardiac gated SPECT image to be denoised is used as the starting point for backdiffusion, encoded again using a VQGAN, and then fed into the backdiffusion model for denoising, guided by the CT image. Backdiffusion denoising utilizes the inverse process of the diffusion model to gradually remove noise from the image, making the resulting image closer to a true static SPECT image. Finally, the reconstruction effect is comprehensively evaluated by comparing it with a true static SPECT image using objective evaluation metrics such as SSIM (structural similarity index), PSNR (peak signal-to-noise ratio), NMSE (normalized mean square error), and motion blur. These metrics measure the similarity and quality difference between the reconstructed image and the true static SPECT image from different perspectives, accurately reflecting the reconstruction performance of the proposed algorithm.
[0078] The modules in the present invention work together to achieve the effect of improving image quality. First, the introduction of the unsupervised strategy enables the generalized diffusion model to be trained without paired data, and the generalized diffusion model improves the accuracy of image denoising through the physically driven attenuation operator and adaptive adjustment of the time step. This combination enables the model to achieve excellent noise reduction and reconstruction performance under limited data conditions. Secondly, VQGAN provides a low-dimensional and compact latent space, which makes the diffusion process more efficient and enhances the ability to restore image details. The diffusion model further performs noise modeling in this latent space to improve the denoising ability and generation quality of the image. Finally, the structural information of the CT image guides the diffusion model to maintain the anatomical consistency of the heart area during the denoising process, ensuring that the generated static SPECT image reduces motion blur while reducing noise, thereby improving image quality and clinical applicability.
[0079] Therefore, the present invention breaks through the limitations of traditional SPECT image denoising methods through the synergistic effect of unsupervised strategies, generalized diffusion models, VQGAN and CT image guidance. It can efficiently remove noise from cardiac gated SPECT images without relying on high-quality paired data, and predict static SPECT images with higher clarity and less motion blur, thereby improving the diagnostic value of SPECT images.
[0080] While cardiac gated SPECT images provide information about cardiac function, they also face problems such as increased noise, decreased signal-to-noise ratio, blurred details, and impaired image quality due to low-dose acquisition conditions, which seriously affect the accuracy of clinical diagnosis. In response to the current technical bottlenecks in the field of cardiac gated SPECT image reconstruction, the present invention focuses on solving the following five key problems: ① The inability to obtain high-quality paired training data limits supervised learning methods. Due to the particularity of SPECT imaging, it is difficult to simultaneously obtain low-noise and high-quality image pairs from the same patient during the same scan. The large-scale paired data that traditional supervised deep learning methods rely on is difficult to construct, which limits model training. ② Traditional methods have limited noise reduction capabilities and can easily lead to over-smoothing of images or loss of details. Although classic filtering methods (such as Gaussian filtering, BM3D, non-local means, etc.) can reduce noise to a certain extent, they often lose image details, resulting in over-smoothing of reconstructed SPECT images that are significantly different from high-quality images. In addition, although deep learning methods based on CNN or GAN can learn complex image distributions, CNN is limited by the local receptive field and has difficulty capturing global information, while GAN training is unstable and prone to artifacts. ③ Cardiac-gated SPECT images lack stable structural prior information, resulting in insufficient anatomical consistency in the reconstructed images. SPECT images have strong functional imaging features but weak structural information. Relying solely on SPECT images for denoising or reconstruction may result in anatomical distortion, affecting clinical usability. ④ Existing diffusion models have high computational costs and long inference times, making them difficult to directly apply to 3D cardiac-gated SPECT image reconstruction. The classic diffusion denoising probabilistic model (DDPM) uses a fixed-step diffusion strategy and only performs denoising using random Gaussian noise, resulting in a lengthy sampling process and slow inference speed, making it difficult to meet the needs of actual clinical applications. Furthermore, traditional diffusion models are primarily used for 2D image processing and are difficult to directly apply to 3D SPECT images. ⑤ Although traditional static SPECT images have low noise, they often suffer from severe motion blur due to the influence of heartbeat. Cardiac-gated SPECT images can alleviate the motion blur problem, but they are noisy under low-dose acquisition conditions.
[0081] To address the above issues, the present invention (1) achieves direct intelligent reconstruction of cardiac gated SPECT images based on an unsupervised learning strategy. This invention employs an unsupervised learning approach, enabling training without the need for matched image data pairs. Through the forward denoising and backward denoising processes of the diffusion model, the intrinsic characteristics of cardiac gated SPECT images are learned, noise is removed while preserving image details, enabling direct intelligent reconstruction of cardiac gated SPECT images. This unsupervised learning strategy not only avoids the reliance of traditional supervised methods on high-quality and low-quality image pairing data, but also improves the robustness of the model, reduces data acquisition costs, and meets the needs of clinical applications. (2) VQGAN is combined with VQGAN for efficient latent variable extraction, optimizing image details and structural information. VQGAN effectively extracts deep features of images through discrete latent variable learning, enabling the diffusion model to perform denoising in a lower-dimensional latent space, reducing computational complexity and improving inference efficiency. Furthermore, VQGAN can better preserve the local and global structure of images during training, avoiding the blurring problem that is prone to traditional generative models, thereby enhancing the clarity and visualization of reconstructed images. (3) CT images are introduced as structural priors to improve the anatomical consistency of reconstructed images. The present invention utilizes CT images to provide stable anatomical constraints. During the back-propagation process of the diffusion model, CT guidance is used to enhance the anatomical consistency of the SPECT images, ensuring that the reconstruction results conform to the structural characteristics of medical images. CT prior information helps optimize the edge clarity of SPECT images, reduce artifacts, and improve the diagnosability of the images. ④ Directly operate on 3D cardiac-gated SPECT images to achieve high-quality image reconstruction in three-dimensional space. Traditional methods mainly process 2D slices or local areas, making it difficult to ensure the integrity of three-dimensional structures. The present invention directly performs VQGAN encoding, diffusion process modeling, and CT prior guidance on 3D cardiac-gated SPECT data, resulting in reconstructed images with better structural consistency and global information capture capabilities in three-dimensional space. ⑤ The present invention uses static SPECT images as target images and denoises cardiac-gated SPECT images to reduce noise while optimizing motion blur. Although traditional static SPECT images have low noise, they often suffer from severe motion blur due to the influence of heartbeats. Cardiac-gated SPECT images can alleviate the motion blur problem, but the noise is relatively high under low-dose acquisition conditions. The present invention uses an adaptive denoising mechanism of a diffusion model to suppress noise while adaptively controlling the degree of blur enhancement, so that the final reconstructed cardiac gated SPECT image has significantly better noise level than the original gated image and lower motion blur than traditional static SPECT images, thereby improving image quality and enhancing the accurate assessment of cardiac function parameters.
[0082] In summary, this method breaks through the traditional supervised learning's reliance on high-quality matching image data. It utilizes an unsupervised diffusion training framework and a backdiffusion inference model to perform image reconstruction in the VQGAN latent space and incorporates CT image information to assist in denoising. Ultimately, this method effectively reduces noise interference in cardiac gated SPECT images, while also minimizing motion blur, improving the quality and diagnostic value of static SPECT images.
[0083] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0084] Based on the same inventive concept, embodiments of the present application also provide a VQGAN and generalized diffusion model-based SPECT image reconstruction device for implementing the aforementioned VQGAN and generalized diffusion model-based SPECT image reconstruction method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the VQGAN and generalized diffusion model-based SPECT image reconstruction device provided below can be found in the above-mentioned limitations of the VQGAN and generalized diffusion model-based SPECT image reconstruction method, and will not be repeated here.
[0085] In one embodiment, a SPECT image reconstruction device based on VQGAN and a generalized diffusion model is provided, comprising: a weight acquisition module for training an unsupervised diffusion training model to obtain target weights;
[0086] The feature extraction module is used to obtain cardiac gated SPECT images and use VQGAN to extract features and obtain latent variables;
[0087] The reconstruction module is used to input latent variables into the back diffusion inference model loaded with target weights, denoise the cardiac gated SPECT images under the guidance of CT images, and obtain reconstructed images.
[0088] Each module in the aforementioned VQGAN and generalized diffusion model-based SPECT image reconstruction device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0089] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in all the above method embodiments when executing the computer program.
[0090] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in all the above method embodiments are implemented.
[0091] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in all the above method embodiments when executed by a processor.
[0092] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0093] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, etc., but are not limited to these.
[0094] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0095] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A SPECT image reconstruction method based on VQGAN and generalized diffusion model, characterized in that: The method comprises: Train the unsupervised diffusion training model to obtain the target weight; Obtain cardiac gated SPECT images and use VQGAN to extract features and obtain latent variables; The latent variables are input into a back-diffusion inference model loaded with the target weights, and the cardiac gated SPECT image is denoised under the guidance of the CT image to obtain a reconstructed image.
2. The method according to claim 1, characterized in that Denoising the cardiac gated SPECT image under the guidance of the CT image includes: Constructing a denoising network including CT image guidance information, and combining the denoising network with the back diffusion inference model to perform denoising; The denoising probability of the denoising network is proportional to the product of the denoising probability of the back diffusion inference model and the prior constraint of the CT image guidance information on the cardiac gated SPECT image.
3. The method according to claim 2, characterized in that The back diffusion inference model includes a first-order neural network and a second-order neural network, the first-order neural network introduces the denoising network, and the second-order neural network introduces the denoising network and a time step adjustment module; The time step adjustment module is used to adaptively adjust the time step of the second-order neural network according to the photon counting situation.
4. The method according to claim 3, characterized in that The time step adjustment module includes a feature adjustment network and a feature extractor; The feature adjustment network is used to estimate the optimization factor based on the Poisson noise predicted by the current time step and the initial input, and the feature extraction is used to extract the time step feature and adjust the time step in combination with the optimization factor.
5. The method according to claim 3, characterized in that The method further comprises: An attenuation coefficient degradation operator driven by patient information is obtained based on the photon attenuation mechanism of SPECT imaging; Using the first-order neural network, the degradation operator, and the second-order neural network to complete a restoration-degradation-restoration denoising process; The attenuation coefficient degradation operator is determined based on the cardiac gated SPECT image, Poisson noise and a set of photon intensities with different training time steps. The set of photon intensities is determined based on the degree of photon attenuation. The degree of photon attenuation is related to the patient information including body fat percentage, lean body mass and human cross-sectional thickness.
6. The method according to claim 1, characterized in that The method further comprises: A global attention module is introduced into the back-diffusion inference model.
7. A SPECT image reconstruction device based on VQGAN and generalized diffusion model, characterized in that: The device comprises: The weight acquisition module is used to train the unsupervised diffusion training model and obtain the target weight; The feature extraction module is used to obtain cardiac gated SPECT images and use VQGAN to extract features and obtain latent variables; A reconstruction module is used to input the latent variables into a back-diffusion inference model loaded with the target weights, denoise the cardiac gated SPECT image under the guidance of the CT image, and obtain a reconstructed image.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.