Sparse view CT reconstruction method based on conditional embedding fusion diffusion model
By using a sparse view CT reconstruction method based on a conditional embedding fusion diffusion model, combined with a Fourier artifact removal module and a U-Net model, the challenges of restoring the global structure and details of images in sparse view CT reconstruction are solved, and high-quality image reconstruction effects are achieved.
Patent Information
- Application Number
- CN202510855428.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing sparse view CT reconstruction technology has difficulty in simultaneously preserving the global structure of the image and restoring fine details during image reconstruction, and there are problems with artifacts and noise.
A sparse view CT reconstruction method based on the conditional embedding fusion diffusion model is adopted. Combined with the Fourier artifact removal module and the U-Net model, a conditional generation model and a conditional embedding fusion diffusion model are designed. Through the adaptive fusion attention generation mechanism and multi-source information fusion strategy, the ability to restore image details and structures is improved.
It significantly improves the quality of sparse view CT reconstructed images, reduces artifacts, enhances image fidelity and detail recovery capabilities, and improves reconstruction accuracy and image quality.
Smart Images

Figure CN120374781B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer imaging, and in particular to a sparse view CT reconstruction method based on a conditional embedding fusion diffusion model. Background Art
[0002] CT images, known for their remarkable ability to clearly depict the internal structure of objects, are widely used in clinical diagnosis, industrial inspection, and other fields. Conventional CT imaging typically relies on dense projection data to ensure high-quality reconstruction. However, this often results in high X-ray radiation doses. Therefore, sparse view CT reconstruction techniques are particularly important to reduce X-ray radiation doses. However, the data incompleteness associated with reducing the number of projections can lead to significant artifacts and loss of detail in the reconstructed images, severely compromising image quality and diagnostic accuracy.
[0003] Existing sparse view CT reconstruction techniques treat sparse view reconstruction as a denoising task, using FBP reconstruction as input to effectively remove artifacts and reduce noise. Additionally, methods combining the U-Net model with residual learning strategies have achieved relatively ideal denoising results. However, these methods often suffer from over-smoothing and loss of detail due to the limitations of pixel-level processing. To address these issues, perceptual losses have been introduced to generate more realistic CT images. Despite this, these post-processing methods still have limitations in recovering subtle structures and features.
[0004] A convolutional neural network (CNN) is used to interpolate missing sinograms, combining residual learning and local training to improve convergence and avoid memory overload. Local linear interpolation is used to synthesize missing data. Furthermore, sinogram interpolation is incorporated into the iterative reconstruction framework, further improving reconstruction performance. However, these methods often produce secondary artifacts during the projection interpolation process, compromising the quality of the reconstructed image.
[0005] Although the above methods have achieved many results, preserving the global structure of the image and recovering fine details simultaneously in sparse view CT reconstruction remains a major challenge. Summary of the Invention
[0006] The purpose of the present invention is to provide a sparse view CT reconstruction method based on the conditional embedding fusion diffusion model. The method aims to construct a robust conditional embedding fusion diffusion model and provide precise conditional guidance for the conditional embedding fusion diffusion model, thereby significantly improving image fidelity and effectively solving the limitations of existing methods.
[0007] To achieve the above functions, the present invention designs a sparse view CT reconstruction method based on a conditional embedded fusion diffusion model. The following steps S1 to S5 are performed to build and train a sparse view CT reconstruction model and complete the reconstruction of sparse view CT images:
[0008] Step S1: Collect sparse view CT images to form a data set, pre-process the data set and divide it into a training set and a test set in proportion;
[0009] Step S2: constructing a sparse view CT reconstruction model, including a conditional generative model and a conditional embedding fusion diffusion model; using the preprocessed sparse view CT image as the initial image, inputting it into the sparse view CT reconstruction model, applying the conditional generative model to the initial image to remove artifacts and generate a preliminary reconstructed image; inputting the preliminary reconstructed image and auxiliary scalar information into the conditional embedding fusion diffusion model to generate residual details; combining the residual details with the preliminary reconstructed image output by the conditional generative model to generate a sparse view CT reconstructed image as the output of the sparse view CT reconstruction model;
[0010] Step S3: Design the loss functions of the conditional generation model and the conditional embedding fusion diffusion model respectively;
[0011] Step S4: Back propagation is used to train the conditional generation model and the conditional embedding fusion diffusion model, and the model parameters are iterated and updated until the preset model convergence conditions are reached, thereby obtaining a trained sparse view CT reconstruction model;
[0012] Step S5: Testing the trained sparse view CT reconstruction model, and deploying the sparse view CT reconstruction model in an actual sparse view CT image reconstruction application environment.
[0013] Beneficial effects: Compared with the prior art, the advantages of the present invention include:
[0014] The present invention designs a sparse view CT reconstruction method based on the conditional embedding fusion diffusion model, combines the Fourier de-artifacting module and the U-Net model to form a conditional generation model. This model effectively suppresses the noise and artifacts in the sparse view CT reconstructed image, and provides precise conditional guidance for the subsequent conditional embedding fusion diffusion model training. In addition, in order to ensure that the conditional information is seamlessly integrated into the conditional embedding fusion diffusion model, a conditional attention embedding module is designed, and an adaptive fusion attention generation mechanism is introduced to improve the spatial adaptability and flexibility of the conditional embedding fusion diffusion model. This mechanism can adaptively generate mixed feature attention of conditions and time steps, so that the conditional embedding fusion diffusion model can dynamically adjust the generation process at each stage, thereby achieving accurate restoration of image details and structures. The multi-source information fusion attention strategy significantly enhances the model's adaptability to noise at different time steps, ultimately improving the spatial adaptability and reconstruction accuracy of the conditional embedding fusion diffusion model. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a flowchart of a sparse view CT reconstruction method based on a conditional embedding fusion diffusion model according to an embodiment of the present invention;
[0016] Figure 2 is a schematic diagram of a sparse view CT reconstruction model provided according to an embodiment of the present invention;
[0017] Figure 3 is a schematic diagram of a conditional generation model provided according to an embodiment of the present invention;
[0018] Figure 4 2 is a schematic diagram of a conditional embedding fusion diffusion model provided according to an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0020] The sparse view CT reconstruction method based on the conditional embedded fusion diffusion model provided in the embodiment of the present invention performs the following steps S1 to S5, referring to Figure 1 , build and train the sparse view CT reconstruction model to complete the reconstruction of sparse view CT images:
[0021] Step S1: Collect sparse view CT images to form a data set, pre-process the data set and divide it into a training set and a test set in proportion;
[0022] The FBP algorithm was used to reconstruct the complete 720 projection views to obtain artifact-free reference images to form a dataset. Initial images of 60, 90, and 120 views were extracted from the dataset as input to the sparse view CT reconstruction model to evaluate the reconstruction effect of the model at different sparsity levels. The dataset was divided into training and test sets in a ratio of 9:1.
[0023] Step S2: Construct a sparse view CT reconstruction model, referring to Figure 2 The proposed framework includes a conditional generative model (FourierNetModule) and a conditional embedding fusion diffusion model (CEF-DM). The preprocessed sparse view CT image is used as the initial image and fed into the sparse view CT reconstruction model. The conditional generative model removes artifacts from the initial image and generates a preliminary reconstructed image. This preliminary reconstructed image and auxiliary scalar information (diffusion time step) are then fed into the conditional embedding fusion diffusion model, enabling better control and guidance of the model's generation of detail increments. This injection not only provides a comprehensive context but also enhances the flexibility of the framework.
[0024] The conditional embedding fusion diffusion model generates residual details, which are combined with the preliminary reconstructed image output by the conditional generation model to generate a sparse view CT reconstructed image as the output of the sparse view CT reconstruction model, increasing the details of the reconstructed image and achieving high-quality reconstruction.
[0025] Reference Figure 3 The conditional generation model is based on the U-Net model structure and introduces the Fourier Domain De-artifacting Module. The Fourier Domain De-artifacting Module first extracts features through two 1×1 convolutional layers and then fuses them to obtain a local feature map. It also extracts features through a 1×1 convolutional layer and a Fourier convolutional layer and then fuses them to obtain a global feature map. Subsequently, the local feature map and the global feature map are normalized and activated respectively, and then feature fused to form a fused feature map.
[0026] The Fourier convolution layer in the Fourier domain artifact removal module includes a Fourier transform and an inverse Fourier transform. The design concept of the Fourier convolution layer is to effectively integrate frequency and spatial domain information by combining Fourier transform and convolution operations. Before performing the Fourier transform and inverse Fourier transform, the feature map passes through a 1×1 convolution layer, a normalization layer, and an activation function layer. The activation function layer uses the ReLU activation function to enhance feature expression.
[0027] Finally, by introducing skip connections to enhance feature propagation, the network is able to retain key spatial information during frequency domain conversion. This design not only improves reconstruction quality and reduces artifacts, but also fully utilizes the global information contained in the Fourier domain. Combining the advantages of convolution in local receptive fields, it achieves more accurate feature extraction and fusion.
[0028] The conditional generation model outputs a preliminary reconstructed image, which serves as the conditional guidance of the conditional embedding fusion diffusion model and is input into the conditional embedding fusion diffusion model.
[0029] The conditional embedding fusion diffusion model includes a denoising module to assist in predicting the noise distribution; the initial reconstructed image output by the conditional generation model is used. L This is used as a conditional input into the denoising module, guiding the model to focus more on refining the reconstruction results and correcting image details and structures. In this way, the conditional embedding fusion diffusion model can more accurately capture residual information, thereby improving image quality, reducing artifacts, and enhancing detail recovery capabilities.
[0030] Reference Figure 4 The conditional embedding fusion diffusion model consists of a backbone module (BM) and a conditional attention embedding module (CAEM). The backbone module includes a normalization layer, a 3×3 convolution layer, an activation function layer, a group convolution layer, and an activation function layer. The activation function layer uses a swish activation function.
[0031] The preliminary reconstructed image and auxiliary scalar information are input into the conditional attention embedding module. The preliminary reconstructed image passes through a 3×3 convolutional layer, an activation function layer, and a 3×3 convolutional layer in sequence to output a feature map X. The activation function layer uses a simplegate activation function. The auxiliary scalar information passes through a linear layer, an activation function layer, and a linear layer in sequence to output a feature map Y. The activation function layer uses a swish activation function. The feature maps X and Y are passed through the Adaptive Fusion Attention Generation Mechanism (AFAGM) to generate a multi-source attention map XY, which is then embedded into the group convolution layer of the backbone module.
[0032] The core idea of the adaptive fusion attention generation mechanism lies in adaptively fusing information from the conditional input and the time-step embedding. By dynamically adjusting and integrating multi-source features, the adaptive fusion attention generation mechanism can enhance the expressiveness of image details, thereby improving the model's generation performance and image reconstruction quality.
[0033] In the process of embedding the multi-source attention map XY into the group convolution layer of the backbone module, the group convolution layer divides the input feature map into n Subgroup characteristics , for each subgroup feature Perform element-wise multiplication with the multi-source attention map XY, and then n The feature maps are summed along the grouping dimension to obtain the final fused feature map, which effectively encodes the conditional information and time step information.
[0034] It is formally expressed as follows:
[0035] ;
[0036] in, N represents the total number of subgroup features, F is the fusion feature map.
[0037] The introduction of group convolution aims to reduce computational complexity and the number of parameters while efficiently extracting local features. By dividing the channels of the feature map into multiple smaller subgroups, group convolution performs independent convolution operations within each group, thereby reducing computational cost while improving the network's feature representation capabilities. In the subsequent feature fusion process, group convolution can also effectively extract and integrate information from different time steps and conditions, thereby optimizing image details and enhancing the model's generative capabilities. The use of the swish activation function and the introduction of skip connections to support residual learning further enhance the model's ability to capture fine-grained details.
[0038] The conditional embedding fusion diffusion model includes a forward diffusion process and a reverse denoising process, where the forward diffusion process is as follows:
[0039] ;
[0040] in, The residual of the preliminary reconstructed image output by the conditional generative model is represents random noise, is the time step t Noise image at ; , is the diffusion noise variance, yes Linear representation, set the initial value , increases linearly to , ,in yes t time The continuous multiplication of
[0041] The reverse denoising process is as follows:
[0042] ;
[0043] in, , represents the variance, which is a linear representation of the variance of the diffusion noise; L represents the preliminary reconstructed image generated using the conditional generation model, represents the noise prediction model, z represents random noise;
[0044] Step S3: Design the loss functions of the conditional generation model and the conditional embedding fusion diffusion model respectively;
[0045] The loss functions designed for the conditional generation model include mean square error (MSE) loss and structural similarity (SSIM) loss, so as to balance pixel-level accuracy and the restoration of structural details during the reconstruction process. MSE loss is used to measure the pixel-by-pixel error between the reconstructed image and the real image, which helps to optimize the overall image quality and thus improve the PSNR index; while SSIM loss focuses on the structure, contrast, and brightness information of the image to ensure that the reconstructed image is visually consistent with the original image, thereby effectively improving the SSIM score. The design of this hybrid loss function enables the model to simultaneously optimize low-level pixel information and high-level structural features, thereby achieving higher quality image reconstruction. By combining MSE loss with SSIM loss, the model can not only reduce artifacts, but also avoid the loss of details, ultimately achieving a more realistic and accurate image reconstruction effect. Specifically as follows:
[0046] ;
[0047] in, represents the loss function of the conditional generation model, represents the mean square error loss, represents the structural similarity loss, Represents the real clean image (Ground Truth), Represents the preliminary reconstructed image generated by the conditional generation model L ;
[0048] The loss function designed for the conditional embedding fusion diffusion model is based on mean squared error (MSE) loss. The MSE loss function can effectively measure the error between predicted noise and actual noise, and has the advantages of simplicity and efficiency. For tasks where the noise is locally random, MSE loss can provide a clear optimization target and maintain good stability during training. In addition, MSE loss is consistent with maximum likelihood estimation theory, which helps the model more accurately approximate the actual noise distribution. The specific formula is as follows:
[0049] ;
[0050] in, represents the loss function of the conditional embedding fusion diffusion model, represents random noise, represents the mean square error loss, , Indicates the maximum value of the image pixel, L represents the preliminary reconstructed image generated using the conditional generation model, H represents the real image, represents the denoising module in the conditional embedding fusion diffusion model, represents the noise variance.
[0051] Step S4: Back propagation is used to train the conditional generation model and the conditional embedding fusion diffusion model, and the model parameters are iterated and updated until the preset model convergence conditions are reached, thereby obtaining a trained sparse view CT reconstruction model;
[0052] For the conditional generation model and the conditional embedding fusion diffusion model, the Adam optimizer was used for 30 rounds of training respectively; and the momentum parameter was set to , represents the exponential decay rate of the first-order moment (i.e. momentum) of the control gradient, Represents the exponential decay rate of the squared (second-order moment) control gradient.
[0053] In step S4, the training steps for the conditional generation model and the conditional embedding fusion diffusion model are as follows:
[0054] Step S4.1: Train the conditional generation model and the conditional embedding fusion diffusion model separately and calculate the loss function value;
[0055] Step S4.2: If the loss function value decreases to a preset value, the conditional generation model and the conditional embedding fusion diffusion model are determined to have converged, and the process proceeds to step S4.3; otherwise, the process returns to step S4.1.
[0056] Step S4.3: Input the sparse view CT image into the trained conditional generative model to complete the initial reconstruction and generate a preliminary reconstructed image. This image is then used as a conditional guide for the conditional embedding fusion diffusion model to generate residual details and complete the final reconstruction of the sparse view CT image.
[0057] Step S5: Testing the trained sparse view CT reconstruction model, and deploying the sparse view CT reconstruction model in an actual sparse view CT image reconstruction application environment.
[0058] To demonstrate the effectiveness of this invention, comparative experiments and ablation experiments were conducted. First, the dataset and training details are introduced. Then, comparative experimental results of different algorithms on the dataset are presented. A series of ablation experiments are conducted to evaluate the effectiveness of the Fourier De-artifacting Module and the Conditional Attention Embedding Module. The initial learning rate of the conditional generative model (FourierNet Module) is set to 1e-4 and gradually decreased to 1e-5 through cosine annealing. The Conditional Embedding Fusion Diffusion Model (CEF-DM) is trained with a fixed learning rate of 1e-4. The diffusion time step T of the Conditional Embedding Fusion Diffusion Model is set to 2000.
[0059] The proposed method was compared with other currently popular advanced methods for sparse view CT reconstruction, and comparative experiments were conducted on the AAPM dataset and the C4KC-KiTS dataset. The results are shown in Tables 1 and 2. As can be seen from the results, the proposed method achieved significant results on various datasets, further verifying the effectiveness of the proposed method.
[0060] Table 1. Comparative experimental results based on the AAPM dataset
[0061]
[0062] Table 2. Comparative experimental results based on the C4KC-KiTS dataset
[0063]
[0064] To verify the effectiveness of each module in our method, we conducted ablation experiments on the AAPM and C4KC-KiTS datasets. We removed the Fourier artifact removal module and the conditional attention embedding module from our method to verify their effectiveness. The ablation results are shown in Table 3, where √ indicates that the corresponding module was retained, and × indicates that the corresponding module was removed.
[0065] Table 3. Ablation experiment results based on AAPM dataset
[0066]
[0067] In order to evaluate the quality of the reconstructed data, the peak signal-to-noise ratio (PSNR) and the structural similarity (SSIM) index were used for quantitative analysis.
[0068] PSNR evaluates image quality by comparing the error between the original image and the reconstructed image. Its value is usually expressed in decibels (dB). The higher the PSNR, the higher the similarity between the reconstructed image and the original image, and the better the quality. and Expressed as the reconstructed image and the original image, the peak signal-to-noise ratio between the reconstructed image and the original image is expressed as:
[0069] ;
[0070] in, represents the peak signal-to-noise ratio between the reconstructed image and the original image, MAX Represents the maximum pixel of the image, MSE stands for mean square error, which is used to measure the difference between the pixel values of two images. The mean square error between the reconstructed image and the original image is expressed as:
[0071] ;
[0072] in, represents the mean square error between the reconstructed image and the original image, W and H are the width and height of the image, i and j is the pixel position in the image.
[0073] SSIM is used to assess the structural similarity between two images. Unlike PSNR, SSIM not only considers the absolute difference in pixels but also comprehensively considers brightness, contrast, and structural information, thus better reflecting the perceived quality of an image. Its value range is between 0 and 1, where 1 indicates that the two images are identical, values close to 1 indicate high similarity, and negative values indicate significant structural differences. The structural similarity between the reconstructed image and the original image is expressed as:
[0074] ;
[0075] in, represents the structural similarity between the reconstructed image and the original image, and are the mean (brightness) of the two images, and are the variance (contrast) of the two images, is the covariance (structure) of the two images, 、 are two constants.
[0076] As shown in Tables 1 and 2, the proposed method achieves higher SSIM and PSNR than other compared methods on the AAPM and C4KC-KiTS datasets. Compared to current mainstream sparse view reconstruction algorithms, the proposed method significantly improves the quality of CT image reconstruction.
[0077] Table 3 shows that for the AAPM dataset, after adding the Fourier artifact removal module, the proposed method improves the SSIM of sparse view CT reconstruction results for 60-degree, 90-degree, and 120-degree views by 1.24% (from 92.64% to 93.88%), 1.37% (from 93.96% to 95.33%), and 1.28% (from 94.35% to 95.63%), respectively. PSNR also increases by 0.76 (from 33.56 to 34.32), 1.33 (from 35.73 to 37.06), and 1.08 (from 37.88 to 38.96), respectively. This demonstrates that the Fourier artifact removal module has a positive impact on reconstruction. After adding the conditional attention embedding module, the proposed method improves the SSIM of sparse view CT reconstruction results for 60-degree, 90-degree, and 120-degree views by 2.08% (from 93.88% to 95.96%), 2.5% (from 95.33% to 97.83%), and 3.01% (from 95.63% to 98.64%), respectively. PSNR also increases by 3.45 (from 34.32 to 37.77), 4.9 (from 37.06 to 41.96), and 4.6 (from 38.96 to 43.56), respectively. The results in Table 3 show that each module in the proposed method brings consistent improvements to the model across different datasets.
[0078] The embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the scope of knowledge possessed by ordinary technicians in this field without departing from the spirit of the present invention.
Claims
1. A sparse view CT reconstruction method based on a conditional embedding fusion diffusion model, characterized by: Perform the following steps S1 to S5 to build and train a sparse view CT reconstruction model to complete the reconstruction of sparse view CT images: Step S1: Collect sparse view CT images to form a data set, pre-process the data set and divide it into a training set and a test set in proportion; Step S2: constructing a sparse view CT reconstruction model, including a conditional generative model and a conditional embedding fusion diffusion model; using the preprocessed sparse view CT image as the initial image, inputting it into the sparse view CT reconstruction model, applying the conditional generative model to the initial image to remove artifacts and generate a preliminary reconstructed image; inputting the preliminary reconstructed image and auxiliary scalar information into the conditional embedding fusion diffusion model to generate residual details; combining the residual details with the preliminary reconstructed image output by the conditional generative model to generate a sparse view CT reconstructed image as the output of the sparse view CT reconstruction model; The conditional generation model is based on the U-Net model structure and introduces a Fourier domain artifact removal module. The Fourier domain artifact removal module first extracts features through two 1×1 convolutional layers and then fuses them to obtain a local feature map. It also extracts features through a 1×1 convolutional layer and a Fourier convolutional layer and then fuses them to obtain a global feature map. Subsequently, the local feature map and the global feature map are normalized and activated respectively, and then feature fusion is performed to form a fused feature map. The Fourier convolution layer in the Fourier domain artifact removal module includes Fourier transform and inverse Fourier transform. Before performing Fourier transform and inverse Fourier transform, the feature map passes through a 1×1 convolution layer, a normalization layer, and an activation function layer respectively; The conditional generation model outputs a preliminary reconstructed image, which serves as the conditional guidance of the conditional embedding fusion diffusion model and is input into the conditional embedding fusion diffusion model; Step S3: Design the loss functions of the conditional generation model and the conditional embedding fusion diffusion model respectively; Step S4: Back propagation is used to train the conditional generation model and the conditional embedding fusion diffusion model, and the model parameters are iterated and updated until the preset model convergence conditions are reached, thereby obtaining a trained sparse view CT reconstruction model; Step S5: Testing the trained sparse view CT reconstruction model, and deploying the sparse view CT reconstruction model in an actual sparse view CT image reconstruction application environment.
2. The sparse view CT reconstruction method based on conditional embedding fusion diffusion model according to claim 1 is characterized in that: In step S1, the FBP algorithm is used to reconstruct the complete 720 projection views to obtain artifact-free reference images to form a data set; initial images of 60, 90, and 120 views are extracted from the data set as the input of the sparse view CT reconstruction model, and the data set is divided into training and test sets in a ratio of 9:
1.
3. The sparse view CT reconstruction method based on conditional embedding fusion diffusion model according to claim 1, characterized in that: In step S2, the conditional embedding fusion diffusion model includes a denoising module, which is composed of a backbone module and a conditional attention embedding module. The backbone module includes a normalization layer, a 3×3 convolution layer, an activation function layer, a group convolution layer, and an activation function layer in sequence; The preliminary reconstructed image and auxiliary scalar information are input into the conditional attention embedding module. The preliminary reconstructed image passes through a 3×3 convolutional layer, an activation function layer, and a 3×3 convolutional layer in sequence to output a feature map X. The auxiliary scalar information passes through a linear layer, an activation function layer, and a linear layer in sequence to output a feature map Y. The feature maps X and Y are subjected to an adaptive fusion attention generation mechanism to generate a multi-source attention map XY, which is then embedded into the group convolution layer of the backbone module.
4. The sparse view CT reconstruction method based on conditional embedding fusion diffusion model according to claim 3 is characterized in that: In the process of embedding the multi-source attention map XY into the group convolution layer of the backbone module, the group convolution layer divides the input feature map into n sub-group features , for each subgroup feature Perform element-wise multiplication with the multi-source attention map XY, and then sum the resulting n feature maps along the grouping dimension to obtain the final fused feature map.
5. The sparse view CT reconstruction method based on conditional embedding fusion diffusion model according to claim 1, characterized in that: The conditional embedding fusion diffusion model includes a forward diffusion process and a reverse denoising process, where the forward diffusion process is as follows: ; in, The residual of the preliminary reconstructed image output by the conditional generative model is represents random noise, is the time step t Noise image at ; , is the diffusion noise variance, yes Linear representation, set the initial value , increases linearly to , ,in yes t time The continuous multiplication of The reverse denoising process is as follows: ; in, , represents the variance, L represents the preliminary reconstructed image generated using the conditional generation model, represents the noise prediction model, and z represents random noise.
6. The sparse view CT reconstruction method based on conditional embedding fusion diffusion model according to claim 1, characterized in that: In step S3, the loss function designed for the conditional generation model includes mean square error loss and structural similarity loss, which are specifically as follows: ; in, represents the loss function of the conditional generation model, represents the mean square error loss, represents the structural similarity loss, represents a real clean image, Represents the preliminary reconstructed image generated by the conditional generation model L ; The loss function designed for the conditional embedding fusion diffusion model is based on the mean square error loss, as shown in the following formula: ; in, represents the loss function of the conditional embedding fusion diffusion model, represents the mean square error loss, represents random noise, L represents the preliminary reconstructed image generated using the conditional generation model, H represents the real image, represents the denoising module in the conditional embedding fusion diffusion model, represents the noise variance.
7. The sparse view CT reconstruction method based on conditional embedding fusion diffusion model according to claim 1, characterized in that: In step S4, the Adam optimizer is used to train the conditional generation model and the conditional embedding fusion diffusion model for 30 rounds respectively, and the momentum parameter is set to , represents the exponential decay rate of the first-order moment of the control gradient, Represents the exponential decay rate of the square of the control gradient.
8. The sparse view CT reconstruction method based on conditional embedding fusion diffusion model according to claim 1, characterized in that: In step S4, the training steps for the conditional generation model and the conditional embedding fusion diffusion model are as follows: Step S4.1: Train the conditional generation model and the conditional embedding fusion diffusion model separately and calculate the loss function value; Step S4.2: If the loss function value decreases to a preset value, the conditional generation model and the conditional embedding fusion diffusion model are determined to have converged, and the process proceeds to step S4.3; otherwise, the process returns to step S4.
1. Step S4.3: Input the sparse view CT image into the trained conditional generative model to complete the initial reconstruction and generate a preliminary reconstructed image. This image is then used as a conditional guide for the conditional embedding fusion diffusion model to generate residual details and complete the final reconstruction of the sparse view CT image.
Citation Information
Patent Citations
Sparse angle CT artifact removal method
CN114596378A
Double-view industrial CT fault online reconstruction method based on hidden space condition diffusion model
CN117635745A