A limited-angle CT reconstruction method based on three-dimensional conditional diffusion model and synchronous iteration
By combining a three-dimensional conditional diffusion model and SIRT iterative reconstruction technology, the artifact and blur problems of CT reconstruction under sparse viewpoints were solved, realizing high-fidelity, physically constrained finite-angle CT reconstruction to meet the high-speed requirements of industrial production lines.
Patent Information
- Application Number
- CN202511440470.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-10-10
AI Technical Summary
Traditional CT reconstruction methods are prone to stripe artifacts, blurring, and loss of detail in reconstruction results under sparse or limited viewpoints, and lack physical mechanism support, resulting in insufficient reliability of the generated tomographic images.
By combining a 3D conditional diffusion model and synchronous iterative reconstruction (SIRT) technology, prior information is obtained through filtered back projection, and local details are enhanced in the frequency domain decoding branch. Combined with SIRT iterative reconstruction, physical constraints are introduced to achieve high-fidelity CT reconstruction.
Achieving high-fidelity CT reconstruction under limited angle conditions meets the high-speed requirements of modern industrial production lines and improves the reconstruction accuracy of local details and edge structures.
Smart Images

Figure CN120912791B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of CT tomographic reconstruction technology and artificial intelligence, and discloses a limited-angle CT reconstruction method based on the combination of a three-dimensional conditional diffusion model and synchronous iteration. BACKGROUND
[0002] Computed Tomography (CT) is a non-destructive imaging technology that irradiates the measured object with X-rays from multiple angles, acquires two-dimensional projection data, and accurately restores the three-dimensional structure using reconstruction algorithms, thereby intuitively presenting the internal details. With the advantages of high resolution, quantification and visualization, CT has been widely used in industrial quality detection, and can identify and locate internal structural defects of workpieces without disassembly. However, traditional CT usually needs to perform 360° full-angle scanning around the object, and hundreds to thousands of X-ray projections are required to complete tomographic reconstruction, which makes it difficult to meet the logistics rhythm and high-speed rhythm requirements of modern industrial production lines.
[0003] To improve the speed of CT to meet the needs of high-speed quality inspection, scholars have proposed methods to reduce the scanning angle and the number of projections, thereby significantly shortening the acquisition time. However, CT reconstruction based on sparse or limited views inevitably faces the problem of insufficient projection information, leading to phenomena such as stripe artifacts, blurring and loss of details in the reconstructed results. To address this problem, patent CN117635745A reports a dual-view reconstruction method based on a hidden space diffusion model, which uses dual-view orthogonal projection as a condition to guide the diffusion model that has learned the tomographic prior distribution to generate, achieving artifact-free CT reconstruction. Patent CN120388093A reports a sparse-view reconstruction method based on a U-Net network, which extracts features through alternating iterations of non-local and local residual optimization modules, thereby effectively suppressing sparse reconstruction artifacts. Patent CN120279209A reports a sparse-view reconstruction method based on discrete Gaussian representation, which uses Gaussian functions to continuously represent target voxels, learns the mapping relationship between projections and voxels, and thereby achieves accurate reconstruction of three-dimensional structures under sparse projection conditions. These methods learn prior information from training data through end-to-end networks to make up for the lack of projection information. However, the inherent geometric constraints and physical mechanisms of CT imaging are ignored during reconstruction, lacking clear theoretical support and interpretability, which may lead to insufficient reliability of the generated tomographic images.
[0004] Unlike existing solutions that simply rely on networks, the present application combines an end-to-end network based on a three-dimensional conditional diffusion model with a simultaneous iterative reconstruction technique (SIRT) to introduce interpretable physical constraints into the deep learning network by taking advantage of the physical mechanism representation of iterative reconstruction. Specifically, on the one hand, the three-dimensional conditional diffusion model is used as the core architecture, and the volume data obtained by filtered back projection (FBP) reconstruction is used as guidance to learn the prior information from the training data, thereby realizing CT tomographic reconstruction under limited angle conditions; on the other hand, the SIRT is used to correct the data consistency of the three-dimensional conditional diffusion model generated results, thereby preserving the details generated by deep diffusion while strictly following the geometric constraints of CT imaging to ensure the reliability of the reconstruction results. In addition, to further improve the reconstruction accuracy of details and edge structures, the network structure is optimized in the latent space decoding part. On the basis of conventional spatial domain decoding, a frequency domain decoding branch based on fast Fourier transform is added. The branch focuses on the extraction and reconstruction of high-frequency features by suppressing redundant information through a low-frequency mask, thereby enhancing the recovery ability of local details and edge structures and building a more solid foundation for subsequent SIRT iteration. In summary, the three-dimensional conditional diffusion model with spatial domain-frequency domain dual decoding combined with SIRT iterative reconstruction of the present application can realize high-fidelity CT reconstruction with physical constraints under limited angle conditions. SUMMARY
[0005] The present application proposes a limited angle CT reconstruction method based on the combination of a three-dimensional conditional diffusion model and a simultaneous iteration, which can realize three-dimensional structure reconstruction by using only 90 projection projections in the range of 0°-90°, thereby meeting the requirements of high-speed rhythm and logistics rhythm of modern industrial production lines. The reconstruction model proposed by the present application includes a three-dimensional conditional diffusion model and a data consistency correction module, and the overall architecture is as shown in Figure 1 .
[0006] The technical scheme of the present application is as follows:
[0007] A limited angle CT reconstruction method based on the combination of a three-dimensional conditional diffusion model and a simultaneous iteration, the steps are as follows:
[0008] Step 1: Construct a CT data set;
[0009] A sufficient number of same kind of workpieces to be inspected are acquired, full-angle CT scanning is performed on each workpiece to be inspected at an interval of 1°, and a tomographic image obtained by FBP reconstruction is taken as a true value to construct a training data set for network parameter training of the three-dimensional conditional diffusion model; meanwhile, 90 projections in a range of 0°-90° are intercepted from the projections obtained by full-angle CT scanning as inputs of the three-dimensional conditional diffusion model; thus, each group of data in the CT data set contains a group of tomographic images obtained by FBP reconstruction of full-angle and 90 corresponding projections in a range of 0°-90°;
[0010] Step two: building the three-dimensional conditional diffusion model;
[0011] (1) Building an encoder: the encoder is used to generate a hidden space representation of a three-dimensional body and encode the three-dimensional body into a hidden space feature map;
[0012] (2) Building a conditional embedding module and a noise prediction network: the conditional embedding module encodes the three-dimensional body into a conditional vector and inputs it into the noise prediction network to guide the denoising process of the hidden space feature map; the noise prediction network is used to estimate the noise component in the hidden space feature map at each time step and gradually remove the noise based on the estimation, so as to generate a target CT body from Gaussian pure noise under the guidance of the conditional vector ;
[0013] (3) Building a decoder: the decoder is responsible for decoding and restoring the hidden space feature map to the target CT body;
[0014] Step three: building a data consistency correction module;
[0015] In order to improve the reliability of the three-dimensional conditional diffusion model in generating the target CT body, a data consistency correction module is built after the three-dimensional conditional diffusion model;
[0016] Step four: training the three-dimensional conditional diffusion model;
[0017] Step five: online running process;
[0018] When deployed online, first, the data set of the workpiece to be inspected is obtained according to step one, and then the three-dimensional conditional diffusion model and the data consistency correction module are built according to steps two and three; subsequently, the training of the three-dimensional conditional diffusion model is completed according to step four, and the three-dimensional conditional diffusion model parameters are fixed after the training is completed; further, the CT device is used to collect 90 projections of the workpiece to be inspected within the range of 0°-90° at intervals of 1°, and the preliminary reconstructed body data is obtained through FBP reconstruction; the preliminary reconstructed body data is input into the conditional embedding module to extract features, and is input into the noise prediction network together with the Gaussian pure noise and time step information, and the target CT body is generated by gradually denoising according to the total denoising time step T; subsequently, the target CT body is input as the initial value into the data consistency correction module, and finally the reconstructed CT body is output.
[0019] The beneficial effects of the present application are: (1) The present application can realize three-dimensional structure reconstruction through limited-angle projection acquisition, and meet the requirements of high-speed rhythm and logistics rhythm of modern industrial production lines. (2) The present application designs a dual-branch architecture of a spatial domain device and a frequency domain decoder, enhances the reconstruction effect of local details and edge structures, and thus improves the problems of detail blur and edge distortion. (3) Unlike the existing scheme which simply relies on a network, the present application combines a three-dimensional conditional diffusion model with SIRT iterative reconstruction, utilizes the advantages of iterative methods in terms of geometric rigor and CT reconstruction mechanism representation, introduces an interpretable physical constraint for a deep learning network, and realizes high-fidelity CT reconstruction under limited-angle conditions. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 is the architecture diagram of the limited-angle CT reconstruction method based on the combination of the three-dimensional conditional diffusion model and the simultaneous iteration of the present application. DETAILED DESCRIPTION
[0021] The specific embodiments of the present application will be further described below in combination with the technical solutions.
[0022] A limited-angle CT reconstruction method based on the combination of a three-dimensional conditional diffusion model and a simultaneous iteration, and the steps are as follows:
[0023] Step one: constructing a CT data set;
[0024] A sufficient number of same type workpieces to be inspected are obtained, full-angle CT scanning is performed on each workpiece to be inspected at intervals of 1°, and the tomographic images obtained through FBP reconstruction are used as true values to construct a training data set for network parameter training of the three-dimensional conditional diffusion model; meanwhile, 90 projections within the range of 0°-90° are intercepted from the projections obtained through the full-angle CT scanning as inputs of the three-dimensional conditional diffusion model; thus, each group of data in the CT data set comprises a group of tomographic images obtained through FBP reconstruction of full-angle and 90 projections within the range of 0°-90°;
[0025] Step 2: Construct a three-dimensional conditional diffusion model;
[0026] (1) Building an encoder: The encoder is used to generate the latent space representation of the three-dimensional volume, encoding the three-dimensional volume into a latent space feature map;
[0027] The encoder consists of a 3D convolution, residual blocks, and downsampling layers. The input single-channel 128³ 3D volume is first expanded from 1 to 32 channels by a 3D convolution. Then, feature extraction is performed sequentially at three scales: at the first two scales, each volume passes through two residual blocks, followed by a downsampling layer with a stride of 2, reducing the size of the 3D volume from 128³ to 64³ and 32³ respectively, while increasing the number of channels from 32 to 64 and 128. At the third scale, the volume passes through two more residual blocks, but no further downsampling is performed. After layer-by-layer extraction and compression, the output is a latent space feature map z containing the 3D volume.
[0028] (2) Constructing a conditional embedding module and a noise prediction network: The conditional embedding module encodes the 3D volume into conditional vectors. The noise prediction network is then fed into the latent space feature map to guide the denoising process. This network estimates the noise components in the latent space feature map at each time step and progressively removes noise based on this estimate, thereby improving the conditional vector. Guided by Gaussian pure noise Generate the target CT image;
[0029] The network structure of the conditional embedding module is the same as that of the encoder;
[0030] The noise prediction network adopts a three-dimensional U-Net structure, mainly composed of convolutional layers, residual blocks, temporal attention, spatial attention, downsampling layers, upsampling layers, and skip connections; the latent space feature map output by the encoder... After passing through a convolutional layer followed by a temporal attention layer, the system proceeds to a downsampling stage with four layers to extract deep features in the planar direction. The first three layers process in the order of two residual blocks, spatial attention, temporal attention, and a downsampling layer. The fourth layer contains only two residual blocks, spatial attention, and temporal attention, without further downsampling. Following the downsampling stage is an intermediate layer, with the order being residual block, spatial attention, temporal attention, and a residual block. Next, the system enters an upsampling stage with three layers, each processing in the order of an upsampling layer, temporal attention, spatial attention, and two residual blocks. Simultaneously, features from the same level saved in the previous downsampling stage are fused through skip connections. The noise prediction network ends with a residual block, spatial attention, temporal attention, a residual block, and a convolutional layer, outputting predicted noise to complete the processing of the latent space feature map. Noise estimation;
[0031] When the noise component in the latent space feature map is estimated After that, the hidden space feature map of each time step is estimated based on the estimation Backward reasoning is performed to realize denoising according to the following formula:
[0032] (1)
[0033] wherein, is a Gaussian noise used to add randomness to the denoising process; is a noise scheduling parameter, indicating the noise intensity injected into the hidden space feature map at the time step , and defines and ; by sequentially iterating formula (1) at each time step, the hidden space feature map is gradually denoised according to the total denoising time step , to generate the hidden space feature map of the target CT volume .
[0034] (3) Build a decoder: the decoder is responsible for decoding and restoring the hidden space feature map to the target CT volume;
[0035] The network structure of the decoder adopts a dual-branch architecture of a spatial domain decoder and a frequency domain decoder; the spatial domain decoder, as the symmetric part of the encoder, is composed of three-dimensional convolution, residual block and up-sampling layer, and is used to restore the hidden space feature map step by step; the frequency domain decoder first performs fast Fourier transform on the hidden space feature map, and then returns to the spatial domain through inverse fast Fourier transform, and further maps the high-frequency enhanced features through the convolution layer; then, the high-frequency enhanced features are fused with the main branch of the spatial domain decoder, and the decoding is completed step by step through the residual block and the three-dimensional convolution.
[0036] Step three: build a data consistency correction module;
[0037] In order to improve the reliability of the three-dimensional conditional diffusion model to generate the target CT volume, a data consistency correction module is built after the three-dimensional conditional diffusion model;
[0038] The data consistency correction module adopts the SIRT iterative method, takes the target CT volume decoded from the hidden space feature map as the initial value , performs forward projection on the target CT volume, calculates the projection residual with the 90 projection images in the range of 0°-90°, and then feeds back the projection residual to the voxel space through back projection to update the voxel value of the target CT volume in an iterative manner; the specific voxel update formula is as follows:
[0039] (2)
[0040] wherein, is the first iteration reconstruction result, is the relaxation factor, is the system matrix, is the projection domain weight matrix, is the projection data.
[0041] Step four: training a three-dimensional conditional diffusion model;
[0042] The specific process of training the three-dimensional conditional diffusion model is as follows:
[0043] (1) Training the encoder and decoder: the encoder is responsible for mapping the input three-dimensional body to the hidden space feature map, and the decoder restores the corresponding three-dimensional body according to the hidden space feature map; the two are jointly optimized as a whole in the training process: the hidden space feature map output by the encoder is directly used as the input of the decoder to realize end-to-end learning;
[0044] Randomly extract the tomographic image obtained by FBP reconstruction from the full-angle in the training data set prepared in step one , input it into the combined network of the encoder and the decoder, and obtain the output ; the tomographic image is used as the supervision signal, and the training target is to make the output consistent with the input as much as possible; the loss function adopts mean square error, which is defined as follows:
[0045] (3)
[0046] Repeat the above training process to jointly train the encoder and the decoder; when the loss function converges, the training process is completed, and the network parameters are frozen for subsequent network training;
[0047] (2) Training the conditional embedding module: the conditional embedding module is trained in cascade with the trained decoder, wherein the trained decoder parameters are fixed, and only the conditional embedding module is trained; 90 projections in the range of 0°-90° are randomly extracted from the training data set prepared in step one , and the tomographic image is obtained by FBP reconstruction as the training input, and the output is obtained through the conditional embedding module and the decoder; the tomographic image is used as the supervision signal, and the loss function adopts mean square error, which is defined as follows:
[0048] (4)
[0049] Repeat the above training process to train the conditional embedding module; when the loss function converges, the training process is completed, and the network parameters are frozen for subsequent network training;
[0050] (3) Training the noise prediction network: keeping the parameters of the trained encoder, conditional embedding module and decoder fixed, only training the noise prediction network: first randomly select a data pair from the training data set prepared in step one, where 90 projections in the range of 0°-90° are input into the conditional embedding module after FBP reconstruction to obtain the conditional feature ; In the data pair, the tomographic image obtained by FBP reconstruction of the full angle is input into the encoder to obtain its corresponding hidden space feature map ; Select the total denoising time step , randomly select the time step , from , add noise to obtain the noisy hidden representation , that is
[0051] (5)
[0052] wherein is the noise sampled from the standard normal distribution ; Then, the hidden space feature map , the conditional feature and the time step are jointly input into the noise prediction network to output the predicted noise ; The training goal is to accurately estimate the noise component contained in the hidden space feature map at any time step, and its loss function is defined as follows:
[0053] (6)
[0054] wherein represents the true noise, is the output of the noise prediction network; In the training process, the time step is randomly selected and iteratively optimized multiple times to ensure that the three-dimensional conditional diffusion model accurately predicts the noise at different time steps; When the loss function converges, the training process is completed.
[0055] Step five: online running process
[0056] When deployed online, first, the data set of the workpiece to be inspected is obtained according to step one, second, the three-dimensional conditional diffusion model and the data consistency correction module are built according to steps two and three; then, the training of the three-dimensional conditional diffusion model is completed according to step four, and the three-dimensional conditional diffusion model parameters are fixed after the training is completed; further, the CT equipment is used to collect 90 projection images of the workpiece to be inspected within the range of 0-90° at intervals of 1°, and the preliminary reconstructed body data is obtained through FBP reconstruction; the preliminary reconstructed body data is input into the conditional embedding module to extract features, and the features are input into the noise prediction network together with the Gaussian pure noise and time step information, and the target CT body is generated by gradually denoising according to the total denoising time step T; then, the target CT body is input as the initial value into the data consistency correction module, and finally, the reconstructed CT body is output.
Claims
1. A limited-angle CT reconstruction method based on the combination of a three-dimensional conditionally diffuse model and synchronous iteration, characterized in that, The steps are as follows: Step one: construct the CT data set; For the same type of workpiece to be inspected, full-angle CT scanning is performed at an interval of 1° for each workpiece to be inspected, and the tomographic image obtained by FBP reconstruction is used as the true value to construct the training data set for network parameter training of the three-dimensional conditional diffusion model; at the same time, 90 projections in the range of 0°-90° are intercepted from the projections obtained from the full-angle CT scanning as the input of the three-dimensional conditional diffusion model; thus, each group of data in the CT data set contains a group of tomographic images obtained by FBP reconstruction of the full-angle and 90 projections in the range of 0°-90° corresponding thereto; Step two: build a three-dimensional conditional diffusion model; (1) Build an encoder: the encoder is used to generate the hidden space representation of the three-dimensional body and encode the three-dimensional body into a hidden space feature map; (2) Build a condition embedding module and a noise prediction network: the condition embedding module encodes the three-dimensional body into a condition vector , which is input into the noise prediction network to guide the denoising process of the hidden space feature map; the noise prediction network is used to estimate the noise component in the hidden space feature map at each time step, and gradually remove the noise based on the estimation, so as to generate the target CT body from the Gaussian pure noise under the guidance of the condition vector ; (3) Build decoder: The decoder is responsible for recovering the target CT volume from the latent space feature map decoder. Step three: build a data consistency correction module; To improve the reliability of the three-dimensional conditional diffusion model in generating the target CT body, a data consistency correction module is built after the three-dimensional conditional diffusion model; The data consistency correction module adopts a SIRT iteration method, and the hidden space feature map The decoded target CT body as an initial value Forward projection is performed on the target CT body, 90 projection calculation projection residuals in the range of 0°-90°, and the projection residuals are fed back to the voxel space by back projection, and the voxel value of the target CT body is updated in an iterative manner; The specific voxel update formula is as follows: (2); wherein, is the reconstruction result of the th iteration, is a relaxation factor, is a system matrix, is a projection domain weight matrix, is projection data; Step four: train the three-dimensional conditional diffusion model; Step five: online running process; When online deployment, first, the data set of the workpiece to be inspected is obtained according to step one, second, the three-dimensional conditional diffusion model and the data consistency correction module are built according to steps two and three; Then, the training of the three-dimensional conditional diffusion model is completed according to step four, and the parameters of the three-dimensional conditional diffusion model are fixed after the training is completed; then, 90 projections of the workpiece to be inspected in the range of 0°-90° are collected by the CT device at an interval of 1°, and the preliminary reconstruction body data is obtained by FBP reconstruction; the preliminary reconstruction body data is input into the conditional embedding module to extract features, and the Gaussian pure noise and time step information are jointly input into the noise prediction network, and the target CT body is generated by gradually denoising according to the total denoising time step T; then, the target CT body is input as the initial value into the data consistency correction module, and the final reconstructed CT body is output.
2. The limited-angle CT reconstruction method based on the combination of the three-dimensional conditional diffusion model and synchronous iteration according to claim 1, characterized in that The encoder is composed of three-dimensional convolution, residual block and down-sampling layer; the input single-channel 128³ three-dimensional body is first expanded from 1 to 32 channels by three-dimensional convolution, and then feature extraction is performed on three scales in turn: on the first two scales, two residual blocks are first passed through, and then a down-sampling layer with a step of 2 is connected to reduce the size of the three-dimensional body from 128³ to 64³ and 32³ in turn, and at the same time, the channel number is increased from 32 to 64 and 128; on the third scale, two residual blocks are continued to be passed through, but no down-sampling is performed; after layer-by-layer extraction and compression, the output contains the hidden space feature map z of the three-dimensional body.
3. The limited-angle CT reconstruction method based on the combination of the three-dimensional conditional diffusion model and synchronous iteration according to claim 2, characterized in that The network structure of the conditional embedding module is the same as that of the encoder; The noise prediction network adopts a three-dimensional U-Net structure, which is composed of convolutional layers, residual blocks, temporal attention, spatial attention, down-sampling layers, up-sampling layers and skip connections; the hidden space feature map z output by the encoder is first processed by a convolutional layer and a layer of temporal attention, and then enters the four levels of the down-sampling stage to extract deep features in the plane direction: the processing order of the first three levels is two layers of residual blocks, spatial attention, temporal attention, and down-sampling layers; the fourth level only contains two layers of residual blocks, spatial attention, and temporal attention, and no longer down-samples; after the down-sampling stage, there is an intermediate layer, which is processed in the order of residual block, spatial attention, temporal attention, and residual block; then, three levels of up-sampling stages are entered, and each level is processed in the order of up-sampling layer, temporal attention, spatial attention, and two layers of residual blocks; at the same time, the same level features saved in the previous down-sampling stage are fused through the skip connection; the tail of the noise prediction network is in the order of residual block, spatial attention, temporal attention, residual block, and convolutional layer, which outputs the predicted noise and completes the noise estimation of the hidden space feature map When estimating the noise component in the latent space feature map After, based on the estimate, the latent space feature map is de-noised by backpropagation, according to the following formula: (1); wherein, is a Gaussian noise, used to add randomness to the denoising process; is a noise scheduling parameter, indicating the noise intensity at time step to the latent space feature map , and defines , and ; by sequentially iterating formula (1) at each time step, the latent space feature map of the target CT volume is gradually denoised according to the total denoising time step , to generate the latent space feature map of the target CT volume .
4. The limited-angle CT reconstruction method based on the combination of the three-dimensional conditional diffusion model and synchronous iteration according to claim 3, characterized in that The network structure of the decoder adopts a dual-branch architecture of a spatial domain decoder and a frequency domain decoder; the spatial domain decoder, as a symmetric part of the encoder, is composed of a three-dimensional convolution, a residual block and an up-sampling layer, and is used to recover the latent space feature map step by step ; the frequency domain decoder firstly performs a fast Fourier transform on the latent space feature map, then returns to the spatial domain through an inverse fast Fourier transform, and further maps the high-frequency enhanced features through a convolution layer. Then, the high-frequency enhancement features are fused with the spatial domain decoder main branch, and the step-by-step decoding is completed through the residual block and three-dimensional convolution. 5.The method of claim 1, wherein, The specific process of training the three-dimensional conditional diffusion model is as follows: (1) training the encoder and the decoder: the encoder is responsible for mapping the input three-dimensional body to the hidden space feature map, and the decoder restores the corresponding three-dimensional body according to the hidden space feature map; the two are jointly optimized as a whole during the training process: the hidden space feature map output by the encoder is directly used as the input of the decoder to realize end-to-end learning; Randomly extract full-angle FBP reconstructed tomographic images from the training dataset prepared in Step 1 , which is input into the combined network of the encoder and the decoder, to obtain an output ; the tomographic image is used as a supervisory signal, and the goal of the training is to make the output as consistent as possible with the input; the loss function uses mean square error, which is defined as follows: (3); Repeat the above training process to jointly train the encoder and decoder; when the loss function After convergence, the training process is complete, and the network parameters are frozen for subsequent network training. (2) training the conditional embedding module: the conditional embedding module is trained in cascade with the trained decoder, wherein the parameters of the trained decoder are fixed, and only the conditional embedding module is trained. Randomly select 90 projections in the range of 0°-90° from the training dataset prepared in step one, and obtain the tomographic image by FBP reconstruction , as a training input, through the conditional embedding module and the decoder output ; take the tomographic image as the supervision signal, and the loss function adopts the mean square error, which is defined as follows: (4); The conditional embedding module is trained by repeating the above training process; when the loss function After convergence, the training process is completed, and the network parameters are frozen for subsequent network training; (3) Training the noise prediction network: keeping the parameters of the trained encoder, conditional embedding module and decoder fixed, only training the noise prediction network: first, randomly select a data pair from the training data set prepared in step one, where 90 projections in the range of 0°-90° are input into the conditional embedding module after FBP reconstruction to obtain the conditional features ; in the data pair, the tomographic image obtained by FBP reconstruction of the full angle is input into the encoder to obtain its corresponding hidden space feature map ; select the total denoising time step , randomly select a time step from , and add noise to obtain the noisy hidden representation , that is (5); wherein, is noise sampled from a standard normal distribution ; then, the latent space feature map , the conditional feature , and the time step are jointly input into a noise prediction network, and a predicted noise is output; the training objective is to accurately estimate the noise component contained in the latent space feature map at any time step, and the loss function is defined as follows: (6); wherein, represents the real noise, is the output of the noise prediction network; in the training process, a time step is randomly selected and iteratively optimized multiple times to ensure that the three-dimensional conditional diffusion model accurately predicts the noise at different time steps; when the loss function converges, the training process is completed.
Citation Information
Patent Citations
CT reconstruction system and method based on discrete Gaussian representation
CN120279209A
Sparse angle CT reconstruction method and model training method based on deep learning
CN120388093A
Double-view industrial CT fault online reconstruction method based on hidden space condition diffusion model
CN117635745A
Unsupervised sparse CT image reconstruction method based on dual-frequency refinement diffusion prior
CN120411291A