Single-view x-ray three-dimensional CT reconstruction method based on self-optimizing double-domain cyclic diffusion probability model

By establishing a cyclical interaction mechanism between the CT volume domain and the projection domain through a self-optimizing dual-domain cyclic diffusion probability model, the problems of insufficient observation information and error accumulation in single-view X-ray 3D CT reconstruction are solved, and 3D CT reconstruction with high accuracy and stability is achieved.

CN122289470APending Publication Date: 2026-06-26CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY
Filing Date
2026-02-25
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing single-view X-ray 3D CT reconstruction methods suffer from uncertainties due to insufficient observation information and error accumulation, resulting in insufficient structural consistency and geometric accuracy of the reconstruction results, making it difficult to meet the rapid imaging needs in low-resource environments.

Method used

A self-optimizing dual-domain cyclic diffusion probability model is adopted. By establishing a cyclic interaction mechanism between the CT volume domain and the projection domain, and combining the perception transformation module and the condition generation module, cross-domain cyclic correction and self-optimization are achieved, thereby enhancing the structural stability and projection physical consistency of the reconstruction process.

Benefits of technology

It significantly improves the accuracy and reliability of reconstructing 3D CT from single-view X-ray images, reduces artifacts and anatomical distortion, and enhances the geometric consistency of the reconstruction results and the ability to express the details of anatomical structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This invention proposes a single-view X-ray 3D CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model, aiming to solve the problem of degraded reconstruction quality caused by severe information loss in single-view CT in traditional CT. This method constructs diffusion denoising processes in both the projection domain and the CT volume domain, and establishes a bidirectional mapping relationship between forward and back projection through a perceptual transformation module. Cyclic consistency constraints are formed during the back diffusion process, achieving cross-domain structural collaborative correction. Simultaneously, a condition generation module based on structural entropy is designed to quantify and dynamically filter uncertainties in multi-view features to enhance model stability. Furthermore, a self-optimizing mechanism is introduced, performing multiple rounds of denoising and error correction within each back diffusion time step to gradually correct prediction bias and suppress error accumulation. By collaboratively optimizing the projection domain and the CT volume domain through a dual-domain joint loss function, a unity of physical consistency and structural realism is achieved, thereby realizing high-precision 3D CT reconstruction under single-view X-ray conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical imaging and deep learning research, specifically involving a single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model. Background Technology

[0002] With the rapid development of modern medical imaging technology, computed tomography (CT) has become an indispensable key imaging tool in clinical diagnosis. CT scans the human body from multiple angles using X-rays and combines this with tomographic reconstruction algorithms to obtain three-dimensional anatomical information with high spatial resolution. This significantly improves doctors' ability to locate and assess the structure of lesions, and has important clinical value, especially in the detailed structural observation and lesion detection of moving organs such as the lungs and liver.

[0003] However, multi-angle X-ray projection acquisition typically relies on complex rotating scanning systems, which significantly increases cumulative radiation dose and imaging time, and is difficult to meet the rapid imaging needs of emergency imaging, bedside detection, and motion-restricted scenarios. Furthermore, the system cost and data redundancy resulting from high-density angle sampling also limit its widespread application in low-resource environments. In contrast, single-angle X-ray reconstruction of 3D CT achieves volumetric structure restoration under extremely underdetermined conditions, providing a novel technical approach for low-dose, high-efficiency imaging, and has significant scientific and clinical value for promoting rapid diagnosis and accurate image reconstruction.

[0004] In recent years, denoising diffusion probabilistic models have made significant progress in 3D shape reconstruction tasks. However, existing methods typically perform the denoising process through a single forward propagation at each sampling time step, lacking an effective error correction mechanism. This easily leads to the accumulation of errors during the progressive generation process, weakening the geometric consistency between the generated results and the input observations, and limiting the ability to accurately recover complex structures and high-frequency details. Furthermore, when only a single image is used as the conditional input during the testing phase, the extremely limited observable information makes it difficult for the model to adequately constrain the spatial distribution of the 3D structure, resulting in a significant increase in the solution space during back-diffusion, thus introducing higher uncertainty. This uncertainty further weakens the structural consistency and geometric accuracy of the generated results, thereby significantly reducing the overall quality and structural reliability of 3D shape reconstruction. Summary of the Invention

[0005] To address the aforementioned problems, this invention proposes a single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model. The specific technical solution adopted is as follows:

[0006] 1. A single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model, characterized in that the system comprises:

[0007] Dual-domain cyclic diffusion strategy: To alleviate the severe ill-posedness caused by insufficient observation information during the reconstruction from single-view X-ray to 3D CT, this method constructs diffusion denoising processes in both the CT volume domain and the projection domain. A cyclic interaction mechanism between the two domains is established through a perceptual transformation module. This allows the deterministic full-view projection image generated in the projection domain to be back-projected into a CT image to guide denoising in the CT volume domain. Simultaneously, the deterministic CT image generated in the CT volume domain can be fed back to the projection domain through forward projection to further optimize the projection generation process. Under this cross-domain cyclic consistency constraint, the results from both domains achieve collaborative correction and gradual convergence at the geometric and anatomical levels of X-ray imaging. This effectively enhances the structural stability and projection physical consistency of the reconstruction process, significantly reduces artifacts and anatomical distortion, and improves the accuracy and reliability of 3D CT reconstruction from single-view X-ray images.

[0008] Conditional Generation Module: During the training phase, multi-scale structural features are first extracted from multi-angle X-ray projection data. Based on structural entropy modeling, the uncertainty of the spatial structure is quantitatively characterized, thereby portraying the structural reliability of different regions. On this basis, a non-parametric degradation and dynamic filtering mechanism is implemented on the conditional features to suppress interference from redundant or highly uncertain information. Conditional variables with both structural awareness and stage-adaptive characteristics are constructed as important prior constraints for the diffusion generation process. In the initial training phase, the projection domain denoising network and the CT volume domain denoising network use single-view projection images respectively. Conditional variables after fusion with multi-view images As a weak condition, the basic mapping relationship for recovering the target structure is progressively learned from limited observation information. As the training process progresses, the full-view projection images and corresponding 3D CT images generated by the two diffusion models at each time step are remapped to the other domain through a perceptual transformation module and fused with the original weak conditions to form a strong condition input containing richer anatomical structural information and imaging geometric constraints. Under this progressive condition enhancement mechanism from weak to strong conditions, the diffusion model can continuously correct the generated results in cross-domain cyclic feedback, significantly reducing structural uncertainty in the generation process, while strengthening the consistency of the reconstruction results at the anatomical structure level and in the physical imaging process, thereby effectively improving the stability and accuracy of single-view 3D CT reconstruction.

[0009] Perception Transformation Module: During the diffusion reconstruction process, to establish a physical consistency between the projection domain and the CT volume domain, this module explicitly constructs the forward and backward projection transformation relationship based on the X-ray cone-beam imaging geometric model, realizing bidirectional mapping and cyclic constraints between the two domains. Specifically, during the backward diffusion process in the projection domain, the full-view projection image generated at each time step is mapped to the three-dimensional volume space through a backward projection operator based on system imaging parameters, forming the corresponding CT image, providing structural reference and spatial guidance for the backward diffusion process in the CT volume domain. Simultaneously, during the backward diffusion process in the CT volume domain, the three-dimensional CT image generated at each time step is remapped to the projection space through a forward projection operator, generating an X-ray projection image consistent with the detector imaging geometry, and fed back to the diffusion process in the projection domain to correct the projection generation results.

[0010] Self-optimization strategy: To effectively address the problem of errors accumulating gradually in a single denoising prediction in traditional diffusion models, a recursive self-optimization mechanism is introduced at each time step of the diffusion inverse process. Specifically, this mechanism uses the previous denoising result as the input for the current stage, performing multiple rounds of recursive optimization on the noise prediction while keeping the conditional information unchanged. This allows the model to gradually correct prediction biases generated in the early stages and continuously approximate the true data distribution. Through this stepwise self-correcting optimization process, not only is the accumulation and propagation of errors during diffusion significantly suppressed, but the model's ability to model complex 3D structural details is also enhanced, thereby generating CT reconstruction results with higher geometric consistency and richer anatomical details.

[0011] 2. A single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model, characterized by comprising the following steps:

[0012] (a) Construct a paired training dataset, with each sample containing labels of real 3D CT images. And its corresponding full-view X-ray projection image labels (360 angles in total). Each real 3D CT image is labeled with its corresponding single real X-ray projection image at any angle. As a single-view condition input, it ensures the physical consistency and spatial correspondence between single-view observation, full-view projection and 3D CT.

[0013] (b) Construct a self-optimizing dual-domain cyclic diffusion model, which includes a projection domain denoising network. and CT volume denoising network and its corresponding condition generation module and ;

[0014] (c) Forward diffusion process: on three-dimensional CT images and full-view X-ray projection images Gaussian noise is added gradually, and its expression is:

[0015]

[0016]

[0017] in, This is a cumulative noise scheduler, which determines the proportion of the original image information remaining at time step t. It is standard Gaussian noise;

[0018] (d) In the projection domain condition generation module In this approach, a single X-ray image from any viewpoint is used as a conditional guide for the projection domain denoising network. This applies to the CT volume domain conditional generation module. First, multi-view X-ray images Normalization preprocessing is performed, and multi-scale features are extracted through a weighted encoding network to form a candidate conditional feature set. Subsequently, the spatial uncertainty of the candidate features is measured by structural entropy modeling, achieving nonparametric degradation and adaptive filtering of the feature set, thereby suppressing interference from unreliable or redundant viewpoint information. Based on this, the filtered features are pooled and fused to obtain a viewpoint-robust global conditional feature representation. Furthermore, the mean and variance of the conditional probability distribution are predicted based on this fused feature, and random conditional embedding vectors are generated through reparameterized sampling. These vectors provide stable and uncertainty-modeling-capable conditional constraints for the CT volume domain denoising network at each time step of the inverse diffusion process, thereby enhancing the geometric consistency and detail representation capability of the 3D structure reconstruction.

[0019] (e) Denoising Network Training: Training the denoising network and This enables it to progressively predict noise based on the noisy image, time step, loop time and conditional features within each time step. .

[0020] 3. The single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model according to claim 2, characterized in that step d specifically includes:

[0021] (d1) Collection of multi-view X-ray images The high-dimensional semantic feature representation is extracted by feeding the data into the ReNet-18 two-dimensional coding network with shared weights. ;

[0022] (d2) For each semantic feature vector Normalized to a probability distribution:

[0023]

[0024] in, For perspective indexing, As an index for the feature dimensions, structural entropy is introduced to quantitatively model this distribution, which is defined as:

[0025]

[0026] (d3) Obtaining the feature vector for each viewpoint structural entropy Subsequently, it can be regarded as a measure of uncertainty of this perspective in the feature representation space. Since the entropy scale may differ between different samples, in order to construct a comparable perspective-level uncertainty representation, the structural entropy is normalized to obtain an uncertainty score:

[0027]

[0028] in, It is a constant to prevent the denominator from being zero, and it is used to construct structural credibility weights based on uncertainty scoring:

[0029]

[0030] (d4) Based on weights, multi-view feature sets Adaptive discarding and retention are employed to suppress the interference of high uncertainty perspectives on the conditional modeling process. Specifically, the discard probability is first constructed based on the weight distribution:

[0031]

[0032] in, This represents the probability that the k-th viewpoint feature is discarded, and it is related to the structural credibility weight. Inversely proportional, that is The lower the value, the greater the probability of it being discarded;

[0033] (d5) Subsequently, random gating variables are introduced to perform sampling screening of multi-view features. Specifically, gating masks are sampled from the Bernoulli distribution:

[0034]

[0035] in, This indicates that the feature vector of the k-th viewpoint is preserved. This indicates that the feature has been discarded. Here, Direct control of the first The probability that a view feature is set to zero. When When the value is larger, this feature is discarded during the sampling process. The higher the probability of ), the better; when The smaller the size, the more it is retained. The higher the probability of ), the more probabilistic feature selection driven by structural uncertainty is achieved;

[0036] (d6) After dynamic filtering, a new feature set is obtained. and to Max pooling is obtained , will integrate features Two independent fully connected blocks are fed into the system to predict the mean and variance of the conditional distribution. Subsequently, the final conditional features are obtained by reparameterizing the sample from this distribution. :

[0037]

[0038] This process can adaptively suppress unreliable structural perspective features, reduce the interference of noisy features on conditional modeling, and retain high-confidence structural information. Furthermore, it obtains stable global features through dynamic screening and max pooling fusion, and generates conditional embeddings with structural prior constraints through distributed modeling and random sampling, thereby significantly improving the robustness and structural consistency of conditional guidance.

[0039] 4. The single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model according to claim 2, characterized in that step e specifically comprises:

[0040] (e1) In the initial stage of reverse diffusion, the generation processes of the projection domain and the CT volume domain are independent of each other, so as to ensure that the model can learn the basic data distribution in their respective domains under weak constraints.

[0041] Specifically, during the inverse diffusion process in the projection domain, the denoising network Full-noise, full-view projection image at the moment the forward diffusion process terminates. As input, it is combined with single-view X-ray projection images at arbitrary angles. As a weakly conditional embedding, it outputs a deterministic full-view X-ray image. The corresponding posterior mean and variance are calculated based on the parameterization formula of the inverse process of the diffusion model, and from... Sampling is performed within a Gaussian distribution to obtain the noise image at the current time step. , It is then passed to the next time step as a new input state, thereby gradually restoring the complete multi-view projection structure.

[0042] In the CT volume domain back diffusion process, the denoising network Full-noise CT image at the moment of termination of the forward diffusion process As input, and in conjunction with the multi-view conditional features generated in step (d). As a weakly conditional embedding, it outputs a deterministic 3D CT image. The corresponding posterior mean and variance are calculated based on the parameterization formula of the inverse process of the diffusion model, and from... Sampling is performed within a Gaussian distribution to obtain the noise image at the current time step. , It is then passed to the next time step as a new input state, thereby gradually recovering the complete 3D CT structure.

[0043] In the initial training phase, since a cross-domain recurrent feedback mechanism had not yet been established, the backdiffusion process in both the projection domain and the CT volume domain always relied solely on single-view projection images. and multi-perspective conditional features Guided as a weak condition;

[0044] (e2) In the mid-to-late stages of the inverse diffusion process, a cross-domain loop and self-optimization mechanism is introduced. By constructing a geometric closed loop between the projection domain and the CT volume domain, recursive correction and consistency constraints of the results in the two domains in the physical imaging space are achieved. Specifically, at each time step, the back-projection operator based on system imaging parameters in the perception transformation module is used to perform deterministic generation of the full-view X-ray projection image in the projection domain. Perform backprojection to reconstruct the corresponding 3D CT image from the projection space. This is used to feed back noise reduction to the CT volume domain to guide the next time step. For the CT volume domain, the cone-beam geometry forward projection operator in the perception transform module is used to denoise the deterministically generated 3D CT volume data. Perform a forward projection transformation to regenerate a full-view projection image consistent with the current volume structure. This information is then fed back to the projection domain to guide the denoising process at the next time step, thereby gradually establishing geometric consistency constraints and cyclic self-correction optimization mechanisms between the projection domain and the CT volume domain during the inverse diffusion process.

[0045] 5. A single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model according to claim 2, characterized in that step e2 specifically comprises:

[0046] (e21) At the t-th time step in the projection domain, the denoising network Full-view noisy image generated at the last time step of the initial stage As input, denoising is performed under strong conditional guidance. Among them, single-view X-ray images... As a weak condition, the deterministic 3D CT image generated in the CT volume domain at the previous time step. The projected image obtained by forward projection using the cone-beam geometric forward projection operator. After being stitched together, they form strong conditions to jointly guide the denoising process at the current time step, resulting in a deterministic projected image for that time step. To ensure the structural accuracy of the generated results, the deterministic full-view projection image generated by the denoising network is used throughout the entire training phase (including the initial phase). With real full-view projection label image Calculate the reconstruction consistency loss to directly constrain the convergence of the network's predicted image to the true projected label image:

[0047]

[0048] Furthermore, on After random sampling, the input image for the next time step is obtained. The entire process can be described as follows:

[0049]

[0050] in, It is a deterministic projected image generated at time step t in the projection domain. The mean, For variance, These are the parameters of the projection domain denoising network. It is a single-view X-ray image from any angle. and the projected image generated by the cone-beam geometric forward projection operator The strong condition formed by splicing is used to guide denoising. Throughout the training phase (including the initial phase), to ensure statistical consistency between the sampling results and the true data distribution, a variational lower bound loss (VLB) is introduced to constrain the backdiffusion distribution learned by the model to approximate the true posterior distribution.

[0051]

[0052] in This represents the true probability distribution at that time step. Given the probability distribution predicted by the denoising network at this time step, the denoising network optimizes the mean using variational lower bound loss. and variance To conform to the true probability distribution. Throughout the mid-to-late stages (excluding the initial stage), to further enhance the consistency between the projection generation results and the actual imaging physics, the CT images reconstructed from the CT volume domain will be... Projected image obtained by forward projection Corresponding real full-view projection label image Construction and reconstruction losses:

[0053]

[0054] (e22) At time step t in the CT volume domain, the denoising network Noisy CT image generated at the last time step of the initial stage As input, denoising is performed under strong conditional guidance. Specifically, the multi-view conditional features generated by the conditional generation module in step (d) are... As a weak condition, and the deterministic full-view projection image generated in the projection domain at the previous time step, Three-dimensional CT images obtained by backprojection using a backprojection operator based on system imaging parameters. After being stitched together, they form strong conditions to jointly guide the denoising process at the current time step, resulting in a deterministic projected image for that time step. To ensure the structural accuracy of the generated results, the deterministic CT images generated by the denoising network are used throughout the entire training phase (including the initial phase). Compared with real CT label images Calculate the consistency loss of CT volumetric reconstruction:

[0055]

[0056] Furthermore, on After random sampling, the input image for the next time step is obtained. The entire process can be described as follows:

[0057]

[0058] in, It is a deterministic CT image generated at time step t in the CT volume region. The mean, For variance, These are the parameters of the CT volume denoising network. The weak condition is generated by step (d). The CT image generated by backprojection in the projection domain at the previous time step The strong condition formed by splicing is used to guide denoising. Throughout the training phase (including the initial phase), to ensure statistical consistency between the sampling results and the true data distribution, a variational lower bound loss (VLB) is introduced to constrain the backdiffusion distribution learned by the model to approximate the true posterior distribution.

[0059]

[0060] in This represents the true probability distribution at that time step. Given the probability distribution predicted by the denoising network at this time step, the denoising network optimizes the mean using variational lower bound loss. and variance To conform to the true probability distribution, throughout the mid-to-late stages (excluding the initial stage), to ensure the physical consistency and structural accuracy of the reconstruction results, the deterministic full-view X-ray image generated by the projection domain denoising network will be used. CT images obtained by back projection Corresponding real CT image labels Construct the reconstruction loss function:

[0061]

[0062] (e23) Throughout the mid-to-late training phase, a recursive self-optimizing diffusion mechanism is designed based on dual-domain cyclic diffusion. This mechanism achieves gradual adaptive optimization of the current reconstruction result by performing multiple recursive denoising and error corrections within each inverse diffusion time step. Specifically, in the standard diffusion model, the denoising network performs only one forward prediction to directly obtain the estimation result for the next time step. This single prediction method easily leads to the gradual accumulation of noise estimation errors, thus compromising the geometric consistency of the reconstruction result. In this paper, at each time step, the current denoising result is re-input into the same denoising network, and denoising and resampling are repeatedly performed under the guidance of input conditional features. This allows the model to use the previous prediction result as feedback to recursively correct its errors and gradually approximate the true structure, thus forming a self-optimizing process. This mechanism enables the model to continuously evaluate and correct its own prediction bias during the inverse diffusion process, gradually strengthening the geometric consistency and detail representation ability of the target structure, effectively suppressing error accumulation, and significantly improving the reconstruction accuracy and stability of complex 3D structures. Specifically, within each time step t in the dual domain, denoising is performed C times in a cyclic manner, where the S-th denoising step can be expressed as:

[0063]

[0064]

[0065] Where S is the S-th cycle time within the t-th time step. and These are the images generated by the projection domain and the CT volume domain at the (S+1)th loop time (the previous loop time) within the t-th time step. The last loop time within the t-th time step is the input for the (t-1)th time step. It is evident that this mechanism does not directly obtain the input from the projection domain and the CT volume domain in one step. Transition to Instead of using a fixed loop, C intermediate loop steps are introduced. This allows the denoising model to utilize intermediate outputs. and It iteratively corrects the errors introduced in the previous step. This differs from the standard diffusion method, where the lack of intermediate correction steps leads to the accumulation of errors and a decrease in the quality of the final output.

[0066] (e24) In the projection domain, the total training loss function is defined as:

[0067]

[0068] In the CT volume domain, the total training loss function is defined as:

[0069]

[0070] in, and where represents the weighting coefficient. To achieve coordinated optimization of the projection domain and the CT volume domain, ensuring that the two domains mutually constrain each other and gradually converge to the true data distribution during structural restoration and probability distribution evolution, a dual-domain joint loss function is used to uniformly optimize the entire cyclic diffusion model. This loss function integrates the supervision information from both the projection domain and the CT volume domain, thereby guaranteeing the structural and physical consistency of the reconstruction results in both the projection space and the volume domain space. Its expression is:

[0071] Attached Figure Description

[0072] Figure 1 Overall framework diagram of the present invention

[0073] Figure 2 A schematic diagram of the S-th time loop in the t-th time step of the method of the present invention.

[0074] Figure 3 Flowchart of Condition Generation Module Detailed Implementation

[0075] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments. (See attached figures) Figures 1 to 3 As shown, a single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model is characterized by the following steps:

[0076] (a) Construct a paired training dataset, with each sample containing labels of real 3D CT images. And its corresponding full-view X-ray projection image labels (360 angles in total). Each real 3D CT image is labeled with its corresponding single real X-ray projection image at any angle. As a single-view condition input, it ensures the physical consistency and spatial correspondence between single-view observation, full-view projection and 3D CT.

[0077] (b) Construct a self-optimizing dual-domain cyclic diffusion model, which includes a projection domain denoising network. and CT volume denoising network and its corresponding condition generation module and ;

[0078] (c) Forward diffusion process: on three-dimensional CT images and full-view X-ray projection images Gaussian noise is added gradually, and its expression is:

[0079]

[0080]

[0081] in, This is a cumulative noise scheduler, which determines the proportion of the original image information remaining at time step t. It is standard Gaussian noise;

[0082] (d) In the projection domain condition generation module In this approach, a single X-ray image from any viewpoint is used as a conditional guide for the projection domain denoising network. This applies to the CT volume domain conditional generation module. First, multi-view X-ray images Normalization preprocessing is performed, and multi-scale features are extracted through a weighted encoding network to form a candidate conditional feature set. Subsequently, the spatial uncertainty of the candidate features is measured by structural entropy modeling, achieving nonparametric degradation and adaptive filtering of the feature set, thereby suppressing interference from unreliable or redundant viewpoint information. Based on this, the filtered features are pooled and fused to obtain a viewpoint-robust global conditional feature representation. Furthermore, the mean and variance of the conditional probability distribution are predicted based on this fused feature, and random conditional embedding vectors are generated through reparameterized sampling. These vectors provide stable and uncertainty-modeling-capable conditional constraints for the CT volume domain denoising network at each time step of the inverse diffusion process, thereby enhancing the geometric consistency and detail representation capability of the 3D structure reconstruction.

[0083] (e) Denoising Network Training: Training the denoising network and This enables it to progressively predict noise based on the noisy image, time step, loop time and conditional features within each time step. .

[0084] 3. The single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model according to claim 2, characterized in that step d specifically includes:

[0085] (d1) Collection of multi-view X-ray images The high-dimensional semantic feature representation is extracted by feeding the data into the ReNet-18 two-dimensional coding network with shared weights. ;

[0086] (d2) For each semantic feature vector Normalized to a probability distribution:

[0087]

[0088] in, For perspective indexing, As an index for the feature dimensions, structural entropy is introduced to quantitatively model this distribution, which is defined as:

[0089]

[0090] (d3) Obtaining the feature vector for each viewpoint structural entropy Subsequently, it can be regarded as a measure of uncertainty of this perspective in the feature representation space. Since the entropy scale may differ between different samples, in order to construct a comparable perspective-level uncertainty representation, the structural entropy is normalized to obtain an uncertainty score:

[0091]

[0092] in, It is a constant to prevent the denominator from being zero, and it is used to construct structural credibility weights based on uncertainty scoring:

[0093]

[0094] (d4) Based on weights, multi-view feature sets Adaptive discarding and retention are employed to suppress the interference of high uncertainty perspectives on the conditional modeling process. Specifically, the discard probability is first constructed based on the weight distribution:

[0095]

[0096] in, This represents the probability that the k-th viewpoint feature is discarded, and it is related to the structural credibility weight. Inversely proportional, that is The lower the value, the greater the probability of it being discarded;

[0097] (d5) Subsequently, random gating variables are introduced to perform sampling screening of multi-view features. Specifically, gating masks are sampled from the Bernoulli distribution:

[0098]

[0099] in, This indicates that the feature vector of the k-th viewpoint is preserved. This indicates that the feature has been discarded. Here, Direct control of the first The probability that a view feature is set to zero. When When the value is larger, this feature is discarded during the sampling process. The higher the probability of ), the better; when The smaller the size, the more it is retained. The higher the probability of ), the more probabilistic feature selection driven by structural uncertainty is achieved;

[0100] (d6) After dynamic filtering, a new feature set is obtained. and to Max pooling is obtained , will integrate features Two independent fully connected blocks are fed into the system to predict the mean and variance of the conditional distribution. Subsequently, the final conditional features are obtained by reparameterizing the sample from this distribution. :

[0101]

[0102] This process can adaptively suppress unreliable structural perspective features, reduce the interference of noisy features on conditional modeling, and retain high-confidence structural information. Furthermore, it obtains stable global features through dynamic screening and max pooling fusion, and generates conditional embeddings with structural prior constraints through distributed modeling and random sampling, thereby significantly improving the robustness and structural consistency of conditional guidance.

[0103] 4. The single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model according to claim 2, characterized in that step e specifically comprises:

[0104] (e1) In the initial stage of reverse diffusion, the generation processes of the projection domain and the CT volume domain are independent of each other, so as to ensure that the model can learn the basic data distribution in their respective domains under weak constraints.

[0105] Specifically, during the inverse diffusion process in the projection domain, the denoising network Full-noise, full-view projection image at the moment the forward diffusion process terminates. As input, it is combined with single-view X-ray projection images at arbitrary angles. As a weakly conditional embedding, it outputs a deterministic full-view X-ray image. The corresponding posterior mean and variance are calculated based on the parameterization formula of the inverse process of the diffusion model, and from... Sampling is performed within a Gaussian distribution to obtain the noise image at the current time step. , It is then passed to the next time step as a new input state, thereby gradually restoring the complete multi-view projection structure.

[0106] In the CT volume domain back diffusion process, the denoising network Full-noise CT image at the moment of termination of the forward diffusion process As input, and in conjunction with the multi-view conditional features generated in step (d). As a weakly conditional embedding, it outputs a deterministic 3D CT image. The corresponding posterior mean and variance are calculated based on the parameterization formula of the inverse process of the diffusion model, and from... Sampling is performed within a Gaussian distribution to obtain the noise image at the current time step. , It is then passed to the next time step as a new input state, thereby gradually recovering the complete 3D CT structure.

[0107] In the initial training phase, since a cross-domain recurrent feedback mechanism had not yet been established, the backdiffusion process in both the projection domain and the CT volume domain always relied solely on single-view projection images. and multi-perspective conditional features Guided as a weak condition;

[0108] (e2) In the mid-to-late stages of the inverse diffusion process, a cross-domain loop and self-optimization mechanism is introduced. By constructing a geometric closed loop between the projection domain and the CT volume domain, recursive correction and consistency constraints of the results in the two domains in the physical imaging space are achieved. Specifically, at each time step, the back-projection operator based on system imaging parameters in the perception transformation module is used to perform deterministic generation of the full-view X-ray projection image in the projection domain. Perform backprojection to reconstruct the corresponding 3D CT image from the projection space. This is used to feed back noise reduction to the CT volume domain to guide the next time step. For the CT volume domain, the cone-beam geometry forward projection operator in the perception transform module is used to denoise the deterministically generated 3D CT volume data. Perform a forward projection transformation to regenerate a full-view projection image consistent with the current volume structure. This information is then fed back to the projection domain to guide the denoising process at the next time step, thereby gradually establishing geometric consistency constraints and cyclic self-correction optimization mechanisms between the projection domain and the CT volume domain during the inverse diffusion process.

[0109] 5. A single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model according to claim 2, characterized in that step e2 specifically comprises:

[0110] (e21) At the t-th time step in the projection domain, the denoising network Full-view noisy image generated at the last time step of the initial stage As input, denoising is performed under strong conditional guidance. Among them, single-view X-ray images... As a weak condition, the deterministic 3D CT image generated in the CT volume domain at the previous time step. The projected image obtained by forward projection using the cone-beam geometric forward projection operator. After being stitched together, they form strong conditions to jointly guide the denoising process at the current time step, resulting in a deterministic projected image for that time step. To ensure the structural accuracy of the generated results, the deterministic full-view projection image generated by the denoising network is used throughout the entire training phase (including the initial phase). With real full-view projection label image Calculate the reconstruction consistency loss to directly constrain the convergence of the network's predicted image to the true projected label image:

[0111]

[0112] Furthermore, on After random sampling, the input image for the next time step is obtained. The entire process can be described as follows:

[0113]

[0114] in, It is a deterministic projected image generated at time step t in the projection domain. The mean, For variance, These are the parameters of the projection domain denoising network. It is a single-view X-ray image from any angle. and the projected image generated by the cone-beam geometric forward projection operator The strong condition formed by splicing is used to guide denoising. Throughout the training phase (including the initial phase), to ensure statistical consistency between the sampling results and the true data distribution, a variational lower bound loss (VLB) is introduced to constrain the backdiffusion distribution learned by the model to approximate the true posterior distribution.

[0115]

[0116] in This represents the true probability distribution at that time step. Given the probability distribution predicted by the denoising network at this time step, the denoising network optimizes the mean using variational lower bound loss. and variance To conform to the true probability distribution. Throughout the mid-to-late stages (excluding the initial stage), to further enhance the consistency between the projection generation results and the actual imaging physics, the CT images reconstructed from the CT volume domain will be... Projected image obtained by forward projection Corresponding real full-view projection label image Construction and reconstruction losses:

[0117]

[0118] (e22) At time step t in the CT volume domain, the denoising network Noisy CT image generated at the last time step of the initial stage As input, denoising is performed under strong conditional guidance. Specifically, the multi-view conditional features generated by the conditional generation module in step (d) are... As a weak condition, and the deterministic full-view projection image generated in the projection domain at the previous time step, Three-dimensional CT images obtained by backprojection using a backprojection operator based on system imaging parameters. After being stitched together, they form strong conditions to jointly guide the denoising process at the current time step, resulting in a deterministic projected image for that time step. To ensure the structural accuracy of the generated results, the deterministic CT images generated by the denoising network are used throughout the entire training phase (including the initial phase). Compared with real CT label images Calculate the consistency loss of CT volumetric reconstruction:

[0119]

[0120] Furthermore, on After random sampling, the input image for the next time step is obtained. The entire process can be described as follows:

[0121]

[0122] in, It is a deterministic CT image generated at time step t in the CT volume region. The mean, For variance, These are the parameters of the CT volume denoising network. The weak condition is generated by step (d). The CT image generated by backprojection in the projection domain at the previous time step The strong condition formed by splicing is used to guide denoising. Throughout the training phase (including the initial phase), to ensure statistical consistency between the sampling results and the true data distribution, a variational lower bound loss (VLB) is introduced to constrain the backdiffusion distribution learned by the model to approximate the true posterior distribution.

[0123]

[0124] in This represents the true probability distribution at that time step. Given the probability distribution predicted by the denoising network at this time step, the denoising network optimizes the mean using variational lower bound loss. and variance To conform to the true probability distribution, throughout the mid-to-late stages (excluding the initial stage), to ensure the physical consistency and structural accuracy of the reconstruction results, the deterministic full-view X-ray image generated by the projection domain denoising network will be used. CT images obtained by back projection Corresponding real CT image labels Construct the reconstruction loss function:

[0125]

[0126] (e23) Throughout the mid-to-late training phase, a recursive self-optimizing diffusion mechanism is designed based on dual-domain cyclic diffusion. This mechanism achieves gradual adaptive optimization of the current reconstruction result by performing multiple recursive denoising and error corrections within each inverse diffusion time step. Specifically, in the standard diffusion model, the denoising network performs only one forward prediction to directly obtain the estimation result for the next time step. This single prediction method easily leads to the gradual accumulation of noise estimation errors, thus compromising the geometric consistency of the reconstruction result. In this paper, at each time step, the current denoising result is re-input into the same denoising network, and denoising and resampling are repeatedly performed under the guidance of input conditional features. This allows the model to use the previous prediction result as feedback to recursively correct its errors and gradually approximate the true structure, thus forming a self-optimizing process. This mechanism enables the model to continuously evaluate and correct its own prediction bias during the inverse diffusion process, gradually strengthening the geometric consistency and detail representation ability of the target structure, effectively suppressing error accumulation, and significantly improving the reconstruction accuracy and stability of complex 3D structures. Specifically, within each time step t in the dual domain, denoising is performed C times in a cyclic manner, where the S-th denoising step can be expressed as:

[0127]

[0128]

[0129] Where S is the S-th cycle time within the t-th time step. and These are the images generated by the projection domain and the CT volume domain at the (S+1)th loop time (the previous loop time) within the t-th time step. The last loop time within the t-th time step is the input for the (t-1)th time step. It is evident that this mechanism does not directly obtain the input from the projection domain and the CT volume domain in one step. Transition to Instead of using a fixed loop, C intermediate loop steps are introduced. This allows the denoising model to utilize intermediate outputs. and It iteratively corrects the errors introduced in the previous step. This differs from the standard diffusion method, where the lack of intermediate correction steps leads to the accumulation of errors and a decrease in the quality of the final output.

[0130] (e24) In the projection domain, the total training loss function is defined as:

[0131]

[0132] In the CT volume domain, the total training loss function is defined as:

[0133]

[0134] in, and where represents the weighting coefficient. To achieve coordinated optimization of the projection domain and the CT volume domain, ensuring that the two domains mutually constrain each other and gradually converge to the true data distribution during structural restoration and probability distribution evolution, a dual-domain joint loss function is used to uniformly optimize the entire cyclic diffusion model. This loss function integrates the supervision information from both the projection domain and the CT volume domain, thereby guaranteeing the structural and physical consistency of the reconstruction results in both the projection space and the volume domain space. Its expression is:

[0135]

Claims

1. A single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model, characterized in that, The system includes: Dual-domain cyclic diffusion strategy: To alleviate the severe ill-posedness caused by insufficient observation information during the reconstruction from single-view X-ray to 3D CT, this method constructs diffusion denoising processes in both the CT volume domain and the projection domain. A cyclic interaction mechanism between the two domains is established through a perceptual transformation module. This allows the deterministic full-view projection image generated in the projection domain to be back-projected into a CT image to guide denoising in the CT volume domain. Simultaneously, the deterministic CT image generated in the CT volume domain can be fed back to the projection domain through forward projection to further optimize the projection generation process. Under this cross-domain cyclic consistency constraint, the results from both domains achieve collaborative correction and gradual convergence at the geometric and anatomical levels of X-ray imaging. This effectively enhances the structural stability and projection physical consistency of the reconstruction process, significantly reduces artifacts and anatomical distortion, and improves the accuracy and reliability of 3D CT reconstruction from single-view X-ray images. Conditional Generation Module: During the training phase, multi-scale structural features are first extracted from multi-angle X-ray projection data. Based on structural entropy modeling, the uncertainty of the spatial structure is quantitatively characterized, thereby portraying the structural reliability of different regions. On this basis, a non-parametric degradation and dynamic filtering mechanism is implemented on the conditional features to suppress interference from redundant or highly uncertain information. Conditional variables with both structural awareness and stage-adaptive characteristics are constructed as important prior constraints for the diffusion generation process. In the initial training phase, the projection domain denoising network and the CT volume domain denoising network use single-view projection images respectively. Conditional variables after fusion with multi-view images As a weak condition, the basic mapping relationship for recovering the target structure is progressively learned from limited observation information. As the training process progresses, the full-view projection images and corresponding 3D CT images generated by the two diffusion models at each time step are remapped to the other domain through a perceptual transformation module and fused with the original weak conditions to form a strong condition input containing richer anatomical structural information and imaging geometric constraints. Under this progressive condition enhancement mechanism from weak to strong conditions, the diffusion model can continuously correct the generated results in cross-domain cyclic feedback, significantly reducing structural uncertainty in the generation process, while strengthening the consistency of the reconstruction results at the anatomical structure level and in the physical imaging process, thereby effectively improving the stability and accuracy of single-view 3D CT reconstruction. Perception Transformation Module: During the diffusion reconstruction process, to establish a physical consistency between the projection domain and the CT volume domain, this module explicitly constructs the forward and backward projection transformation relationship based on the X-ray cone-beam imaging geometric model, realizing bidirectional mapping and cyclic constraints between the two domains. Specifically, during the backward diffusion process in the projection domain, the full-view projection image generated at each time step is mapped to the three-dimensional volume space through a backward projection operator based on system imaging parameters, forming the corresponding CT image, providing structural reference and spatial guidance for the backward diffusion process in the CT volume domain. Simultaneously, during the backward diffusion process in the CT volume domain, the three-dimensional CT image generated at each time step is remapped to the projection space through a forward projection operator, generating an X-ray projection image consistent with the detector imaging geometry, and fed back to the diffusion process in the projection domain to correct the projection generation results. Self-optimization strategy: To effectively address the problem of errors accumulating gradually in a single denoising prediction in traditional diffusion models, a recursive self-optimization mechanism is introduced at each time step of the diffusion inverse process. Specifically, this mechanism uses the previous denoising result as the input for the current stage, performing multiple rounds of recursive optimization on the noise prediction while keeping the conditional information unchanged. This allows the model to gradually correct prediction biases generated in the early stages and continuously approximate the true data distribution. Through this stepwise self-correcting optimization process, not only is the accumulation and propagation of errors during diffusion significantly suppressed, but the model's ability to model complex 3D structural details is also enhanced, thereby generating CT reconstruction results with higher geometric consistency and richer anatomical details.

2. A single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model, characterized in that, Includes the following steps: (a) Construct a paired training dataset, where each sample contains labels of real 3D CT images. And its corresponding full-view X-ray projection image labels (360 angles in total). Each real 3D CT image is labeled with its corresponding single real X-ray projection image at any angle. As a single-view condition input, it ensures the physical consistency and spatial correspondence between single-view observation, full-view projection, and 3D CT. (b) Construct a self-optimizing dual-domain cyclic diffusion model, which includes a projection domain denoising network. and CT volume denoising network and its corresponding condition generation module and ; (c) Forward diffusion process: on three-dimensional CT images and full-view X-ray projection images Gaussian noise is added gradually, and its expression is: in, This is a cumulative noise scheduler, which determines the proportion of the original image information remaining at time step t. It is standard Gaussian noise; (d) In the projection domain condition generation module In this approach, a single X-ray image from any viewpoint is used as a conditional guide for the projection domain denoising network. This applies to the CT volume domain conditional generation module. First, multi-view X-ray images Normalization preprocessing is performed, and multi-scale features are extracted through a weighted encoding network to form a candidate conditional feature set. Subsequently, the spatial uncertainty of the candidate features is measured by structural entropy modeling, achieving nonparametric degradation and adaptive filtering of the feature set, thereby suppressing interference from unreliable or redundant viewpoint information. Based on this, the filtered features are pooled and fused to obtain a viewpoint-robust global conditional feature representation. Furthermore, the mean and variance of the conditional probability distribution are predicted based on this fused feature, and random conditional embedding vectors are generated through reparameterized sampling. These vectors provide stable and uncertainty-modeling-capable conditional constraints for the CT volume domain denoising network at each time step of the inverse diffusion process, thereby enhancing the geometric consistency and detail representation capability of the 3D structure reconstruction. (e) Denoising Network Training: Training the denoising network and This enables it to progressively predict noise based on the noisy image, time step, loop time and conditional features within each time step. .

3. The single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model according to claim 2, characterized in that, Step d specifically includes: (d1) Collection of multi-view X-ray images The high-dimensional semantic feature representation is extracted by feeding the data into the ReNet-18 two-dimensional coding network with shared weights. ; (d2) For each semantic feature vector Normalized to a probability distribution: in, For perspective indexing, As an index for the feature dimensions, structural entropy is introduced to quantitatively model this distribution, which is defined as: (d3) Obtaining the feature vector for each viewpoint structural entropy Subsequently, it can be regarded as a measure of uncertainty of this perspective in the feature representation space. Since the entropy scale may differ between different samples, in order to construct a comparable perspective-level uncertainty representation, the structural entropy is normalized to obtain an uncertainty score: in, It is a constant to prevent the denominator from being zero, and it is used to construct structural credibility weights based on uncertainty scoring: (d4) Based on weights, multi-view feature sets Adaptive discarding and retention are employed to suppress the interference of high uncertainty perspectives on the conditional modeling process. Specifically, the discard probability is first constructed based on the weight distribution: in, This represents the probability that the k-th viewpoint feature is discarded, and it is related to the structural credibility weight. Inversely proportional, that is The lower the value, the greater the probability of it being discarded; (d5) Subsequently, random gating variables are introduced to perform sampling screening of multi-view features. Specifically, gating masks are sampled from the Bernoulli distribution: in, This indicates that the feature vector of the k-th viewpoint is preserved. This indicates that the feature has been discarded. Here, Direct control of the first The probability that a view feature is set to zero. When When the value is larger, this feature is discarded during the sampling process. The higher the probability of ), the better; when The smaller the size, the more it is retained. The higher the probability of ), the more probabilistic feature selection driven by structural uncertainty is achieved; (d6) After dynamic filtering, a new feature set is obtained. and to Max pooling is obtained , will integrate features Two independent fully connected blocks are fed into the system to predict the mean and variance of the conditional distribution. Subsequently, the final conditional features are obtained by reparameterizing the sample from this distribution. : This process can adaptively suppress unreliable structural perspective features, reduce the interference of noisy features on conditional modeling, and retain high-confidence structural information. Furthermore, it obtains stable global features through dynamic screening and max pooling fusion, and generates conditional embeddings with structural prior constraints through distributed modeling and random sampling, thereby significantly improving the robustness and structural consistency of conditional guidance.

4. The single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model according to claim 2, characterized in that, Step e specifically includes: (e1) In the initial stage of reverse diffusion, the generation processes of the projection domain and the CT volume domain are independent of each other, so as to ensure that the model can learn the basic data distribution in their respective domains under weak constraints. Specifically, during the inverse diffusion process in the projection domain, the denoising network Full-noise, full-view projection image at the moment the forward diffusion process terminates. As input, it is combined with single-view X-ray projection images at arbitrary angles. As a weakly conditional embedding, it outputs a deterministic full-view X-ray image. The corresponding posterior mean and variance are calculated based on the parameterization formula of the inverse process of the diffusion model, and from... Sampling is performed within a Gaussian distribution to obtain the noise image at the current time step. , It is then passed to the next time step as a new input state, thereby gradually restoring the complete multi-view projection structure. In the CT volume domain back diffusion process, the denoising network Full-noise CT image at the moment of termination of the forward diffusion process As input, and in conjunction with the multi-view conditional features generated in step (d). As a weakly conditional embedding, it outputs a deterministic 3D CT image. The corresponding posterior mean and variance are calculated based on the parameterization formula of the inverse process of the diffusion model, and from... Sampling is performed within a Gaussian distribution to obtain the noise image at the current time step. , It is then passed to the next time step as a new input state, thereby gradually recovering the complete 3D CT structure. In the initial training phase, since a cross-domain recurrent feedback mechanism had not yet been established, the backdiffusion process in both the projection domain and the CT volume domain always relied solely on single-view projection images. and multi-perspective conditional features Guided as a weak condition; (e2) In the mid-to-late stages of the inverse diffusion process, a cross-domain loop and self-optimization mechanism is introduced. By constructing a geometric closed loop between the projection domain and the CT volume domain, recursive correction and consistency constraints of the results in the two domains in the physical imaging space are achieved. Specifically, at each time step, the back-projection operator based on system imaging parameters in the perception transformation module is used to perform deterministic generation of the full-view X-ray projection image in the projection domain. Perform backprojection to reconstruct the corresponding 3D CT image from the projection space. This is used to feed back noise reduction to the CT volume domain to guide the next time step. For the CT volume domain, the cone-beam geometry forward projection operator in the perception transform module is used to denoise the deterministically generated 3D CT volume data. Perform a forward projection transformation to regenerate a full-view projection image consistent with the current volume structure. This information is then fed back to the projection domain to guide the denoising process at the next time step, thereby gradually establishing geometric consistency constraints and cyclic self-correction optimization mechanisms between the projection domain and the CT volume domain during the inverse diffusion process.

5. The single-view X-ray three-dimensional CT reconstruction method based on a self-optimizing dual-domain cyclic diffusion probability model according to claim 2, characterized in that, The specific content of step e2 is as follows: (e21) At the t-th time step in the projection domain, the denoising network Full-view noisy image generated at the last time step of the initial stage As input, denoising is performed under strong conditional guidance. Among them, single-view X-ray images... As a weak condition, the deterministic 3D CT image generated in the CT volume domain at the previous time step. The projected image obtained by forward projection using the cone-beam geometric forward projection operator. After being stitched together, they form strong conditions to jointly guide the denoising process at the current time step, resulting in a deterministic projected image for that time step. To ensure the structural accuracy of the generated results, the deterministic full-view projection image generated by the denoising network is used throughout the entire training phase (including the initial phase). With real full-view projection label image Calculate the reconstruction consistency loss to directly constrain the convergence of the network's predicted image to the true projected label image: Furthermore, on After random sampling, the input image for the next time step is obtained. The entire process can be described as follows: in, It is a deterministic projected image generated at time step t in the projection domain. The mean, For variance, These are the parameters of the projection domain denoising network. It is a single-view X-ray image from any angle. and the projected image generated by the cone-beam geometric forward projection operator The strong condition formed by splicing is used to guide denoising. Throughout the training phase (including the initial phase), to ensure statistical consistency between the sampling results and the true data distribution, a variational lower bound loss (VLB) is introduced to constrain the backdiffusion distribution learned by the model to approximate the true posterior distribution. in This represents the true probability distribution at that time step. Given the probability distribution predicted by the denoising network at this time step, the denoising network optimizes the mean using variational lower bound loss. and variance To conform to the true probability distribution. Throughout the mid-to-late stages (excluding the initial stage), to further enhance the consistency between the projection generation results and the actual imaging physics, the CT images reconstructed from the CT volume domain will be... Projected image obtained by forward projection Corresponding real full-view projection label image Construction and reconstruction losses: (e22) At time step t in the CT volume domain, the denoising network Noisy CT image generated at the last time step of the initial stage As input, denoising is performed under strong conditional guidance. Specifically, the multi-view conditional features generated by the conditional generation module in step (d) are... As a weak condition, and the deterministic full-view projection image generated in the projection domain at the previous time step, Three-dimensional CT images obtained by backprojection using a backprojection operator based on system imaging parameters. After being stitched together, they form strong conditions to jointly guide the denoising process at the current time step, resulting in a deterministic projected image for that time step. To ensure the structural accuracy of the generated results, the deterministic CT images generated by the denoising network are used throughout the entire training phase (including the initial phase). Compared with real CT label images Calculate the consistency loss of CT volumetric reconstruction: Furthermore, on After random sampling, the input image for the next time step is obtained. The entire process can be described as follows: in, It is a deterministic CT image generated at time step t in the CT volume region. The mean, For variance, These are the parameters of the CT volume denoising network. The weak condition is generated by step (d). The CT image generated by backprojection in the projection domain at the previous time step The strong condition formed by splicing is used to guide denoising. Throughout the training phase (including the initial phase), to ensure statistical consistency between the sampling results and the true data distribution, a variational lower bound loss (VLB) is introduced to constrain the backdiffusion distribution learned by the model to approximate the true posterior distribution. in This represents the true probability distribution at that time step. Given the probability distribution predicted by the denoising network at this time step, the denoising network optimizes the mean using variational lower bound loss. and variance To conform to the true probability distribution, throughout the mid-to-late stages (excluding the initial stage), to ensure the physical consistency and structural accuracy of the reconstruction results, the deterministic full-view X-ray image generated by the projection domain denoising network will be used. CT images obtained by back projection Corresponding real CT image labels Construct the reconstruction loss function: (e23) Throughout the mid-to-late training phase, a recursive self-optimizing diffusion mechanism is designed based on dual-domain cyclic diffusion. This mechanism achieves gradual adaptive optimization of the current reconstruction result by performing multiple recursive denoising and error corrections within each inverse diffusion time step. Specifically, in the standard diffusion model, the denoising network performs only one forward prediction to directly obtain the estimation result for the next time step. This single prediction method easily leads to the gradual accumulation of noise estimation errors, thus compromising the geometric consistency of the reconstruction result. In this paper, at each time step, the current denoising result is re-input into the same denoising network, and denoising and resampling are repeatedly performed under the guidance of input conditional features. This allows the model to use the previous prediction result as feedback to recursively correct its errors and gradually approximate the true structure, thus forming a self-optimizing process. This mechanism enables the model to continuously evaluate and correct its own prediction bias during the inverse diffusion process, gradually strengthening the geometric consistency and detail representation ability of the target structure, effectively suppressing error accumulation, and significantly improving the reconstruction accuracy and stability of complex 3D structures. Specifically, within each time step t in the dual domain, denoising is performed C times in a cyclic manner, where the S-th denoising step can be expressed as: Where S is the S-th cycle time within the t-th time step. and These are the images generated by the projection domain and the CT volume domain at the (S+1)th loop time (the previous loop time) within the t-th time step. The last loop time within the t-th time step is the input for the (t-1)th time step. It is evident that this mechanism does not directly obtain the input from the projection domain and the CT volume domain in one step. Transition to Instead of using a fixed loop, C intermediate loop steps are introduced. This allows the denoising model to utilize intermediate outputs. and It iteratively corrects the errors introduced in the previous step. This differs from the standard diffusion method, where the lack of intermediate correction steps leads to the accumulation of errors and a decrease in the quality of the final output. (e24) In the projection domain, the total training loss function is defined as: In the CT volume domain, the total training loss function is defined as: in, and where represents the weighting coefficient. To achieve coordinated optimization of the projection domain and the CT volume domain, ensuring that the two domains mutually constrain each other and gradually converge to the true data distribution during structural restoration and probability distribution evolution, a dual-domain joint loss function is used to uniformly optimize the entire cyclic diffusion model. This loss function integrates the supervision information from both the projection domain and the CT volume domain, thereby guaranteeing the structural and physical consistency of the reconstruction results in both the projection space and the volume domain space. Its expression is: