4D-CBCT dynamic reconstruction method

By integrating ConvLSTM and dual attention mechanism into a 4D-CBCT dynamic reconstruction method, and combining PCA model and deep learning network, the real-time performance and accuracy issues of 4D-CBCT dynamic reconstruction in existing technologies are solved, achieving high-quality 4D-CBCT image reconstruction and improving the visual quality and motion artifact reduction of reconstructed images.

CN121392141APending Publication Date: 2026-01-23AFFILIATED HOSPITAL OF JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511537269.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing 4D-CBCT dynamic reconstruction methods have problems such as insufficient real-time performance, inaccurate motion modeling, sensitivity to noise, and difficulty in taking into account both local details and global motion patterns.

Method used

A dynamic reconstruction method for 4D-CBCT is employed, integrating ConvLSTM and a dual-attention mechanism. Combining principal component analysis (PCA) with prior 4D-CBCT data, this method accurately and rapidly reconstructs high-quality 4D-CBCT image sequences using a finite number of 2D DRR projection images. The method includes steps such as constructing a prior motion model, training a deep learning network, and reconstructing the 4D-CBCT image sequence. ConvLSTM is used to capture spatiotemporal dynamic evolution, and CBAM and SPA modules are used for feature selection and compensation.

Benefits of technology

It significantly improves the accuracy of predicting motion parameters from limited projection data, resulting in higher visual quality and fewer motion artifacts in the reconstructed images. It demonstrates high accuracy and robustness, with significant advantages in key metrics such as PSNR, SSIM, and MAPE.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121392141A_ABST
    Figure CN121392141A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and discloses a 4D-CBCT dynamic reconstruction method which comprises the following steps: constructing a priori motion model based on principal component analysis in an off-line manner for describing a main mode of organ respiratory motion; designing and training a deep learning network fusing a convolutional long short-term memory network and a double attention mechanism, wherein the network can learn nonlinear mapping from a two-dimensional projection image sequence to a PCA coefficient representing three-dimensional motion; in the online reconstruction stage, real-time projection is input into a network to predict a PCA coefficient, and a high-quality 4D-CBCT image sequence is quickly reconstructed through PCA inverse transformation and image deformation. According to the invention, the CBAM module enhances local features and inhibits noise in an encoder, and the SPA module carries out global context compensation after ConvLSTM. According to the method, the precision and robustness of 4D-CBCT dynamic reconstruction are improved, and complex respiratory movement can be accurately captured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image processing, and particularly relates to a 4D-CBCT dynamic reconstruction method. BACKGROUND

[0002] Stereotactic Body Radiation Therapy (SBRT) has become an effective treatment for early lung cancer and other malignant tumors due to its high dose rate characteristics. However, respiratory motion causes changes in target position, posing a significant challenge to high-precision radiotherapy. Therefore, accurate image guidance of moving targets during treatment is crucial.

[0003] Currently, integrated three-dimensional cone beam CT (3D-CBCT) is widely used in clinical image guidance, but traditional static 3D-CBCT cannot capture respiratory motion information of organs, limiting its application in moving organ treatment. To solve this problem, four-dimensional cone beam CT (4D-CBCT) technology has emerged, which can provide image sequences containing time dimension, thereby achieving dynamic tracking of moving organs such as the lungs.

[0004] Traditional analytical 4D-CBCT reconstruction methods, such as improved versions of the FDK (Feldkamp-Davis-Kress) algorithm, have been applied on commercial linear accelerators, but the quality of the reconstructed images is often limited by insufficient projection data, leading to problems such as reduced contrast and motion blur. Another type of reconstruction method based on image deformation reconstructs images by calculating deformation vector fields (DVFs) between the reference phase and other phases, but the optimization process is usually very time-consuming, making it difficult to meet the real-time requirements of clinical practice.

[0005] In recent years, methods based on deep learning, particularly those using limited single-direction X-ray projections for online real-time reconstruction, have received considerable attention. This approach not only enables rapid generation of 4D-CBCT images but also significantly reduces the radiation dose received by patients. Existing related research includes: 1. Traditional iterative methods: attempt to minimize the intensity difference between digital reconstructed radiographs (DRRs) and actual X-ray projections through iterative optimization. Such methods are computationally expensive and time-consuming.

[0006] 2. Regression-based methods: establish a linear model to construct the mapping relationship between DRRs and 4D-CBCT images. However, the complexity of respiratory motion far exceeds the scope that can be described by a linear model.

[0007] 3. Traditional convolutional neural network (CNN) based methods: These methods are mainly used for static reconstruction and cannot effectively model the dependence in the time dimension, and are not sensitive enough to motion information.

[0008] 4. Methods based on temporal networks: Some methods have begun to use temporal networks such as ConvLSTM, but they usually only capture local spatiotemporal features through convolutional gating, and their performance decreases in the presence of noise, and they lack the ability to focus on key motion areas.

[0009] In summary, the prior art has the problems of insufficient real-time performance, inaccurate motion modeling, sensitivity to noise, and difficulty in balancing local details and global motion patterns when using limited projections for 4D-CBCT dynamic reconstruction. SUMMARY

[0010] The present application aims to at least partially solve the above technical problems. To this end, the present application aims to provide a 4D-CBCT dynamic reconstruction method that combines ConvLSTM and a dual attention mechanism, which combines a principal component analysis (PCA) model with prior 4D-CBCT data, and can accurately and quickly reconstruct a high-quality 4D-CBCT image sequence from a limited number of two-dimensional DRR projection images.

[0011] The technical solution adopted by the present application is as follows: A 4D-CBCT dynamic reconstruction method that can achieve dynamic reconstruction of 4D-CBCT images through a limited number of two-dimensional projections; the method comprises the following steps: Step a): Constructing a prior motion model. This step is an offline stage, and a low-dimensional model that can represent organ respiratory motion is established using a set of high-quality prior 4D-CBCT image sequences. Specifically, first, the three-dimensional deformation vector field (3D-DVF) between adjacent phases in the prior 4D-CBCT data is calculated one by one, and finally an accurate 4D-DVF containing organ motion information throughout the respiratory cycle is obtained. Subsequently, principal component analysis (PCA) is applied to the high-dimensional 4D-DVF data to extract its main motion patterns, thereby constructing a mapping relationship from high-dimensional 3D-DVF to low-dimensional PCA coefficient space, i.e. a PCA model.

[0012] Step b) : Training the deep learning network. This step aims to establish an end-to-end deep learning network whose core task is to learn the complex nonlinear mapping relationship from the two-dimensional DRR projection image sequence to the latent variable (i.e. PCA coefficient) representing the three-dimensional motion. To achieve this goal, it is necessary to first generate paired training data. By interpolating in the PCA coefficient space, a set of PCA coefficients capable of representing continuous motion is generated. Then, the interpolated PCA coefficients are subjected to inverse operation through the PCA model established in step 1 to obtain a set of 3D-DVF groups containing continuous motion information. Next, use this set of 3D-DVF to deform the 3D-CBCT image of the reference phase (such as 0% phase) to generate a set of continuously changing 3D-CBCT image groups. Finally, the 3D-CBCT image group is subjected to single-direction digital simulation projection to obtain the corresponding two-dimensional DRR projection image group. These DRR image groups and their corresponding PCA coefficient groups together constitute the paired data set for training the deep learning network.

[0013] The innovation of the present application lies in the architecture design of the deep learning network. The network combines ConvLSTM and dual attention mechanism. The dual attention mechanism adopts a phased collaborative design: The first attention mechanism, Convolutional Block Attention Module (CBAM), is embedded in the image encoder of the network. CBAM processes the feature maps extracted by the encoder in turn through channel attention and spatial attention. Its function is to optimize the features within a single frame in the early stage of feature extraction, amplify the feature channels carrying important motion information, and focus on the spatial regions with significant motion, thereby achieving local noise suppression and detail enhancement.

[0014] The ConvLSTM network receives the preprocessed feature sequence through CBAM. It captures the spatiotemporal dynamic evolution law in the DRR image sequence through internal input gate, forget gate, output gate and cell state in the time dimension, and establishes a mapping from two-dimensional image sequence to hidden motion state.

[0015] The second attention mechanism, Spatial Attention (SPA) module, acts on the high-level hidden state output by the ConvLSTM network. The SPA module generates a spatial weight map that emphasizes key information by analyzing the high-level feature map that has fused spatiotemporal information, performs final global feature selection and compensation at the high-level semantic level, and further strengthens the regions important for motion modeling.

[0016] Step c): reconstruction of 4D-CBCT image sequence. This is the online application stage. The limited number of two-dimensional projection images acquired in real time in the clinic are input into the deep learning network trained in step 2, and the network will predict the corresponding PCA coefficient group. Then, the predicted PCA coefficient group is input into the prior PCA model established in step 1, and the corresponding 3D-DVF image group is obtained through inverse operation. Finally, using this group of 3D-DVF, the 3D-CBCT image of the reference phase (such as 0% phase) is sequentially deformed, and the final high-quality 4D-CBCT dynamic image sequence is obtained.

[0017] The training process is also optimized in the application, and the weighted mean square error (MSE) is used as the loss function to balance the importance of different PCA principal components. At the same time, gradient clipping and mixed precision training techniques are introduced to ensure the stability of the training process and improve the computational efficiency. In addition, in the training data generation stage, by adding simulated Poisson-Gaussian mixed noise to the DRR image, the robustness of the model to clinical real noise data is enhanced.

[0018] The beneficial effects of the application are: The application proposes a fusion architecture of double attention mechanism and ConvLSTM. CBAM enhances local features and suppresses noise at the input end, providing high-quality input for time series modeling; ConvLSTM effectively captures the dynamic evolution law of the sequence; SPA performs advanced feature screening and compensation based on global context at the output end. The three work together to significantly improve the accuracy of predicting motion parameters from limited projection data. Compared with the baseline model, the method performs significantly better in key indicators such as PSNR, SSIM and MAPE, and the visual quality of the reconstructed image is higher and the motion artifacts are fewer, proving the high precision and robustness of the method in the 4D-CBCT dynamic reconstruction task. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 The deep learning model structure diagram constructed by the application.

[0020] Figure 2 The 4D-CBCT dynamic reconstruction technology roadmap of the application.

[0021] Figure 3 The performance comparison table of different model configurations on the test set.

[0022] Figure 4 The reconstructed image result visualization comparison chart of different models. DETAILED DESCRIPTION

[0023] The technical solutions of the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0024] It should be understood that it should also be noted that the functions / actions appearing in the embodiments may not appear in the order shown in the drawings. For example, depending on the functions / actions involved, the two figures shown successively can actually be performed substantially concurrently, or sometimes in the opposite order.

[0025] The present application proposes a 4D-CBCT dynamic reconstruction method fusing ConvLSTM and double attention mechanism, and the complete technical process is as shown in Figure 2 The main parts include data preprocessing, prior motion model construction, deep learning network training and online reconstruction.

[0026] 1. 4D-CBCT data acquisition and preprocessing Firstly, in the 4D-CBCT imaging process, the respiratory motion of the patient is synchronously monitored through a camera-based respiratory monitoring system (such as Varian RPM system), and the amplitude and phase waveform of the respiration are recorded in real time. At the same time of CBCT projection acquisition, each frame of projection data is marked with the corresponding time stamp to associate the current respiratory state. After the acquisition is completed, all projection data are phase-grouped according to the respiratory signal, and a 4D-CBCT projection data set containing 10 phases (0% to 90%) is formed.

[0027] To simulate the quantum noise existing in the CBCT image in the clinical environment, a linear noise combination model is constructed in the present application and applied to the generated DRR image. The model is as follows:

[0028] Wherein is the noise-free signal line integral, N is the X-ray projection intensity with noise, is set to 10 5 , the background electronic noise variance is set to 10. By adding such noise to the DRR image of the training data, the diversity and authenticity of the data can be enhanced, and the model has better generalization ability to real clinical noise data.

[0029] In constructing the training sample, a continuous 4-frame DRR image is used to form an input sequence, with a dimension of 4x1xHxW (time sequence x channel x height x width), and the true PCA coefficient corresponding to the 4th frame DRR is taken as the regression label of the sequence. The data set is strictly divided into a training set (80%), a validation set (10%) and a test set (10%) in chronological order to avoid data leakage.

[0030] 2. Construction of prior motion model The goal of this step is to establish a low-dimensional model that can describe the main mode of respiratory motion. First, calculate the 3D-DVF between adjacent phases of 4D-CBCT. Read the 4D-CBCT three-dimensional image sequence (10 phases) of the entire respiratory cycle, and designate the 0% phase image at the end of inspiration as the fixed reference image. To improve computational efficiency, all phase images are downsampled to 1 / 2 of the original size. Then, using the Demons non-rigid registration algorithm, the deformation of the moving image of the other 9 phases (10% to 90%) to the 0% phase reference image is calculated in turn. Each registration generates a three-dimensional deformation vector field (3D-DVF), which consists of three matrices (Fx, Fy, Fz), representing the displacement of each voxel in the X, Y, Z three directions. The three displacement fields are spliced to form a complete 3D-DVF file. Finally, 9 accurate 3D-DVFs are obtained, which completely describe the motion trajectory of the organ relative to the reference phase during the respiratory cycle.

[0031] Then, a PCA dimension reduction model is constructed. Load all 3D-DVF files obtained in the above step, flatten each DVF into a long vector, and form a data matrix X with a dimension of NxD, where N is the number of phases (9) and D is the dimension of each DVF (number of voxels x 3). Use principal component analysis (PCA) to reduce the dimension of the data matrix, retaining the first 3 principal components, as they are usually sufficient to explain more than 95% of the variation in respiratory motion. After the PCA model is learned, a mean vector and 3 principal component vectors (component matrix) are obtained. At this point, each original 3D-DVF can be represented as a 3-dimensional PCA coefficient vector. In this way, a mapping relationship from high-dimensional 3D-DVF to low-dimensional 3-D PCA coefficient is established.

[0032] 3. Generation of training data To train the deep learning network, a large amount of paired data is needed. The present application generates large-scale continuous motion data by interpolating in the PCA coefficient space.

[0033] First, linear interpolation is performed between the PCA coefficient vectors p i and p i+1 corresponding to the adjacent two original phases, generating k interpolation points between each pair of phases. The interpolation formula is: p interp (j)=p i +j / k*(p i+1 -p i ), where j ranges from 1 to k.

[0034] This method can generate a large number of smoothly transitioning PCA coefficient sets to characterize continuous organ motion.

[0035] Next, each PCA coefficient vector generated by interpolation is inversely transformed using the established PCA model to reconstruct the corresponding 3D-DVF vector. The inverse transformation formula is: DVF reconstructed =mean vector +PCA coeffs *component matrix This yields a set of 3D-DVFs containing information about the continuous motion of the image.

[0036] Then, the 3D-CBCT reference image at 0% phase is deformed using this reconstructed 3D-DVF. For each voxel position x' in the deformed image, its intensity value is obtained by bilinear interpolation of the reference image at the x'-DVF(x') position. This generates a continuously varying set of 3D-CBCT images.

[0037] Finally, this set of 3D-CBCT images is digitally simulated and projected in one direction to generate a two-dimensional DRR image set. The projection process simulates the geometry of cone-beam CT, and the attenuation integral of X-rays passing through the 3D-CBCT volume data is calculated using a ray tracing algorithm (such as the DDA algorithm) to generate DRR images. As mentioned earlier, a noise model can be selectively added at this stage. Ultimately, the resulting two-dimensional DRR projected image set and its corresponding PCA coefficient set constitute the paired dataset used for training.

[0038] 4. Design and Training of Deep Learning Networks The core of this invention is an end-to-end deep learning network, the overall architecture of which is as follows: Figure 1 As shown, it includes an image encoder, a CBAM module, a ConvLSTM network, a SPA module, and a regressor.

[0039] Image Encoder and CBAM Module (First Attention Mechanism): The input is a DRR image sequence (shape B×T×C×H×W). Each frame in the sequence first undergoes spatial feature extraction via a shared-weight 2D convolutional encoder. This encoder contains four cascaded convolutional layers, each with a stride of 2, achieving spatial downsampling. After each convolutional layer, a CBAM module is embedded. The CBAM performs channel attention and spatial attention sequentially. Channel Attention: The spatial information is aggregated by global average pooling and global max pooling, and then a shared multi-layer perceptron (MLP) and Sigmoid function are used to calculate the weight of each channel, which is used to calibrate the feature channels and enhance the motion-related information.

[0040] Spatial Attention: The feature map is averaged and maximized in the channel dimension to generate a two-dimensional spatial attention map, which is then processed by convolution and Sigmoid function to highlight the key spatial regions of motion significance. CBAM works at the bottom of the encoder, optimizing the intra-frame features of single-frame images, and providing high-quality and high signal-to-noise ratio input features for subsequent time series modeling.

[0041] ConvLSTM Network: The feature sequence output by the encoder (shape ) is sent to the ConvLSTM layer. ConvLSTM replaces the fully connected operation in traditional LSTM with convolution, allowing it to process spatial and temporal information simultaneously. It controls the flow of information through input gates, forget gates, output gates, and cell states, capturing the spatiotemporal dynamic evolution rules in the DRR sequence, and finally outputting a high-level spatiotemporal feature representation that integrates the context of the entire historical sequence.

[0042] SPA Module (Second Attention Mechanism): The hidden state feature map of the last time step of ConvLSTM (shape B hidden ×(H / 16)×(W / 16)) is sent to the SPA module. This module aims to compensate and filter the global context information. It first performs multi-scale pooling operations on the input feature map, including 1x1, 3x3, and 5x5, to capture spatial context at different granularities. Then these multi-scale features are upsampled and fused, and a 3x3 convolution and Sigmoid function are used to generate the final spatial attention map. This attention map is used to reweight the output features of ConvLSTM, further strengthening the regions that are crucial for the final PCA coefficient regression.

[0043] Regressor: The feature map optimized by the SPA module is compressed in the spatial dimension by a global average pooling layer, and then flattened into a feature vector. This vector is then passed through two fully connected layers: the first fully connected layer reduces the dimension and uses a LeakyReLU activation function; the second fully connected layer maps the features to 3 dimensions and directly outputs the predicted PCA coefficients.

[0044] Staged Collaborative Design: The dual-attention mechanism of the present application works collaboratively in stages. CBAM is located in the encoder and performs local noise suppression and detail enhancement in the early stage of feature extraction; SPA is located after ConvLSTM and performs global context compensation and screening at the high-level semantic feature level. The two processes different stages and different levels of abstract features, forming an effective synergistic effect.

[0045] Loss Function: The Weighted MSE (Mean Squared Error) loss function is used, wherein is a preset weight used to balance the importance of different PCA principal components.

[0046] Optimizer: The Adam optimizer is used.

[0047] Training Techniques: To stabilize training and accelerate convergence, gradient clipping (limiting the gradient norm to prevent gradient explosion) and mixed precision training (AMP, using FP16 half-precision calculation to save memory and improve speed, while retaining FP32 main parameters to ensure accuracy) are used.

[0048] Early Stopping Mechanism: The loss on the validation set is monitored, and when the validation loss does not decrease for multiple epochs in a row, training is terminated early to prevent overfitting.

[0049] Final Reconstruction of 4D-CBCT Image Sequence: In the clinical application stage, a sequence of real-time acquired limited times (e.g. 4 frames) of two-dimensional projection images is input into the trained deep learning network. The network quickly predicts the corresponding 3D PCA coefficients. Then, the predicted PCA coefficients are subjected to inverse operation through the prior PCA model established in step 2 to reconstruct the corresponding 3D-DVF image. Finally, using this 3D-DVF to spatially deform the 3D-CBCT image at the reference phase (0% phase), the high-quality 3D-CBCT image at the current respiratory phase can be obtained in real time and accurately. Repeating this process, a complete 4D-CBCT dynamic image sequence can be obtained.

[0050] To verify the effectiveness of the method of the present application, real patient data from the TrueBeam STx platform was used. The 4D-CBCT data of this patient contains 10 phases, and the original CBCT image size is 512x512x102, with a voxel size of 1.172x1.172x3mm 3 .

[0051] First, the patient data set was uniformly resampled to have a voxel size of 2x2x2mm 3, the adjusted CBCT size is 300x300x153. The generated DVF size is 300x300x459 (as it contains three components of x, y, z). The DRR image size generated by single-direction projection is set to 512x384.

[0052] According to the description of step c), the PCA coefficients corresponding to 9 original phases are interpolated, 120 points are interpolated between each phase, a total of 8*120=960 sets of interpolated PCA coefficients are obtained. Then through inverse PCA transformation, image deformation and DRR projection, 960 DRR images and their corresponding PCA coefficient labels are finally obtained, which constitute the paired training data. The input sequence is composed of 4 consecutive DRR images, and the label is the PCA coefficient of the 4th frame. The data set is divided into training set, validation set and test set in the ratio of 8:1:1.

[0053] Model training is performed on a server equipped with NVIDIA GeForce RTX 4090 GPU, using PyTorch2.5.1 framework and CUDA 12.4. The key hyperparameter settings are as follows: batch size is 8, learning rate is 1e-5, and the optimizer is Adam. The model is trained for a total of 200 epochs, and the early stopping strategy is adopted, and the model weight with the lowest loss on the validation set is saved.

[0054] Three evaluation indicators are used to quantify the reconstruction effect: structural similarity index (SSIM), peak signal-to-noise ratio (PSNR) and mean absolute percentage error (MAPE). In order to verify the effectiveness of the dual attention mechanism in the present application, a series of ablation experiments are conducted, and four model configurations are compared: 1. ConvLSTM (baseline model) 2. ConvLSTM+CBAM 3. ConvLSTM+SPA 4. ConvLSTM+CBAM+SPA (complete model proposed in the present application) The quantitative results are shown in Figure 3 From the figure, it can be seen that the fusion model (ConvLSTM+CBAM+SPA) proposed in the present application achieves the best performance in all three indicators. Its SSIM reaches 0.9998, PSNR reaches 61.40 dB, and MAPE is as low as 0.0157. Compared with the baseline model ConvLSTM, the PSNR is improved by 7.47 dB, and the MAPE is reduced by 0.0234, with significant performance improvement. This proves the effectiveness of the cooperative work of CBAM and SPA dual attention mechanisms.

[0055] At the same time, the ConvLSTM+CBAM model also has obvious improvement compared with the baseline model (PSNR is improved by 5.04 dB), indicating that the CBAM module can independently bring performance gain. The performance of the ConvLSTM+SPA model does not see significant improvement, which may be because without the pre-feature enhancement of CBAM, SPA is difficult to play the maximum role on the high-level features full of noise. However, when CBAM and SPA are used in combination, a 1+1>2 synergistic effect is produced, which is much better than the simple addition of a single module, fully proving the correctness and superiority of the phased collaborative design of the application.

[0056] The visualization results are shown in FIG. 4. Figure 4 As shown in the figure. The figure shows the images reconstructed by different models and the difference maps between them and the real target image. All models can restore image details to a certain extent. However, from the comparison of the error maps, it can be clearly seen that the error map of the method of the application (ConvLSTM+CBAM+SPA) has the lowest overall brightness and the closest color to black, indicating that its reconstruction result has the smallest error with the real image. Compared with the original ConvLSTM, the reconstruction result of the method is more accurate in vision and has fewer artifacts, further verifying the effectiveness of the ConvLSTM model fused with double attention mechanisms.

[0057] The application is not limited to the above optional embodiments, and anyone can derive other various forms of products under the inspiration of the application, but regardless of any changes in shape or structure, any technical solutions falling within the scope defined by the claims of the application fall within the protection scope of the application.

Claims

1. A 4D-CBCT dynamic reconstruction method, characterized in that, Includes the following steps: a) Constructing a priori motion model: Based on a set of prior 4D-CBCT image sequences, calculate the three-dimensional deformation vector field between adjacent phases, and use principal component analysis to construct a PCA model that maps from 3D-DVF to low-dimensional PCA coefficients. b) Training a deep learning network: Build and train an end-to-end deep learning network to establish a mapping relationship from two-dimensional digitally reconstructed radiographic projection image sequences to PCA coefficients representing three-dimensional motion; the deep learning network integrates a convolutional long short-term memory network and a dual attention mechanism; c) Reconstructing the 4D-CBCT image sequence: Input the real-time two-dimensional projection image sequence to be reconstructed into the trained deep learning network to predict the corresponding PCA coefficient set; perform inverse operation on the PCA model established in step a) to obtain the three-dimensional deformation vector field set; use the 3D-DVF set to deform a reference phase 3D-CBCT image to obtain the final 4D-CBCT image sequence. The dual attention mechanism includes a first attention mechanism located in the image encoder and a second attention mechanism located after the output of the ConvLSTM network, which work together in stages.

2. The method according to claim 1, characterized in that, The specific steps for constructing the prior motion model in step a) include: For a priori 4D-CBCT image sequence, the 3D-DVF between adjacent phases is calculated one by one to obtain the 4D-DVF characterizing the motion information of the entire respiratory cycle. Principal component analysis was performed on the 4D-DVF data to extract the main motion patterns and construct a PCA model from high-dimensional 3D-DVF to low-dimensional PCA coefficients.

3. The method according to claim 1, characterized in that, Before training the deep learning network in step b), there is also a step of generating training data, which specifically includes: Interpolation is performed in the PCA coefficient space to generate a set of PCA coefficients that can characterize continuous motion; The interpolated PCA coefficient set is inversely operated through the PCA model to obtain a 3D-DVF set containing continuous motion information. The 3D-DVF group is applied to the 3D-CBCT image of the reference time phase for deformation, generating a set of continuously changing 3D-CBCT images; The 3D-CBCT image set is projected in one direction to generate a corresponding two-dimensional DRR projection image set. This image set, together with the interpolated PCA coefficient set, constitutes the training dataset.

4. The method according to claim 1, characterized in that, The first attention mechanism is a convolutional block attention module, which is embedded in the image encoder of the deep learning network. It sequentially performs channel attention and spatial attention processing on the feature maps extracted by the encoder to enhance motion-related features and suppress noise.

5. The method according to claim 1 or 4, characterized in that, The ConvLSTM network receives the feature sequence processed by the first attention mechanism and captures the spatiotemporal dynamic evolution of the DRR image sequence in the time dimension through its internal input gate, forget gate and output gate structure, and outputs a high-level spatiotemporal feature representation that integrates historical sequence context information.

6. The method according to claim 1, characterized in that, The second attention mechanism is a spatial attention module, which operates on the high-level hidden state feature map output by the ConvLSTM network. It generates a spatial weight map through multi-scale pooling and feature fusion, and performs global context compensation and filtering on high-level semantic features to strengthen the spatial regions that are crucial to the final regression task.

7. The method according to claims 1, 4, and 6, characterized in that, The phased collaborative design of the dual attention mechanism is as follows: the CBAM module performs local optimization of single-frame image features in the early stage of feature extraction, and the SPA module performs global optimization of high-level semantic features after temporal modeling. The two modules process features at different levels of abstraction respectively.

8. The method according to claim 1, characterized in that, The specific steps for reconstructing the 4D-CBCT image sequence in step c) are as follows: the predicted PCA coefficient set is inversely transformed through the prior PCA model to reconstruct the corresponding 3D-DVF image set, and then the 3D-DVF image set is used to spatially deform the 0% phase three-dimensional CBCT image in sequence to obtain the final high-quality 4D-CBCT image sequence.

9. The method according to claim 1, characterized in that, In step b), the process of training the deep learning network uses weighted mean square error as the loss function, and employs gradient pruning and mixed precision training techniques to stabilize the training process and improve training efficiency.

10. The method according to claim 3, characterized in that, In the process of generating the two-dimensional DRR projection image set, in order to simulate the real clinical situation, a linear combined noise model containing Poisson noise and Gaussian noise was applied to the generated DRR images to enhance the diversity of training data and the robustness of the model.