A four-dimensional CT reconstruction method driven by real-time kV images
By employing a real-time kV image-driven four-dimensional CT reconstruction method, combined with data preprocessing and a deep learning model, the problem of high temporal resolution and low radiation dose in radiotherapy has been solved, achieving efficient and accurate four-dimensional CT reconstruction and dose calculation, and supporting real-time target area monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING UNIV CANCER HOSPITAL
- Filing Date
- 2026-03-12
- Publication Date
- 2026-06-02
AI Technical Summary
Current radiotherapy imaging technology cannot simultaneously achieve high temporal resolution, low radiation dose, and high-precision three-dimensional/four-dimensional dose calculation, resulting in target localization errors and dose distribution uncertainties.
A four-dimensional CT reconstruction method based on real-time kV image-driven reconstruction is adopted. High-resolution three-dimensional CT volume data is reconstructed through data preprocessing (respiratory phase synchronization, spatial resampling, geometric consistency registration and clipping) and deep learning model (three-dimensional U-Net architecture). Combined with the dose calculation module, real-time target area monitoring and dynamic dose optimization are realized.
It achieves efficient four-dimensional CT reconstruction under low radiation dose, improves reconstruction speed and spatial accuracy, and has the feasibility and practical value for clinical application.
Smart Images

Figure CN122134872A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image reconstruction and radiotherapy technology, specifically to a four-dimensional CT reconstruction method based on real-time kV image-driven reconstruction. Background Technology
[0002] In high-precision radiotherapy for tumors (such as lung cancer, where respiratory motion is significantly a factor), the patient's respiratory motion and dynamic organ deformation are the main factors leading to target localization errors and dose distribution uncertainties. To achieve accurate image guidance, current clinical practice mainly employs the following imaging techniques, but all have significant technical limitations:
[0003] Conventional cone-beam computed tomography (CBCT) can provide three-dimensional structural information, but its temporal resolution is severely insufficient, and a single scan will bring patients a high additional radiation dose; in addition, it is very easy to produce severe motion artifacts when facing periodic movements such as breathing, which affects the positioning accuracy.
[0004] Magnetic resonance imaging (MRI): Although it has excellent soft tissue resolution, it takes a long time to image and lacks CT values (HU values) for tissue electron density mapping, making it difficult to use directly and efficiently for accurate dose calculation in radiotherapy.
[0005] Airborne kV Imaging System for Linear Accelerators: Modern medical linear accelerators are typically equipped with orthogonal kV imaging systems, which offer extremely high temporal resolution and can capture patient dynamics in real time. However, these systems only acquire two-dimensional projection images. Lacking depth information, they cannot be directly used to reconstruct the patient's three-dimensional / four-dimensional anatomy, let alone for complex three-dimensional and four-dimensional dose calculations and assessments.
[0006] In summary, existing conventional radiotherapy imaging techniques cannot simultaneously achieve high temporal resolution, low radiation dose, and high-precision 3D / 4D dose calculation. Therefore, there is an urgent need in this field for a novel technology that can fully utilize the high temporal resolution advantage of real-time 2D kV images from accelerators as a driving force to quickly and accurately reconstruct high-fidelity 4D CT images. This, combined with a dose calculation module, can then enable real-time target monitoring and dynamic dose optimization during radiotherapy. Summary of the Invention
[0007] To address the technical problem that existing conventional radiotherapy imaging techniques cannot simultaneously achieve optimal temporal resolution, imaging radiation dose, and applicability for dose assessment, this invention provides a four-dimensional CT reconstruction method based on real-time kV image-driven imaging.
[0008] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0009] A four-dimensional CT reconstruction method based on real-time kV image-driven reconstruction includes the following steps:
[0010] S1. Data Preprocessing:
[0011] S11. Synchronized Respiratory Phase: For each patient Real-time acquisition of timestamp tags associated with respiratory movements using an existing linear accelerator-borne kV imaging system. The original kV projection image is recorded as Raw 4D CT scan data taken during the radiotherapy planning phase This represents the CT imaging of the patient's thoracic and pulmonary anatomy as a function of respiration. The changes; this invention combines real-time respiratory monitoring signals acquired during radiotherapy to record the patient's respiratory status at each time frame. Collected Image sequence allocation and respiratory phases corresponding to 4D CT After this processing, the kV projection image sequence, originally arranged according to clinical acquisition time, was... It was converted into a timeline synchronized with the CT scan and based on the respiratory phase. Indexed four-dimensional sequence data This constitutes the original dataset used in this invention. :
[0012]
[0013] S12, Spatial Resampling: For each patient raw CT body data Perform three-dimensional resampling to unify its voxel space and obtain the resampled CT volume data. Simultaneously, the serialized kV projection images Two-dimensional resampling is performed to unify the pixel space to the same scale, resulting in a sampled projected image. ;
[0014] S13. Geometric Consistency Registration Clipping: This includes between-group difference alignment and within-group difference registration. For between-group differences, this invention introduces an between-group geometric consistency clipping strategy for the resampled data. For each patient... The phase with the largest lung volume in the respiratory cycle, i.e., the end-expiratory phase, is selected. The CT images are used as a reference standard, denoted as This ensures that the cropping intervals for each time phase cover the entire lung region; subsequently, the lung region is automatically segmented to generate a binary mask for the lung region. Based on this segmentation mask, the geometric center point of the lung region is calculated. This center point As the spatial anchor point for the cutting operation, a fixed-size three-dimensional space is cut; in the coronal plane, i.e. In the axial direction, with Extending upwards and downwards from the center. Layer slices, in cross-section On a plane, with Extending forward, backward, left, and right from the center. Individual unit, the formal definition of this clipping operation is:
[0015]
[0016] Then perform this cropping operation Consistently applied to all respiratory phases of this patient. Data of a uniform spatial scale can be obtained from image data:
[0017]
[0018] in, Model training for 4D CT reconstruction This indicates the height of the CT scan data after uniform cropping. This indicates the width of the CT scan data after uniform cropping. This indicates the depth of the CT volume data after uniform cropping.
[0019] To address intra-group differences, this invention introduces a spatial alignment registration strategy for the resampled data, the specific process of which is as follows:
[0020] Digital reconstruction of radiographic images is generated. This step defines a simulated fluoroscopic projection function based on the physical geometry parameters of the treatment accelerator kV imaging system. For patients During a specific respiratory phase Below, CT body data after the aforementioned processing Digital reconstructed images are generated through this projection function. :
[0021]
[0022] in, The 3D CT volumetric data is virtually projected onto a spatial coordinate system consistent with the kV image to simulate the data input used for model inference in a real radiotherapy scenario. Simultaneously, it provides sufficient data support for the training of the 4D CT reconstruction model, forming a system where each patient has spatiotemporally correlated 3D CT volumetric data and 2D... Data input pairs:
[0023]
[0024] kV image preprocessing is a key step in transforming clinically acquired, synchronized, and sampled real two-dimensional kV images into digital images. Through cropping and geometric transformation, the model is adjusted to align with the data obtained during the planning and training phases in terms of lung anatomy and spatial features, thereby providing consistent and comparable input for the four-dimensional CT reconstruction model. Specifically:
[0025] Perform two-dimensional registration on the kV projection image, to For fixed images, generated by in-phase CT For floating images, optimize the spatial deformation field Maximize the image similarity between the two :
[0026]
[0027] in, Indicates the deformation field Apply to floating images ;
[0028] The optimized deformation field obtained by optimization Applied to The image is initially aligned with the kV projection image:
[0029]
[0030] Subsequently, the region of interest determined by this alignment process is directly applied to the original kV image to perform a cropping operation:
[0031]
[0032] After the above registration and cropping preprocessing, a dataset with spatiotemporal correlation and domain adaptation preprocessing was formed for each patient in the model testing and clinical inference stages:
[0033]
[0034] S2, 4D CT reconstruction:
[0035] Four-dimensional CT reconstruction is achieved using a deep learning model based on a 3D U-Net architecture. The model training and validation phases utilize 3D CT volumetric data. kV projection image of the target time phase Using the input image as input, a deep learning model predicts and outputs a high-resolution 3D CT volumetric image corresponding to the same respiratory phase as the input kV image. This allows for real-time, precise dynamic modeling of changes in the patient's lung anatomy during respiratory movements, simulating the process of radiotherapy.
[0036] The model training employs a supervised learning paradigm, using multiple sets of 3D CT volumetric data collected from the same patient at different time phases and data from the same time phase. Images are used as training sample pairs. By optimizing the composite loss function, the network learns to reconstruct the accurate three-dimensional structure of the target phase from the mapping between the two-dimensional projection and the prior three-dimensional reference image. Finally, an optimal model that can be stably reconstructed using prior data in real radiotherapy scenarios is obtained.
[0037] The model uses normalized prior 3D CT volume data and corresponding 2D kV projection images as joint inputs. It learns the mapping relationship between 2D projection images and 3D volume data through multi-scale feature encoding and decoding structures, thereby predicting and outputting 3D CT volume data at the current respiratory phase.
[0038] Furthermore, in step S2 of the four-dimensional CT reconstruction, the normalization process includes: for CT volume data expressed in HU values... Utilize window width and window position The range of values in the region of interest is linearly mapped to the same value range as that of a single-channel kV image. The data normalization process involves truncating voxel values that are too high or too low within a certain range to reduce their impact on the prediction results. This normalization is denoted as [missing information]. , means as follows:
[0039]
[0040] Predicted image generated by the network The value range remains the single-channel grayscale value. Intervals, through inverse normalization operations This is then mapped back to a preset window width HU value range to fit clinical radiation dose estimation. The specific inverse normalization operation is as follows:
[0041] .
[0042] Furthermore, in step S2 of the four-dimensional CT reconstruction, the three-dimensional U-Net consists of symmetrical encoder and decoder paths connected by skip connections. The encoder extracts and compresses multi-scale features progressively through 3D convolution, activation, and pooling operations. The decoder upsamples and fuses contextual information based on features from the encoder through transposed convolution and skip connections, ultimately outputting the predicted single-channel three-dimensional CT volume data through a 1×1×1 convolutional layer. .
[0043] Furthermore, in step S2, the four-dimensional CT reconstruction, the composite loss function... It is composed of multiple loss functions, specifically represented as follows:
[0044]
[0045] in, The mean squared error loss function is used to measure the numerical difference between the predicted CT and the real CT at the voxel level, forcing the reconstructed image to approximate the real value at the voxel level and ensuring the accuracy of the HU value. The structural similarity loss function measures the similarity between the reconstructed image and the real image in terms of brightness, contrast, and local structure. Minimizing the structural similarity loss helps the network generate visually more natural and structurally clearer images, thereby improving the credibility of the reconstructed CT in terms of anatomical structures. It is a first-order spatial gradient loss function used to measure the difference between the reconstructed image and the real image in the spatial gradient domain. It constrains the model to learn high-frequency information of organ contours and vascular nodules at the spatial gradient level, ensuring that the reconstruction results have clear anatomical boundaries and retain texture details, thereby improving the accuracy of radiotherapy dose calculation and evaluation. To define the hyperparameters for the weights of the structural similarity loss function; This defines the hyperparameters for the weights of the first-order spatial gradient loss function.
[0046] Compared with existing technologies, the four-dimensional CT reconstruction method based on real-time kV image driving provided by this invention realizes an end-to-end closed loop from two-dimensional projection to four-dimensional CT to dose assessment, and has the following advantages: 1. It reduces the radiation dose required for image reconstruction; 2. It improves the reconstruction speed; 3. It ensures spatial accuracy and temporal continuity; 4. It has the feasibility and practical value for clinical application. Attached Figure Description
[0047] Figure 1 This is a schematic diagram of the data preprocessing process in the four-dimensional CT reconstruction method provided by the present invention.
[0048] Figure 2 This is a case study illustrating the differences between the present invention and traditional CT reconstruction methods.
[0049] Figure 3 This is a visualization demo script diagram of the present invention for reconstructing CT scans. Detailed Implementation
[0050] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below with reference to specific illustrations.
[0051] This invention provides a four-dimensional CT reconstruction method based on real-time kV image-driven reconstruction, comprising the following steps:
[0052] S1. Data Preprocessing:
[0053] The real-time acquired kV projection images differ significantly from the pre-acquired 3D / 4D CT data in terms of grayscale distribution, imaging geometry, and temporal sampling methods. If these images are directly used for subsequent 4D reconstruction and dose calculation, it will lead to unstable reconstruction, temporal mismatch, and dose assessment errors.
[0054] like Figure 1 As shown, this invention performs data preprocessing on the input historical CT data and real-time kV image data, including respiratory phase synchronization, spatial resampling, and geometric consistency registration and cropping. This constructs a unified input data format suitable for deep learning reconstruction models and radiotherapy dose calculation, providing a reliable data foundation for subsequent four-dimensional CT reconstruction and dose assessment. Specifically, the data preprocessing includes:
[0055] S11. Synchronized respiratory phase:
[0056] Real-time acquired kV projection images exhibit significant temporal continuity, and their frame sequence is closely related to the patient's respiratory movements. The core of four-dimensional CT reconstruction lies in accurately depicting the temporal information of anatomical structures as they change with respiration. Therefore, this invention introduces a respiratory phase synchronization processing mechanism to establish a temporal correspondence between kV projection images and four-dimensional CT scans, providing crucial temporal priors for subsequent high-precision four-dimensional reconstruction, dose calculation, and assessment.
[0057] For each patient Real-time acquisition of timestamp tags associated with respiratory movements using an existing linear accelerator-borne kV imaging system. The original kV projection image is recorded as Raw 4D CT scan data taken during the radiotherapy planning phase This represents the CT imaging of the patient's thoracic and pulmonary anatomy as a function of respiration. The changes; this invention combines real-time respiratory monitoring signals acquired during radiotherapy to record the patient's respiratory status at each time frame. Collected Image sequence allocation and respiratory phases corresponding to 4D CT After this processing, the kV projection image sequence, originally arranged according to clinical acquisition time, was... It was converted into a timeline synchronized with the CT scan and based on the respiratory phase. Indexed four-dimensional sequence data This constitutes the original dataset used in this invention. :
[0058]
[0059] Includes respiratory phases The original kV projection images and original CT images are used for subsequent processing, registration and reconstruction.
[0060] S12, Spatial Resampling:
[0061] CT images and kV projection images differ in spatial dimension, voxel (pixel) size, and resolution. Inconsistent inputs lead to inconsistent model inputs, affecting network structure design, parameter sharing, and the feasibility of batch training. They also hinder the accurate mapping of physical quantities in subsequent dose calculations. Therefore, this invention performs spatial resampling on the original image data, specifically:
[0062] For each patient raw CT body data Perform 3D resampling to unify the voxel space (e.g., to 2mm×2mm×2mm) to obtain the resampled CT volume data. Simultaneously, the serialized kV projection images Two-dimensional resampling is performed to unify the pixel space to the same scale (e.g., 2mm × 2mm) to obtain the sampled projection image. This processing ensures that image data from different sources and patients are consistent in spatial resolution, laying the data foundation for subsequent geometrical registration and spatial cropping.
[0063] S13, Geometric Consistency Registration Clipping:
[0064] The training data involved in this invention faces geometric inconsistencies at both the inter-group and intra-group levels. The noise and bias introduced by geometric inconsistencies force the model to learn irrelevant spatial transformations instead of focusing on the core motion laws, thereby reducing the model's efficiency in focusing on key anatomical structures and severely reducing the accuracy and robustness of four-dimensional reconstruction. Therefore, geometric consistency registration and clipping includes inter-group difference alignment and intra-group difference registration.
[0065] (1) Alignment of differences between groups
[0066] At the intergroup level, due to differences in body position, scanning range and anatomical structure among different patients during CT scans, there are significant inconsistencies in the spatial location and area of interest of their CT data, with large differences in the location distribution and center offset of the lung regions.
[0067] To address inter-group differences, this invention introduces an inter-group geometric consistency pruning strategy for the resampled data, for each patient... The phase with the largest lung volume in the respiratory cycle, i.e., the end-expiratory phase, is selected. The CT images are used as a reference standard, denoted as This ensures that the cropping intervals for each time phase cover the entire lung region; subsequently, the lung region is automatically segmented to generate a binary mask for the lung region. Based on this segmentation mask, the geometric center point of the lung region is calculated. This center point As the spatial anchor point for the cutting operation, a fixed-size three-dimensional space is cut; in the coronal plane, i.e. In the axial direction, with Extending upwards and downwards from the center. Layer slices, in cross-section On a plane, with Extending forward, backward, left, and right from the center. Individual unit, the formal definition of this clipping operation is:
[0068]
[0069] Then perform this cropping operation Consistently applied to all respiratory phases of this patient. Data of a uniform spatial scale can be obtained from image data:
[0070]
[0071] in, Model training for 4D CT reconstruction This indicates the height of the CT scan data after uniform cropping. This indicates the width of the CT scan data after uniform cropping. This indicates the depth of the CT volume data after uniform cropping. During implementation, if the cropping range exceeds the original image boundary, an air value of HU=-1000 is used for filling.
[0072] This standardization process unifies the spatial scale of CT data: regardless of the original CT image size, after this fixed-range cropping, the spatial dimension of all input data is forced to be uniform, resolving the issue of inconsistent original data sizes. Secondly, the anatomical centers of different patients are aligned: cropping is performed based on the geometric center of each patient's lung region, concentrating the main anatomical structures of the lungs in the cropped data blocks to the same spatial location, significantly eliminating inter-group differences caused by overall organ positional shifts. This reduction in inter-group variation helps deep learning models focus more stably and efficiently on learning the inherent anatomical features of the lungs and their movement patterns during respiration, thereby improving the model's generalization ability and analytical accuracy.
[0073] (2) Within-group difference registration
[0074] At the intra-group level, this invention also faces the problem of inherent geometric inconsistencies between images due to the different imaging systems used. Due to deviations in physical coordinate systems, imaging geometry, and patient positioning between fractionated treatments, a direct and accurate correspondence cannot be established between the simulated positioning of the same patient, the CT body data obtained from the radiotherapy plan, and the kV projection data obtained in the radiotherapy room.
[0075] To address intra-group differences, this invention introduces a spatial alignment and registration strategy for the resampled data. The strategy balances clinically acceptable computational efficiency with algorithm robustness and provides geometrically pre-aligned data input for subsequent deep learning models. The specific process is as follows:
[0076] Digital reconstruction of radiographic images is generated. This step defines a simulated fluoroscopic projection function based on the physical geometry parameters of the treatment accelerator kV imaging system. For patients During a specific respiratory phase Below are the CT body data after the aforementioned preprocessing. Digital reconstructed images are generated through this projection function. :
[0077]
[0078] in, The 3D CT volumetric data is virtually projected onto a spatial coordinate system consistent with the kV image to simulate the data input used for model inference in a real radiotherapy scenario. Simultaneously, it provides sufficient data support for the training of the 4D CT reconstruction model, forming a system where each patient has spatiotemporally correlated 3D CT volumetric data and 2D... Data input pairs:
[0079]
[0080] This invention addresses the interdomain discrepancy problem faced by models during training and inference phases—specifically, training uses simulated DRR images and 4DCT data, while clinical inference requires processing real-time acquired kV images. Therefore, a crucial kV image preprocessing step is introduced before performing formal 4D reconstruction.
[0081] kV image preprocessing is a key step in transforming clinically acquired, synchronized, and sampled real two-dimensional kV images into digital images. Through cropping and geometric transformation, the model is adjusted to align with the data obtained during the planning and training phases in terms of lung anatomy and spatial features, thereby providing consistent and comparable input for the four-dimensional CT reconstruction model. Specifically:
[0082] Perform two-dimensional registration on the kV projection image, to For fixed images, generated by in-phase CT For floating images, optimize the spatial deformation field Maximize the image similarity between the two :
[0083]
[0084] in, Indicates the deformation field Apply to floating images ;
[0085] The optimized deformation field obtained by optimization Applied to The image is initially aligned with the kV projection image:
[0086]
[0087] Subsequently, the region of interest determined by this alignment process (i.e., the registered region) The pixel spatial range of the image is directly applied to the original kV image to perform a cropping operation:
[0088]
[0089] Ensure the kV image after cropping Compared with the training phase The data achieves geometric coarse registration in terms of lung regional anatomical perspective and pixel size, thereby effectively reducing the inter-domain distribution differences caused by different imaging modes.
[0090] Finally, after the above registration and pruning preprocessing, a dataset with spatiotemporal relevance and domain adaptation preprocessing was formed for each patient during the model testing and clinical inference stages:
[0091] .
[0092] S2, 4D CT reconstruction:
[0093] Four-dimensional CT reconstruction is the core processing step of this invention. In the actual radiotherapy stage, the model uses the patient's historical three-dimensional CT volume data and real-time two-dimensional kV projection images under a single respiratory phase to dynamically generate three-dimensional CT volume data with clinical standard CT values (Henry's units, HU) that are synchronized with the real-time images.
[0094] Specifically, this invention employs a deep learning model based on a 3D U-Net architecture to achieve 4D CT reconstruction. During the model training and validation phases, 3D CT volume data is used... kV projection image of the target time phase Using the input image as input, a neural network model predicts and outputs a high-resolution 3D CT volumetric data corresponding to the same respiratory phase as the input kV image. This allows for real-time, precise dynamic modeling of changes in the patient's lung anatomy during respiratory movements, simulating the process of radiotherapy. This provides data support for real-time dose calculation, accumulation, and evaluation, assisting physicians in making real-time treatment decisions in image-guided radiotherapy.
[0095] The model training employs a supervised learning paradigm, using multiple sets of 3D CT volumetric data collected from the same patient at different time phases and data from the same time phase. Images are used as training sample pairs. By optimizing the composite loss function, the network learns to reconstruct the accurate three-dimensional structure of the target phase from the mapping between the two-dimensional projection and the prior three-dimensional reference image. Finally, an optimal model that can be stably reconstructed using prior data in real radiotherapy scenarios is obtained.
[0096] The model uses normalized prior 3D CT volume data and corresponding 2D kV projection images as joint inputs. It learns the mapping relationship between 2D projection images and 3D volume data through multi-scale feature encoding and decoding structures, thereby predicting and outputting 3D CT volume data at the current respiratory phase.
[0097] As a specific embodiment, in step S2 of the four-dimensional CT reconstruction, since the voxel values of three-dimensional CT volume data are represented by HU values, while two-dimensional kV projection images are usually single-channel grayscale values, there are significant differences between the two in terms of numerical range and physical meaning. To ensure that images of different modalities can be used as unified model inputs for joint modeling, this invention first performs normalization (including data normalization and denormalization) processing on the input images in the four-dimensional CT reconstruction.
[0098] Specific standardization processes include: for CT volume data expressed in terms of HU values. Utilize window width and window position The range of values in the region of interest is linearly mapped to the same value range as that of a single-channel kV image. The data normalization process involves truncating voxel values that are too high or too low within a certain range to reduce their impact on the prediction results. This data normalization operation is denoted as [insert operation here]. Specifically, it is expressed as follows:
[0099]
[0100] Predicted image generated by the network The value range remains the single-channel grayscale value. Intervals, through inverse normalization operations This is then mapped back to a preset window width HU value range to fit clinical radiation dose estimation. The specific inverse normalization operation is as follows:
[0101] .
[0102] As a specific embodiment, in step S2 of the four-dimensional CT reconstruction, the two-dimensional kV projection image is... Extending along its projection direction, it forms a three-dimensional "pseudo-data". In the depth direction The depth is the same as that of the 3D CT volume data. Provide three-dimensional spatial structure priors, Then the target phase is encoded on the projection path. The spatiotemporal distribution information of anatomical structures. and By splicing along the channel dimension, a two-channel 4D tensor is formed. As network input.
[0103] The 3D U-Net consists of a symmetrical encoder and decoder path connected by skip connections. The encoder extracts and compresses multi-scale features progressively through a series of operations including 3D convolution, activation, and pooling. The decoder upsamples and fuses contextual information from the encoder features through transposed convolution and skip connections, ultimately outputting the predicted single-channel 3D CT volume data through a 1×1×1 convolutional layer. .
[0104] As a specific embodiment, in step S2 of the four-dimensional CT reconstruction, during the model training process, in order to constrain the performance of the reconstructed CT in terms of voxel HU value accuracy, structural consistency, and edge continuity, this invention constructs a joint optimization objective composed of multiple loss functions, namely a composite loss function. It is composed of multiple loss functions, specifically represented as follows:
[0105]
[0106] in, To define the hyperparameters for the weights of the structural similarity loss function; To define the hyperparameters of the weights of the first-order spatial gradient loss function; The mean squared error (MSE) loss function measures the numerical difference between the predicted CT and the true CT at the voxel level, forcing the reconstructed image to approximate the true value at the voxel level and ensuring the accuracy of the HU value.
[0107]
[0108] in, Representing the true time phase In the CT body data below, the first The size of the voxel value at each position. This indicates that among the CT volume data reconstructed at the same time phase, the first... The size of the voxel value at each position. This indicates the total number of voxels.
[0109] As a structural similarity loss function, a structural similarity (SSIM) loss is introduced as a regularization term to measure the similarity between the reconstructed image and the real image in terms of brightness, contrast, and local structure. Minimizing the structural similarity loss helps the network generate visually more natural and structurally clearer images, improving the reliability of reconstructed CT images in terms of anatomical structures. This invention designs a local cubic window with a size of 5×5×5 voxel spacing, and sets... The structural similarity index represents the set of voxel values in the reconstructed image patch and the real image patch, respectively. The calculation is as follows:
[0110]
[0111] in, windows respectively and The voxel mean represents the brightness comparison component; and For window and voxel variance The covariance of the two radii represents the contrast and structural comparison components; to avoid a denominator of 0 and to enhance computational stability, a... and Two small positive numbers.
[0112] For the entire 3D CT volume data, the overall structural similarity loss is defined as the mean of the 1-SSIM values of all sliding local windows:
[0113]
[0114] in, This represents the total number of windows calculated from the entire volume of data. and Reconstructed CT and real CT at the 1st A voxel block within a window.
[0115] The first-order spatial gradient loss function is introduced as a second regularization term to measure the difference between the reconstructed image and the real image in the spatial gradient domain. This constrains the model to learn high-frequency information of organ contours and vascular nodules at the spatial gradient level, ensuring that the reconstruction results have clear anatomical boundaries and preserve texture details, thereby improving the accuracy of radiotherapy dose calculation and assessment.
[0116]
[0117] in, Indicates the image in Spatial gradient operators in three directions.
[0118] Figure 2 This is a case study illustrating the differences between the present invention and traditional CT reconstruction methods. The reconstruction performance is as follows:
[0119] RMSE PSNR SSIM optimal 40.88 40.02 0.995 worst 96.72 32.53 0.967 average 60.16±16.81 36.98±2.33 0.988±0.007
[0120] Figure 3 This is a visualization demo script diagram of the present invention for reconstructing CT scans.
[0121] Compared with existing technologies, the four-dimensional CT reconstruction method based on real-time kV image driving provided by this invention realizes an end-to-end closed loop from two-dimensional projection to four-dimensional CT to dose assessment, and has the following advantages: 1. It reduces the radiation dose required for image reconstruction; 2. It improves the reconstruction speed; 3. It ensures spatial accuracy and temporal continuity; 4. It has the feasibility and practical value for clinical application.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A four-dimensional CT reconstruction method based on real-time kV image-driven reconstruction, characterized in that, Includes the following steps: S1. Data Preprocessing: S11. Synchronized Respiratory Phase: For each patient Real-time acquisition of timestamp tags associated with respiratory movements using an existing linear accelerator-borne kV imaging system. The original kV projection image is recorded as Raw 4D CT scan data taken during the radiotherapy planning phase This represents the CT imaging of the patient's thoracic and pulmonary anatomy as a function of respiration. The changes; this invention combines real-time respiratory monitoring signals acquired during radiotherapy to record the patient's respiratory status at each time frame. Collected Image sequence allocation and respiratory phases corresponding to 4D CT After this processing, the kV projection image sequence, originally arranged according to clinical acquisition time, was... It was converted into a timeline synchronized with the CT scan and based on the respiratory phase. Indexed four-dimensional sequence data This constitutes the original dataset used in this invention. : S12, Spatial Resampling: For each patient raw CT body data Perform three-dimensional resampling to unify its voxel space and obtain the resampled CT volume data. Simultaneously, the serialized kV projection images Two-dimensional resampling is performed to unify the pixel space to the same scale, resulting in a sampled projected image. ; S13. Geometric Consistency Registration Clipping: This includes between-group difference alignment and within-group difference registration. For between-group differences, this invention introduces an between-group geometric consistency clipping strategy for the resampled data. For each patient... The phase with the largest lung volume in the respiratory cycle, i.e., the end-expiratory phase, is selected. The CT images are used as a reference standard, denoted as This ensures that the cropping intervals for each time phase cover the entire lung region; subsequently, the lung region is automatically segmented to generate a binary mask for the lung region. Based on this segmentation mask, the geometric center point of the lung region is calculated. This center point As the spatial anchor point for the cutting operation, a fixed-size three-dimensional space is cut; in the coronal plane, i.e. In the axial direction, with Extending upwards and downwards from the center. Layer slices, in cross-section On a plane, with Extending forward, backward, left, and right from the center. Individual unit, the formal definition of this clipping operation is: Then perform this cropping operation Consistently applied to all respiratory phases of this patient. Data of a uniform spatial scale can be obtained from image data: in, Model training for 4D CT reconstruction This indicates the height of the CT scan data after uniform cropping. This indicates the width of the CT scan data after uniform cropping. This indicates the depth of the CT volume data after uniform cropping. To address intra-group differences, this invention introduces a spatial alignment registration strategy for the resampled data, the specific process of which is as follows: Digital reconstruction of radiographic images is generated. This step defines a simulated fluoroscopic projection function based on the physical geometry parameters of the treatment accelerator kV imaging system. For patients During a specific respiratory phase Below, CT body data after the aforementioned processing Digital reconstructed images are generated through this projection function. : in, The 3D CT volumetric data is virtually projected onto a spatial coordinate system consistent with the kV image to simulate the data input used for model inference in a real radiotherapy scenario. Simultaneously, it provides sufficient data support for the training of the 4D CT reconstruction model, forming a system where each patient has spatiotemporally correlated 3D CT volumetric data and 2D... Data input pairs: kV image preprocessing is a key step in transforming clinically acquired, synchronized, and sampled real two-dimensional kV images into digital images. Through cropping and geometric transformation, the model is adjusted to align with the data obtained during the planning and training phases in terms of lung anatomy and spatial features, thereby providing consistent and comparable input for the four-dimensional CT reconstruction model. Specifically: Perform two-dimensional registration on the kV projection image, to For fixed images, generated by in-phase CT For floating images, optimize the spatial deformation field Maximize the image similarity between the two : in, Indicates the deformation field Apply to floating images ; The optimized deformation field obtained by optimization Applied to The image is initially aligned with the kV projection image: Subsequently, the region of interest determined by this alignment process is directly applied to the original kV image to perform a cropping operation: After the above registration and cropping preprocessing, a dataset with spatiotemporal correlation and domain adaptation preprocessing was formed for each patient in the model testing and clinical inference stages: S2, 4D CT reconstruction: Four-dimensional CT reconstruction is achieved using a deep learning model based on a 3D U-Net architecture. The model training and validation phases utilize 3D CT volumetric data. kV projection image of the target time phase Using the input image as input, a deep learning model predicts and outputs a high-resolution 3D CT volumetric image corresponding to the same respiratory phase as the input kV image. This allows for real-time, precise dynamic modeling of changes in the patient's lung anatomy during respiratory movements, simulating the process of radiotherapy. The model training employs a supervised learning paradigm, using multiple sets of 3D CT volumetric data collected from the same patient at different time phases and data from the same time phase. Images are used as training sample pairs. By optimizing the composite loss function, the network learns to reconstruct the accurate three-dimensional structure of the target phase from the mapping between the two-dimensional projection and the prior three-dimensional reference image. Finally, an optimal model that can be stably reconstructed using prior data in real radiotherapy scenarios is obtained. The model uses normalized prior 3D CT volume data and corresponding 2D kV projection images as joint inputs. It learns the mapping relationship between 2D projection images and 3D volume data through multi-scale feature encoding and decoding structures, thereby predicting and outputting 3D CT volume data at the current respiratory phase.
2. The four-dimensional CT reconstruction method based on real-time kV image-driven reconstruction according to claim 1, characterized in that, In step S2 of the four-dimensional CT reconstruction, the normalization process includes: for CT volume data expressed in HU values... Utilize window width and window position The range of values in the region of interest is linearly mapped to the same value range as that of a single-channel kV image. The data normalization process involves truncating voxel values that are too high or too low within a certain range to reduce their impact on the prediction results. This normalization is denoted as [missing information]. , means as follows: Predicted image generated by the network The value range remains the single-channel grayscale value. Intervals, through inverse normalization operations This is then mapped back to a preset window width HU value range to fit clinical radiation dose estimation. The specific inverse normalization operation is as follows: 。 3. The four-dimensional CT reconstruction method based on real-time kV image-driven reconstruction according to claim 1, characterized in that, In step S2, the four-dimensional CT reconstruction, the three-dimensional U-Net consists of a symmetrical encoder and decoder path connected by skip connections. The encoder extracts and compresses multi-scale features progressively through 3D convolution, activation, and pooling operations. The decoder upsamples and fuses contextual information based on features from the encoder through transposed convolution and skip connections, ultimately outputting the predicted single-channel three-dimensional CT volume data through a 1×1×1 convolutional layer. .
4. The four-dimensional CT reconstruction method based on real-time kV image-driven reconstruction according to claim 1, characterized in that, In step S2, the four-dimensional CT reconstruction, the composite loss function... It is composed of multiple loss functions, specifically represented as follows: in, The mean squared error loss function is used to measure the numerical difference between the predicted CT and the real CT at the voxel level, forcing the reconstructed image to approximate the real value at the voxel level and ensuring the accuracy of the HU value. The structural similarity loss function measures the similarity between the reconstructed image and the real image in terms of brightness, contrast, and local structure. Minimizing the structural similarity loss helps the network generate visually more natural and structurally clearer images, thereby improving the credibility of the reconstructed CT in terms of anatomical structures. It is a first-order spatial gradient loss function used to measure the difference between the reconstructed image and the real image in the spatial gradient domain. It constrains the model to learn high-frequency information of organ contours and vascular nodules at the spatial gradient level, ensuring that the reconstruction results have clear anatomical boundaries and retain texture details, thereby improving the accuracy of radiotherapy dose calculation and evaluation. To define the hyperparameters for the weights of the structural similarity loss function; This defines the hyperparameters for the weights of the first-order spatial gradient loss function.