A Four-Dimensional Cone-Beam CT Reconstruction Method Based on Differential Transformer Motion Compensation

By employing the differential Transformer motion compensation method, the image blurring problem caused by respiratory motion in four-dimensional cone-beam CT reconstruction was solved, achieving high-precision four-dimensional cone-beam CT image reconstruction, which is suitable for high-precision radiotherapy and other medical imaging scenarios.

CN122134877APending Publication Date: 2026-06-02SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
Filing Date
2026-01-29
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing four-dimensional cone-beam CT reconstruction methods struggle to achieve high-precision image reconstruction in radiotherapy for thoracic and abdominal tumors due to image blurring and artifacts caused by respiratory motion. This is especially true in image-guided radiotherapy for lung and liver cancer, where conventional methods cannot effectively distinguish anatomical structures at different respiratory phases, leading to errors in target delineation and inaccurate organ irradiation.

Method used

A method based on differential Transformer motion compensation is adopted. By acquiring respiratory motion signals in cone-beam CT scan data, the data is divided into multiple temporal phase groups. Feature encoding is performed using differential feature vectors and scan geometric parameters. Feature modeling is performed by combining a Transformer network with spatial and temporal attention modules. Motion-aware regularization constraints are introduced to distinguish between moving and stationary regions for image reconstruction.

Benefits of technology

It achieves high-quality four-dimensional cone-beam CT image reconstruction, eliminates motion blur and artifacts, improves signal-to-noise ratio and spatial consistency of images, provides high-precision image support, and provides reliable image guidance for high-precision radiotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134877A_ABST
    Figure CN122134877A_ABST
Patent Text Reader

Abstract

This invention provides a four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation, comprising: dividing acquired continuous projection data into multiple time-phase groups according to respiratory motion signals; selecting at least one frame in each group as a base reference frame, dividing each group of images into multiple local image blocks and performing feature encoding to obtain the original feature vector of each block; calculating the difference feature vector between the original feature vector of the non-base reference frame and the original feature vector of the corresponding base reference frame; jointly encoding the original feature vector, the difference feature vector, and the angle embedding vector to form an enhanced feature vector, which is input into a Transformer network for feature modeling and motion compensation; introducing motion-aware regularization constraints to distinguish between moving and stationary regions; decoding each group of features, reconstructing the corresponding three-dimensional images of each group, and combining them to form a complete four-dimensional cone-beam CT image sequence. The method of this invention achieves high signal-to-noise ratio, clear, and morphologically accurate 4D-CBCT image reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of medical image processing and analysis, and particularly relates to a four-dimensional cone-beam CT reconstruction method based on differential Transform motion compensation. BACKGROUND

[0002] CBCT has become a standard imaging modality in image-guided radiation therapy (IGRT) due to its small device volume, high spatial resolution, and ability to be integrated into a medical linear accelerator (LINAC). In the process of radiotherapy, CBCT is mainly used to correct patient positioning errors and ensure that the treatment beam is accurately projected onto the tumor target. However, in the image-guided radiotherapy of thoracic and abdominal tumors, the imaging quality of cone-beam computed tomography (CBCT) faces severe challenges.

[0003] However, due to patient respiratory motion and other physiological and pathological processes, as well as the displacement and deformation of thoracic and abdominal organs such as the lungs and liver caused thereby, the projection data collected during the conventional scanning process often contains complex motion information. Conventional 3D-CBCT imaging is based on the assumption of "static anatomical structure", that is, it is assumed that the patient's anatomical structure remains static during the approximately 60-120 seconds of scanning during data acquisition. However, in the radiotherapy of thoracic and abdominal tumors (such as lung cancer and liver cancer), the patient inevitably has respiratory motion. This non-rigid motion of internal organs makes it impossible for conventional 3D-CBCT to distinguish the patient's respiratory motion state during the acquisition process, and the projection data collected is inconsistent in the time dimension, mixing anatomical structures at different respiratory phases.

[0004] This inconsistency in the data will produce a "motion averaging effect" in the image reconstruction process, and the 3D-CBCT image reconstructed therefrom will have severe motion blur and artifacts. Specifically, the tumor volume is elongated or deformed in the image, resulting in morphological distortion and failing to reflect the true anatomical morphology; at the same time, the contrast between the tumor and the surrounding normal tissues such as lung parenchyma and blood vessels is reduced, showing a state of unclear boundary and blurred edge. This low-quality image seriously hinders the accurate delineation of the radiotherapy target area by doctors, easily causing delineation errors, leading to missed irradiation of the tumor or excessive irradiation of the surrounding critical organs, and the potential clinical risks make it impossible for conventional techniques to meet the stringent needs of motion management in high-precision radiotherapy such as stereotactic body radiotherapy. Conventional three-dimensional cone-beam CT (3D-CBCT) is difficult to cope with the non-rigid respiratory motion of internal organs, resulting in mixing of projection data of different phases of anatomical structures, causing severe image blur and tumor morphological distortion, making it difficult to identify the target area boundary.

[0005] To address these issues, 4D-CBCT technology emerged. This technology introduces a time dimension, namely respiratory signals, to classify continuously acquired projection data according to respiratory phase or amplitude, thereby reconstructing 3D image sequences under different respiratory states. Existing 4D-CBCT reconstruction methods can be broadly categorized into three types: traditional analytical and iterative reconstruction methods, deep learning methods based on convolutional neural networks (CNNs), and emerging methods based on Transformers. All of these methods utilize attention mechanisms to model intra- and inter-group relationships in both spatial and temporal dimensions, and combine prior information or motion compensation strategies, providing an effective technical foundation for high-quality 4D-CBCT reconstruction under sparse projection conditions. However, while existing four-dimensional cone-beam computed tomography (4D-CBCT) reconstruction methods can classify projection data according to respiratory phase, the sparse sampling of data for each phase leads to problems such as motion blur, undersampling artifacts, intra-group motion residues, and insufficient inter-group motion compensation. Furthermore, insufficient utilization of geometric priors such as scanning angles results in difficulty in accurately aligning images across phases and loss of detail. Summary of the Invention

[0006] In view of this, the present invention provides a four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation to solve the above problems.

[0007] This invention provides a four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation, comprising: acquiring continuous projection data and scanning geometric parameters during cone-beam CT scanning; extracting respiratory motion signals based on the projection data; dividing the continuous projection data into multiple time phase groups according to the respiratory motion signals; selecting at least one frame in each time phase group as a base reference frame; dividing the image in each time phase group into multiple local image blocks and performing feature encoding to obtain the original feature vector of each block; for non-base reference frames in each time phase group, calculating the difference between the original feature vector of each local image block and the original feature vector of the corresponding base reference frame. The original feature vector, the differential feature vector, and the angle embedding vector generated based on the scanning geometry parameters are jointly encoded to form an enhanced feature vector. The enhanced feature vector is then input into a Transformer network containing spatial attention and temporal attention modules for feature modeling and motion compensation. During the reconstruction process, motion-aware regularization constraints based on the differential feature vector are introduced to distinguish between moving and stationary regions and to apply structure-preserving constraints to stationary regions. The features processed by the Transformer network and guided by regularization constraints are decoded to reconstruct the three-dimensional images corresponding to each temporal phase group and combine them to form a complete four-dimensional cone-beam CT image sequence.

[0008] In another implementation of the present invention, the scanning geometric parameters include at least the projection angle corresponding to each frame projection, the relative spatial position relationship between the X-ray source and the detector, the distance from the source to the isocenter, the distance from the isocenter to the detector, and the pixel size and spatial arrangement of the detector.

[0009] In another implementation of the present invention, the base reference frame is an image representing the average spatial position of the moving organ in a statistical sense within the time phase, used to describe the relatively stable anatomical structure under the respiratory phase, and serves as a unified reference benchmark for motion compensation and feature difference calculation within the time phase group in subsequent processing.

[0010] In another implementation of the present invention, the differential feature vector is represented as:

[0011] in, The original feature vector, The feature vector corresponding to the base reference frame.

[0012] In another implementation of the present invention, the method further includes: obtaining scanning angle information based on geometric parameters during cone-beam CT scanning; mapping the scanning angle information into angles and embedding vectors to characterize the geometric prior information corresponding to different viewpoints during projection acquisition, thereby enhancing the model's ability to perceive multi-view projection distribution and spatial geometric relationships.

[0013] In another implementation of the present invention, the spatial attention module is used to model the spatial correlation between different local patches within the same time phase group to constrain the spatial consistency of anatomical structures and reduce or eliminate local motion residuals caused by irregular patient breathing. It is also used to model the spatial correlation between corresponding patches in different time phase groups to capture cross-phase changes in local anatomical structures. The temporal attention module is used to model the temporal sequence relationship between corresponding patches in different time phase groups to characterize the temporal changes caused by respiratory motion or organ motion, and to achieve cross-phase feature alignment and motion compensation.

[0014] In another implementation of the present invention, the motion-sensing regularization constraint is based on the magnitude of the differential feature vector and its temporal variation pattern to determine the motion state of the reconstructed region, so as to distinguish between the motion region caused by respiratory motion and the relatively static region; for the region determined to be relatively static, a structure-preserving regularization constraint is applied to maintain the spatial consistency of the anatomical structure at different time phases; for the region determined to be a motion region, it is allowed to deform within a reasonable range, thereby avoiding excessive smoothing of the motion region or structural distortion introduced by strong constraints while ensuring the overall anatomical consistency, and improving the comprehensive performance of the four-dimensional reconstruction results in terms of spatiotemporal continuity and local structural realism.

[0015] In another aspect, the present invention provides a four-dimensional cone-beam CT reconstruction system based on differential Transformer motion compensation, comprising: a data preprocessing module: acquiring continuous projection data and scanning geometric parameters during cone-beam CT scanning, extracting respiratory motion signals based on the projection data, and dividing the continuous projection data into multiple time phase groups according to the respiratory motion signals; a feature extraction module: selecting at least one frame as a base reference frame within each time phase group, dividing the image within each time phase group into multiple local image blocks and performing feature encoding to obtain the original feature vector of each block; and a feature enhancement module: for non-base reference frames within each time phase group, calculating the original feature vector of each local image block and the original feature vector of the corresponding base reference frame. The original feature vector, the differential feature vector, and the angle embedding vector generated based on the scanning geometry parameters are jointly encoded to form an enhanced feature vector. The feature reconstruction module inputs the enhanced feature vector into a Transformer network containing spatial and temporal attention modules for feature modeling and motion compensation. During reconstruction, motion-aware regularization constraints based on the differential feature vector are introduced to distinguish between moving and stationary regions and apply structure-preserving constraints to the stationary regions. The result output module decodes the features processed by the Transformer network and guided by regularization constraints, reconstructs the three-dimensional images corresponding to each temporal phase group, and combines them to form a complete four-dimensional cone-beam CT image sequence.

[0016] In another aspect, the present invention provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of a four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation as described in any of the preceding claims. In another aspect, the present invention provides a computer storage medium storing a computer program that, when executed by a processor, implements the steps of a four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation as described in any of the preceding claims.

[0017] The present invention provides a four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation, which accurately compensates for intra-group and inter-group motion, eliminating afterimages and blurring; utilizes patch differential enhancement and full projection data to reduce sparse sampling artifacts and improve the signal-to-noise ratio; introduces scanning angle embedding and static region regularization to achieve local and global motion alignment; and reconstructs high-quality, morphologically accurate 4D-CBCT images, providing reliable imaging support for high-precision radiotherapy. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. By reading the detailed description of the embodiments below, the advantages and benefits of the solutions will become clear to those skilled in the art. The accompanying drawings are only for illustrating preferred embodiments and are not intended to limit the present invention. In the accompanying drawings: Figure 1 This is a schematic diagram of a four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation, according to an embodiment of the present invention.

[0019] Figure 2 This is a schematic diagram of an end-to-end 4D-CBCT image reconstruction framework according to an embodiment of the present invention. Detailed Implementation

[0020] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art should fall within the protection scope of the present invention.

[0021] Figure 1 This is a schematic diagram of a four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation provided in an embodiment of the present invention, as shown below. Figure 1 As shown, this embodiment mainly includes: S101. Acquire continuous projection data and scanning geometric parameters during cone-beam CT scanning, extract respiratory motion signals based on the projection data, and divide the continuous projection data into multiple time phase groups according to the respiratory motion signals.

[0022] For example, a respiratory motion extraction method based on the projection data itself is used to automatically extract the patient's respiratory signal from the continuous projection data. This respiratory signal is obtained using the Amsterdam Shroud respiratory signal extraction algorithm, which utilizes the time-varying projection intensity distribution characteristics in the projection sequence to estimate the respiratory motion cycle and phase information. This method does not rely on external respiratory monitoring equipment; the patient's respiratory motion information can be obtained solely based on the raw projection data acquired during the scanning process and the corresponding geometric information. The continuously acquired projection data is divided into several time phase groups based on the respiratory signal. (Groups), each time phase group corresponds to a specific respiratory state.

[0023] S102. Select at least one frame as the base reference frame within each time phase group, divide the image within each time phase group into multiple local image blocks and perform feature encoding to obtain the original feature vector of each block.

[0024] For example, the reconstructed image or intermediate features in each time phase group are segmented into blocks, and the reconstructed image or intermediate features are divided into multiple local image blocks (Patch) according to a preset spatial scale. Feature encoding is performed on each local image block to obtain the corresponding original feature vector.

[0025] S103. For each non-base reference frame in each time phase group, calculate the difference feature vector between the original feature vector of each local image block and the original feature vector of the corresponding base reference frame.

[0026] S104. The original feature vector, the differential feature vector, and the angle embedding vector generated based on the scanning geometric parameters are jointly encoded to form an enhanced feature vector.

[0027] For example, concatenating or linearly mapping the original feature vector, the difference feature vector, and the angle embedding vector to form an enhanced feature vector significantly improves the model's ability to perceive multi-view projection distributions and spatial geometric relationships.

[0028] S105. Input the enhanced feature vector into a Transformer network containing spatial attention and temporal attention modules for feature modeling and motion compensation.

[0029] For example, by cascading computation of multiple spatial-temporal attention units, the network can simultaneously capture subtle motion changes in local areas and temporal evolution features at the global scale within a unified framework, thereby providing a stable and accurate feature representation foundation for subsequent high-quality four-dimensional image reconstruction.

[0030] S106. In the reconstruction process, motion-aware regularization constraints based on differential feature vectors are introduced to distinguish between moving regions and stationary regions and to apply structure-preserving constraints to stationary regions.

[0031] S107. Decode the features after processing by the Transformer network and guided by regularization constraints, reconstruct the three-dimensional images corresponding to each time phase group, and combine them to form a complete four-dimensional cone-beam CT image sequence.

[0032] For example, the feature vectors output by the Transformer are mapped back to the image space to reconstruct the three-dimensional CBCT images corresponding to each time phase, and then combined in chronological order to form a complete four-dimensional CBCT image sequence.

[0033] The present invention provides a four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation, which accurately compensates for intra-group and inter-group motion, eliminating afterimages and blurring; utilizes patch differential enhancement and full projection data to reduce sparse sampling artifacts and improve the signal-to-noise ratio; introduces scanning angle embedding and static region regularization to achieve local and global motion alignment; and reconstructs high-quality, morphologically accurate 4D-CBCT images, providing reliable imaging support for high-precision radiotherapy.

[0034] In another implementation of the present invention, the scanning geometric parameters include at least the projection angle corresponding to each frame projection, the relative spatial position relationship between the X-ray source and the detector, the distance from the source to the isocenter, the distance from the isocenter to the detector, and the pixel size and spatial arrangement of the detector.

[0035] In another implementation of the present invention, the base reference frame is an image representing the average spatial position of the moving organ in a statistical sense within the time phase, used to describe the relatively stable anatomical structure under the respiratory phase, and serves as a unified reference benchmark for motion compensation and feature difference calculation within the time phase group in subsequent processing.

[0036] In another implementation of the present invention, the differential feature vector is represented as:

[0037] in, The original feature vector, The feature vector corresponding to the base reference frame.

[0038] For example, suppose the first In the time phase group, the The first frame corresponding to The original feature vectors of each local image patch are The feature vector corresponding to the base reference frame is .

[0039] For each non-base frame within a temporal phase group, the feature difference between it and the base reference frame at the corresponding local patch is calculated to explicitly represent local anatomical changes caused by respiration or organ movement. This allows the motion compensation process to proactively provide motion information instead of relying on implicit learning by the network. By calculating the feature difference between non-base frames and the base reference frame, local motion information is explicitly represented, enabling the space-time attention transformer to accurately capture subtle changes caused by respiratory or organ movement. This achieves intra-group micro-motion compensation and cross-phase group alignment, significantly reducing motion ghosting and blurring effects.

[0040] In another implementation of the present invention, the method further includes: obtaining scanning angle information based on geometric parameters during cone-beam CT scanning; mapping the scanning angle information into angles and embedding vectors to characterize the geometric prior information corresponding to different viewpoints during projection acquisition, thereby enhancing the model's ability to perceive multi-view projection distribution and spatial geometric relationships.

[0041] For example, the scanning angle and geometric parameters are mapped to angle embedding vectors and jointly encoded with the original features and differential features, so that the network can make full use of the spatial geometric prior of the projection acquisition, enhance the reconstruction stability and accuracy of the model under sparse projection conditions, and improve the image signal-to-noise ratio and structural integrity.

[0042] In another implementation of the present invention, the spatial attention module is used to model the spatial correlation between different local patches within the same time phase group to constrain the spatial consistency of anatomical structures and reduce or eliminate local motion residuals caused by irregular patient breathing. It is also used to model the spatial correlation between corresponding patches in different time phase groups to capture cross-phase changes in local anatomical structures. The temporal attention module is used to model the temporal sequence relationship between corresponding patches in different time phase groups to characterize the temporal changes caused by respiratory motion or organ motion, and to achieve cross-phase feature alignment and motion compensation.

[0043] For example, spatial attention modeling is performed on patches within the same time frame, and spatial and temporal attention modeling is performed on corresponding patches between different time phase groups. This achieves unified optimization of local anatomical structure consistency and cross-phase motion compensation, thereby improving the spatial resolution and temporal continuity of the four-dimensional image.

[0044] In another implementation of the present invention, the motion-sensing regularization constraint is based on the magnitude of the differential feature vector and its temporal variation pattern to determine the motion state of the reconstructed region, so as to distinguish between the motion region caused by respiratory motion and the relatively static region; for the region determined to be relatively static, a structure-preserving regularization constraint is applied to maintain the spatial consistency of the anatomical structure at different time phases; for the region determined to be a motion region, it is allowed to deform within a reasonable range, thereby avoiding excessive smoothing of the motion region or structural distortion introduced by strong constraints while ensuring the overall anatomical consistency, and improving the comprehensive performance of the four-dimensional reconstruction results in terms of spatiotemporal continuity and local structural realism.

[0045] For example, the amplitude and time-series variation patterns of the differential feature vectors distinguish between moving regions and relatively static regions. Structural preservation constraints are applied to static regions, while reasonable deformation is allowed to moving regions. This achieves the realistic preservation of local anatomical structures and consistency with the overall space, avoids excessive smoothing or structural distortion, and improves the anatomical realism and local detail fidelity of the reconstructed image.

[0046] Optionally, for continuously acquired cone-beam CT projection data or corresponding reconstructed images, considering the large amount of data input to the deep learning model, consecutive frames within a temporal phase group or frames across groups can be divided into adjacent frame pairs. Only two adjacent frames are fed into the Transformer network for processing at a time, and the entire time series is processed sequentially to gradually complete motion compensation within and between groups. This approach effectively reduces network memory usage and computational pressure without significantly decreasing compensation accuracy, making it suitable for clinical environments with limited hardware resources.

[0047] Optionally, the patch can be divided into variable-scale local image blocks based on the anatomical complexity or motion amplitude of the scanned area. For example, areas with large motion amplitude can be divided into smaller patches to capture fine motion, while areas with small motion amplitude can be divided into larger patches to reduce computational load. This approach can flexibly adapt to the reconstruction needs of different body parts.

[0048] This invention achieves effective compensation for minute intra-group and cross-group motions by using intra-group and inter-group differential tokens, spatial and temporal attention Transformer, scanning angle embedding, and motion-aware regularization constraints. It eliminates motion blur and undersampling artifacts without increasing radiation dose, thereby obtaining high signal-to-noise ratio, clear and morphologically accurate 4D-CBCT images.

[0049] Example 1 like Figure 2As shown, an end-to-end 4D-CBCT (four-dimensional cone-beam computed tomography) image reconstruction framework is presented, aiming to address the limitations of traditional reconstruction methods in handling respiratory motion artifacts. The core idea is to capture subtle motion changes by explicitly extracting "differential features" and leveraging the powerful sequence modeling capabilities of Transformer to simultaneously optimize image quality in both spatial and temporal dimensions, ultimately outputting a high-precision dynamic 4D image sequence.

[0050] Detailed stage descriptions: 1. First stage: Data acquisition and preprocessing Input data: The system receives continuously acquired X-ray projection data and corresponding scanning geometric parameters.

[0051] Respiratory signal extraction: The Amsterdam Shroud algorithm is used to automatically extract respiratory motion signals from the raw projection data without the need for additional hardware monitoring.

[0052] Phase segmentation: Based on the periodicity of the respiratory signal, the continuous projection data is divided into G different time phases to prepare for subsequent dynamic reconstruction.

[0053] 2. Second Stage: Differential Feature Construction and Geometric Embedding Base reference frame selection: Select a base reference frame t0.

[0054] Difference calculation This is the core innovation of the algorithm: it not only focuses on the current feature vector f, but also calculates the difference between the current feature and the feature of the reference frame. .

[0055] Geometric embedding: Geometric parameters such as scanning angle are mapped into vectors and injected into the network as prior knowledge to help the model understand spatial relationships under different projection angles.

[0056] Feature fusion: Jointly encode the original features, differential features, and geometric embedding vectors to generate an "enhanced feature vector" containing rich spatiotemporal information.

[0057] 3. Third stage: Space-time Transformer modeling Dual attention mechanism: Enhanced features are fed into the Transformer network, which contains two key components: Spatial attention: Within the same temporal phase, it learns the correlation between pixels at different spatial locations to ensure the integrity and consistency of the anatomical structure.

[0058] Temporal attention: Learns the evolution of features across different time phases, and performs cross-phase feature alignment and motion compensation.

[0059] The two work in cascade, effectively utilizing the global contextual information of 4D data.

[0060] 4. Fourth Stage: Regularization Constraints and Image Reconstruction Motion perception regularization constraint: The system is based on difference features Based on amplitude and timing patterns, it intelligently determines the state of image regions: Relatively static regions: Apply structure-preserving constraints to prevent spurious movements in static tissues (such as bones).

[0061] Motion zone: Allows reasonable non-rigid deformation, avoiding excessive smoothing that can blur the edges of tumors or organs.

[0062] Decoding and Output: Guided by the regularization term, the decoder maps the optimized features back to the image space to reconstruct 3D-CBCT images of each phase.

[0063] Final output: Images from all phases are combined in chronological order to form a complete 4D-CBCT image sequence, clearly showing the dynamic changes in anatomical structures during the respiratory cycle.

[0064] The method of this invention generally includes steps such as temporal phase partitioning, differential feature construction, geometric prior embedding, spatial-temporal attention modeling, motion-aware regularization constraints, and four-dimensional image reconstruction. Its core lies in explicitly introducing differential feature vectors to represent motion information and combining this with the Transformer's self-attention mechanism to achieve multi-scale, multi-temporal motion compensation. This invention does not allow the network to "implicitly guess where the movement is," but rather actively calculates the "motion changes" explicitly in the form of differential features, and then passes them to the Transformer for unified modeling across multiple spatial and temporal scales through the attention mechanism, thereby achieving more accurate motion compensation. By organically combining the base reference frame, differential features, enhanced feature vectors, spatial-temporal attention modeling, and motion-aware regularization constraints, it achieves intra-group micro-motion compensation and cross-phase inter-group alignment, thus obtaining high signal-to-noise ratio, low artifact, and structurally accurate 4D-CBCT image sequences under sparse projection conditions.

[0065] Through the above technical solution, this invention can fully utilize temporal dimension information and scanning geometric priors under sparse projection conditions to achieve coordinated motion compensation within and between groups. In this process, through explicit representation of differential feature vectors, joint modeling of the space-time attention Transformer, and guidance from motion-aware regularization constraints, this invention can effectively suppress blurring effects and sparse sampling artifacts caused by respiratory or organ motion, while maintaining the spatial consistency and temporal continuity of anatomical structures, thereby significantly improving the spatial resolution, temporal consistency, and structural accuracy of 4D-CBCT images. This method is applicable to image guidance and motion management in high-precision radiotherapy, providing clear and reliable four-dimensional image information support for clinical practice.

[0066] The method of this invention is not only applicable to image guidance and motion management in high-precision radiotherapy, but also to other four-dimensional medical imaging scenarios, such as dynamic angiography, cardiac motion CT, and respiratory motion assessment. In these scenarios, by explicitly representing motion information through differential features and combining it with a space-time attention Transformer for modeling, high-quality four-dimensional dynamic image reconstruction and motion analysis can be achieved.

[0067] Another aspect of the present invention provides a four-dimensional cone-beam CT reconstruction system based on differential Transformer motion compensation, comprising: Data preprocessing module: acquires continuous projection data and scanning geometric parameters during cone-beam CT scanning, extracts respiratory motion signals based on the projection data, and divides the continuous projection data into multiple time phase groups according to the respiratory motion signals.

[0068] Feature extraction module: Select at least one frame as the base reference frame within each time phase group, divide the image within each time phase group into multiple local image blocks and perform feature encoding to obtain the original feature vector of each block.

[0069] Feature enhancement module: For each non-base reference frame in each time phase group, calculate the difference feature vector between the original feature vector of each local image block and the original feature vector of the corresponding base reference frame; jointly encode the original feature vector, the difference feature vector and the angle embedding vector generated based on the scanning geometry parameters to form the enhanced feature vector.

[0070] Feature Reconstruction Module: The enhanced feature vector is input into a Transformer network containing spatial attention and temporal attention modules for feature modeling and motion compensation; during the reconstruction process, motion-aware regularization constraints based on differential feature vectors are introduced to distinguish between moving and stationary regions and apply structure-preserving constraints to stationary regions.

[0071] The output module decodes the features processed by the Transformer network and guided by regularization constraints, reconstructs the three-dimensional images corresponding to each time phase group, and combines them to form a complete four-dimensional cone-beam CT image sequence.

[0072] The present invention provides a four-dimensional cone-beam CT reconstruction system based on differential Transformer motion compensation, which accurately compensates for intra- and inter-group motion, eliminating afterimages and blurring; utilizes patch differential enhancement and full projection data to reduce sparse sampling artifacts and improve the signal-to-noise ratio; introduces scanning angle embedding and static region regularization to achieve local and global motion alignment; and reconstructs high-quality, morphologically accurate 4D-CBCT images, providing reliable imaging support for high-precision radiotherapy.

[0073] In another aspect of the present invention, the electronic device includes: a processor, a memory, and a communication bus and a communication interface.

[0074] in: The processor, memory, and communication interface communicate with each other via a communication bus.

[0075] A communication interface is used to communicate with other electronic devices or servers.

[0076] The processor is used to execute programs, specifically, to perform any of the steps of the four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation in the above embodiments.

[0077] Specifically, the program may include program code, which includes computer operation instructions.

[0078] The processor may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0079] Memory is used to store programs. Memory may include high-speed RAM, and may also include non-volatile memory, such as at least one disk drive.

[0080] Specifically, the program can be used to cause the processor to execute the steps of any of the four-dimensional cone-beam CT reconstruction methods based on differential Transformer motion compensation described in the embodiments. The specific implementation of each step in the program can be found in the corresponding descriptions of the steps and units executed in any of the four-dimensional cone-beam CT reconstruction methods based on differential Transformer motion compensation described above, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments.

[0081] An exemplary embodiment of this application also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods of various embodiments of this application.

[0082] The methods described above according to embodiments of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded via a network. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0083] Specific embodiments of the present invention have now been described. Other embodiments are within the scope of the appended claims. In some cases, the actions described in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result.

[0084] It should be noted that all directional indications (such as up, down, left, right, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship between the components in a certain order (as shown in the figure). If the specific order changes, the directional indication will also change accordingly.

[0085] In the description of this invention, the terms "first" and "second" are used only for convenience in describing different components or names, and should not be construed as indicating or implying a sequential relationship, relative importance, or implicitly specifying the number of technical features indicated. Thus, a feature defined with "first" and "second" may explicitly or implicitly include at least one of that feature.

[0086] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0087] It should be noted that although specific embodiments of the present invention have been described in detail with reference to the accompanying drawings, this should not be construed as limiting the scope of protection of the present invention. Various modifications and variations that can be made by those skilled in the art without inventive effort within the scope described in the claims still fall within the scope of protection of the present invention.

[0088] The examples of the embodiments of the present invention are intended to concisely illustrate the technical features of the embodiments of the present invention, so that those skilled in the art can intuitively understand the technical features of the embodiments of the present invention, and are not intended to be an improper limitation of the embodiments of the present invention.

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A four-dimensional cone-beam computed tomography (CBCT) reconstruction method based on differential Transformer motion compensation, characterized in that, include: Acquire continuous projection data and scanning geometric parameters during cone-beam CT scanning, extract respiratory motion signals based on the projection data, and divide the continuous projection data into multiple time phase groups according to the respiratory motion signals; At least one frame is selected as the base reference frame within each time phase group. The image within each time phase group is divided into multiple local image blocks and feature encoding is performed to obtain the original feature vector of each block. For each non-base reference frame within a time phase group, calculate the difference feature vector between the original feature vector of each local image patch and the original feature vector of the corresponding base reference frame. The original feature vector, the differential feature vector, and the angle embedding vector generated based on the scanning geometry parameters are jointly encoded to form an enhanced feature vector; The enhanced feature vectors are input into a Transformer network containing spatial attention and temporal attention modules for feature modeling and motion compensation. During the reconstruction process, motion-aware regularization constraints based on the differential feature vectors are introduced to distinguish between moving regions and stationary regions and to apply structure-preserving constraints to stationary regions. The features processed by the Transformer network and guided by the regularization constraints are decoded to reconstruct the three-dimensional images corresponding to each time phase group, and combined to form a complete four-dimensional cone-beam CT image sequence.

2. The method according to claim 1, characterized in that, The scanning geometry parameters include at least the projection angle corresponding to each frame projection, the relative spatial position relationship between the X-ray source and the detector, the distance from the source to the isocenter, the distance from the isocenter to the detector, and the pixel size and spatial arrangement of the detector.

3. The method according to claim 1, characterized in that, The base reference frame is an image representing the statistically average spatial position of the moving organs within that time phase. It is used to describe the relatively stable anatomical structure under that respiratory phase and serves as a unified reference benchmark for motion compensation and feature difference calculation within that time phase group in subsequent processing.

4. The method according to claim 3, characterized in that, The difference feature vector is represented as: in, The original feature vector, The feature vector corresponding to the base reference frame.

5. The method according to claim 1, characterized in that, Also includes: Scanning angle information is obtained based on geometric parameters during cone-beam CT scanning; The scanning angle information is mapped to an angle and embedded as a vector to represent the geometric prior information corresponding to different viewpoints during the projection acquisition process, thereby enhancing the model's ability to perceive multi-view projection distribution and spatial geometric relationships.

6. The method according to claim 1, characterized in that, The spatial attention module is used to model the spatial relationships between different local patches within the same time phase group to constrain the spatial consistency of anatomical structures and reduce or eliminate local motion residuals caused by irregular patient breathing. It is also used to model the spatial relationships between corresponding patches between different time phase groups to capture cross-phase changes in local anatomical structures. The time attention module is used to model the time series relationship between corresponding patches in different time phase groups to characterize the temporal changes caused by respiratory movements or organ movements, and to achieve cross-phase feature alignment and motion compensation.

7. The method according to claim 1, characterized in that, The motion-sensing regularization constraint is based on the magnitude of the differential feature vector and its temporal change pattern to determine the motion state of the reconstructed region, so as to distinguish between the motion region caused by respiratory movement and the relatively static region. For regions determined to be relatively static, structure-preserving regular constraints are applied to maintain the spatial consistency of anatomical structures at different time phases. For regions identified as motion areas, deformation is allowed within a reasonable range. This ensures overall anatomical consistency while avoiding excessive smoothing of motion areas or structural distortions introduced by strong constraints, thereby improving the overall performance of four-dimensional reconstruction results in terms of spatiotemporal continuity and local structural realism.

8. A four-dimensional cone-beam CT reconstruction system based on differential Transformer motion compensation, characterized in that, include: Data preprocessing module: acquires continuous projection data and scanning geometric parameters during cone-beam CT scanning, extracts respiratory motion signals based on the projection data, and divides the continuous projection data into multiple time phase groups according to the respiratory motion signals; Feature extraction module: Select at least one frame as the base reference frame within each time phase group, divide the image within each time phase group into multiple local image blocks and perform feature encoding to obtain the original feature vector of each block; Feature enhancement module: For each non-base reference frame in each time phase group, calculate the difference feature vector between the original feature vector of each local image block and the original feature vector of the corresponding base reference frame; jointly encode the original feature vector, the difference feature vector and the angle embedding vector generated based on the scanning geometry parameters to form the enhanced feature vector; Feature Reconstruction Module: The enhanced feature vector is input into a Transformer network containing spatial attention and temporal attention modules for feature modeling and motion compensation; during the reconstruction process, motion-aware regularization constraints based on differential feature vectors are introduced to distinguish between moving and stationary regions and apply structure-preserving constraints to stationary regions; The output module decodes the features processed by the Transformer network and guided by regularization constraints, reconstructs the three-dimensional images corresponding to each time phase group, and combines them to form a complete four-dimensional cone-beam CT image sequence.

9. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which, when executed by a processor, implements the steps in the four-dimensional cone-beam CT reconstruction method based on differential Transformer motion compensation as described in any one of claims 1 to 7.