Method for generating three-dimensional CT image from cross-modal dual-view DR image based on structure guidance
The method of generating three-dimensional CT images from cross-modal dual-view DR images based on structure guidance solves the problem of generating high-fidelity standing three-dimensional CT images from full-length DR anteroposterior and lateral radiographs of the spine in the existing technology. It realizes high-precision, low-radiation three-dimensional CT image reconstruction, which is suitable for the diagnosis and surgical planning of scoliosis.
Patent Information
- Application Number
- CN202511104379.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies cannot effectively generate high-fidelity standing 3D CT images from conventional full-length spinal DR anteroposterior and lateral radiographs, leading to errors and radiation risks in surgical planning, as well as insufficient equipment accessibility.
A structure-guided method for generating 3D CT images from cross-modal dual-view DR images is proposed. This method utilizes structural priors, multi-angle information fusion, voxel-level reconstruction mechanisms, and multi-target loss optimization strategies to generate high-quality 3D CT images from dual-view DR images. The process includes data preprocessing, feature extraction, feature fusion, layer-by-layer voxel growth reconstruction, and multi-target loss optimization.
It achieves high-precision, low-radiation 3D CT image reconstruction, reduces the risk of radiation exposure for patients, improves surgical safety and equipment versatility, and is suitable for the diagnosis and surgical planning of adolescent scoliosis.
Smart Images

Figure CN120997395A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, specifically to a method for generating three-dimensional CT images from structure-guided cross-modal dual-view DR images. Background Technology
[0002] Scoliosis is a common disease characterized by three-dimensional spinal deformity, with an incidence rate of approximately 2%-4% in adolescents. This disease causes abnormal curvature in the sagittal, coronal, and axial planes of the spine, leading not only to postural imbalance and cardiopulmonary dysfunction but also potentially progressive nerve compression, severely impacting adolescent growth and development. Currently, standing full-length digital radiography (DR) of the spine and conventional 3D CT are core tools for clinical diagnosis and surgical navigation. Doctors assess the degree of deformity and develop corrective treatment plans by measuring parameters such as the Cobb angle and apical vertebral offset. However, this technology has the following inherent defects: (1) Lack of three-dimensional structural information in DR: The two-dimensional projection characteristics of full-length DR anteroposterior and lateral radiographs of the spine cannot accurately quantify the vertebral rotation angle and the three-dimensional orientation of the pedicle, resulting in an experience-dependent error in the screw implantation path during surgical planning; (2) Poor repeatability of DR measurements: Studies have shown that the Cobb angle measurement of full-length DR anteroposterior and lateral radiographs of the spine can vary by 5°-10° within the group; (3) Distortion of biomechanical state in conventional three-dimensional CT: Due to hardware limitations, conventional CT scans can only be performed in the supine position, which cannot simulate the standing weight-bearing state, resulting in a systematic underestimation of key parameters (such as the mean difference of Cobb angle between the standing and supine positions of 7.2°±2.1°).
[0003] The full-length standing 3D CT scan of the spine obtains a high-resolution 3D model of the spine by scanning the patient while standing and bearing weight. It can completely replace the standing full-length DR of the spine and conventional 3D CT. Its advantages are: (1) spatial resolution of 1mm, accurately presenting the pedicle morphology and facet joint spatial relationship; (2) restoring the real biomechanical state and supporting personalized orthopedic force line design.
[0004] However, the clinical application of standing 3D CT technology faces severe challenges: (1) High radiation dose: The effective dose of a single full-length spinal CT is about 4-8 mSv, which is more than 200 times that of DR, and the risk of ionizing radiation exposure to patients is significantly increased; (2) Insufficient equipment accessibility: Standing CT is a high-end specialized CT, with only 12 units installed worldwide and only two units in China, which are still in the clinical validation stage, resulting in high cost for patients per examination.
[0005] Currently, some international studies have attempted to use single-view DR images for 3D reconstruction. However, due to limitations in 2D information and the performance of reconstruction algorithms, the resulting 3D models still have significant shortcomings in terms of spatial accuracy and structural integrity. Furthermore, traditional 3D reconstruction methods based on physical models or geometric assumptions are computationally intensive and complex, making it difficult to achieve real-time intraoperative 3D reconstruction and failing to meet the needs of rapid clinical response. Therefore, generating high-fidelity standing 3D CT images from conventional full-length DR anteroposterior and lateral views of the spine, while ensuring low radiation dose and universal equipment availability, has become a critical technical bottleneck that spinal surgery urgently needs to overcome.
[0006] To address the aforementioned issues, this invention proposes a method for generating 3D CT images from cross-modal dual-view DR images based on structure guidance. This method can synthesize high-quality 3D CT images in the preoperative stage using only conventionally acquired dual-view DR images, thereby obtaining complete 3D structural information of the standing spine without traditional CT scans. It is particularly suitable for the diagnosis and follow-up of adolescent scoliosis, surgical path planning, and intraoperative navigation, effectively reducing the risk of ionizing radiation exposure and improving surgical safety. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention aims to provide a method for generating three-dimensional CT images from cross-modal dual-view DR images based on structure guidance. By combining structural priors, multi-angle information fusion, voxel-level reconstruction mechanisms, and multi-objective loss optimization strategies, an effective conversion from two-dimensional DR images to high-fidelity three-dimensional CT images is achieved, making it particularly suitable for the diagnosis, treatment, and follow-up of adolescent scoliosis patients.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] A method for generating 3D CT images from cross-modal dual-view DR images based on structure guidance includes the following steps:
[0010] Step S1: Acquire dual-view DR images and paired 3D CT images of the patient from two different projection angles, perform data preprocessing, and construct a training dataset;
[0011] Step S2: Input the dual-view DR images into the structure prior-guided dual-branch encoder network respectively, and extract their respective image feature information;
[0012] Step S3: The image features from the two perspectives are fused through the orientation difference perception module to obtain an enhanced three-dimensional structural representation;
[0013] Step S4: Input the fused features into the layer-by-layer voxel growth and reconstruction module to gradually generate three-dimensional CT image data with consistent spatial structure.
[0014] Step S5: Introduce a global-local collaborative attention mechanism during image decoding to enhance the structural representation of key anatomical regions;
[0015] Step S6: Train and optimize the network based on a multi-objective loss function driven by structural consistency to improve the density accuracy and anatomical consistency of the generated images.
[0016] Step S7: Train the model constructed in steps S2, S3, S4, S5, and S6 to generate three-dimensional CT volume data from dual-view DR images.
[0017] Furthermore, in step S1, acquiring dual-view DR images of the patient from two different projection angles and paired 3D CT images, performing data preprocessing, and constructing a training dataset specifically includes the following steps:
[0018] (1) Acquire anteroposterior and lateral DR images of patients using clinical DR equipment to ensure short imaging time intervals and uniform posture;
[0019] (2) Use high-resolution CT scanning equipment to acquire three-dimensional CT data that are consistent with the anatomical region of the DR image;
[0020] (3) Denoise and normalize the dual-view DR images, unify the image size and resolution, crop the three-dimensional CT data, remove the non-interest regions, and unify the spatial resolution and coordinate system to obtain the matching dual-view DR images and three-dimensional CT data as training datasets.
[0021] Furthermore, in step S2, the dual-view DR images are input into a structure-prior-guided dual-branch encoder network to extract their respective image feature information, specifically including the following steps:
[0022] (1) Input the DR images of the two projection angles into a dual-branch convolutional encoder with a shared structure prior module to enhance cross-view feature consistency by sharing anatomical guidance parameters;
[0023] (2) Extract multi-scale image features in each coding branch, including edge, texture and structural information, and enhance the expressive power of key regions through structural attention units;
[0024] (3) A parameter alignment mechanism is adopted to maintain the semantic coordination between the two encoding paths, providing prior support for structural alignment for subsequent feature fusion.
[0025] Furthermore, in step S3, the image features from the two perspectives are fused through the orientation difference perception module to obtain an enhanced three-dimensional structural representation, specifically including the following steps:
[0026] (1) Establish an angle difference perception module, including an angle attention sub-module and structural contrast loss, to simulate the angle transformation relationship between viewpoints;
[0027] (2) Construct a directional encoding tensor to participate in feature weighted fusion in order to improve the structural consistency representation;
[0028] (3) Early detailed features are called back through a jump connection mechanism to enhance the ability to perceive 3D structures.
[0029] Furthermore, in step S4, the fused features are input into the layer-by-layer voxel growth and reconstruction module to gradually generate three-dimensional CT image data with consistent spatial structure, specifically including the following steps:
[0030] (1) Design a layer-by-layer voxel growth and reconstruction module, which consists of a voxel candidate generator and a spatial consistency filter;
[0031] (2) A bottom-up approach is used to generate three-dimensional CT image voxels layer by layer. The voxel filling order is guided by a probability map, and the spatial structure is reconstructed progressively from low resolution to high resolution.
[0032] (3) Introduce a structural sparsity regularization term to constrain the continuity and density consistency between the reconstructed voxels in each layer, so as to prevent artifacts or structural drift in the reconstructed image.
[0033] Furthermore, in step S5, the fused features are input into the layer-by-layer voxel growth and reconstruction module to gradually generate three-dimensional CT image data with consistent spatial structure, specifically including the following steps:
[0034] (1) Construct a global-local collaborative attention mechanism, in which the global branch generates structural context representation based on the global response of the image, and the local branch focuses on key anatomical regions such as the spine, ribs and other detailed parts;
[0035] (2) By using fused attention maps to guide the image decoding process, the model can more accurately recover the boundary and density changes of complex tissues during the generation process;
[0036] (3) By using a dual-scale supervision mechanism, structural similarity loss is applied to both the intermediate layer and the output layer to improve the detail fidelity of the decoded output;
[0037] Furthermore, in step S6, the network is trained and optimized based on a multi-objective loss function driven by structural consistency to improve the density accuracy and anatomical consistency of the generated images. This specifically includes the following steps:
[0038] (1) Construct a multi-objective joint loss function, which includes three sub-items: density regression loss, structural similarity loss, and anatomical consistency loss;
[0039] (2) The density regression loss uses a weighted mean square error function to measure the difference in gray density between the predicted CT image and the real CT image;
[0040] (3) Structural similarity loss is used to preserve the overall structural information of the image and improve the restoration of image texture and boundaries;
[0041] (4) Anatomical consistency loss: Combine prior segmentation mask or key point label to apply position and shape consistency constraints to important organ regions to ensure the anatomical accuracy of the generated image.
[0042] (5) Set dynamic weighting coefficients for each of the above sub-loss functions and adjust them automatically according to the training stage to optimize the overall learning process.
[0043] Furthermore, in step S7, the model constructed in steps S2, S3, S4, S5, and S6 is trained to generate three-dimensional CT volume data from dual-view DR images, specifically including the following steps:
[0044] (1) Input the dual-view DR images of the patient to be reconstructed into the trained and optimized structure-guided network model, and extract the image feature information of the two views through the dual-branch encoder module respectively;
[0045] (2) Input the extracted dual-view image features into the orientation difference perception module, fuse the multi-view spatial structure information, and generate a unified three-dimensional structural feature representation;
[0046] (3) Input the fusion features into the layer-by-layer voxel growth module to generate preliminary three-dimensional CT image voxel data;
[0047] (4) During the decoding process, a global-local collaborative attention mechanism is automatically applied to enhance features and restore details in key anatomical regions;
[0048] (5) Output the volumetric data of the three-dimensional CT image after structural reconstruction as the three-dimensional CT reconstruction result of the corresponding dual-view DR image.
[0049] Furthermore, the method also includes system deployment operations:
[0050] (1) After the model training is completed, it is deployed in the inference system to realize the prediction of three-dimensional CT images of unknown dual-view DR images;
[0051] (2) During the input prediction stage, the image preprocessing, structural feature extraction, feature fusion and three-dimensional reconstruction modules are automatically executed to complete the integrated processing flow.
[0052] (3) Provides a GPU-based accelerated inference framework with an average inference time of less than 10 seconds, meeting the efficiency requirements of clinical preoperative planning.
[0053] (4) The system supports the visualization and export of slices of key anatomical areas, which facilitates doctors to conduct further diagnostic analysis or preoperative assessment.
[0054] Compared with the closest existing technology, the technical solution provided by this invention has the following beneficial effects:
[0055] This invention significantly improves image reconstruction accuracy: by using a structure prior-guided dual-branch encoder network and a direction difference perception module, complementary information from dual-view DR images is effectively fused to construct a highly consistent and accurate representation of three-dimensional structural features. Compared with traditional single-view or geometrically hypothetical reconstruction methods, it can more accurately reconstruct complex anatomical structures.
[0056] This invention achieves high-fidelity CT image reconstruction at the voxel level: it adopts a layer-by-layer voxel growth reconstruction mechanism and multi-target structural sparsity regularization to ensure that the generated CT images meet clinically usable standards in terms of spatial structure, density continuity and anatomical consistency, and solves the problems of poor image quality, blurred texture and obvious artifacts generated by existing algorithms.
[0057] This invention enhances the ability to express key anatomical regions: by introducing a global-local collaborative attention mechanism, it effectively improves the ability to identify the boundaries and density of key regions such as the spine (pedicles, spinous processes, and intervertebral foramina), which is especially suitable for the high sensitivity requirements of anatomical details in preoperative planning.
[0058] This invention balances structural reconstruction and inference efficiency: it constructs a multi-objective loss function driven by structural consistency, which improves the fidelity of image structure and the accuracy of density prediction, while controlling the average inference time to less than 10 seconds through GPU acceleration and end-to-end architecture design, thus meeting the needs of rapid clinical response.
[0059] This invention reduces the risk of radiation exposure for patients and optimizes the preoperative process: This method achieves three-dimensional CT image synthesis based on preoperative dual-view DR images, which is suitable for patients who are not suitable for repeated conventional CT scans due to trauma, scoliosis, etc., or who do not have the conditions to undergo standing CT scans. It avoids the risks of radiation accumulation and transport, and has important value in orthopedic preoperative planning.
[0060] This invention has good system deployment capabilities: the proposed model can be embedded into the existing imaging system of the hospital to realize automated DR-to-CT image generation, and supports functions such as three-dimensional visualization, slice export, and preoperative navigation, and has broad clinical adaptability and potential for promotion and application. Attached Figure Description
[0061] Figure 1 This invention provides a method for generating three-dimensional CT images from cross-modal dual-view DR images based on structure guidance.
[0062] Figure 2 A schematic diagram of a dual-branch prior encoder provided in an embodiment of the present invention;
[0063] Figure 3 A schematic diagram of a direction difference sensing and fusion module provided in an embodiment of the present invention;
[0064] Figure 4 A schematic diagram of a KAN-based diffusion voxel reconstruction module provided in an embodiment of the present invention;
[0065] Figure 5 This is a schematic diagram of a global-local attention decoder provided in an embodiment of the present invention. Detailed Implementation
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0067] Step S1: Acquire dual-view DR images and paired 3D CT images of the patient from two different projection angles, perform data preprocessing, and construct a training dataset;
[0068] Step S2: Input the dual-view DR images into the structure prior-guided dual-branch encoder network respectively, and extract their respective image feature information;
[0069] Step S3: The image features from the two perspectives are fused through the orientation difference perception module to obtain an enhanced three-dimensional structural representation;
[0070] Step S4: Input the fused features into the layer-by-layer voxel growth and reconstruction module to gradually generate three-dimensional CT image data with consistent spatial structure.
[0071] Step S5: Introduce a global-local collaborative attention mechanism during image decoding to enhance the structural representation of key anatomical regions;
[0072] Step S6: Train and optimize the network based on a multi-objective loss function driven by structural consistency to improve the density accuracy and anatomical consistency of the generated images.
[0073] Step S7: Train the model constructed in steps S2, S3, S4, S5, and S6 to generate three-dimensional CT volume data from dual-view DR images.
[0074] Example 1
[0075] Figure 1 A flowchart of a method for generating three-dimensional CT images from cross-modal dual-view DR images based on structure-guided imaging is provided in this embodiment of the invention. The following refers to... Figure 1 Each step is explained in detail.
[0076] Step S110: Acquire dual-view DR images and paired 3D CT images of the patient from two different projection angles, perform data preprocessing, and construct a training dataset;
[0077] Furthermore, use clinical DR equipment to acquire anteroposterior (AP) and lateral (left-right) images of the patient. Ensure that the time interval between the acquisition of the two images is short and that the patient's posture is consistent. Pay special attention to maintaining image alignment during acquisition to ensure that both views correspond to the same patient.
[0078] Furthermore, a high-resolution CT scanner is used to acquire three-dimensional CT volumetric image data of the patient, ensuring that the CT images can cover the anatomical regions corresponding to the DR images and are spatially aligned with the DR images.
[0079] Furthermore, the collected dual-view DR and CT volumetric data underwent data preprocessing, specifically as follows:
[0080] (1) Denoise the DR image, such as by using a Gaussian filter to reduce noise;
[0081] (2) Normalize the DR and CT images to [-1, 1]. The size of the DR images is uniformly 1024×1024 and the size of the CT images is uniformly 128×512×512 to ensure that they are suitable for input into the deep learning model.
[0082] (3) Crop the CT image to remove irrelevant areas, ensuring that only the region of interest is retained, and perform spatial alignment and coordinate system unification.
[0083] Step S120: Input the dual-view DR images into the structure prior-guided dual-branch encoder network respectively, and extract their respective image feature information.
[0084] For details, see Figure 2 The diagram illustrates a dual-branch prior encoder provided in an embodiment of the present invention. DR images from both perspectives are input into a dual-branch convolutional encoder network. Each branch utilizes a shared structural prior module and anatomically guided parameters to enhance the consistency of cross-view image features. Within each encoding branch, a convolutional neural network (CNN) is used to extract multi-scale image features, such as edges, textures, and structural information.
[0085] Furthermore, structural attention units are used to enhance the expressive power of specific key regions in the image, ensuring that the network can focus on anatomically important areas.
[0086] Furthermore, semantic consistency is maintained between the two encoding paths by using a KL divergence minimization strategy for feature alignment, providing prior support for structural alignment for subsequent feature fusion.
[0087] Step S130: The image features from the two perspectives are fused through the orientation difference perception module to obtain an enhanced three-dimensional structural representation;
[0088] For details, see Figure 3 The diagram shown is a schematic of a directional difference sensing fusion module provided in an embodiment of the present invention. The module includes:
[0089] (1) Angle encoder: By learning the viewpoint direction embedding vector, the two viewpoint images are encoded with angle guidance;
[0090] (2) Directional difference calculation: Extract the high-dimensional feature maps of the two branches, perform difference modeling and directional fusion;
[0091] (3) Feature fusion module: The channel attention mechanism and skip connection strategy are adopted to realize context enhancement and output the fused spatial structure representation tensor (128×128×128×64).
[0092] Step S140: Input the fused features into the layer-by-layer voxel growth and reconstruction module to gradually generate three-dimensional CT image data with consistent spatial structure.
[0093] For details, see Figure 4 The diagram shown is a schematic of a voxel growth and reconstruction module provided in an embodiment of the present invention. This module consists of the following components:
[0094] (1) Model initialization: UNet based on KAN (Kolmogorov-Arnold Networks) is used as the backbone of the diffusion model to enhance nonlinear modeling capabilities;
[0095] (2) Forward process: Gaussian noise is added to the structural features in T diffusion steps to generate a blurred voxel map;
[0096] (3) Reverse diffusion process: KAN-UNet is used for stepwise denoising and reconstruction, combined with dual-view structural features as conditions to guide the reconstruction from random noise to high-quality CT volume;
[0097] (4) Training objective: Minimize the mean square error between the predicted noise and the actual noise at each step, while applying the structural consistency loss L. structThe final output is a 3D CT image with a size of 256×256×128.
[0098] Step S150: Introduce a global-local collaborative attention mechanism during image decoding to enhance the structural representation of key anatomical regions; for details, see... Figure 5 The diagram shown is a schematic of a global-local attention decoder provided in an embodiment of the present invention. The module structure is as follows:
[0099] (1) Global attention path: Use a 3-layer Transformer model to model the overall spatial information;
[0100] (2) Local attention path: The features of the vertebral structure (such as pedicle, spinous process and transverse process) are refined by a 3-layer attention enhancement ResNet module;
[0101] (3) Feature fusion and decoding: The global and local attention maps are fused and then input into the 4-level upsampling path, and finally the CT image volume of 256×256×128 is output.
[0102] Step S160: Train and optimize the network based on a multi-objective loss function driven by structural consistency to improve the density accuracy and anatomical consistency of the generated images. The loss function is constructed as follows: L total =λ1L density +λ2L struct +λ3L anatomy +λ4L spatial Among them: L density : Weighted mean square error of gray-level density between three-dimensional voxels; L struct Structural similarity loss (SSIM) measures the structural consistency between the output and the real CT image; L anatomy : Displacement and shape constraint loss based on prior knowledge of key region anatomy; L spatial : 3D spatial consistency regularization term. Weights λ are dynamically adjusted during training. i In the early stage, structural constraints are strengthened, and in the later stage, density fitting is strengthened.
[0103] Step S170: Train the model constructed in steps S120, S130, S140, S150, and S160 to generate three-dimensional CT volume data from dual-view DR images.
[0104] Specifically, the trained model is used to generate a corresponding three-dimensional CT volume image from the input dual-view DR image;
[0105] (1) Input the dual-view DR image to the encoder branch and extract features;
[0106] (2) The directional difference module performs fusion encoding to generate a spatial structure feature tensor;
[0107] (3) The voxel reconstruction module generates a three-dimensional low-resolution initial CT image;
[0108] (4) The decoder module performs decoding and structural detail restoration, and finally outputs a three-dimensional CT image volume (256×256×128);
[0109] (5) The average time for inference on the NVIDIA RTX 3080 platform is 8.9 seconds.
[0110] Through the above implementation method, we can generate three-dimensional CT images from the original dual-view DR images.
[0111] Any aspects of this invention not described in detail are well-known to those skilled in the art.
[0112] The technical means disclosed in this invention are not limited to those disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications are also considered within the scope of protection of this invention.
Claims
1. A method for generating three-dimensional CT images from structure-guided cross-modal dual-view DR images, characterized in that, Includes the following steps: Step S1: Acquire dual-view DR images and paired 3D CT images of the patient from two different projection angles, perform data preprocessing, and construct a training dataset; Step S2: Input the dual-view DR images into the structure prior-guided dual-branch encoder network respectively, and extract their respective image feature information; Step S3: The image features from the two perspectives are fused through the orientation difference perception module to obtain an enhanced three-dimensional structural representation; Step S4: Input the fused features into the layer-by-layer voxel growth and reconstruction module to gradually generate three-dimensional CT image data with consistent spatial structure. Step S5: Introduce a global-local collaborative attention mechanism during image decoding to enhance the structural representation of key anatomical regions; Step S6: Train and optimize the network based on the structural consistency-driven multi-objective loss function to improve the density accuracy and anatomical consistency of the generated images; Step S7: Train the model constructed in steps S2, S3, S4, S5, and S6 to generate three-dimensional CT volume data from dual-view DR images.
2. The method for generating three-dimensional CT images based on structure-guided cross-modal dual-view DR images according to claim 1, characterized in that, In step S1, the specific steps are as follows: Step 1: Acquire DR images of the patient from two different projection angles (e.g., anteroposterior and lateral views) using clinical DR equipment, ensuring that the two sets of images correspond to the same patient and that the acquisition time interval is short, thus ensuring accurate image pairing. Step 2: Use a high-resolution CT scanner to acquire three-dimensional CT volumetric image data of the patient, covering the anatomical area corresponding to the DR image, and ensure that the CT data and DR image are spatially aligned. Step 3: Denoise and normalize the dual-view DR images, unify the image size and resolution, crop the 3D CT data to remove non-interest regions, and unify the spatial resolution and coordinate system to obtain matching dual-view DR images and 3D CT data as training datasets.
3. The method for generating three-dimensional CT images based on structure-guided cross-modal dual-view DR images according to claim 1, characterized in that, Step S2 specifically includes: Step 1: Input the DR images of the two projection angles into a dual-branch convolutional encoder with a shared structure prior module to enhance cross-view feature consistency by sharing anatomical guidance parameters. Step 2: Extract multi-scale image features, including edge, texture and structural information, from each coding branch, and enhance the expressive power of key regions through structural attention units; Step 3: Employ a parameter alignment mechanism to maintain semantic consistency between the two encoding paths, providing prior support for structural alignment for subsequent feature fusion.
4. The method for generating three-dimensional CT images based on structure-guided cross-modal dual-view DR images according to claim 1, characterized in that, In step S3, the specific steps are as follows: Step 1: Construct an orientation difference perception module, which includes an angle attention submodule and a structural contrast loss function, used to explicitly model the projection angle difference between two viewpoints; Step 2: Introduce directional encoding tensors during the fusion process to guide the weighted aggregation of multi-scale features and improve the spatial consistency and stability of structural representation; Step 3: Fuse detailed information from the early layers of the encoder through a skip connection mechanism to enrich the low-level texture features required for 3D structure reconstruction.
5. The method for generating three-dimensional CT images based on structure-guided cross-modal dual-view DR images according to claim 1, characterized in that, In step S4, the specific steps are as follows: Step 1: Design a layer-by-layer voxel growth and reconstruction module, which consists of a voxel candidate generator and a spatial consistency filter; Step 2: Generate 3D CT image voxels layer by layer in a bottom-up manner, and use probability maps to guide the voxel filling order to progressively reconstruct the spatial structure from low resolution to high resolution. Step 3: Introduce a structural sparsity regularization term to constrain the continuity and density consistency between the reconstructed voxels in each layer, preventing artifacts or structural drift in the reconstructed image.
6. The method for generating three-dimensional CT images based on structure-guided cross-modal dual-view DR images according to claim 1, characterized in that, In step S5, the specific steps are as follows: Step 1: Construct a global-local collaborative attention mechanism, in which the global branch generates a structural context representation based on the global response of the image, and the local branch focuses on key anatomical regions such as the spine, ribs and other detailed parts; Step 2: Use fused attention maps to guide the image decoding process, enabling the model to more accurately recover the boundaries and density changes of complex tissues during the generation process; Step 3: By applying a dual-scale supervision mechanism, structural similarity loss is applied to both the intermediate layer and the output layer to improve the detail fidelity of the decoded output.
7. The method for generating three-dimensional CT images based on structure-guided cross-modal dual-view DR images according to claim 1, characterized in that, In step S6, the specific steps are as follows: Step 1: Construct a multi-objective joint loss function, which includes three sub-items: density regression loss, structural similarity loss, and anatomical consistency loss; Step 2: The density regression loss uses a weighted mean square error function to measure the difference in gray density between the predicted CT image and the real CT image. Step 3: Structural similarity loss is used to preserve the overall structural information of the image and improve the restoration of image texture and boundaries; Step 4: Anatomical consistency loss combines prior segmentation masks or keypoint labels to apply positional and shape consistency constraints to important organ regions, ensuring the anatomical accuracy of the generated images. Step 5: Set dynamic weighting coefficients for each of the above sub-loss functions, and automatically adjust them according to the training phase to optimize the overall learning process.
8. The method for generating three-dimensional CT images based on structure-guided cross-modal dual-view DR images according to claim 1, characterized in that, In step S7, the specific steps are as follows: Step 1: Input the dual-view DR images of the patient to be reconstructed into the trained and optimized structure-guided network model, and extract the image feature information of the two views through the dual-branch encoder module respectively; Step 2: Input the extracted dual-view image features into the orientation difference perception module, fuse the multi-view spatial structure information, and generate a unified three-dimensional structural feature representation; Step 3: Input the fusion features into the layer-by-layer voxel growth module to generate preliminary 3D CT image voxel data; Step 4: During the decoding process, a global-local collaborative attention mechanism is automatically applied to enhance features and restore details in key anatomical regions; Step 5: Output the volumetric data of the 3D CT image after structural reconstruction, as the 3D CT reconstruction result of the corresponding dual-view DR image.
9. A method for generating three-dimensional CT images based on structure-guided cross-modal dual-view DR images according to any one of claims 1 to 7, characterized in that, The method further includes the following steps: Step 1: After the model training is completed, deploy it in the inference system to achieve 3D CT image prediction of unknown dual-view DR images; Step 2, during the input prediction stage, automatically executes various modules such as image preprocessing, structural feature extraction, feature fusion, and 3D reconstruction to complete the integrated processing flow; Step 3: Provide a GPU-based accelerated inference framework with an average inference time of less than 10 seconds to meet the efficiency requirements of clinical preoperative planning. Step 4: The system supports visual annotation and slice export of key anatomical areas, which facilitates doctors to conduct further diagnostic analysis or preoperative assessment.