System and method for synthesizing 3D volumetric medical images using a probabilistic-based machine learning model
A generative probabilistic model reconstructs three-dimensional CT images from two-dimensional radiographs, addressing the limitations of conventional radiographs by providing detailed volumetric data without additional radiation or specialized equipment.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-04-02
AI Technical Summary
Conventional radiographs lack depth information and volumetric detail, necessitating additional radiation exposure or specialized equipment for three-dimensional imaging, which is not always feasible.
A method using a generative probabilistic model, such as a diffusion model, to reconstruct three-dimensional CT images from two-dimensional radiographic images, incorporating pose estimation and spatial alignment to generate high-resolution volumetric data without additional radiation.
Enables three-dimensional imaging from two-dimensional radiographs, reducing the need for conventional CT imaging and maintaining diagnostic quality for surgical planning and fracture evaluation.
Smart Images

Figure US2025048292_02042026_PF_FP_ABST
Abstract
Description
MGH 2024-020-02 125141.04890 SYSTEM AND METHOD FOR SYNTHESIZING 3D VOLUMETRIC MEDICAL IMAGES USING A PROBABILISTIC-BASED MACHINE LEARNING MODEL CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 700,192, filed on September 27, 2024, and entitled “System and Method for Synthesizing 3D Volumetric Medical Images Using a Diffusion-Based Machine Learning Model,” which is herein incorporated by reference in its entirety. BACKGROUND
[0002] Medical imaging plays an advantageous role in modern healthcare, with computed tomography (CT) and radiography serving as complementary diagnostic modalities. While conventional radiographs provide rapid, cost-effective two-dimensional projections of anatomical structures, they inherently lack the depth information and volumetric detail that three-dimensional CT imaging offers. CT scans enable detailed visualization of complex anatomical relationships, cross-sectional views, and precise spatial measurements that are valuable for surgical planning, fracture assessment, and treatment decision-making.
[0003] However, CT imaging presents several practical limitations including higher radiation exposure, increased cost, longer acquisition times, and limited availability in certain clinical settings. These constraints often result in delayed diagnosis, additional imaging appointments, or reliance on incomplete anatomical information from radiographs alone. The ability to derive three-dimensional volumetric information from standard two-dimensional radiographic images would address these limitations by providing enhanced diagnostic capabilities without requiring additional radiation exposure or specialized equipment. Recent advances in artificial intelligence and generative modeling techniques have created new opportunities to bridge the gap between two-dimensional and three-dimensional medical imaging modalities. SUMMARY OF THE DISCLOSURE
[0004] According to an aspect of the present disclosure, a method for generating three- dimensional computed tomography image data from two-dimensional radiographic images is provided. The method includes accessing two-dimensional radiographic image data with a computer system. The method includes processing the two-dimensional radiographic image data to estimate pose data indicating relative orientations between radiographic views. The method includes spatially aligning the two-dimensional radiographic image data based on the pose data to generate aligned radiographic images. The method includes generating coarse 1 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 three-dimensional volumetric data via backprojection of the aligned radiographic images into three-dimensional space using the pose data. The method includes applying the aligned radiographic images and the coarse three-dimensional volumetric data to a trained generative probabilistic model to generate three-dimensional computed tomography image data.
[0005] According to another aspect of the present disclosure, a method for training a generative probabilistic model to synthesize three-dimensional computed tomography volumes from two-dimensional radiographic projections is provided. The method includes accessing training data including three-dimensional computed tomography volumes and corresponding digitally reconstructed radiographs with a computer system. The method includes generating backprojected volumetric guidance data by projecting the digitally reconstructed radiographs into three-dimensional space. The method includes training a generative probabilistic model on the training data and backprojected volumetric guidance data. The method includes storing the trained generative probabilistic model for generating three-dimensional computed tomography image data from two-dimensional radiographic images. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] FIG. 1 is a flowchart of an example method for generating 3D CT image data from 2D image data (e.g., radiographs) using a pose-conditioned generative probabilistic model.
[0007] FIG. 2 is a flowchart of an example method for training a pose-conditioned generative probabilistic model.
[0008] FIGS. 3A–3C illustrate a workflow for an example process for 3D U-Net training with DRRs. FIG. 3A illustrates an example patch extraction process. FIG. 3B illustrates an example training framework. FIG.3C illustrates an implementation of an example patch-wise loss function.
[0009] FIG.4 illustrates an example reverse process for 3D CT volume synthesization.
[0010] FIG.5 illustrates an example process of using a 3D diffusion model for 3D CT image data synthesis. (A) The forward processing of the diffusion model involves patch-wise training of the Denoiser 3D U-Net, which takes conditional input in the form of patches consisting of the training CT volume, anterior-posterior (AP) view, and lateral (LAT) view x- ray images, along with patch coordinates. (B) To synthesize the entire 3D CT volume using the patch-wisely trained denoiser 3D U-Net, the 3D Gaussian noise and AP / LAT x-ray images are iteratively denoised, following the probability flow ordinary differential equation. 2 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0011] FIG. 6 shows representative wrist CT images; The synthesized wrist CT (Bottom) accurately represents bone anatomy and matches the original CT (Top), as seen in both 3D volume rendering and slice images.
[0012] FIG. 7 shows an example clinical workflow comparing the standard pathway (top) with the proposed Virtual CT–augmented pathway (bottom) for 3D wrist imaging.
[0013] FIG.8 shows an example inference pipeline for using a generative probabilistic model to generate 3D CT image data.
[0014] FIG.9 shows an example training pipeline for training a generative probabilistic model to generate 3D CT image data.
[0015] FIG. 10 shows an example pose estimation module trained on synthetic DRR pairs to predict relative rotations (Rest) with a supervised loss against ground-truth (Rgt), improving alignment robustness for real radiographs.
[0016] FIG. 11 shows quantitative evaluation from a majority level consensus from three hand surgeons (≤2 of 3 readers agreeing).
[0017] FIG. 12 shows an example of in-plane alignment of wrist radiographs. Left panels show the radiography (top) and CT (bottom) acquisition poses, both performed without external positioning jigs, reflecting the inherent variability and non-uniformity of real-world wrist radiography and CT datasets. Right panels display posteroanterior (PA), oblique (OBL), and lateral (LAT) views before (top row) and after (bottom row) manual in-plane alignment. Alignment was performed using a custom graphical user interface by rotating each image to align the ulnar shaft (colored vertical lines) with the proximal–distal axis and translating the image to center the base of the lunate bone (green circles), chosen as a stable anatomical reference consistently visible across all three views. Following alignment, all radiographs were adjusted to have the same vertical position, with the lunate base aligned to a common height from the top of the image, as indicated by the red horizontal line in the bottom row.
[0018] FIG. 13 shows representative examples of fracture-negative wrist cases comparing Original CT (ground truth) and Virtual CT (synthesized from radiographs) across axial, coronal, and sagittal planes, demonstrating preservation of cortical smoothness, trabecular consistency, and joint congruity in the synthesized volumes.
[0019] FIG. 14 illustrates fracture-positive wrist cases with multiple displaced fragments, showing how Virtual CT reconstructions capture diagnostic features including cortical breaches, fragment contours, and intra-articular involvement while maintaining sufficient precision for expert interpretation despite minor smoothing at fragment boundaries. 3 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0020] FIG. 15 presents additional fracture-positive examples demonstrating Virtual CT's ability to reproduce comminution patterns and joint deformity with diagnostic quality, preserving key morphologic cues required for fracture classification including radial inclination and articular step-off.
[0021] FIG.16 shows complex intra-articular fracture patterns comparing Original and Virtual CT reconstructions, illustrating how the synthesized volumes maintain major topological relationships including fragment displacement and spatial orientation without artificial bone formation or erasure of critical disruptions.
[0022] FIG. 17 shows a block diagram of an example 3D CT image synthesis system for generating 3D CT image data.
[0023] FIG.18 shows a block diagram of example components that can implement the system of FIG.17.
[0024] FIGS.19A and 19B illustrate an example CT system. DETAILED DESCRIPTION
[0025] Described here are systems and methods for generating three-dimensional computed tomography (CT) images from two-dimensional radiographic projections using generative probabilistic models, such as diffusion models or flow matching models. The systems may utilize a pose-conditioned, score-based generative probabilistic model framework that reconstructs high-resolution volumetric CT data from standard radiographic views without requiring additional radiation exposure or specialized imaging equipment. In some aspects, the methods may incorporate a convolutional neural network-based pose estimation module that infers radiographic view orientations, enabling spatial alignment and geometric conditioning of the diffusion process. The systems may employ patch-wise training techniques to synthesize high-resolution three-dimensional volumes within computational memory constraints while preserving fine anatomical details such as cortical boundaries and trabecular patterns.
[0026] In some cases, the methods may enable point-of-care volumetric assessment by transforming routine two-dimensional radiographs into diagnostic-quality three-dimensional reconstructions, potentially reducing the need for conventional CT imaging in certain clinical scenarios while maintaining clinically relevant structural information for surgical planning and fracture evaluation.
[0027] The generation of three-dimensional volumetric images from two-dimensional radiographic projections presents a challenging computational problem due to the inherent loss of depth information during the projection process. Traditional reconstruction methods 4 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 typically require multiple X-ray projections from different angles to reconstruct three- dimensional volumes, necessitating specialized scanning equipment and controlled imaging protocols. Recent advances in machine learning and artificial intelligence have opened new possibilities for addressing this reconstruction challenge through data-driven approaches that can learn complex mappings between two-dimensional and three-dimensional medical image domains.
[0028] Diffusion models and other generative probabilistic models (e.g., flow matching models) represent a class of generative artificial intelligence techniques that have demonstrated capabilities in synthesizing high-quality images across various domains. Diffusion models operate through a two-stage process involving forward and reverse diffusion processes. During the forward process, noise may be progressively added to training data, while the reverse process involves learning to remove this noise iteratively to generate new samples. The application of diffusion models to medical imaging reconstruction tasks offers potential advantages in terms of image quality and the ability to capture complex anatomical variations present in training datasets.
[0029] The present disclosure describes systems and methods for generating three- dimensional CT volumes from two-dimensional radiographic images using diffusion-based approaches. The disclosed techniques may address computational and memory constraints associated with three-dimensional volume generation while maintaining high spatial resolution in the reconstructed images. The approach may incorporate patch-wise training strategies and conditioning mechanisms that enable the diffusion model to learn mappings between limited radiographic views and corresponding three-dimensional anatomical structures. The systems may be configured to process various types of radiographic inputs and generate volumetric outputs suitable for clinical evaluation and treatment planning applications.
[0030] Referring now to FIG. 1, a flowchart is illustrated as setting forth the steps of an example method for generating 3D CT image data using a suitably trained generative probabilistic model. As will be described, the generative probabilistic model takes 2D image data as an input and generates 3D CT image data as an output.
[0031] The method includes accessing 2D image data with a computer system, as indicated at step 102. Accessing the 2D image data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the 2D image data may include acquiring such data with a medical imaging system (e.g., an x-ray system) and transferring or otherwise communicating the data to the computer 5 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 system, which may be a part of the medical imaging system. By way of example, the 2D image data may include radiographs.
[0032] Various types of input radiographic views may be processed to accommodate different clinical scenarios and imaging protocols. In some cases, single x-ray images may be processed, which may be particularly useful in emergency or resource-limited settings where only one radiographic view is available. Additionally or alternatively, bi-directional x-ray images may be processed, including anterior-posterior (AP) and lateral (LAT) views, which provide complementary geometric information for three-dimensional reconstruction. In still other examples, tri-planar radiographs including posterior-anterior (PA), oblique (OBL), and lateral (LAT) views may be processed, offering enhanced spatial coverage and improved reconstruction accuracy through multiple viewing angles.
[0033] In some cases, the 2D image data may be preprocessed to enhance the quality and consistency of radiographic data before three-dimensional reconstruction. As one example, body-part segmentation may be performed to identify and remove external objects that could interfere with the reconstruction process. In some cases, casts, splints, jewelry, or other non- anatomical objects are automatically detected and excluded from the input data. The segmentation process may utilize software tools such as ITK-SNAP, which provides manual and semi-automatic segmentation functionality for medical images. Alternatively, a Segment Anything Model (SAM) may be used, which offers automated segmentation capabilities using machine learning techniques to identify and isolate anatomical structures from background elements.
[0034] The preprocessing pipeline may further include image normalization and standardization procedures to ensure consistent input formatting across different imaging systems and protocols. In some cases, the system applies intensity normalization to account for variations in x-ray exposure parameters and detector characteristics. The preprocessing may also involve spatial alignment operations to establish consistent anatomical orientation across multiple radiographic views. Background removal techniques may be applied to eliminate non- relevant image regions and focus computational resources on anatomically relevant areas. These preprocessing steps may collectively improve the robustness and accuracy of the subsequent three-dimensional reconstruction process by providing clean, standardized input data to the generative probabilistic model.
[0035] In some cases, the method may dynamically adapt the reconstruction approach based on the number and type of available radiographic views. When processing single-view 6 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 inputs, the method may rely more heavily on learned anatomical priors to compensate for limited geometric information. For bi-directional configurations, the method may leverage the complementary information from orthogonal or near-orthogonal views to resolve depth ambiguities inherent in two-dimensional projections. Tri-planar configurations may provide the most comprehensive geometric coverage, enabling the method to generate three- dimensional reconstructions with enhanced accuracy and detail preservation across all anatomical planes. In still other examples, additional radiographic views beyond three (e.g., multi-oblique or angled acquisitions) may be incorporated to further improve geometric fidelity.
[0036] A trained generative probabilistic model (e.g., a diffusion model, a flow model) is also accessed with the computer system, as indicated at step 104. In general, the generative probabilistic model is trained, or has been trained, on training data in order to generate 3D CT image data from 2D image data.
[0037] Accessing the trained model may include accessing model parameters (e.g., weights, biases, or both) that have been optimized or otherwise estimated by training the model on training data. In some instances, retrieving the model can also include retrieving, constructing, or otherwise accessing the particular model architecture to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.
[0038] In general, the generative probabilistic model can implement any number of different model architectures. As one example, the generative probabilistic model can implement a diffusion model. As another example, the generative probabilistic model can implement a flow model. Other generative probabilistic models may also be used.
[0039] By way of example, the generative probabilistic model may be a diffusion model that employs a score-based generative modeling approach, which operates through forward and reverse diffusion processes. During the forward process, Gaussian noise may be progressively added to clean 3D CT volumes according to a predefined noise schedule, while the reverse process learns to iteratively denoise random noise to generate synthetic volumetric data. The model architecture may center around various neural network designs that serve as the denoising network. In some cases, a modified 3D U-Net architecture may be employed, incorporating encoder-decoder pathways with skip connections to preserve spatial information across different resolution levels. Alternatively, the architecture may utilize a 3D ResNet 7 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 backbone with dense connections, a vision transformer (ViT) adapted for volumetric data processing, a hybrid convolutional-transformer architecture that combines local feature extraction with global attention mechanisms, or the like. The network may include multiple processing levels, with each level handling volumetric features at progressively different spatial resolutions while maintaining semantic information through various connection strategies, such as skip connections, dense connections, or attention-based feature aggregation.
[0040] The network architecture incorporates feature-wise affine modulation layers that enable conditioning on both noise levels and input radiographic data. These modulation layers may be distributed throughout the neural network at multiple resolution scales, allowing the network to adapt its processing based on the current noise level during the diffusion process and the specific radiographic conditioning information. In U-Net implementations, the modulation may occur at corresponding encoder and decoder stages, while in ResNet architectures, the modulation may be applied at residual block boundaries. For transformer- based architectures, the conditioning may be integrated through cross-attention mechanisms or through modified self-attention layers that incorporate conditioning tokens. The conditioning mechanism may involve concatenating back-projected radiographic volumes with positional encoding information as additional input channels to the denoising network. The noise level conditioning may be implemented through learned embeddings that are processed through fully connected layers and subsequently used to modulate feature maps via affine transformations, attention weights, or adaptive normalization parameters at various network depths.
[0041] Patch-wise training methodology addresses the computational constraints associated with processing full-resolution 3D volumes during model training. The training process may utilize overlapping patches extracted from the full volumetric data, with patch dimensions that can vary based on available GPU memory and computational requirements. In some cases, patches of size 64×64×64 may be employed for memory-constrained training scenarios, while larger patches of size 256×64×160 may be used when greater memory capacity allows for processing extended spatial regions. The overlapping nature of patch sampling ensures that boundary regions receive adequate training coverage and helps maintain spatial consistency across patch boundaries during inference. Alternatively, non-patch-based training methods may be employed when sufficient computational resources are available to process full-resolution volumes directly, which may provide enhanced global spatial context at the cost of increased memory requirements. 8 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0042] The patch-wise loss function may be formulated as a sum of individual patch losses, leveraging the shift-invariant properties of convolutional operations. Each patch contributes to the overall denoising score matching objective, with the network learning to predict the noise component added to each volumetric patch. The overlapping patch strategy may provide multiple training contexts for shared voxel locations, encouraging consistent predictions across different spatial neighborhoods and improving the model's ability to generate coherent full-volume reconstructions during inference. Positional encoding may be incorporated to help the network distinguish between patches sampled from different spatial locations within the volume, providing additional context that aids in maintaining anatomical consistency. In other examples, non-patch-based training approaches may operate on complete volumes using standard loss functions computed across the entire three-dimensional space, potentially offering improved global coherence.
[0043] During inference, the trained diffusion model may employ various sampling methods to generate 3D CT volumes from the learned data distribution. The sampling process may utilize ordinary differential equation (ODE) solvers that implement the reverse-time diffusion process through numerical integration schemes. Different solver implementations may be employed depending on the desired balance between generation quality and computational efficiency. The Heun solver may provide second-order accuracy for the numerical integration, potentially offering improved sample quality compared to first-order methods. Alternatively, the Euler method may be used as a first-order solver that provides faster sampling at the potential cost of some generation fidelity.
[0044] The number of function evaluations (NFE) during sampling may be adjusted to control the trade-off between generation quality and inference time. Lower NFE values such as 10 evaluations may enable rapid sampling for time-sensitive applications, while higher NFE values such as 50, 100, or 1000 evaluations may provide progressively improved sample quality through more refined denoising steps. The sampling process begins with pure Gaussian noise matching the dimensions of the target CT volume and iteratively applies the learned denoising function to gradually transform the noise into a coherent anatomical structure. The conditioning information from input radiographs may be maintained throughout the sampling process, guiding the generation toward anatomically consistent reconstructions that align with the provided 2D projections.
[0045] The 2D image data are processed to estimate the pose data, as indicated at step 106. By way of example, the pose estimation may include implementing a convolutional neural 9 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 network configured to predict relative rotation angles between paired radiographic views. The pose estimation may determine the radiographic view orientation of input x-ray images with respect to a canonical view, such as a canonical posteroanterior view. In some cases, multi- channel radiographic input data are processed to extract geometric cues implicit in the radiographic projection and learn rotational features from training data synthesized via digitally reconstructed radiographs.
[0046] The convolutional neural network architecture may include a feature extractor backbone followed by a fully connected regression head. In some cases, the feature extractor backbone comprises a ResNet-18 network pretrained on ImageNet and adapted for multi- channel radiographic input processing. The ResNet-18 backbone may process 256×256 digitally reconstructed radiographs through grouped convolutions to capture geometric features relevant to rotation estimation. Alternative convolutional neural network backbones may be employed, including other ResNet variants, DenseNet architectures, or custom convolutional architectures designed for medical image analysis. The fully connected regression head may regress rotation parameters in a tangent space to provide geometrically consistent angular predictions.
[0047] The pose estimation module may be trained using a geodesic loss function that operates to minimize discrepancies between predicted and ground-truth rotations on the manifold of 3D rotations. This geodesic loss formulation may penalize angular deviations in a manner that respects the underlying geometry of the rotation group, providing more stable training compared to Euclidean distance metrics applied directly to rotation matrices.
[0048] Training of the pose estimation module may utilize synthetic training data generated by sampling azimuth angles relative to posteroanterior views during each trainingepoch. The azimuth angles may be sampled as ^ i ^ ^ ^ ^ ^ ^ i , where Δα represents discreteangular offsets corresponding to differenti represents random perturbationterms. In some cases, Δα may be set to values including 0, ^ , or ^ for generating 4 2posteroanterior, oblique, or lateral views respectively, withterms ^^ isampled from uniform distributions to introduce variability in the training data. The pose estimation module may be trained using Adam optimization with learning rates on the order of 1×10−4and cosine learning rate decay schedules over multiple training epochs. 10 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0049] The pose estimation process may extract clinically interpretable projection angles from the estimated rotation matrices for use in downstream processing. The inter-plane projection angle may be computed as, ^^proj ^arctan 2 ^ r32 , r 33 ^
[0050] where rijdenotes elements of the estimated rotation matrix. This projected angle may correspond to observable rotations about a proximal-distal axis on standard radiographs and may serve as conditioning parameters for subsequent diffusion model processing. Separate pose estimation models may be trained for different view pair combinations, including posteroanterior-oblique and posteroanterior-lateral view estimation, with each model specialized for the geometric relationships inherent to the specific view pairing.
[0051] Following pose estimation, the 2D image data are spatially aligned based on the pose data, as indicated at step 108. By way of example, spatial alignment includes applying geometric transformations to ensure anatomical consistency across different input views.
[0052] Spatial alignment of radiographic images may be performed to standardize positioning and orientation across different patient anatomies and imaging protocols. In some cases, the system incorporates manual alignment capabilities through a graphical user interface (GUI) that allows operators to adjust radiographic positioning based on anatomical landmarks. The GUI-based alignment process may involve identification of specific anatomical structures such as the ulnar shaft and lunate bone to establish consistent coordinate systems across multiple radiographic views. Manual alignment may provide precise control over positioning parameters, allowing for correction of patient positioning variations that occur during clinical imaging procedures.
[0053] Automated alignment approaches may also be implemented to reduce operator dependency and improve workflow efficiency. In some cases, automated alignment algorithms analyze anatomical landmarks within the radiographic images to determine optimal positioning parameters. The automated alignment process may utilize image processing techniques to identify key anatomical structures and calculate transformation matrices for standardizing image orientation. Automated alignment may incorporate machine learning algorithms trained on anatomically aligned datasets to recognize consistent positioning patterns across different patient populations.
[0054] Domain transfer capabilities may be incorporated to address differences between synthetic training data and real-world clinical radiographs. In some cases, a CycleGAN model may perform domain adaptation between digitally reconstructed radiographs 11 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 (DRRs) and actual x-ray images. The CycleGAN architecture may include two generator networks and two discriminator networks that learn bidirectional mappings between the synthetic and real image domains. Domain transfer may help bridge the gap between the controlled synthetic training environment and the variable characteristics of clinical radiographic imaging systems.
[0055] The CycleGAN model may be trained using unpaired datasets of real x-rays and DRRs, allowing the system to learn domain-specific characteristics without requiring exact correspondence between synthetic and real images. In some cases, the domain transfer process preserves anatomical structure while adapting image appearance characteristics such as noise patterns, contrast variations, and imaging artifacts typical of clinical x-ray systems. The generator networks within the CycleGAN may learn to transform DRRs to appear more similar to real x-rays while maintaining the underlying anatomical information necessary for accurate 3D reconstruction. Cycle consistency losses may ensure that the domain transfer process remains reversible and preserves anatomical fidelity throughout the transformation process.
[0056] The aligned radiographic images in the 2D image data are subsequently processed to generate coarse 3D volumetric data, as indicated at step 110. By way of example, the initial three-dimensional volumetric estimates may be generated based on the determined pose parameters and spatial relationships. Positional encoding schemes that provide depth information and coordinate references for the subsequent generative probabilistic model processing stages may be incorporated into a backprojection framework to estimate the coarse 3D volumetric data. The initial volumetric estimates serve as conditioning inputs for the generative probabilistic model, which refines the volumetric data through iterative operations. As described in more detail below, the generative probabilistic model applies learned anatomical priors and geometric constraints to enhance the structural accuracy and detail resolution of the generated volumes.
[0057] The backprojection process transforms two-dimensional radiographic images into three-dimensional volumetric guidance data that conditions the generative probabilistic model during CT image volume synthesis. In some cases, the backprojection process receives the aligned radiographic views and generates initial volumetric estimates by distributing pixel intensities along projection rays through the three-dimensional space. The backprojection operation may utilize geometric parameters derived from the pose data to determine the correct spatial orientation and positioning of each radiographic view relative to the target CT coordinate system. The resulting volumetric guidance data provide spatial context that helps 12 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 constrain the diffusion process, or other generative probabilistic process, to anatomically plausible reconstructions.
[0058] Positional encoding schemes may be incorporated into the backprojection process to provide additional spatial information that disambiguates depth relationships across different regions of the reconstructed volume. In some cases, one-dimensional z-axis positional maps are generated that encode the superior-inferior coordinate position for each voxel location within the target CT volume. These positional maps may be computed as normalized coordinate values ranging from zero to one along the z-axis direction, providing explicit spatial reference information that helps the diffusion model understand the relative positioning of anatomical structures. The z-axis positional encoding may be particularly useful for distinguishing between overlapping bone structures that appear similar in the two-dimensional radiographic projections but occupy different depths within the three-dimensional anatomy.
[0059] The conditioning mechanisms can integrate multiple sources of guidance information into the generative probabilistic process through feature-wise affine modulation and multi-scale injection pathways. In some cases, the backprojected volumetric guidance data are concatenated with a noisy CT volume input and fed into a neural network architecture of the generative probabilistic model at the encoder input level. The positional encoding maps may be processed through separate convolutional layers and injected at multiple resolution levels throughout the encoder-decoder pathway using affine transformation parameters that modulate the feature representations. This multi-scale conditioning approach allows the generative probabilistic model to incorporate spatial guidance information at both fine-grained and coarse-grained levels of the reconstruction process.
[0060] Coordinate encoding schemes may extend beyond simple z-axis positioning to include full three-dimensional spatial coordinates or relative positioning information between different anatomical landmarks. In some cases, the coordinate encoding incorporates normalized x, y, and z position values that are embedded as additional input channels alongside the back-projected volumetric guidance. These coordinate embeddings may be processed through learned embedding layers that transform the raw positional information into higher- dimensional feature representations suitable for integration with the diffusion model architecture. The coordinate encoding may also include relative distance measurements from key anatomical reference points such as joint centers or bone landmarks to provide additional spatial context for the reconstruction process. 13 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0061] The integration of conditioning information into the generative probabilistic model may occur through multiple pathways that operate at different stages of the sampling process (e.g., diffusion sampling process, flow sampling process). In some cases, the backprojected guidance and positional encoding are combined through element-wise addition or concatenation operations before being processed by the denoising network. The conditioning information may also be injected through cross-attention mechanisms that allow the generative probabilistic model to selectively attend to relevant spatial regions based on the guidance provided by the radiographic projections. These attention-based conditioning mechanisms enable the model to focus computational resources on anatomically important regions while maintaining global spatial coherence throughout the reconstruction process.
[0062] The 2D image data and coarse 3D volumetric data are then input to the generative probabilistic model, generating output as 3D CT image data, as indicated at step 112. During the refinement process, the generative probabilistic model may operate on patch- based segments of the volumetric data to accommodate computational memory limitations while maintaining spatial coherence across the entire volume. The patch-wise processing approach allows the 3D CT image synthesis process to handle high-resolution volumetric outputs while distributing computational load across available hardware resources. Alternatively, non-patch-based approaches may be employed when sufficient computational resources are available to process full-resolution volumes directly, which may provide enhanced global spatial context at the cost of increased memory requirements. The generative probabilistic model may be a diffusion model that applies score-based denoising techniques that progressively remove noise artifacts and enhance anatomical features based on training data patterns. The iterative refinement continues through multiple sampling steps until the volumetric output reaches the desired quality and resolution specifications.
[0063] Advantageously, the model generates 3D CT image data with configurable spatial resolutions and dimensional parameters. In some cases, the generated 3D CT image data may have isotropic spatial resolutions of 0.2 mm³, providing high-resolution anatomical detail for precise clinical assessment. Alternative configurations may produce volumes at 0.5 mm³ resolution, which may balance computational efficiency with diagnostic quality. The model may accommodate other isotropic resolutions based on specific clinical requirements or hardware constraints. The flexibility in resolution selection allows adaptation to different diagnostic scenarios, where higher resolutions may be selected for detailed fracture analysis and lower resolutions may be used for general anatomical assessment. 14 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0064] Volume dimensions may vary according to anatomical coverage requirements and computational resources. In some cases, the model may generate 3D CT image data with dimensions of 256×256×160 voxels, providing comprehensive coverage of wrist anatomy. Alternative dimensional configurations may include 256×224×160 voxels, which may optimize memory usage while maintaining adequate anatomical representation. In still other cases, the dimensions may be tailored to the desired clinical application. As such, the model may support other volumetric sizes to accommodate different anatomical regions or specific imaging protocols. These dimensional variations allow the model to adapt to different clinical scenarios while managing computational load and memory requirements during the 3D CT image data generation process.
[0065] Inference timing characteristics may depend on hardware configuration and sampling parameters selected during the generation process. In some cases, the model may complete 3D CT image data generation in approximately 5 minutes per patient when operating on suitable graphics processing units. The inference time may vary based on the number of sampling steps employed during the diffusion process, flow process, or other probabilistic process modeled by the generative probabilistic model, where higher numbers of function evaluations may increase generation time while potentially improving output quality. Hardware specifications, including memory capacity and processing power, may influence the overall generation speed. The model may be optimized for different hardware configurations, allowing deployment on various computational platforms while maintaining acceptable generation times for clinical workflows.
[0066] The 3D CT image data generation process may incorporate adaptive sampling strategies that balance output quality with computational efficiency. In some cases, the model may employ different numbers of function evaluations during the reverse diffusion process (e.g., for a diffusion model), reverse mapping (e.g., for a flow model), or the like, ranging from 10 to 1000 steps depending on the desired quality-speed trade-off. Lower numbers of function evaluations may reduce generation time while maintaining acceptable anatomical accuracy for many clinical applications. Higher numbers of function evaluations may produce enhanced detail and reduced artifacts at the cost of increased computational time. The system may automatically select appropriate sampling parameters based on the specific clinical application or allow manual configuration by the operator. 15 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0067] The 3D CT image data generated by inputting the 2D image data to the trained generative probabilistic model can then be displayed to a user, stored for later use or further processing, or both, as indicated at step 114.
[0068] Referring now to FIG. 2, a flowchart is illustrated as setting forth the steps of an example method for training a generative probabilistic model on training data, such that the generative probabilistic model is trained to receive 2D image data as an input in order to generate 3D CT image data as an output. In general, the generative probabilistic model can implement any number of different model architectures. As one example, the generative probabilistic model can implement a diffusion model. As another example, the generative probabilistic model can implement a flow model. Other generative probabilistic models may also be used.
[0069] The method includes accessing training data with a computer system, as indicated at step 202. Accessing the training data may include retrieving such data from a memory or other suitable data storage device or medium. Alternatively, accessing the training data may include acquiring such data with one or more medical imaging systems (e.g., x-ray imaging systems, CT imaging systems) and transferring or otherwise communicating the data to the computer system.
[0070] In general, the training data can include three-dimensional CT image volumes and corresponding DRRs derived from clinical imaging datasets. In some cases, the CT volumes may be collected from patients who have undergone both standard plain radiographs and three-dimensional CT scans within a specified time interval, such as a 7-day period, to ensure anatomical consistency across modalities. The CT volumes may be resampled to a consistent spatial resolution, such as 0.5 mm × 0.5 mm × 0.5 mm, and manually aligned along anatomical axes. Data preprocessing may include body-part segmentation to eliminate extraneous objects such as splints, casts, or other non-anatomical materials that could interfere with the training process. In some implementations, left-sided anatomical structures may be mirrored to match the coordinate system of right-sided structures, thereby standardizing anatomical orientation across the dataset.
[0071] The DRRs may be generated by simulating x-ray projections from the aligned CT volumes at anatomically defined rotation angles corresponding to standard radiographic views. The simulation process may involve applying rigid in-plane rotations along the proximal-distal axis of the training CT volumes, followed by parallel x-ray beam projection with detector spacing and field of view parameters configured to mimic routine clinical 16 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 radiographs. The training dataset may be assembled by extracting overlapping three- dimensional patches from the CT volumes and generating corresponding back-projected volumetric guidance data from the simulated radiographs. Additional conditioning information, such as one-dimensional positional maps encoding spatial coordinates, may be incorporated to provide depth information and coordinate references. Data augmentation techniques may be applied during training iterations, including random rotations within specified angular ranges and spatial translations, to artificially expand the effective dataset size and improve model generalization capabilities across different anatomical variations and imaging conditions.
[0072] The generative probabilistic model is then trained on the training data, as indicated at step 204. In general, the generative probabilistic model can be trained by optimizing model (e.g., weights, biases, or both). A workflow of an example process for 3D U-Net training with DRRs is illustrated in FIGS.3A–3C, and a corresponding reverse process for 3D CT volume synthesization is illustrated in FIG.4. FIG.3A illustrates an example patch extraction process. FIG. 3B illustrates an example training framework. FIG. 3C illustrates an implementation of an example patch-wise loss function.
[0073] By way of example, the diffusion model training process may utilize a score- based generative modeling approach that operates through forward and reverse diffusion processes to learn the mapping between two-dimensional radiographic projections and three- dimensional computed tomography volumes. During the forward diffusion process, Gaussian noise may be progressively added to clean three-dimensional CT volumes according to a predefined noise schedule, transforming the original data distribution into a prior noise distribution over a series of time steps. The forward process may be characterized by sampling a noise level σ from a log-normal distribution p(σ) and forming noisy samples. The denoising network may be trained to predict the clean volume through score matching, where the network learns to estimate a score function given conditioning information that includes backprojected radiographic volumes and positional encoding data. The training process is described in more detail below.
[0074] The reverse diffusion process may enable the generation of three-dimensional CT volumes by iteratively denoising random Gaussian noise through learned score functions until coherent anatomical structures emerge. During inference, the trained model may employ various sampling methods including ordinary differential equation solvers that implement the reverse-time diffusion process through numerical integration schemes. The sampling process 17 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 may utilize different numbers of function evaluations, such as 10 to 1000 steps, where higher numbers of sampling steps may provide progressively improved sample quality through more refined denoising iterations at the cost of increased computational time. The model may be trained using patch-wise loss functions that operate on overlapping three-dimensional patches extracted from the full CT volumes, allowing the system to handle high-resolution volumetric data within computational memory constraints while maintaining spatial coherence across patch boundaries through the overlapping sampling strategy and positional encoding mechanisms. Alternatively, non-patch-based training methods may be employed when sufficient computational resources are available to process full-resolution volumes directly, which may provide enhanced global spatial context at the cost of increased memory requirements.
[0075] By way of example, the training methodology may accommodate various dataset configurations to support robust model development across different clinical scenarios. In an example study, the training dataset included 61 patients that served as a foundational cohort for initial model development and proof-of-concept validation. Alternative configurations may utilize larger patient populations, which may offer enhanced diversity in anatomical variations and pathological presentations and enable more comprehensive training across a broader spectrum of anatomies and fracture patterns.
[0076] Patient cohort configurations may vary based on anatomical laterality and demographic distributions. In some cases, the dataset may include both right-hand and left- hand CT scans, with proportions such as 32 right-hand cases and 44 left-hand cases in smaller cohorts. The patient population may span various age groups and fracture severities, ensuring representative sampling of clinical presentations encountered in orthopedic practice. Data preprocessing may include body-part segmentation procedures to eliminate extra-body objects such as splints, casts, or other external materials that may interfere with model training. Left- hand CT scans may undergo spatial transformation to achieve anatomically identical positioning relative to right-hand scans, standardizing the training data orientation.
[0077] Data splitting strategies may employ various train-test ratios to optimize model performance and validation reliability. In some cases, an 80 / 20 split may be implemented, allocating the majority of data for training while reserving a portion for independent testing. For a 76-patient dataset, this configuration may result in 61 training cases and 15 testing cases. Alternative splitting approaches may utilize different ratios based on dataset size and validation requirements. The training subset may undergo further subdivision to create validation sets for 18 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 hyperparameter tuning and model selection during the training process. Cross-validation techniques may be employed to maximize data utilization and assess model generalizability across different patient subsets.
[0078] Training data augmentation strategies may enhance dataset diversity without requiring additional patient acquisitions. Random rotations within specified angular ranges, such as ±10°, may be applied consistently to both CT volumes and corresponding guidance images. Translation augmentations may introduce spatial variations of ±5 voxels to improve model robustness to positioning variations. These augmentation techniques may be applied during training iterations to artificially expand the effective dataset size and improve model generalization capabilities. The augmentation parameters may be adjusted based on anatomical constraints and clinical relevance to maintain realistic data distributions while maximizing training diversity.
[0079] The trained generative probabilistic model is then stored for later use, as indicated at step 206. Storing the generative probabilistic model may include storing model parameters (e.g., weights, biases, or both), which have been computed or otherwise estimated by training the generative probabilistic model on the training data. Storing the generative probabilistic model may also include storing the particular model architecture to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be stored.
[0080] In an example study, a 3D diffusion model was used to synthesize high- resolution 3D CT volumes from 2D X-ray images. The 3D diffusion model was designed as a score-based generative model (FIG.5). To enable training the 3D diffusion model on a single GPU, the sub-regions extracted from the x-ray images in the superior and inferior direction and regional coordinates were used as conditional inputs to generate the corresponding 3D sub- volume. The trained model samples a high-resolution 3D CT volume (0.5 mm3) using bidirectional 2D images and coordinates of the entire image. The proposed 3D diffusion model- based CT synthesis was trained and validated using 3D wrist CT and digitally reconstructed radiography (DRR).
[0081] The diffusion model is a generative model including two processes: forward and reverse processes. In the forward process, the conditional data distribution, p(x|y), is diffused by adding Gaussian noise to generate the series of noisy samples, p^ x 2t y ^^ N ^ x t ;^ t x , ^ t ^ 2t I ^ ,QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0082] where the condition y is provided. The forward process can be interpreted as a continuous mapping from data distribution to prior distribution with time variable t ∈ [0, T] such that x(0) ∼ p0, which equals a data distribution, and x(T) ∼ pT equals a prior distribution. In the view of score-based generative modeling, the forward processing could be characterized by a stochastic differential equation (SDE), dx ^ f ^ x, t ^ dt ^ g ^ t ^ dw ,
[0083] f(・, t) : Rd→ Rdis the drift function, and g(t) :R → R is a
[0084] In the reverse process, the data distribution x(0) ∼ p0started from the prior distribution x(T) ∼ pT, can be sampled while the backward processing is also given by the reverse time SDE. The closed form of reverse time SDE is provided as, dx^^^ f ^ x, t ^ ^ g ^ t ^2 ^x log p t ^ x y ^ ^^ dt ^ g ^ t ^ dw ,
[0085] starting from t =T to t = 0. Here, the gradient of the log density of the data datasets, ^xlog p t ^ x y ^ , is a scorefunction. The SDE in the reverse process can be converted into the differentialequation (ODE) named probability flow ODE. This is given as, dx^^^ ^^^ ^ t ^ ^ ^ t ^ ^xlog p t ^ x y ^ ^ ^ dt ,
[0086] over time. In order to solve thereverse process, the score function ^xlog p t ^ x y ^ has to be determined. Therefore, theneural network can be trained toby score matching with a denoiser function D(x; σ, y), ^^ D ^ x ;^, y ^^x ^^ ^ .
[0087] is given as, LE^2 D ^ x ^ n ; ^ , y ^ ^ x2 2.
[0088] significant memory and is made even more challenging by the need to gather extensive 3D CT data. To reduce memory usage and the need for a large dataset, the 3D diffusion model can be trained using patch-wise loss. The loss function of the denoiser function can be decomposed as the sum of patch-wise losses 20 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 because convolution layers are shift-invariant. A uniformly sampled of the 3D CT subvolume, x, can be denoted as x^. Then, x^ ^N 1x ^
[0089] a total number of sub-divisions. Then, the patch-wise loss is given as,2 LP^^Ep dataEN ^ 0,^2 I ^ D ^ x^t^n ;^, y ^^x^ 2.
[0090] right-hand and 44 with left-hand 3D CT scans. prepare was performed to eliminate extra-body objects, such as splints or casts, using ITK-snap. Additionally, the left-hand CT scans were flipped to have identical anatomical positions as the right-hand scan. The segmented 3D CT volumes have the size of 256 × 256 × 160, which corresponds to a spatial resolution of 0.5 mm3.
[0091] The dataset was randomly split into 80% / 20% for model training (n=61) and testing(n=15), respectively. The neural network is a modified 3D DDPM++ U-net model by changing convolutional layers and attention modules to 3D. For the patch-wise training, a 3D CT volume was randomly sampled using a 64-slice window in the superior-inferior direction, which can be seen in both AP and LAT view images. The 3D sub-volume has the size of 256 × 64 × 160. Subsequently, the 3D sub-volume was projected in the AP view (256 × 64) and LAT view (64 × 160) to create DRR patch images while preserving resolution in the superior- inferior direction.
[0092] The dimensions of AP and LAT images were expanded by repeating images in the LAT and AP directions, respectively. Therefore, the inputs have identical spatial sizes as the target 3D sub-volume. Similarly, the superior-inferior coordinates of the sampled 3D CT sub-volume were incorporated as input by adding additional channels. The batch size was set to 1, and the network was trained for 170k steps. The data augmentation was not used. Training and testing were performed on a single DGX-A100 GPU with 40GB memory (NVIDIA Santa Clara, California, USA).
[0093] The performance of the 3D diffusion model was quantitatively evaluated using peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) in comparison to the original CT scans. The evaluation included comparisons with 2-view and 36-view filtered back projection reconstruction (FBP), as well as variations dependent on the number of function 21 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 evaluations (NFE) for diffusion sampling. The reported metrics are the mean and standard deviation of 15 test cases.
[0094] Representative images demonstrate the potential of the 3D diffusion model to synthesize high-resolution wrist CT scans from AP and LAT image inputs. The 3D volume rendering and slices of the synthesized CT scans show anatomical agreement with the original CT images, while variations in trabecular bone structure are observed (FIG.6). The quantitative evaluation showed that CT generated with the diffusion model was superior in PSNR and SSIM compared to FBP CT reconstructed with 2 views. In the comparison of the performance based on NFE, synthesized CTs with over 50 NFE had higher SSIM than FBP CTs with 36 views.
[0095] A 3D diffusion model is presented for synthesizing high-resolution 3D CT images from bidirectional X-ray images. The diffusion model is conditioned on images and their corresponding patch coordinates. Therefore, the denoiser 3D U-net of the diffusion model can be trained using a patch loss function, which was executed on a single GPU. Further research is warranted to assess the performance of the 3D diffusion model in a group of patients with wrist fractures.
[0096] In another example study, a diffusion-based generative framework for synthesizing 3D wrist CT volumes from standard tri-planar radiographs in patients with distal radius fractures was developed and evaluated. The framework incorporates a convolutional neural network (CNN)–based pose estimator that infers wrist orientation from radiographs and conditions a 3D score-based diffusion model. Radiographs are back-projected into volumetric space using the estimated pose, enabling the model to encode view-specific spatial constraints throughout the generative process. Patchwise training enables high-resolution (0.5 mm) volumetric synthesis within routine GPU memory limits, preserving local structural detail such as cortical step-offs and fracture lines that are critical for clinical assessment.
[0097] The clinical applicability of the method was evaluated through a reader-blinded study and quantitative experiments. Using a dataset of 209 patients with matched radiographs and CT scans, image usability and the fidelity of fracture morphology in synthesized volumes was assessed relative to ground-truth CT. Three fellowship-trained hand surgeons, blinded to case identity and imaging types, independently counted the number of major articular fragments on both Original and Virtual CT scans. In addition, binary endpoints—including fragment presence and usability for surgical planning—were analyzed to quantify inter-rater agreement and correspondence with conventional CT. 22 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0098] This study demonstrated that high-resolution 3D wrist anatomy can be reliably synthesized from routine radiographs using a pose-conditioned diffusion framework. By bridging the information gap between planar and volumetric imaging, the method enables CT- like visualization without requiring cross-sectional acquisition, thus minimizing radiation exposure and improving access to 3D anatomical data across care settings. Beyond enhancing surgical planning, this approach may reduce downstream CT burden, shorten emergency department dwell times, and streamline outpatient workflows. These findings underscore the potential of generative imaging to expand diagnostic capacity using standard hardware, offering scalable augmentation for musculoskeletal care.
[0099] Data was collected from 688 patients diagnosed with distal radius fractures, each of whom underwent both standard plain radiographs and 3D wrist CT scans. All data acquisition was approved by the institutional review board, and informed consent was obtained from all participants. Radiographs and CT scans were acquired within a 7-day interval to ensure anatomical consistency across modalities.
[0100] Following quality control, 163 patients were excluded due to insufficient CT resolution exceeding 1.0 mm voxel spacing (n=62), presence of extraneous body parts or external devices in the CT field-of-view (n=57), or incorrect or missing anatomical orientation during radiograph acquisition (n=44). The final cohort included 525 patients, which were stratified into a model development set (n=316; mean age 52 ± 19 years; 62% female) and an independent test set (n=209; mean age 51 ± 20 years; 60% female), maintaining balance in age, sex, and fracture characteristics.
[0101] To focus on bone structures and ensure consistency across samples, any splints or casts present in the CT images were manually segmented and removed using ITK-SNAP. The CT volumes were resampled to a consistent spatial resolution of 0.5 mm×0.5 mm×0.5 mm and manually aligned along the proximal-distal and anteroposterior axes. Left wrist CT volumes were mirrored in the sagittal plane to match the coordinate system of right wrists, standardizing the anatomical orientation across the dataset.
[0102] FIG. 7 shows a clinical workflow comparing the standard pathway (top) with the proposed Virtual CT–augmented pathway (bottom) for 3D wrist imaging. From tri-planar radiographs (PA, OBL, LAT), a CNN-based pose estimator aligns views for 3D back- projection with positional encoding. The resulting volume is processed by a conditional 3D diffusion model to synthesize diagnostic-quality CT, enabling faster decision-making, triage of non-fracture cases, and reduced CT utilization without additional radiation. FIG. 8 shows an 23 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 example inference pipeline with: (1) CNN-based radiography pose estimation; (2) geometric alignment and back-projection into coarse 3D volumes; and (3) conditional 3D diffusion synthesis via iterative score-based denoising. FIG.9 shows an example training pipeline using synthetic data. High-resolution 3D wrist CTs undergo forward projection at clinical PA, OBL, and LAT angles to generate DRRs. DRRs are back-projected into 3D, fused with positional encoding, cropped along the Z-axis to focus on the ROI, and sampled into patches. Ground- truth CT patches, perturbed with Gaussian noise, are denoised by a 3D U-Net trained via score matching. FIG. 10 shows an example pose estimation module trained on synthetic DRR pairs to predict relative rotations (Rest) with a supervised loss against ground-truth (Rgt), improving alignment robustness for real radiographs. FIG. 11 shows an example quantitative evaluation from a majority level consensus from three hand surgeons (≤2 of 3 readers agreeing). For C1 (image usability), Virtual CT was rated diagnostic in 200 / 209 cases. For C2 (fracture presence), Virtual CT achieved sensitivity 1.00 and specificity 0.24 (172 / 172 fractures correctly identified, no missed fractures). For C3 (fragment count), Virtual CT showed lower inter-reader reliability (AC₂ 0.716 vs.0.863) with a tendency for overestimation (p=0.001), yet maintained directional concordance with Original CT.
[0103] Digitally reconstructed radiographs (DRRs) were created from the aligned CT volumes (FIG. 9) by simulating x-ray projections at anatomically defined rotation angles corresponding to the standard radiographic views: posteroanterior (PA), lateral (LAT), and oblique (OBL). DRRs were produced by applying rigid in-plane rotations along the proximal- distal axis of the training CT volumes, followed by parallel X-ray beam projection. The detector spacing and field of view were set to mimic routine clinical radiographs, so the simulated images follow realistic geometry. These DRRs served two purposes in training: first, each view was back-projected to build per-view volumetric guidance that conditions the denoising score-matching 3D U-Net used in the diffusion framework; second, each DRR was paired with its known rotation to train a CNN that estimates view angles, which is then applied to real radiographs at inference.
[0104] For real radiographs used during test-time inference, background structures were removed using the Segment Anything Model (SAM), which automatically segmented the visible anatomical region. Radiographs were then resampled to a resolution of 0.5 mm × 0.5 mm to match the scale of the CT-derived DRRs. In-plane alignment of radiographs was performed manually using a custom graphical user interface. As shown in FIG.12, each image was rotated to align the ulnar shaft with the vertical axis, corresponding to the anatomical 24 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 proximal–distal direction. Subsequently, the image was translated such that the base of the lunate bone was centered both vertically and horizontally. The lunate was selected as the centering reference because it is consistently visible across PA, OBL, LAT views and remains largely unaffected by distal radius fractures, making it a reliable and anatomically stable landmark.
[0105] A 3D score-based diffusion model is trained to synthesize high-resolution wrist CT volumes directly from radiographs, conditioned on view-wise back-projections (FIG.9). For each CT, DRRs are rendered at arbitrary wrist rotations and at canonical views (e.g., PA: ^െ15,15^∘, OBL: ^30,60^∘, LAT: ^75,105^∘), and volumetric guidance is obtained by back-projecting each DRR, yielding ^^^^^,^^, ^^^^,^^^, ^^^^,^^^^. A 1D z-axis positional map ^^௭ ∈^െ1,1^^ൈுൈ^ൈ^is added for condition. The conditioning set is, ^^:ൌ ^^^^^,^^, ^^^^,^^^, ^^^^,^^^,^^௭^. (1)
[0106] Forward and reverse diffusion (EDM-style).
[0107] EDM parameterization was followed. Let ^^ ∈ ℝுൈ^ൈ^ be the clean CT and ^^the conditioning. The forward corruption samples a noise level ^^ ∈ ^^^୫୧୬,^^୫ୟ^^ from a log-normal ^^^^^^ and forms ^^ఙ ൌ ^^ ^ ^^^^, ^^ ∼ ^^^0, ^^^.
[0108] The denoiser predicts the clean volume via EDM preconditioning ^^^ ^^^ ,^^, ^^^ ൌ ^^ ^^^^ ^^ ^ ^^ ^^^^ ^^^^^ ^^^^ ^^ , ^^, ^^^, (2)^ ఙ ^୩୧୮ ఙ ୭^^ ఏ ୧୬ ఙ
[0109] ^ ఙ మ ^^ ^^^^ ൌ , ^^ ^^^^ ൌ ^^౪^୧୬ ^୩୧୮ఙమାమ , ^^୭^^^^^^ ൌఙ ఙ^^౪^, (3) ^ ఙ
[0110] By Tweedie’s relation, ^ ௫^௫ ,ఙ,^^ି௫బ ^ ^ ∇ log^^^^^ |^^^ ^. (4)௫ ఙమ^ఙ
[0111] ^^:^ ௗ௫ఙ^ ^௧^ௗ௫ ௫ ି௫^௫ ,ఙ,^^^ ^ ^ బ ^ ൌ ^^^ െ ^^^ ^^^ ,^^, ^^^^ ^ ൌ. (5)ఙ ^ ఙ
[0112] ^^ / ఘ ^ / ఘ ^ / ఘ ఘ ^^ ൌ ^ ^^ ^^^^െ ^^ ^^ , ^^ ൌ 0, … ,^^ െ 1, (6)^ ୫ୟ^ ୫ୟ^୫୧୬
[0113] ODE. Noise level ^^ is embedded and fused into the 3D U-Net via feature-wise affine 25 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 modulation at multiple scales; ^^ (the three back-projections and ^^௭) is concatenated and injected through the same modulators.
[0114] Denoising score matching and the score identity
[0115] Let ^^ ∼ ^^^ୟ^ୟ be a clean CT, ^^ ∼ ^^^0,^^ଶ^^^, and ^^ ൌ ^^ ^ ^^. Givenconditioning ^^, a denoiser ^^ఏwas trained with
[0116] ℒୈୗ^^^^,^^^ ൌ ^^௫∼^^^౪^ ^^^∼^^^^,ఙమூ^‖^^ఏ^^^; ^^, ^^^ െ ^^‖ଶଶ. (7) ^^^ ^^^ ൌఙమ. In practice the plug-in denoising score estimator ^^ఏ^^^,^^, ^^^:ൌ ^^^ఏ^^^;^^, ^^^ െused as the conditional score. Training objective (patch-wise)
[0119] DSM was optimized while sampling ^^ ∼ ^^^^^^ at each step:ଶℒୈୗ^ ൌ ^^௫∼^^^౪^ ^^ఙ∼^^ఙ^ ^^^∼^^^^,ఙమூ^‖^^ఏ^^^ ^ ^^^^;^^, ^^^ െ ^^‖ଶ. (8)
[0120] performed ,the corresponding crops (including ^^ ). The patch loss is௭ℒൌ ∑ே^^ ^^ ^^ ^^ ^ ^^^ ^^, ^ଶ ୮ୟ^ୡ୦ ^ୀ^ ௫∼^ ఙ∼^^ఙ^ మ ฮ^^ఏ൫ ஐ ^ ; ^ ஐ ൯ െ ^^ஐ ฮ.^∼^^^^,ఙ ூ^^^౪^ ^ ^ ^ ଶ
[0121] local context. Minimizing ℒ therefore forces consistent predictions for shared voxels୮ୟ^ୡ୦across contexts, aligning patch boundaries during training. The z-axis positional map ^^ further௭disambiguates depth across adjacent patches and stabilizes cross-patch alignment.
[0122] The denoising score matching network is a 3D U-Net with encoder–decoderpaths and skip connections. Inputs are ^^^ ^ ^^,^^^^^^, ^^^; the three back-projections and ^^ are௭concatenated as conditioning and injected at multiple resolutions via affine modulation. ିସ ିହ Training was done with Adam (learning rate 2 ൈ 10 , weight decay 1 ൈ 10 ^, batch size 1,for 65,000 iterations on four NVIDIA A100 (40 GB) GPUs with mixed precision. Z-axis ROI ∘ cropping focuses the wrist. Augmentations include random rotations (േ10 ) and translations (േ5 voxels), applied consistently to CT and guidance volumes.
[0123] To estimate the radiographic view orientation, a convolutional neural network (CNN) was trained to predict the relative along the proximal-distal anatomical axis rotation angle of each input radiograph with respect to the canonical PA view (FIG. 10). The model architecture consists of a feature extractor backbone (a ResNet-18 pretrained on ImageNet and 26 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 adapted for multi-channel radiographic input) followed by a fully connected regression head. This design enables the model to capture geometric cues implicit in the radiographic projection and learn rotational features from training data synthesized via DRRs.
[0124] To formulate the angular prediction task in a geometrically consistent manner, the estimated and ground-truth rotation matrices in homogeneous coordinates were defined as, ^^ ൌ ^^^^^^^^ ^^ ^ ^^^^ୃ1൨ , ^^^ ൌ ^ ^^^ ^ୃ൨^^ 1
[0125] where ^^^^^,^^^^ ∈ SO^3^ represent the estimated and ground-truth rotations,respectively.
[0126] A geodesic loss in the Lie algebra ^^^^^3^ can be minimized: ℒ^^୭ ൌ ฮlog൫^^ୃ^^^ ^^^^൯ฮி (10)
[0127] thematrix , norm. discrepancies in the predicted rotation relative to the true orientation on the manifold of 3D rotations.
[0128] For interpretability in clinical settings, the inter-plane projection angle ^^^୮୰୭୨was also extracted from ^^^^^as: ^^^୮୰୭୨ ൌ arctan2^^^ଷଶ, ^^ଷଷ^
[0129] , where ^^^^denotes the element in the ^^-th row and ^^-th column of ^^^^^. This projected angle corresponds to the observable rotation about the horizontal axis on standard radiographs and serves as the primary conditioning parameter for the downstream diffusion model.
[0130] Dynamic generation of synthetic training data was implemented by samplingazimuth angles ^^^ ൌ Δ^^ ^ ^^^^ relative to posteroanterior views during each epoch. The angularstep was set as Δ^^ ∈ ^0, గ , గସ ଶ^ for generating posteroanterior views, oblique views, or lateralviews, respectively, with the random perturbation term ^^^^ ∼ ^^^െగ, గ^ଶ ^ଶ^.
[0131] To train the view estimation model, 171(voxel size: 0.5ଷ mm,randomly drawn from model development) were used, from which digitally reconstructed radiographs (DRRs) were generated at various projection angles. For each scan, a maximum of 500 individual projections were produced, and view pairs ^^^^,^^^^ were sampled to form up to 250 paired combinations per patient, resulting in 42,750 training pairs for each view class(PA–LAT and PA–OBL). Angular separation between projections |^^^ െ ^^^| was constrained27 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 to ensure sufficient geometric variation within each pair while maintaining realistic anatomical alignment.
[0132] Adam optimization was used with learning rate 1 ൈ 10ିସ, batch size 8, andcosine learning rate decay over 20 epochs. Separate models were trained for posteroanterior- oblique (PA-OBL) and posteroanterior-lateral (PA-LAT) view estimation. The network backbone processes 256ൈ256 DRRs through grouped convolutions, followed by a fully- connected head regressing rotation parameters in the ^^^^^3^ tangent space. Train was done on a NVIDIA A100 GPU with 40GB memory, implementing early stopping after 5 epochs of validation loss plateau.
[0133] The model was evaluated on an independent set of 10 patients (randomly drawn from model development), each contributing 50 DRRs, totaling 500 unseen projections. Validation included Bland–Altman and linear regression analyses against ground-truth rotation angles. As shown in the Supplementary 0.11, the PA–LAT estimation exhibited a bias ofെ4.6∘ with limits of agreement from െ8.7∘ to ^17.9∘, while the PA–OBL estimation showeda bias of ^2.0∘ and limits from െ12.3∘ to ^16.3∘. Regression slopes were 0.78 (^^ଶ ൌ 0.71)for PA–LAT and 0.89 (^^ଶ ൌ 0.73) for PA–OBL, indicating strong linear correspondencebetween estimated and true inter-view rotation.
[0134] In the testing phase (FIG. 8), standard radiographs from PA, LAT, and OBL views were collected for each patient. The pose estimator model estimated the rotation angles from these radiographs. An internally developed GUI was used to align the radiographs along the proximal-distal axis and to adjust in-plane rotations based on the ulna bone axis. The aligned radiographs were backprojected using the estimated rotation angles from CNN-base Pose Estimation model to generate an initial volumetric estimate. The 3D diffusion model then refined this estimate to produce the final synthesized CT volume. The number of reverse-time sampling steps was set to 100, using the deterministic second-order Heun method. Inference required approximately 5 minutes on a single NVIDIA A100 GPU with 40 GB of memory. The generated virtual CT volumes had a size of 256 ൈ 224 ൈ 160 voxels with an isotropic spatial resolution of 0.5 mm.
[0135] A generative diffusion pipeline that synthesizes high-resolution 3D wrist CT volumes directly from tri-planar wrist radiographs (Posteroanterior, Lateral, and Oblique) in patients with suspected distal radius fractures (FIG. 7) was developed. The system includes three modules (FIG. 8): (1) a CNN-based pose estimation module that infers global wrist orientation from the input radiographs; (2) a geometric alignment module that spatially 28 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 reorients in-plane and voxelizes the back projections into a 3D volume; and (3) a conditional 3D score-based diffusion module trained in a patch-wise fashion to accommodate high- resolution CT synthesis under GPU memory constraints. All modules were integrated into a unified inference framework enabling end-to-end CT volume generation from input radiographs.
[0136] The framework was trained using CT volumes and digitally reconstructed radiographs (DRRs) on a cohort of 316 patients (mean age 52 ± 19 years; 62% female), and the final model was validated on a cohort of 209 patients (mean age 51 ± 20 years; 60% female) with suspected distal radius fractures, all of whom underwent both imaging modalities within a 7-day interval.
[0137] Representative wrist trauma cases (FIGS.13–16) were qualitatively assessed to examine the visual fidelity and diagnostic sufficiency of Virtual CT reconstructions relative to their paired Original CTs. These examples were selected to span the full clinical spectrum, from intact anatomy to complex intra-articular fracture patterns. In fracture-negative cases (FIG. 13), both Original and Virtual CTs preserved cortical smooth-ness, trabecular consistency, and joint congruity across all planes. In fracture-positive examples (FIGS.14–16), Virtual CT consistently captured diagnostic features such as cortical breaches, fragment contours, and intra-articular involvement, with increasing visibility in more severe injuries. While subtle reductions in edge sharpness and trabecular granularity were noted in Virtual CT, major topological relationships, including fragment displacement and spatial orientation, remained intact.
[0138] 3D renderings further corroborated structural alignment, showing preservation of metaphyseal contour, articular surface integrity, and fragment positioning. In particular, Virtual CT reconstructions did not exhibit artificial bone formation or erasure of critical disruptions in radius bone region. Even in cases with multiple displaced fragments (FIG. 14), the model reproduced comminution patterns and joint deformity with sufficient precision for expert interpretation.
[0139] Virtual CTs were found to preserve key morphologic cues required for fracture classification, including radial inclination, lunate fossa contour, and articular step-off. Minor smoothing at fragment boundaries did not hinder recognition of clinically relevant findings. These observations support the use of Virtual CT as a viable surrogate in diagnostic workflows, particularly in contexts requiring fracture detection, characterization, and surgical planning. 29 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0140] The Virtual CT is therefore advantageous for screening and prioritization: promptly flagging potentially complex, multifragmentary injuries for expedited formal CT, while allowing straight-forward cases to proceed without immediate cross-sectional imaging, optimizing resource allocation without compromising patient safety.
[0141] Distal radius fractures are among the most common orthopedic injuries, yet their evaluation of-ten requires computed tomography (CT) to assess fragment configuration and joint involvement. While CT provides detailed anatomical information, its routine use is limited by concerns related to radiation exposure, cost, and access, particularly in low-resource settings. Standard radiographs remain the first-line modality worldwide but lack volumetric detail and suffer from overlapping structures and projectional ambiguity, limiting their utility in complex fracture assessment.
[0142] To address this diagnostic gap, a diffusion model–based framework that synthesizes high-resolution 3D CT volumes from standard radiographs without requiring external fiducials or mechanical stabilization was proposed. The approach aims to bridge 2D and 3D imaging by generating volumetric anatomical reconstructions from widely available radiographs, offering a scalable and radiation-free alternative for comprehensive fracture evaluation.
[0143] Traditional methods for reconstructing 3D structures from 2D radiographs have faced significant challenges due to the loss of depth information and overlapping anatomical features inherent in x-ray images. These methods often require extensive computational resources, specialized equipment, or may lack accuracy and generalizability, limiting their practical use in clinical settings. Previous attempts to bridge this gap have not fully leveraged the advancements in deep learning, particularly in generative models like diffusion models, which have shown promising results in other imaging domains. Generation of 3D images from radiographs of rigidly fixed wrists with external constraints is technically feasible but clinically impractical in patients with a painful fracture. In addition, the generation of 3D images from unfractured wrists using generative techniques are also described, but the heterogeneity of distal radius fracture patterns adds unique challenges especially in view estimation and standardization of axes.
[0144] To overcome these limitations, a generative pipeline that directly learns anatomical structure from radiographic input without requiring external constraints or specialized equipment was designed. The approach incorporates diffusion model architecture in generating high-quality images while effectively managing the high-dimensionality of 3D 30 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 data through patch-wise diffusion model training. This strategy enables efficient high- resolution image synthesis without exceeding hard-ware limitations. The integration of CNN- based pose estimation addresses the geometric ambiguities by providing precise conditioning parameters for the diffusion model. By aligning radiographs using anatomical landmarks and estimating rotation angles, it can be ensured that the synthesized 3D images accurately represent the patient’s anatomy.
[0145] The clinical utility of the method is notable. By generating diagnostic-quality 3D CT volumes from routine radiographs, the approach offers an alternative to conventional CT in assessing distal radius fractures. This could reduce radiation exposure, lower imaging costs, and increase access, particularly in low-resource environments where CT scanners are unavailable. Notably, once radiographs are acquired, nearly all distal radius fractures can be volumetrically characterized without further imaging, enabling more consistent and timely fracture evaluation. These synthetic CTs could be directly integrated into radiology reading platforms and orthopedic surgical planning workflows, and extended to other joints with high fracture prevalence.
[0146] FIG.17 shows an example of a system 1700 for generating 3D CT image data in accordance with some embodiments described in the present disclosure. As shown in FIG. 17, a computing device 1750 can receive one or more types of data (e.g., 2D image data, pose data, etc.) from data source 1702. In some embodiments, computing device 1750 can execute at least a portion of a 3D CT image synthesis system 1704 to generate 3D CT image data from 2D images received from the data source 1702.
[0147] Additionally or alternatively, in some embodiments, the computing device 1750 can communicate information about data received from the data source 1702 to a server 1752 over a communication network 1754, which can execute at least a portion of the 3D CT image synthesis system 1704. In such embodiments, the server 1752 can return information to the computing device 1750 (and / or any other suitable computing device) indicative of an output of the 3D CT image synthesis system 1704.
[0148] The 3D CT image synthesis system 1704 operates through an integrated workflow that combines multiple processing modules to transform two-dimensional radiographic images into three-dimensional computed tomography volumes. The workflow begins when input radiographic images are received by a pose estimation module, which analyzes the geometric relationships between multiple views to determine relative positioning and orientation angles. By way of example, the pose estimation module processes the 31 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 radiographic data using convolutional neural network architectures that extract spatial features and predict rotation matrices in three-dimensional space. Following pose estimation, a spatial alignment module receives the orientation data and applies geometric transformations to ensure anatomical consistency across different input views.
[0149] The aligned radiographic images are subsequently processed by a backprojection module, which generates initial three-dimensional volumetric estimates based on the determined pose parameters and spatial relationships. The backprojection module incorporates positional encoding schemes that provide depth information and coordinate references for the subsequent generative probabilistic model processing stages. The initial volumetric estimates serve as conditioning inputs for the three-dimensional generative probabilistic model, which refines the volumetric data through iterative operations. The generative probabilistic model applies learned anatomical priors and geometric constraints to enhance the structural accuracy and detail resolution of the generated volumes.
[0150] During the refinement process, the three-dimensional generative probabilistic model may operate on patch-based segments of the volumetric data to accommodate computational memory limitations while maintaining spatial coherence across the entire volume. The patch-wise processing approach allows the 3D CT image synthesis system 1704 to handle high-resolution volumetric outputs while distributing computational load across available hardware resources. Alternatively, non-patch-based approaches may be employed when sufficient computational resources are available to process full-resolution volumes directly, which may provide enhanced global spatial context at the cost of increased memory requirements. The generative probabilistic model may be a diffusion model that applies score- based denoising techniques that progressively remove noise artifacts and enhance anatomical features based on training data patterns. The iterative refinement continues through multiple sampling steps until the volumetric output reaches the desired quality and resolution specifications.
[0151] The 3D CT image synthesis system 1704 architecture may be adapted for different anatomical regions beyond wrist applications, including ankle, elbow, and other joints with high fracture prevalence. For ankle applications, the pose estimation module may be retrained using ankle-specific radiographic datasets to learn the geometric relationships between tibial, fibular, and tarsal bone structures. Moreover, the disclosed systems and methods may process images obtained with or without contrast-agent administration, as well as dynamic fluoroscopic sequences, thereby extending applicability across a broader range of 32 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 interventional and diagnostic workflows. The spatial alignment procedures may be modified to account for the different anatomical landmarks and orientation references specific to ankle anatomy. The three-dimensional generative probabilistic model may be adapted through transfer learning techniques that leverage pre-trained wrist models while fine-tuning the network parameters for ankle-specific bone morphology and tissue characteristics.
[0152] Similarly, for elbow applications, the 3D CT image synthesis system 1704 may incorporate specialized preprocessing steps that account for the complex joint articulations between the humerus, radius, and ulna bones. The pose estimation algorithms may be configured to recognize elbow-specific anatomical features and geometric constraints that differ from wrist anatomy. The backprojection module may be adjusted to handle the different spatial relationships and projection angles commonly used in elbow radiography. The generative probabilistic model training may incorporate elbow-specific computed tomography datasets to learn the appropriate anatomical priors and structural patterns for accurate volume generation.
[0153] The modular design of the 3D CT image synthesis system 1704 architecture facilitates adaptation to various anatomical regions through component-specific modifications rather than complete system redesign. The pose estimation module may be retrained with region-specific datasets while maintaining the same underlying convolutional neural network architecture. In some implementations, rather than estimating inter-plane rotations from image data alone, the disclosed systems may also utilize rotation encodings provided directly by the imaging system (e.g., gantry or C-arm angle readouts) as conditioning inputs, thereby improving geometric accuracy and reducing estimation ambiguity. The spatial alignment procedures may be updated with new anatomical landmark definitions and coordinate systems appropriate for different joint structures. The generative probabilistic model may leverage transfer learning approaches that adapt pre-trained models to new anatomical domains while preserving learned geometric and structural relationships that are common across different skeletal regions.
[0154] In some embodiments, computing device 1750 and / or server 1752 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable computer, a server computer, a virtual machine being executed by a physical computing device, and so on. The computing device 1750 and / or server 1752 can also reconstruct images from the data. 33 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0155] In some embodiments, data source 1702 can be any suitable source of data (e.g., measurement data, images reconstructed from measurement data, processed image data), such as an x-ray imaging system, another computing device (e.g., a server storing measurement data, images reconstructed from measurement data, processed image data), and so on. In some embodiments, data source 1702 can be local to computing device 1750. For example, data source 1702 can be incorporated with computing device 1750 (e.g., computing device 1750 can be configured as part of a device for measuring, recording, estimating, acquiring, or otherwise collecting or storing data). As another example, data source 1702 can be connected to computing device 1750 by a cable, a direct wireless link, and so on. Additionally or alternatively, in some embodiments, data source 1702 can be located locally and / or remotely from computing device 1750, and can communicate data to computing device 1750 (and / or server 1752) via a communication network (e.g., communication network 1754).
[0156] In some embodiments, communication network 1754 can be any suitable communication network or combination of communication networks. For example, communication network 1754 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, a 4G network, etc., complying with any suitable standard, such as CDMA, GSM, LTE, LTE Advanced, WiMAX, etc.), other types of wireless network, a wired network, and so on. In some embodiments, communication network 1754 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Communications links shown in FIG.17 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, and so on.
[0157] Referring now to FIG. 18, an example of hardware 1800 that can be used to implement data source 1702, computing device 1750, and server 1752 in accordance with some embodiments of the systems and methods described in the present disclosure is shown.
[0158] As shown in FIG.18, in some embodiments, computing device 1750 can include a processor 1802, a display 1804, one or more inputs 1806, one or more communication systems 1808, and / or memory 1810. In some embodiments, processor 1802 can be any suitable hardware processor or combination of processors, such as a central processing unit (“CPU”), a graphics processing unit (“GPU”), and so on. In some embodiments, display 1804 can include any suitable display devices, such as a liquid crystal display (“LCD”) screen, a light-emitting 34 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 diode (“LED”) display, an organic LED (“OLED”) display, an electrophoretic display (e.g., an “e-ink” display), a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 1806 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0159] In some embodiments, communications systems 1808 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1754 and / or any other suitable communication networks. For example, communications systems 1808 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1808 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0160] In some embodiments, memory 1810 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1802 to present content using display 1804, to communicate with server 1752 via communications system(s) 1808, and so on. Memory 1810 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1810 can include random-access memory (“RAM”), read-only memory (“ROM”), electrically programmable ROM (“EPROM”), electrically erasable ROM (“EEPROM”), other forms of volatile memory, other forms of non-volatile memory, one or more forms of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1810 can have encoded thereon, or otherwise stored therein, a computer program for controlling operation of computing device 1750. In such embodiments, processor 1802 can execute at least a portion of the computer program to present content (e.g., images, user interfaces, graphics, tables), receive content from server 1752, transmit information to server 1752, and so on. For example, the processor 1802 and the memory 1810 can be configured to perform the methods described herein.
[0161] In some embodiments, server 1752 can include a processor 1812, a display 1814, one or more inputs 1816, one or more communications systems 1818, and / or memory 1820. In some embodiments, processor 1812 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, display 35 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 1814 can include any suitable display devices, such as an LCD screen, LED display, OLED display, electrophoretic display, a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 1816 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0162] In some embodiments, communications systems 1818 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1754 and / or any other suitable communication networks. For example, communications systems 1818 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1818 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0163] In some embodiments, memory 1820 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1812 to present content using display 1814, to communicate with one or more computing devices 1750, and so on. Memory 1820 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1820 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1820 can have encoded thereon a server program for controlling operation of server 1752. In such embodiments, processor 1812 can execute at least a portion of the server program to transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 1750, receive information and / or content from one or more computing devices 1750, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone), and so on.
[0164] In some embodiments, the server 1752 is configured to perform the methods described in the present disclosure. For example, the processor 1812 and memory 1820 can be configured to perform the methods described herein.
[0165] In some embodiments, data source 1702 can include a processor 1822, one or more data acquisition systems 1824, one or more communications systems 1826, and / or memory 1828. In some embodiments, processor 1822 can be any suitable hardware processor 36 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, the one or more data acquisition systems 1824 are generally configured to acquire data, images, or both, and can include an x-ray imaging system. Additionally or alternatively, in some embodiments, the one or more data acquisition systems 1824 can include any suitable hardware, firmware, and / or software for coupling to and / or controlling operations of an x-ray imaging system. In some embodiments, one or more portions of the data acquisition system(s) 1824 can be removable and / or replaceable.
[0166] Note that, although not shown, data source 1702 can include any suitable inputs and / or outputs. For example, data source 1702 can include input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, a trackpad, a trackball, and so on. As another example, data source 1702 can include any suitable display devices, such as an LCD screen, an LED display, an OLED display, an electrophoretic display, a computer monitor, a touchscreen, a television, etc., one or more speakers, and so on.
[0167] In some embodiments, communications systems 1826 can include any suitable hardware, firmware, and / or software for communicating information to computing device 1750 (and, in some embodiments, over communication network 1754 and / or any other suitable communication networks). For example, communications systems 1826 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1826 can include hardware, firmware, and / or software that can be used to establish a wired connection using any suitable port and / or communication standard (e.g., VGA, DVI video, USB, RS-232, etc.), Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0168] In some embodiments, memory 1828 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1822 to control the one or more data acquisition systems 1824, and / or receive data from the one or more data acquisition systems 1824; to generate images from data; present content (e.g., data, images, a user interface) using a display; communicate with one or more computing devices 1750; and so on. Memory 1828 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1828 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1828 can have encoded thereon, or otherwise stored 37 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 therein, a program for controlling operation of data source 1702. In such embodiments, processor 1822 can execute at least a portion of the program to generate images, transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 1750, receive information and / or content from one or more computing devices 1750, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone, etc.), and so on.
[0169] In some embodiments, any suitable computer-readable media can be used for storing instructions for performing the functions and / or processes described herein. For example, in some embodiments, computer-readable media can be transitory or non-transitory. For example, non-transitory computer-readable media can include media such as magnetic media (e.g., hard disks, floppy disks), optical media (e.g., compact discs, digital video discs, Blu-ray discs), semiconductor media (e.g., RAM, flash memory, EPROM, EEPROM), any suitable media that is not fleeting or devoid of any semblance of permanence during transmission, and / or any suitable tangible media. As another example, transitory computer- readable media can include signals on networks, in wires, conductors, optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.
[0170] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms “component,” “system,” “module,” “framework,” and the like are intended to encompass part or all of computer-related systems that include hardware, software, a combination of hardware and software, or software in execution. For example, a component may be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a computer and the computer can be a component. One or more components (or system, module, and so on) may reside within a process or thread of execution, may be localized on one computer, may be distributed between two or more computers or other processor devices, or may be included within another component (or system, module, and so on).
[0171] In some implementations, devices or systems disclosed herein can be utilized or installed using methods embodying aspects of the disclosure. Correspondingly, description herein of particular features, capabilities, or intended purposes of a device or system is generally intended to inherently include disclosure of a method of using such features for the intended purposes, a method of implementing such capabilities, and a method of installing 38 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 disclosed (or otherwise known) components to support these purposes or capabilities. Similarly, unless otherwise indicated or limited, discussion herein of any method of manufacturing or using a particular device or system, including installing the device or system, is intended to inherently include disclosure, as embodiments of the disclosure, of the utilized features and implemented capabilities of such device or system.
[0172] Referring particularly now to FIGS. 19A and 19B, an example of an x-ray computed tomography (“CT”) imaging system 1900 is illustrated. The CT system includes a gantry 1902, to which at least one x-ray source 1904 is coupled. The x-ray source 1904 projects an x-ray beam 1906, which may be a fan-beam or cone-beam of x-rays, towards a detector array 1908 on the opposite side of the gantry 1902. The detector array 1908 includes a number of x-ray detector elements 1910. Together, the x-ray detector elements 1910 sense the projected x-rays 1906 that pass through a subject 1912, such as a medical patient or an object undergoing examination, that is positioned in the CT system 1900. Each x-ray detector element 1910 produces an electrical signal that may represent the intensity of an impinging x-ray beam and, hence, the attenuation of the beam as it passes through the subject 1912. In some configurations, each x-ray detector 1910 is capable of counting the number of x-ray photons that impinge upon the detector 1910. During a scan to acquire x-ray projection data, the gantry 1902 and the components mounted thereon rotate about a center of rotation 1914 located within the CT system 1900.
[0173] The CT system 1900 also includes an operator workstation 1916, which typically includes a display 1918; one or more input devices 1920, such as a keyboard and mouse; and a computer processor 1922. The computer processor 1922 may include a commercially available programmable machine running a commercially available operating system. The operator workstation 1916 provides the operator interface that enables scanning control parameters to be entered into the CT system 1900. In general, the operator workstation 1916 is in communication with a data store server 1924 and an image reconstruction system 1926. By way of example, the operator workstation 1916, data store sever 1924, and image reconstruction system 1926 may be connected via a communication system 1928, which may include any suitable network connection, whether wired, wireless, or a combination of both. As an example, the communication system 1928 may include both proprietary or dedicated networks, as well as open networks, such as the internet.
[0174] The operator workstation 1916 is also in communication with a control system 1930 that controls operation of the CT system 1900. The control system 1930 generally 39 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 includes an x-ray controller 1932, a table controller 1934, a gantry controller 1936, and a data acquisition system 1938. The x-ray controller 1932 provides power and timing signals to the x-ray source 1904 and the gantry controller 1936 controls the rotational speed and position of the gantry 1902. The table controller 1934 controls a table 1940 to position the subject 1912 in the gantry 1902 of the CT system 1900.
[0175] The DAS 1938 samples data from the detector elements 1910 and converts the data to digital signals for subsequent processing. For instance, digitized x-ray data is communicated from the DAS 1938 to the data store server 1924. The image reconstruction system 1926 then retrieves the x-ray data from the data store server 1924 and reconstructs an image therefrom. The image reconstruction system 1926 may include a commercially available computer processor, or may be a highly parallel computer architecture, such as a system that includes multiple-core processors and massively parallel, high-density computing devices. Optionally, image reconstruction can also be performed on the processor 1922 in the operator workstation 1916. Reconstructed images can then be communicated back to the data store server 1924 for storage or to the operator workstation 1916 to be displayed to the operator or clinician.
[0176] The CT system 1900 may also include one or more networked workstations 1942. By way of example, a networked workstation 1942 may include a display 1944; one or more input devices 1946, such as a keyboard and mouse; and a processor 1948. The networked workstation 1942 may be located within the same facility as the operator workstation 1916, or in a different facility, such as a different healthcare institution or clinic.
[0177] The networked workstation 1942, whether within the same facility or in a different facility as the operator workstation 1916, may gain remote access to the data store server 1924 and / or the image reconstruction system 1926 via the communication system 1928. Accordingly, multiple networked workstations 1942 may have access to the data store server 1924 and / or image reconstruction system 1926. In this manner, x-ray data, reconstructed images, or other data may be exchanged between the data store server 1924, the image reconstruction system 1926, and the networked workstations 1942, such that the data or images may be remotely processed by a networked workstation 1942. This data may be exchanged in any suitable format, such as in accordance with the transmission control protocol (“TCP”), the internet protocol (“IP”), or other known or suitable protocols. 40 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890
[0178] The present disclosure has described one or more preferred embodiments, and it should be appreciated that many equivalents, alternatives, variations, and modifications, aside from those expressly stated, are possible and within the scope of the disclosure. 41 QB\125141.04890\98542490.2
Claims
MGH 2024-020-02 125141.04890 CLAIMS 1. A method for generating three-dimensional computed tomography image data from two-dimensional radiographic images, comprising: accessing two-dimensional radiographic image data with a computer system; processing the two-dimensional radiographic image data to estimate pose data indicating relative orientations between radiographic views; spatially aligning the two-dimensional radiographic image data based on the pose data to generate aligned radiographic images; generating coarse three-dimensional volumetric data by backprojecting the aligned radiographic images into three-dimensional space using the pose data; and applying the aligned radiographic images and the coarse three-dimensional volumetric data to a trained generative probabilistic model to generate three-dimensional computed tomography image data.
2. The method of claim 1, wherein processing the two-dimensional radiographic image data to estimate pose data comprises applying a convolutional neural network trained to predict relative rotation angles between paired radiographic views.
3. The method of claim 2, wherein the convolutional neural network comprises a pretrained ResNet-18 backbone adapted for multi-channel radiographic input processing.
4. The method of claim 2, wherein the convolutional neural network is trained using a geodesic loss function.
5. The method of claim 1, wherein spatially aligning the two-dimensional radiographic image data comprises identifying anatomical landmarks within the radiographic images to determine optimal positioning parameters.
6. The method of claim 1, wherein generating coarse three-dimensional volumetric data comprises distributing pixel intensities along projection rays through three- dimensional space based on the pose data. 42 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 7. The method of claim 6, further comprising incorporating positional encoding schemes that provide depth information and coordinate references for the three-dimensional space.
8. The method of claim 7, wherein the positional encoding schemes comprise one-dimensional z-axis positional maps that encode superior-inferior coordinate positions for each voxel location.
9. The method of claim 1, wherein the generative probabilistic model comprises a diffusion model that employs a score-based generative modeling approach.
10. The method of claim 9, wherein the diffusion model comprises a modified three-dimensional U-Net architecture with encoder-decoder pathways and skip connections.
11. The method of claim 10, wherein the three-dimensional U-Net architecture incorporates feature-wise affine modulation layers that enable conditioning on noise levels and input radiographic data.
12. The method of claim 1, wherein the generative probabilistic model comprises a flow model.
13. The method of claim 1, wherein the generative probabilistic model is trained using patch-wise loss functions on overlapping three-dimensional patches extracted from training computed tomography volumes.
14. The method of claim 13, wherein the patch-wise loss functions are formulated as a sum of individual patch losses leveraging shift-invariant properties of convolutional operations.
15. The method of claim 1, wherein the three-dimensional computed tomography image data has an isotropic spatial resolution of 0.5 mm³. 43 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 16. The method of claim 1, wherein applying the aligned radiographic images and the coarse three-dimensional volumetric data to the trained generative probabilistic model comprises iterative denoising through multiple sampling steps using a second-order Heun solver.
17. The method of claim 1, wherein the two-dimensional radiographic image data comprises at least one of posteroanterior, lateral, or oblique radiographic views of a wrist.
18. A method for training a generative probabilistic model to synthesize three- dimensional computed tomography volumes from two-dimensional radiographic projections, comprising: accessing training data including three-dimensional computed tomography volumes and corresponding digitally reconstructed radiographs with a computer system; generating backprojected volumetric guidance data by projecting the digitally reconstructed radiographs into three-dimensional space; training a generative probabilistic model on the training data and backprojected volumetric guidance data; and storing the trained generative probabilistic model for generating three-dimensional computed tomography image data from two-dimensional radiographic images.
19. The method of claim 18, wherein the digitally reconstructed radiographs are generated by simulating x-ray projections at anatomically defined rotation angles corresponding to posteroanterior, lateral, and oblique views.
20. The method of claim 19, wherein the posteroanterior views correspond to rotation angles within a range of -15° to 15°.
21. The method of claim 19 or 20, wherein the oblique views correspond to rotation angles within a range of 30° to 60°.
22. The method any one of claims 19, 20, or 21, wherein the lateral views correspond to rotation angles within a range of 75° to 105°. 44 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 23. The method of claim 18, wherein generating back-projected volumetric guidance data comprises incorporating one-dimensional z-axis positional maps that encode superior-inferior coordinate positions.
24. The method of claim 23, wherein the z-axis positional maps have values ranging from -1 to 1 across spatial dimensions.
25. The method of claim 18, wherein the generative probabilistic model comprises a three-dimensional denoising neural network having a modified three-dimensional U-Net architecture with encoder-decoder pathways and skip connections.
26. The method of claim 25, wherein the three-dimensional U-Net architecture incorporates feature-wise affine modulation layers that enable conditioning on noise levels and back-projected volumetric guidance.
27. The method of claim 26, wherein the backprojected volumetric guidance data and z-axis positional maps are concatenated and injected at multiple resolution levels through the affine modulation layers.
28. The method of claim 18, further comprising: extracting overlapping three-dimensional patches from the three-dimensional computed tomography volumes; and training the generative probabilistic model using a patch-wise loss function that operates on the overlapping three-dimensional patches, wherein the patch-wise loss function comprises a sum of individual patch losses computed across multiple overlapping spatial regions.
29. The method of claim 18, wherein the overlapping three-dimensional patches are extracted using a slice window in a superior-inferior direction. 45 QB\125141.04890\98542490.2MGH 2024-020-02 125141.04890 30. The method of claim 28, wherein the patch-wise loss function leverages shift- invariant properties of convolutional operations to ensure consistent predictions for shared voxels across different spatial contexts.
31. The method of claim 18, further comprising applying data augmentation including random rotations and translations to computed tomography volumes and corresponding guidance volumes.
32. The method of claim 18, wherein the generative probabilistic model comprises a score-based diffusion model that employs forward and reverse diffusion processes.
33. The method of claim 32, wherein the forward diffusion process adds Gaussian noise according to a predefined noise schedule and the reverse diffusion process learns to iteratively denoise random noise to generate synthetic volumetric data.
34. The method of claim 33, wherein the reverse diffusion process utilizes a second-order Heun solver with a Karras ρ-schedule for sampling three-dimensional computed tomography volumes from the learned data distribution. 46 QB\125141.04890\98542490.2
Citation Information
Patent Citations
Method and software for shape representation with curve skeletons
US20080018646A1
A System and Computer-Implemented Method for Segmenting an Image
US20200167930A1
Systems and methods of using three-dimensional image reconstruction to aid in assessing bone or soft tissue aberrations for orthopedic surgery
US20230005232A1