Navigation of medical instrument based on real-time intraoperative CT imaging using conventional x-ray device
The method addresses navigation challenges in surgeries by generating real-time 3D CT imaging from 2D X-ray images, improving accuracy and reducing radiation, enabling clear spatial visualization and expanding procedure access to outpatient settings.
Patent Information
- Application Number
- PCT/US2025/030620
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-22
- Filing Date
- 2025-05-22
- Publication Date
- 2025-11-27
AI Technical Summary
Current methods for navigating medical instruments during surgeries using 2D X-ray images face challenges such as registration errors, high radiation exposure, costly equipment, and limited access to advanced imaging, particularly in outpatient settings, with Axial and Sagittal views often being blurred or unclear.
A method and system for navigating medical instruments using real-time high-quality 3D CT imaging generated from standard 2D X-ray images, eliminating the need for 2D to 3D registration and reducing radiation exposure by employing a deep learning-based approach with MedNeRF or residual DNN to reconstruct clear spatial visualizations on all anatomical planes.
Enhances navigation accuracy, reduces costs and radiation exposure, and expands the availability of advanced procedures to outpatient settings by providing clear spatial visualization of medical instruments in real-time 3D CT imaging.
Smart Images

Figure US2025030620_27112025_PF_FP_ABST
Abstract
Description
Navigation Of Medical Instrument Based On Real-Time Intraoperative CT Imaging Using Conventional X-Ray DeviceCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This Application claims priorityto related U.S. provisional application serial number 63 / 650,726, filed May 22, 2024, by inventor Dorian Averbuch, entitled Navigation Of Medical Instrument Based On Real-Time Intraoperative CT Imaging Using Conventional X-Ray Device, the contents of which are incorporated by reference herein in their entirety.FIELD OF THE INVENTION
[0002] The present invention pertains to the field of medical imaging, specifically to a method and system for navigating medical instrument using real-time high-quality three-dimensional (3D) computed tomography (CT) imaging from 2D X-ray images acquired using conventional imaging equipment such as 2D X-Ray or C-Arm.BACKGROUND OF THE INVENTION
[0003] Three-dimensional (3D) imaging has become an essential tool in a multitude of medical intervention procedures such as spine, dental, orthopedic, and cardiovascular surgeries. The ability to accurately visualize and navigate complex 3D anatomical structures from two- dimensional (2D) X-ray images is a common and critical challenge faced by physicians. Particularly in spinal surgeries, discerning intricate details like the spinal cord and pedicle pathways can be exceedingly difficult from 2D images alone.
[0004] The use of 3D imaging necessitates access to CT scanners during interventional procedures by doctors of multiple specialties. This requirement involves several challenges, such as high radiation doses to the patient and staff, regulatory and training requirements around the facility using CT, and costs associated with the acquisition, installation, and maintenance of these complex equipment items.
[0005] Radiation dose remains one of the limiting factors for using CT intraoperatively. Some approaches aim to redesign CT devices to emit lower radiation and use machine learning (ML) to enhance the accuracy of tomographic reconstructions (see [8]).
[0006] In the realm of navigating surgical procedures, the two conventional approaches have traditionally leaned on the integration of CT scans with intraoperative 2D X-rays while 2D X-rays remains the main real-time navigation modality.
[0007] One solution involves the registration of preoperative CT scans with intraoperative 2D X- rays, allowing for the navigation of surgical tools within a virtual representation of the preoperative CT scan (see
[0001] , [4]). While this technique aids in visualizing the surgical field, it is fraught with potential for registration errors, lacks the capability for real-time updates to account for intraoperative changes, and depends on costly proprietary disposable navigation tools. Post-procedure CT verification is needed, further increasing radiation exposure and costs.
[0008] The alternative, which employs intraoperative CT imaging devices, while providing realtime feedback (see [2]), introduces its own set of issues. Such systems are expensive, lead to higher levels of radiation exposure for both patient and staff, and require operations within shielded rooms by specially trained professionals. The registration between intraoperative CT and X-Ray is still needed and remains an essential part of the process. This not only increases the financial burden but also limits the accessibility of advanced imaging to well-equipped hospital settings, thereby restricting the ability to perform certain procedures in outpatient facilities.
[0009] Yet, another alternative solution involves radiation dose reduction as a combination of complete redesign of CT hardware and the use of Al model to enhance tomographic reconstruction through number of iterations on the raw data (see [8]).
[0010] Yet, another alternative solution was developed to create tomographic reconstruction of pulmonary nodules from C-Arm that is limited to lung application and requires to use a preoperative CT as a part of the workflow. The reconstruction of limited angle tomography is performed using traditional back-face projection algorithm for CT reconstruction and then the tomography quality is refined through multiple iterations using ML model (see
[0013] ). Refining tomographic reconstruction with ML model was also proposed by additional patent applications (see [9],
[0010] ). Navigation of flexible catheters in the airways allows to use the map of bronchial airways to improve the localization of this instrument inside the lung (see
[0014] ). However, this approach is limited to natural cavities of the body and not applicable to spine or orthopedic procedures.
[0011] Although prior art methods have permitted the creation of limited angle tomography, which offers tomosynthesis of spinal anatomy, these methods are constrained to lung applications. While some can present good quality oblique views such as Anterior-Posterior (AP), Left Anterior Oblique (LAO), and Right Anterior Oblique (RAO) projections, the Axial and Sagittal views remain blurred, unclear, and are consequently of limited utility to physicians. This limitation represents a significant shortfall, as real-time access to sharp Axial and Sagittal views is critical for performing intricate procedures on complex bony structures like the spine or knee. Moreover, the existing applications do not allow using CT imaging as real-time intraoperative modality.
[0012] Recent advancement in Deep Neural Network (DNN) development enabled alternative approaches to traditional computed tomography reconstruction that allow reconstruction of CT imaging directly from a few or even a single-view X-ray. The examples of such are presented in open access Prior Art articles using MedNeRF architecture, such as MedNeRF: Medical Neural Radiance Fields for Reconstructing 3D-aware CT-Projections from a Single X-ray, Abril Corona-Figueroa and Co, Durham, UK, which Proposes a MedNeRF Deep Learning model that learns to reconstruct CT imaging from a few or even a single-view X-ray. This is based on a novel architecture that builds from neural radiance fields (NeRF), which learns a continuous representation of CT scans by disentangling the shape and volumetric depth of surface and internal anatomical structures from 2D images. Additionally, prior art such as US 20210393229 A1 , entitled, Single Or A Few Views Computed Tomography Imaging With Deep Neural Network, to Shen Liyue, Zhao Wei, and Xing Lei, discusses reconstructing a 3D volumetric image from the set of one or more 2D projection images using a residual deep learning network comprising an encoder network, a transform module and a decoder network. Thus it can be seen that the field of ML is rapidly evolving, with researchers globally developing novel architectural solutions regularly.
[0013] This invention describes the navigation and tracking of medical instrument on the realtime high-quality 3D CT imaging continuously and iteratively generated from a 2D X-ray images obtained using standard, widely accessible imaging equipment like a simple X-Ray devices or C-Arms. The proposed method stands to revolutionize the current practice by eliminating the need for both navigation systems based on registration of X-Ray with intraoperative or preoperative CT and, also, the use of CT scanners intraoperatively. The benefits of such an innovation include reduced costs, simplified workflow, minimized radiation exposure, and the expansion of interventional procedures to outpatient settings, ultimately democratizing theavailability of advanced medical procedures and broadening the scope of minimally invasive procedures.SUMMARY OF THE INVENTION
[0014] The scope of this invention encompasses a novel method and system for navigation of medical instrument using high-definition 3D images acquired from standard 2D X-ray images with standard X-Ray devices without the need for registration between intraoperative 2D to 3D imaging. This system is designed to overcome the need for 2D to 3D registration increasing the navigation accuracy, deficiencies of blurred or unclear views in Axial and Sagittal planes, which are critical for the real-time navigation and decision-making processes during surgeries. By improving upon existing technologies, this invention ensures that physicians can spatially localize their instruments during intervention procedures with great accuracy and high visual quality that facilitate better patient outcomes and allow for a broader range of procedures to be performed in various settings, including outpatient facilities.
[0015] One aspect of the invention provides a solution that transcends the limitations of existing art by enabling instrument navigation using clear spatial visualization on all anatomical planes, including Axial and Sagittal views, while the 3D image is reconstructed from 2D X-ray data. The invention aims to equip physicians with the ability to interactwith a real-time 3D representation of the patient's anatomy using their common instruments during interventional procedures, thereby greatly enhancing the precision and safety of surgical and interventional procedures.
[0016] Another aspect of the invention introduces a method of Medical Instrument Navigation using real-time 3D CT images generated from a set of 2D X-ray images.DEFINITIONS
[0017] For clarity and to ensure proper understanding of the terms used throughout this document, the following definitions are provided:
[0018] Medical Instrument: specialized tools or devices designed specifically for navigating, diagnosing, and treating conditions within complex anatomical areas, such as spine, knee, elbow, etc. These instruments include catheters, guidewires, implants, and surgical tools that allow for precise manipulation, delivery of therapeutic agents, or correction of anatomical structures.
[0019] Computed Tomography (CT): A medical imaging technique that uses computer- processed combinations of multiple X-ray measurements taken from different angles to produce cross-sectional images of specific areas of a scanned object, allowing the user to see inside the object without cutting.
[0020] C-Arm: A medical imaging device named for its C-shaped arm or Fluoroscope, used to connect the X-ray source and X-ray detector to one another, which allows for movement horizontally, vertically, and around the swivel axes, thus providing a variety of imaging angles.
[0021] Isocenter: The fixed point in space around which the X-ray tube and image intensifier (detector) of a C-arm rotate during imaging.
[0022] Isocentering: The technique of aligning the isocenter of the imaging equipment with the region of interest within the patient's body to maintain a consistent focal point throughout different imaging angles.
[0023] Artificial Intelligence (Al): A branch of computer science that aims to create systems capable of performing tasks that usually require human intelligence. This includes the ability to learn, reason, solve problems, perceive, and understand natural language.
[0024] Machine Learning (ML): A subset of artificial intelligence that involves the development of algorithms that can learn and make predictions or decisions based on data. Machine learning enables computers to identify patterns and learn from past experiences without being explicitly programmed.
[0025] Neural Network: A computational model inspired by the way neural networks in the human brain process information. It consists of interconnected nodes, or neurons, that process input data and can adapt and learn by adjusting the connections (weights) between these nodes.
[0026] Deep Neural Network (DNN): An advanced type of neural network that contains multiple layers of interconnected nodes. Each layer transforms the input data with the aim to extract increasingly higher-level features of the data fortasks such as image and speech recognition.
[0027] Convolutional Neural Network(CNN): A class of deep neural networks, most commonly applied to analyzing visual imagery. CNNs are modeled on the organization of the animal visual cortex and are designed to automatically and adaptively learn spatial hierarchies of features from input images. They achieve this through a multilayered architecture composedof convolutional layers that filter inputs for useful information, pooling layers that reduce dimensionality, and fully connected layers that interpret the feature data to make predictions or classifications. This structure enables CNNs to transform raw pixel data into incremental levels of abstract representation, making them highly efficient for tasks such as image recognition, object detection, and more.
[0028] Neural Radiance Fields (NeRF): A computational framework for rendering highly detailed 3D scenes from a set of 2D images. It uses a deep neural network to model the volumetric scene function, encoding both the color and density of points in 3D space as viewed from any angle. NeRF has been revolutionary in computer graphics and computer vision for its ability to produce photorealistic reconstructions and novel viewpoints of complex scenes with unprecedented detail and realism.
[0029] Medical Neural Radiance Fields (MedNeRF): An adaptation of the Neural Radiance Fields (NeRF) framework specifically tailored for medical imaging applications. It leverages the principles of NeRF to reconstruct and visualize high-fidelity 3D models of anatomical structures from medical imaging data, such as CT or MRI scans. MedNeRF aims to improve the accuracy and detail of medical visualizations, facilitating better diagnosis, surgical planning, and understanding of complex biological mechanisms by rendering intricate details of internal organs and tissues from various perspectives.
[0030] Pose of the C-Arm: The spatial configuration of the X-ray source in a C-Arm imaging system at a specific point in time, defined by its three-dimensional position and orientation relative to a fixed coordinate system. The pose comprises six degrees of freedom: three translational components (x, y, z) describing the position of the source in space, and three rotational components (0x, 0y, 0z) describing its angular orientation about each principal axis. Accurate estimation of the C-Arm source pose is critical for reconstructing volumetric (3D) images from multiple 2D fluoroscopic projections.
[0031] These terms will be used in the context of their definitions in the ensuing descriptions of the invention.BRIEF DESCRIPTION OF DRAWINGS
[0032] Fig. 1 is a flowchart of an embodiment of a method of the invention;
[0033] Fig. 2 is a perspective view of a C-arm being isocentered around the area of interest without preoperative CT according to an embodiment of a method of the invention;
[0034] Fig. 3 is an example of a multilayered system architecture based on residual DNN accordin to an embodiment of the invention; and,
[0035] Fig. 4 is an example of a multilayered system architecture based on MedNeRF accordingto an embodiment of the invention.DETAILED DESCRIPTION OF THE INVENTION
[0036] Referring now to the Figures, and first to Fig. 1 , there is shown a typical imaging suite 100 used to practice the method of the invention. Generally, the imaging suite 100 includes at least a C-arm 102, under which a patient 104 is placed on an operating table 106, which allows a region of i nterest 108 of the patient 104 to be placed at the isocenter 110 (denoted by a dotted line) of the C-arm 102. The patient 104 is lying supine on the operating table 106 with a skeletal overlay, indicating the Intraoperative 3D / CT Model 118 of the patient 104. This CT Model 118 is an adaptable, generic 3D representation of patient anatomy that may be used as a reference during different steps of a procedure such as setup, surgical navigation, treatment, and others.
[0037] A C-arm monitor 112 displays, as a result of the method of the invention, what appears to be a live X-ray image 113, offering real-time visual feedback during a procedure. Additionally, X-ray snapshots 114 are displayed next to the live X-ray image 113 on the c-arm monitor 112, showing different radiographic views captured by the C-arm 102.
[0038] METHOD 120 OF THE INVENTION
[0039] Fig. 2 shows an example of a method 120 of the invention that includes steps 122-128, each discussed in detail.
[0040] Step 122 - Isocentering
[0041] The method 120 begins, at step 122, with placing a region of interest of the patient at the isocenter 110 of the C-arm 102 (isocentering) using application guidance on the intraoperative CT model to ensure the targeted area is at the focal point of the imaging process.
[0042] Generally, isocentering a C-Arm during surgical interventions is a technique aimed at precisely positioning imaging equipment to capture the necessary views of a patient's anatomy with minimal radiation exposure and time. In complex procedures, such as those involving the spine or in trauma surgery, it is crucial to visualize anatomical structures relative to surgical tools and implants. Heretofore, achieving the correct radiographic views often required a skilled operator to maneuver the C-Arm through various angles, which could be both timeconsuming and expose patients and staff to additional radiation.
[0043] Methods have been developed by others in attempts to address these challenges that utilize preoperative CT scans to assist in the C-Arm positioning process. One example of such a method, which is discussed in Virtual Fluoroscopy For Intraoperative C-Arm Positioning And Radiation Dose Reduction. Tharindu De Silva, Journal of Medical Imaging, Jan 2018, employs an image-based registration technique that leverages the gantry's position encoders to guide the C-Arm into the desired position without additional hardware. While this approach integrates with standard surgical procedures and potentially reduces the time and radiation dose typically associated with manual C-Arm positioning, the method is image-based and thus relies upon an additional, preoperative image. An alternative approach utilizes virtual fluoroscopy, which is also generated from preoperative CT (id.).
[0044] The innovative approach of current invention completely obviates the need for a preoperative CT scan. Rather, the method provides the ability to guide the isocenter process using an intraoperative CT model. This approach involves using an adaptable generic anatomic model that can be created using at least one X-ray snapshot of the area of interest with known pose, or positions and orientation of the radiation source. Based on the at least one X-ray image, the anatomic model will be generated and automatically registered with the imaged object. Once established, the known image-based registration technique is used for isocentering.
[0045] Another innovative step of the current invention is the ability to change the isocenter virtually in the future within the field of view and the area of interest before the 3D tomographic reconstruction to achieve certain level of details and higher image quality of the anatomy reconstructed in the area of interest.
[0046] In summary, the use of simulated preoperative CT data in the current invention, based on the generic Al anatomy model, for isocentering aids in the effective and efficient positioning of the C-Arm, offering a solution that fits within the existing surgical workflow while aiming to improve outcomes and reduce risks associated with radiation exposure.
[0047] Step 124 - Acquire At Least One X-ray Image
[0048] After the isocentering step 122 is completed, at step 124 the X-ray 116 of the C-arm 102 is used to acquire at least one X-ray image of a known or roughly estimated pose that preferably includes six degrees of freedom (DOF) defining the angular and positional relationship of the imaging equipment relative to the imaged area or object. The C-Arm captures at least one 2D X- ray image or, alternatively, a series of 2D X-ray images from multiple directions. This task canbe completed by taking a single snapshot, distinct snapshots from various, preferably predetermined angles, or through continuous radiation exposure as the C-Arm moves around the target area. For each captured image, the C-Arm's pose is roughly estimated or provided. This information includes the position and orientation of the radiation source relative to the patient's body and, depending on reconstruction method, maybe pivotal for the accurate reconstruction of CT images. This data influences the transformation of two-dimensional images for integration into the 3D model generation process.
[0049] As background, a standard C-Arm device does not inherently provide pose information; therefore, supplementary methods are needed. These may include external tracking systems, mechanical boards equipped with metal markers, algorithms for image registration, or the use of positional encoders attached to the C-Arm. External tracking may use optical, electromagnetic, or electromechanical sensors. Mechanical boards are designed to use 2D or 3D patterns of markers that can be automatically detected on X-ray images for further processing. These methods for pose calculation are well-documented in existing literature and prior art. The advantage of the present invention is that continuous instrument navigation and tracking becomes an integral part of the process.
[0050] Analog C-Arms, which typically use image intensifiers, necessitate distortion calibration to address inherent pincushion or S-distortions. Distortion calibration involves capturing images of a known grid or pattern to map distortions and correct new images, a critical step for accurately depicting the imaged area.
[0051] The calibration grid serves a dual purpose in medical imagingwith C-Arm fluoroscopy. It is used for distortion calibration and as an aid for C-Arm pose estimation. The grid's known geometry and pattern provide reference points in fluoroscopic images that software algorithms can use to deduce the C-Arm's pose during imaging, streamlining the workflow by combining calibration and pose estimation into a single step.
[0052] By calculating the C-Arm's pose, the imaging system can align each 2D image with its spatial orientation, foundational for the Al-driven reconstruction process. These calculated poses allow for the precise merging of images to form a cohesive 3D visualization critical for both diagnostic and interventional use.
[0053] An innovation of this invention lies in determining the optimal number and orientation of 2D X-ray images required for Al-driven CT reconstruction, depending on the precision needs of each medical application when reconstructed CT image and 3D models will be used. Thisinnovation is pivotal in enhancingthe efficiency of the imaging process by reducing the number of images and thus exposure to radiation, while ensuring the adequate quality of the reconstructed CT per specific medical application requirements.
[0054] Another alternative non-limiting solution involves generating limited angle tomography from a series of 2D X-ray images using a traditional approach and such is used as an input for the ML model in the following step for certain or all layers of the ML model.
[0055] Step 126 -2D to 3D Transformation
[0056] Next, at step 126, the method 120 involves performing a 2D to 3D transformation on the image of step 124 using an ML Pipeline, which is and end-to-end system that handles the full flow of data from input to final output. Figs. 3 and 4 each depict an ML Pipeline - 3000 and 4000, respectfully. The process of transforming 2D X-ray data into 3D CT images in this invention involves a hierarchical ML model specially designed for CT imaging generated from a single 2D X-ray snapshot or from a small or limited number of 2D X-ray images or projection views. This deep learning approach operates within an encoder-decoder framework, a representation-generation model that learns and applies the complex relationship between the 2D X-ray and 3D CT data.
[0057] One innovative aspect of the present invention lies in the multiple processing layers within the system as shown in Fig. 3 and Fig. 4. The advantage of utilizing multiple layers is the ability to recognize different types of instances from the input image. The non-limiting examples of such instances are static patient anatomy such as vertebrae or rib, blood vessels that are temporarily highlighted by the injected contrast and moving or static medical instruments within the body. Each layer is methodically tailored to recognize, enhance, and reconstruct these different instances. This specialization is achieved through separate configuration and architecture of each layer, ensuring that they are finely tuned to the unique characteristics of the respective instances including their static or dynamic nature. Some medical applications may use just a single layer of representation network while others may use multiple layers. The architecture of each layer is detailed below.
[0058] For example, attention is drawn to the ML Pipelines 3000, 4000 shown in Figs. 3 and 4. Fig. 3 shows the progression of a multilayer system architecture based on Residual DNN. Fig. 4 shows the progression of a multilayer system architecture based on MedNeRF. First, in each instance layer (e.g. static anatomy layer, dynamic anatomy layer, surgical tools layer, etc.), a 2D X-ray image 300 or 400, respectively, is pre-processed at 310 or 410 respectively, toemphasize the instances of the input image and / or reduce noise. The processed images then pass through a tailored ML model 320 or 420, respectively that allows an accurate 3D reconstruction of the instance. The details of these models are discussed in more detail below.
[0059] In the final stage, the instance representations 330 and 430 from Figs. 3 and 4, respectively, of each layer, specifically the 3D Models 332, 432 and the CTVolumes 334, 434, are merged using a novel model-weighing approach 350, 450 that allows the reconstruction of a high-resolution CT-like volume as well as a combined 3D view of the underlying instances.
[0060] The innovative aspects of the ML model 320, 420 include its ability to decipher higherdimensional information from multiple 2D projections, whereas a single X-ray image does not necessarily have the information required to reconstruct a 3D volume because it may be obscured to vital features. The ML model 320, 420 overcomes this limitation by employing an iterative approach, where the output generated at each instance layer is refined with the inclusion of additional X-ray images. As previously mentioned, the nonlimiti ng examples of the instance layer model may entail an encoder-decoder architecture, where the representation network serves as the encoder, compressing the input into a dense feature representation, where each layer of the generation network acts as the decoder for the correspondent layer of the representation network (Fig. 3) or, alternatively, use MedNeRf to apply DDN to model the volumetric scene function, encoding both the color and density of points in 3D space (Fig. 4), reconstructing the data into the 3D Model 332, 432 and the corresponding CT image 334, 434.
[0061] The detailed architecture of residual DNN architecture ML pipeline 3000 instance layer is shown in Fig. 3 and described as follows:
[0062] The layer-specific X-ray image preprocessing aims to highlight the relevant instances in the acquired images per the layer requirements at 310. Such non-limiting examples of processing may include thresholding, low or high-pass filtering, histogram equalization, etc. These images form the raw data that the ML model 320 will process.
[0063] The representation network (CNN) 322 of the ML model receives processed images as inputs. It then undergoes multiple convolutional layers that extract the semantic features necessary for the subsequent reconstruction of the instance at 324.
[0064] The features identified in step 324, remain in 2D space. The transformation module, at 326, reshapes the features into a representation that is conducive to the instance generation phase.
[0065] The generation network 328 takes as input the transformed images from 326 and uses them to reconstruct a 3D model 332 of the instance at 330. The non-limiting examples of 3D model representation could be voxel-based, solid, polynomial, NURBS, and others. Also at 330, the generation network may use a series of 3D deconvolutional layers to upscale the feature maps and output a volume 334. It may also include elements of residual learning or skip connections to retain fine-grained details necessary for high-quality reconstruction. Other examples of such generative models can be Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs).
[0066] The output of the generation network is a 3D model 332 and a CT volume 334 that progressively improves with each iteration.
[0067] The detailed architecture of MedNeRf DNN architecture ML pipeline 4000 instance layer is shown in Fig. 4 and described as follows:
[0068] The layer-specific X-ray image preprocessing aims to highlight the relevant instances in the acquired images per the layer requirements 410. Such non-limiting examples of processing may include thresholding, low or high-pass filtering, histogram equalization, etc. These images form the raw data thatthe ML model 420 will process.
[0069] Positional encoding 422 is applied to enable the neural network to better capture high- frequency details by transforming the input coordinates (e.g., spatial positions) into a higherdimensional space. This is achieved by applying a set of sinusoidal functions with different frequencies to each input coordinate, thereby allowing the model to differentiate between positions more effectively with subsequent input of this information into MedNeRF DNN 424.
[0070] At its core, MedNeRF DNN 424 represents a continuous volumetric scene. This neural network predicts the color and density of points in 3D space, given their coordinates and viewing direction. This representation is particularly suited for medical images, where capturing the intricate details of anatomical structures is crucial.
[0071] A Structure Activation Function 426 is applied to emphasize anatomically relevant regions during the volumetric representation process. This function acts as a spatially aware gating mechanism that selectively enhances features within the neural network based on their anatomical salience, such as bone, soft tissue, or implants. By weighting internal feature activations using learned structural relevance cues, the network more effectively captures the fine-grained variations criticalfor medical reconstruction.
[0072] MedNeRF employs volume rendering techniques 428 to synthesize 2D projections (views) from the 3D neural representation. This involves casting rays through the volume and accumulating color and opacity values along each ray based on the neural network's predictions. This process is key for generating views from any desired perspective, not present in the original dataset.
[0073] The output 430 of the volume rendering 428 is a 3D model 432 and a CT volume 434 that progressively improve with each iteration.
[0074] The ML model 420 is trained using a large number of paired X-ray images and instance 3D models, as well as simulated data. Since the training of such a network entails a large number of X-ray images some may be simulated computationally.
[0075] In this invention, each instance layer may have a different architecture and the generation of a Combined 3D Model 360, 460 and Combined CT volume 370, 470 from multiple 3D Models 332, 432 and CT volumes 334, 434 is achieved through a sophisticated process of Weighted 3D Model and CT Volume Composition 350 / 450 that involves a weighted averaging technique at the voxel level. Here's a conceptual explanation of the process applied for CT Volume:
[0076] A separate CT volume 334, 434 is generated for each layer using the ML model's multilayer structure. Each volume provides a three-dimensional representation of specific instances.
[0077] The CT volumes 334, 434 are composed of numerous tiny units called voxels, analogous to pixels in a 2D image but extending into three dimensions. Each voxel in the volume contains data about the structure at that point in space.
[0078] To merge these volumes, a weight is assigned to each voxel in each volume. The weights are determined based on the relevance and accuracy of the information that a particular voxel represents, which could be influenced by factors such as tissue density, the clarity of the feature, the confidence of the ML model 320, 420 in its reconstruction, movement, and the clinical significance of the anatomical structure.
[0079] For each voxel position in the combined volume, the corresponding voxels from each separate CT volume 334, 434 are taken, and a weighted average is calculated. This means that the final value of a voxel in the combined CT volume is the sum of the values of that voxel in all the individual volumes, each multiplied by its assigned weight and then divided by the sum of the weights.
[0080] The process is repeated for every voxel in the 3D space, resulting in a single, combined realistic CT volume. This volume now contains a harmonious blend of the anatomical features and instances, with the contribution of each feature to the final image determined by its assigned weights.
[0081] Finally, the combined volume may be further refined to enhance the visibility of critical structures, and parts of medical instruments or to suppress artifacts. This involves additional post-processing steps such as smoothing, thresholding, or applying advanced imaging filters.
[0082] Combining separately generated 3D models within the same coordinate system for display in a single scene involves importing each model into a shared environment and positioning them relative to one another before rendering the composite scene. This process is managed by tools provided in 3D graphics software and standard graphic engines.
[0083] Repeat Steps 124 and 126
[0084] Steps 124 and 126 are repeated several times, each iteration being either a refinement of the previous iteration or a new transformation depending on the application needs.
[0085] Step 128 - Display Instrument on CT Volume and 3D Model
[0086] Step 128 involves displaying a medical instrument being used to perform a medical procedure on the patient 104 on the CT Volume and 3D Model on the c-arm monitor 112 to facilitate user interaction with the 3D anatomy. While the user is using the C-arm generated X- ray images to navigate the medical instrument and access the anatomy of the spine (for example) in a clinically appropriate position, the ML Model will continue to update the 3D model progressively and continuously. This visual format can be easily analyzed by medical professionals. For instance, the software presents Axial, Sagittal, and Coronal views, Multiple Intensive projection images or a three-dimensional image that can be manipulated on-screen, allowing for rotation, zooming, and slicing through different planes to examine various angles and layers of the anatomy. This visualization depends on the specific medical application and aids clinicians in diagnosis, surgical planning, and patient education. If desired, the tip of the instrument may be used by doctor as an interaction mean with CT Volume and 3D Model.
Claims
CLAIMS1 . A method of generating three-dimensional (3D) computed tomography (CT) images, useable for navigating medical instruments, comprising: isocentering a c-arm around a region of interest using application guidance on an intraoperative CT model to ensure a targeted area is at a focal point of an imager of the c- arm; acquiring at least one X-ray image of a known or estimated pose; performing a two-dimensional (2D) to 3D transformation using a machine learning (ML) pipeline to create at least one of a CT volume or a 3D model; repeating the acquiring and performing steps until a desired image accuracy of the at least one of the CT volume or the 3D model is generated; displaying the at least one of the CT volume or the 3D model on a monitor along with a realtime image of a medical instrument being used to perform a procedure in the region of interest.
2. A method for reconstructing a three-dimensional (3D) computed tomography (CT)-like volume from intraoperative two-dimensional (2D) X-ray images, comprising:(a) acquiring at least one intraoperative 2D X-ray image of a patient’s anatomical region of interest using a conventional X-ray imaging device;(b) pre-processing the at least one 2D X-ray image to emphasize anatomical and instrument instances and reduce noise;(c) inputting the pre-processed 2D X-ray images into a machine learning (ML) model comprising multiple instance-specific layers, wherein each layer is tailored to reconstruct a partial 3D volume for distinct anatomical or instrument instances;(d) generating, using each instance-specific layer, an instance-specific 3D model and CT volume representation;(e) merging the instance-specific 3D volumes from the multiple layers using voxel-level weighted averaging, wherein weights assigned to each voxel are determined based on one or more of tissue density, feature clarity, reconstruction confidence, movement artifacts, or clinical significance;(f) iteratively refining the reconstructed 3D CT-like volume in real-time by integrating additional intraoperative 2D X-ray images into the ML model, continuously updating the 3D CT-like volume;(g) outputting a high-resolution combined 3D CT-like volume suitable for intraoperative visualization without requiring any preoperative CT or external reference system.
3. The method of claim 2, wherein the pre-processing step comprises at least one of thresholding, low-pass filtering, high-pass filtering, or histogram equalization.
4. The method of claim 2, wherein at least one instance-specific layer comprises a residual deep neural network (DNN) architecture, said architecture including: a convolutional neural network (CNN) for extracting semantic features from pre-processed images; a feature transformation module configured to reshape extracted features; and a generation network configured to reconstruct the 3D model and CT volume from reshaped features.
5. The method of claim 4, wherein the generation network comprises one or more of 3D deconvolutional layers, residual learning blocks, skip connections, generative adversarial networks (GANs), or variational autoencoders (VAEs).
6. The method of claim 2, wherein at least one instance-specific layer comprises a Medical Neural Radiance Fields (MedNeRF) architecture configured to encode density and color information of anatomical structures in continuous 3D space, using positional encoding and network optimization emphasizing anatomically relevant regions.
7. The method of claim 2, further comprising refining the combined CT-like volume by applying additional post-processing including one or more of smoothing, thresholding, or advanced image filtering techniques.
8. A method of navigating a surgical instrument, comprising: (a) reconstructing a 3D CT- like volume according to the method of claim 1 ; (b) tracking a position and orientation of the surgical instrumentwithin the patient’s anatomical region of interest; (c) continuously integrating the tracked position of the surgical instrumentwithin the reconstructed 3D CT- like volume; and (d) displaying the surgical instrument’s position in real-time within the reconstructed 3D CT-like volume to guide an interventional procedure.
9. The method of claim 8, wherein the surgical instrument itself serves as an interactive tool, enabling the physician to manipulate or interact with the displayed 3D CT-like volume, further enhancing intraoperative decision-making.
10. A surgical navigation system comprising: (a) a conventional intraoperative X-ray imaging device; (b) a computing unit comprising at least one processor and memory configured to perform the method according to claim 1 to reconstruct a high-resolution 3D CT-like volume from intraoperative X-ray images; (c) an instrument tracking subsystem configured to track and output real-time position data of a surgical instrument; and (d) a display unit configured to visualize the reconstructed 3D CT-like volume and the tracked surgical instrument’s position in real-time for intraoperative surgical navigation.11 . The surgical navigation system of claim 10, wherein the computing unit further comprises software instructions to iteratively update the 3D reconstruction in real-time as additional X-ray images are acquired, without requiring a preoperative CT scan or external spatial registration.
Citation Information
Patent Citations
Method and apparatus for instrument tracking on a scrolling series of 2D fluoroscopic images
US20050169510A1
Method for reconstructing a 3D image from 2d x-ray images
US20170164919A1
Systems and methods for reconstruction of 3D anatomical images from 2d anatomical images
US20200334897A1
A machine learning model to adjust c-arm cone-beam computed tomography device trajectories
US20220323025A1
System and method for local three dimensional volume reconstruction using a standard fluoroscope
US20230157652A1