Systems and methods for robotic endoscopy systems utilizing tomosynthesis and enhanced fluoroscopy
The tomosynthesis-based method addresses CT2BD in robotic bronchoscopy by generating a 3D image and augmented fluoroscopy overlay, improving tool placement accuracy and reducing procedure time.
Patent Information
- Application Number
- JP2025527079
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-18
- Filing Date
- 2023-11-13
- Publication Date
- 2025-12-16
AI Technical Summary
Conventional robotic bronchoscopy systems face challenges with CT-to-body discrepancy (CT2BD) during navigation, leading to increased procedure length and non-diagnostic outcomes due to inadequate intraoperative real-time correction, particularly in determining tool position relative to lesions using tomosynthesis with poor depth resolution.
A tomosynthesis-based method that provides improved accuracy by generating a reconstructed 3D image and augmented fluoroscopy overlay, utilizing markers and pose estimation to confirm tool position within lesions, enhancing navigation and visualization.
Improves the accuracy and efficiency of determining tool placement within lesions by providing quantitative spatial relationships, reducing procedural time and enhancing diagnostic yield.
Smart Images

Figure 2025540630000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims priority to U.S. Provisional Patent Application No. 63 / 384,312, filed November 18, 2022, which is incorporated herein by reference in its entirety. [Background technology]
[0002] Early diagnosis of lung cancer is crucial. Lung cancer remains the most lethal form of cancer, resulting in over 150,000 deaths annually. Compared with computed tomography (CT)-guided transthoracic needle aspiration (CT-TTNA), navigation bronchoscopy has a better safety profile (lower risk of pneumothorax, life-threatening bleeding, and length of stay) and the ability to stage the mediastinum, but is associated with a lower diagnostic yield. Endoscopy (e.g., bronchoscopy) can involve accessing and visualizing the inside of a patient's internal cavity (e.g., the trachea) for diagnostic or therapeutic purposes. During the procedure, a flexible tubular tool, such as an endoscope, can be inserted into the patient's body, and instruments can be passed through the endoscope to the tissue site identified for diagnosis or treatment.
[0003] Robotic bronchoscopy systems have garnered interest for biopsy of peripheral pulmonary lesions. Robotic platforms offer superior stability, distal articulation, and visualization compared to conventional pre-curved catheters. Some conventional robotic bronchoscopy systems utilize shape-sensing technology (SS) for guidance. SS catheters may have embedded fiber-optic sensors that measure the catheter's shape hundreds of times per minute. Other conventional robotic bronchoscopy systems incorporate direct visualization, optical pattern recognition, and geographic location sensing (OPRGPS) for guidance. Both SS and OPRGPS systems utilize pre-planned CT scans to create electronically generated virtual targets. However, SS and OPRGPS systems can be prone to CT-to-body discrepancy (CT2BD). CT2BD is the discrepancy between the electronic virtual target and the actual anatomical location of peripheral pulmonary lesions. CT2BD can occur for a variety of reasons, including atelectasis, anesthesia-induced neuromuscular weakness, tissue distortion from the catheter system, bleeding, ferromagnetic interference, and perturbations in the anatomy, such as pleural effusion. Neither the SS system nor the OPRGPS platform has intraoperative real-time correction of CT2BD, which, in particular, can increase the length of the procedure, disturb the operator, and ultimately result in a non-diagnostic procedure. Summary of the Invention
[0004] Digital tomosynthesis algorithms have been introduced in recent years to correct for CT2BD. Tomosynthesis (also referred to as "tomo") is limited-angle tomography, as opposed to full-angle (e.g., 180-degree) tomography. However, tomosynthesis reconstructions do not have uniform resolution. For example, resolution is often lowest in the depth direction. The standard method of displaying a 3D volume dataset through three orthogonal planes (e.g., axial, sagittal, and coronal) can be ineffective because two of the planes have lower resolution. A common method for viewing a tomosynthesis volume is to scroll through the depth direction, where each slice has good resolution. In the case of pulmonology, the coronal plane is viewed and then scrolled through the anterior-posterior (AP) direction. However, this has led to difficulties in determining the spatial relationships of structures in the depth direction. It can be difficult to determine whether a tool (e.g., a biopsy needle) is inside a lesion in the AP direction of a chest tomosynthesis reconstruction.
[0005] There is a need for methods and systems that can determine whether a tool is within a target (e.g., a lesion) with improved accuracy or efficiency. The present disclosure addresses the above need by providing a tomosynthesis-based intra-lesion tool determination method with improved accuracy and efficiency. In particular, the provided method can provide a user with quantitative information about the spatial relationship between a thin tool and a target region (e.g., a lesion) in the depth direction. The methods, systems, computer-readable media, and techniques herein can identify the (depth-wise) positional relationship between the tool and the lesion by separately identifying their depths, and quantitatively determine whether the (thin) tool is within the lesion.
[0006] The methods herein may be applied after the robotic platform is set up, target lesions are identified and / or segmented, airway registration is performed, and individual target lesions are selected. The methods herein may be applied during or after the navigation process to identify the position of a portion of the tool relative to the target. Endoscopy navigation systems may use different sensing modalities (e.g., camera imaging data, electromagnetic (EM) position data, robot position data, etc.). In some cases, the navigation approach may rely on an initial estimate of where the tip of the endoscope is relative to the airway to begin tracking the tip of the endoscope. Some endoscopy techniques may involve a three-dimensional (3D) model of the patient's anatomy (e.g., CT images) and guided navigation using EM field and position sensors.
[0007] In some cases, 3D images of a patient's anatomy may be captured one or more times for various purposes. For example, prior to a medical procedure, a 3D model of the patient's anatomy may be created to identify a target location. In some cases, the precise alignment (e.g., registration) between the virtual space of the 3D model, the physical space of the patient's anatomy represented by the 3D model, and the EM field may be unknown. Therefore, the endoscope position within the patient's anatomy cannot be accurately mapped to the corresponding position in the 3D model before generating the registration. In another example, during a surgical procedure, 3D imaging may be performed to update / confirm the location of the target (e.g., lesion) in case of target issues or lesion migration.
[0008] In some cases, fluoroscopic imaging systems may be used to determine the position and orientation of medical instruments and patient anatomical structures within the coordinate system of the surgical environment via fluoroscopy (also referred to as "fluoro"). Fluoroscopy is a method that provides real-time X-ray imaging. For imaging data to aid in the precise localization of medical instruments, the imaging system's coordinate system may be required to reconstruct a 3D model. For example, to better visualize and provide the 3D coordinates of anatomical structures, multiple 2D fluoroscopic images can be used to create a tomosynthesis or cone-beam CT (CBCT) reconstruction. During a CBCT scan, the CBCT scanner can acquire projections along a 180° to 360° angle of rotation (i.e., a full rotation of the x-ray source and detector) across the region of interest to obtain a volumetric data set. Scanning software collects the data, reconstructs it, and generates a digital volume composed of 3D voxels of anatomical data, which can then be manipulated and visualized. Tomosynthesis is similar to CBCT scanning but uses a limited rotation angle (e.g., 15-60 degrees) and therefore has a reduced scan time compared to CBCT. Tomosynthesis has an additional benefit over CBCT in that the limited range of motion required for tomosynthesis allows it to be used in more constrained patient environments where full 360° access around the patient is difficult to achieve during a procedure. Tomosynthesis can be performed to determine the position and orientation of medical instruments and patient anatomy. However, traditional tomosynthesis has poor depth resolution (AP direction), creating difficulties in determining whether a tool is within a target region (e.g., a lesion) or the position of a thin tool relative to the target region. The systems, methods, and techniques herein beneficially provide tool intrusion confirmation in a quantitative manner, thereby improving the accuracy and precision of locating a tool (e.g., a needle) relative to a target region. As used herein, the term CBCT may also refer to tomosynthesis, which is used interchangeably throughout this specification unless the context suggests otherwise.
[0009] As described above, tomosynthesis or CBCT reconstruction of an anatomical structure involves combining data from images of 2D projections taken at multiple angles relative to the anatomical structure and combining the multiple 2D images to reconstruct a 3D view of the anatomical structure. The mathematical process of combining the 2D projections to create the 3D view requires as input the relative pose (angle and position) of the camera from which each of the 2D projections is recorded. In some cases, the methods herein may employ pose estimation methods to obtain the relative pose of the camera. For example, the relative pose of the camera may be obtained by using features within the image itself. In some instances, when markers (e.g., an array of artificial markers with known positions, or natural features such as bones) are captured in the image, the relative positions of the markers relative to each other in the 2D projections may be processed using computer vision methods to estimate the pose of the camera in a 3D world reference frame. In other cases, the pose of the camera from which each of the 2D projections is recorded may be obtained from independent measurements of the camera position and orientation (e.g., an accelerometer, IMU, separate imaging device, or other orientation sensor). The present disclosure can utilize the methods described above to generate 3D diagram constructions from combinations of 2D projections.
[0010] In some cases, features identified from tomosynthesis or CBCT images acquired after intubation of the patient but before the start of bronchoscopy may be utilized to generate augmented fluoroscopy images. Augmented reality has previously been associated with improvements in diagnostic accuracy, procedure time, and radiation dose in biopsies. Specifically, augmented fluoroscopy can be utilized to reduce radiation exposure without compromising diagnostic accuracy. Augmented fluoroscopy can display an enhanced layer of information on top of a live fluoroscopic view.
[0011] In one aspect of the present disclosure, a computer-implemented method for an endoscopic device is provided, the method including: (a) providing a first graphical user interface (GUI) for a tomosynthesis mode and a second GUI for a fluoroscopy view mode for viewing a portion of an endoscopic device and a target within an object; (b) receiving a sequence of fluoroscopy image frames including a portion of the endoscopic device, a marker, and the target, the sequence of fluoroscopy image frames corresponding to various orientations of an imaging system acquiring the sequence of fluoroscopy image frames; (c) upon switching to the tomosynthesis mode, i) performing a uniqueness check on the sequence of fluoroscopy image frames and ii) generating a reconstructed 3D tomosynthesis image based at least in part on the imaging system orientation estimated using the marker; and (d) upon switching to the fluoroscopy view mode, i) generating an estimated pose of the imaging system associated with the fluoroscopy image frames from the sequence of fluoroscopy image frames based at least in part on the markers included in the fluoroscopy image frames, and ii) generating an overlay of the target displayed on the fluoroscopy image frames based at least in part on the estimated pose.
[0012] In a related yet separate aspect, a non-transitory computer-readable medium is provided that stores instructions that, when executed by at least one processor, cause the at least one processor to (a) provide a first graphical user interface (GUI) for a tomosynthesis mode and a second GUI for a fluoroscopy view mode for viewing a portion of an endoscopic device and a target within a subject, (b) receive a sequence of fluoroscopy image frames including a portion of the endoscopic device, a marker, and the target, the sequence of fluoroscopy image frames corresponding to various orientations of an imaging system that acquire the sequence of fluoroscopy image frames, and (c) switch to the tomosynthesis mode. When the fluoroscopy view mode is selected, the system performs operations including: i) performing a uniqueness check on the sequence of fluoroscopic image frames; and ii) generating a reconstructed 3D tomosynthesis image based at least in part on the imaging system pose estimated using the markers; and (d) when switching to a fluoroscopy view mode, i) generating an estimated imaging system pose associated with the fluoroscopic image frames from the sequence of fluoroscopic image frames based at least in part on the markers included in the fluoroscopic image frames; and ii) generating a target overlay to be displayed on the fluoroscopic image frames based at least in part on the estimated pose.
[0013] In some embodiments, the uniqueness check is not performed in fluoroscopic view mode. In some embodiments, the uniqueness check includes determining whether a fluoroscopic image frame from a sequence of fluoroscopic image frames is unique based at least in part on an intensity comparison.
[0014] In some embodiments, the marker has a 3D pattern. Optionally, the marker includes a plurality of features arranged on at least two different planes. In some embodiments, the marker has a plurality of features of different sizes arranged in a coding pattern. Optionally, the coding pattern includes a plurality of subareas, each having a unique pattern. Optionally, in a tomosynthesis mode, the pose of the imaging system is estimated by matching a patch of a plurality of features in a sequence of fluoroscopic image frames to the coding pattern. In some examples, the method further includes identifying one or more fluoroscopic image frames having a high pattern match score. Optionally, in a fluoroscopic view mode, the estimated pose of the imaging system is generated by matching a patch of a plurality of features in the fluoroscopic image frames to the coding pattern.
[0015] In some embodiments, the first GUI is configured to display a reconstructed 3D tomosynthesis image and receive user input on the reconstructed 3D tomosynthesis image indicating a location of the target. In some cases, the second GUI displays a fluoroscopic image frame with an overlay of the target, the location of the target displayed on the fluoroscopic image frame being based at least in part on the location of the target.
[0016] In some embodiments, the shape of the overlay is based at least in part on a 3D model of the target projected onto the fluoroscopic image frame based on the estimated pose. In some cases, the 3D model is generated based on a computed tomography image. In some embodiments, the second GUI provides a graphical element for enabling or disabling the display of the overlay.
[0017] In another aspect, the systems, methods, and computer-readable media of the present disclosure may implement operations including: (a) navigating an endoscopic device toward a target within a subject in a navigation mode of a graphical user interface (GUI), the GUI displaying a virtual view having visual elements to guide navigation of the endoscopic device; (b) upon switching to a tomosynthesis mode of the GUI, i) receiving a sequence of fluoroscopic image frames corresponding to various orientations of an imaging system that includes a portion of the endoscopic device and the target and that acquires the sequence of fluoroscopic image frames, ii) generating a reconstructed 3D tomosynthesis image based at least in part on the orientation of the imaging system, and iii) determining a position of the target based at least in part on the reconstructed 3D tomosynthesis image; and (c) upon switching to a fluoroscopic view mode of the GUI, i) obtaining an imaging system orientation associated with the fluoroscopic image frames acquired in the fluoroscopic view mode, and ii) generating a target overlay to be displayed on the fluoroscopic image frames based at least in part on the orientation of the imaging system and the position of the target determined in (b).
[0018] In some embodiments, the virtual view in the navigation mode includes rendering a graphical representation of the target and an indicator showing the angle of the target relative to the exit axis of the working channel of the endoscopic device when the distal tip of the endoscopic device is determined to be within a predetermined proximity of the target. In some embodiments, the position of the target displayed in the navigation mode is updated based on the position of the target determined in (b). In some embodiments, the pose of the imaging system in the tomosynthesis mode is estimated using markers included in the sequence of fluoroscopic image frames. In some embodiments, the pose of the imaging system in the tomosynthesis mode is measured by one or more sensors.
[0019] In some embodiments, the pose of the imaging system relative to the fluoroscopic image frames in a fluoroscopic view mode is estimated using markers included in the fluoroscopic image frames. In some cases, the markers have a 3D pattern. In some examples, the markers include multiple features arranged on at least two different planes. In some cases, the markers include multiple features of different sizes arranged in a coding pattern. In some examples, the coding pattern includes multiple subareas, each with a unique pattern. In some examples, the pose of the imaging system is estimated by aligning patches of multiple features in the fluoroscopic image frames to the coding pattern.
[0020] In some embodiments, a pose of the imaging system relative to the fluoroscopic image frames in a fluoroscopic view mode is measured by one or more sensors. In some embodiments, in a tomosynthesis mode, the sequence of fluoroscopic image frames is processed by performing a uniqueness check on the sequence of fluoroscopic image frames. In some embodiments, the uniqueness check further includes determining whether a fluoroscopic image frame from the sequence of fluoroscopic image frames is unique based at least in part on an intensity comparison.
[0021] In some cases, the systems, methods, and computer-readable media of the present disclosure may implement operations including (a) receiving instructions to present, on one or more graphical displays, one or both of one or more tomosynthesis reconstructions or one or more enhanced fluoroscopy overlays. The tomosynthesis reconstructions may be generated by (i) acquiring, via a first imaging device of the one or more imaging devices, one or more tomosynthesis images over a region of interest of a patient, where at least a portion of the tomosynthesis images over the region of interest includes first image data corresponding to a plurality of markers, and the tomosynthesis images include a plurality of depth-stacked tomosynthesis slices, and (ii) generating a tomosynthesis reconstruction based on the tomosynthesis images and the plurality of markers, where the tomosynthesis reconstruction includes the tomosynthesis images. The augmented fluoroscopy overlay is generated by (i) acquiring one or more fluoroscopy images across a region of interest of a patient, at least a portion of the fluoroscopy images across the region of interest including second image data corresponding to a plurality of markers, the fluoroscopy images including a plurality of depth-stacked fluoroscopy slices; (ii) generating an augmented fluoroscopy overlay based on the fluoroscopy images and the plurality of markers, the augmented fluoroscopy overlay including the augmented fluoroscopy image; and (b) in response to receiving a command, causing one or more graphical displays to present one or both of the tomosynthesis reconstruction or the augmented fluoroscopy overlay.
[0022]
[0013] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0023] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In the event that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede or supersede any such conflicting material. [Brief explanation of the drawings]
[0024] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figure" and "FIG."). [Figure 1] FIG. 1 illustrates an exemplary process for tomosynthesis image reconstruction. [Figure 2] FIG. 2 illustrates an exemplary process for generating an enhanced fluoroscopic overlay. [Figure 3] FIG. 3 is a diagram illustrating an exemplary system of various state machines. [Figure 4] FIG. 4 shows an example of a configuration state machine. [Figure 5] FIG. 5 shows exemplary state machine logic. [Figure 6A] 6A-6C show exemplary tomosynthesis board marker designs. [Figure 6B] 6A-6C show exemplary tomosynthesis board marker designs. [Figure 6C] 6A-6C show exemplary tomosynthesis board marker designs. [Figure 7A] Figure 7A shows an example of blob detection of markers on an image of the tomosynthesis board. [Figure 7B]FIG. 7B shows examples of candidate points on an image of the tomosynthesis board. [Figure 7C] Figure 7C shows an example of marker extraction on an image of the tomosynthesis board. [Figure 8] FIG. 8 illustrates an exemplary process for robust tomosynthesis marker alignment. [Figure 9] FIG. 9 shows an example result of marker tracking across a sequence of tomosynthesis frames on an image of the tomosynthesis board. [Figure 10] FIG. 10 shows an example of camera pose estimation. [Figure 11] FIG. 11 shows an example of an extended fluoroscopic projection. [Figure 12] FIG. 12 illustrates an example of a robotic bronchoscopy system according to some embodiments of the present invention. [Figure 13] FIG. 13 shows an example of a fluoroscopy (tomosynthesis) imaging system. [Figure 14] 14 and 15 show examples of flexible endoscopes. [Figure 15] 14 and 15 show examples of flexible endoscopes. [Figure 16] FIG. 16 shows an example of an instrument drive mechanism that provides a mechanical interface to the handle portion of a robotic bronchoscope. [Figure 17] FIG. 17 shows an example of the distal tip of an endoscope. [Figure 18] FIG. 18 shows an exemplary distal portion of a catheter with integrated imaging and illumination devices. [Figure 19] FIG. 19 shows an example of a user interface with a tomosynthesis dashboard. [Figure 20] FIG. 20 shows an example of a user interface with a C-arm settings dashboard. [Figure 21] FIG. 21 shows an example of a user interface with a scope selection dashboard. [Figure 22]FIG. 22 shows an example of a user interface with a selection crosshair panel. [Figure 23] FIG. 23 shows an example of a user interface with a lesion selection dashboard. [Figure 24] FIG. 24 shows an example of a user interface with an enhanced fluoroscopy panel. [Figure 25] FIG. 25 shows an example of a user interface for driving or navigating an endoscope. [Figure 26] FIG. 26 shows an example of a virtual endoluminal view displaying a target. [Figure 27] FIG. 27 illustrates a computer system that is programmed or otherwise configured to implement the methods provided herein. [Figure 28] FIG. 28 shows an example of a method for presenting either or both of a tomosynthesis reconstruction or an enhanced fluoroscopic overlay. DETAILED DESCRIPTION OF THE INVENTION
[0025] Although the exemplary embodiments are primarily directed to tomosynthesis, enhanced fluoroscopy, bronchoscopy, and the like, those skilled in the art will understand that this is not intended to be limiting and that the systems, methods, and techniques described herein may be used for other therapeutic or diagnostic procedures and in other anatomical regions of a patient's body, such as the digestive system, including but not limited to the esophagus, liver, stomach, colon, urinary tract, or the respiratory system, including but not limited to the bronchi, lungs, and various others.
[0026] The embodiments disclosed herein can be combined in one or more of many ways to provide improved diagnosis and treatment to patients. The disclosed embodiments can be combined with existing methods and devices to provide improved treatment, such as by combining with known methods of pulmonary diagnosis, surgery, and surgery of other tissues and organs. It should be understood that any one or more of the structures and steps described herein can be combined with any one or more additional structures and steps of the methods and devices described herein, and the figures and supporting text provide an explanation according to the embodiments.
[0027] Although the treatment plans and definitions of diagnostic or surgical procedures as described herein are presented in the context of pulmonary diagnosis or surgery, the methods and devices described herein can be used to treat any tissue of the body and any organ and vessel of the body, such as the brain, heart, lungs, intestines, eyes, skin, kidneys, liver, pancreas, stomach, uterus, ovaries, testes, bladder, ear, nose, mouth, soft tissue such as bone marrow, adipose tissue, muscle, glandular and mucosal tissue, spinal cord and nerve tissue, cartilage, hard biological tissue such as teeth, bones, etc., and body cavities and passageways such as the sinuses, ureters, colon, esophagus, pulmonary passageways, blood vessels and throat.
[0028] As used herein, a processor encompasses one or more processors, e.g., a single processor, or multiple processors, e.g., in a distributed processing system. A controller or processor described herein generally includes a tangible medium for storing instructions for implementing process steps, and the processor may include, for example, one or more of a central processing unit, programmable array logic, gate array logic, or field programmable gate array. In some cases, the one or more processors may be a programmable processor (e.g., a central processing unit (CPU) or microcontroller), a digital signal processor (DSP), a field programmable gate array (FPGA), or one or more advanced RISC machine (ARM) processors. In some cases, the one or more processors may be operably coupled to a non-transitory computer-readable medium. The non-transitory computer-readable medium may store logic, code, or program instructions executable by one or more processor units to perform one or more steps. The non-transitory computer-readable medium may include one or more memory units (e.g., removable media or external storage such as an SD card or random access memory (RAM)). One or more of the methods or operations disclosed herein may be implemented in hardware components or a combination of hardware and software, such as, for example, an ASIC, a special purpose computer, or a general purpose computer.
[0029] As used herein, the terms distal and proximal may generally refer to a location referenced from the device, as opposed to an anatomical reference. For example, a distal location of a bronchoscope or catheter may correspond to a proximal location of an elongated member on a patient, and a proximal location of a bronchoscope or catheter may correspond to a distal location of an elongated member on a patient.
[0030] The systems described herein include an elongated portion or member, such as a catheter. The terms "elongated member," "catheter," and "bronchoscope" are used interchangeably throughout this specification unless the context suggests otherwise. The elongated member can be placed directly within a body cavity or lumen. In some embodiments, the system may further include a support device, such as a robotic manipulator (e.g., a robotic arm), for driving, supporting, positioning, or controlling the movement or motion of the elongated member. Alternatively, or in addition, the support device may be a handheld device or other control device that may or may not include a robotic system. In some embodiments, the system may further include peripheral devices and subsystems, such as an imaging system, that may assist or facilitate navigation of the elongated member to a target site within the subject's body. Such navigation may require a registration process, as described later in this specification.
[0031] Some embodiments of the present disclosure provide a robotic bronchoscopy system for performing surgery or diagnosis with improved performance at low cost. For example, the robotic bronchoscopy system may include a steerable catheter, which may be completely disposable. This may beneficially reduce the need for sterilization, which may be costly or difficult to operate, but sterilization or disinfection may not be effective. Furthermore, one challenge in bronchoscopy is navigating the airways to reach the upper lobes of the lungs. In some cases, the provided robotic bronchoscopy system may be designed with the ability to navigate through airways with small curvatures autonomously or semi-autonomously. Autonomous or semi-autonomous navigation may require a registration process. Alternatively, the robotic bronchoscopy system may be navigated by an operator through a control system with visual guidance.
[0032] The typical lung cancer diagnosis and surgical treatment process can vary dramatically depending on the technology, clinical protocols, and clinical setting used by healthcare providers. Inconsistent processes can cause delays in diagnosing lung cancer early, lead to high costs for the healthcare system for patients to diagnose and treat lung cancer, and result in a high risk of clinical and procedural complications. The robotic bronchoscopy system herein utilizes integrated tomosynthesis to improve lesion visibility and intralesional tool confirmation, and utilizes enhanced fluoroscopy to enable real-time navigation updates and guidance in all regions of the lung, thus enabling standardized early lung cancer diagnosis and treatment.
[0033] FIG. 1 illustrates an exemplary process 100 for tomosynthesis image reconstruction. In some cases, the tomosynthesis image reconstruction of process 100 may include generating a 3D volume using a combination of X-ray projection images (acquired by any type of C-arm system) acquired at different angles. FIG. 2 illustrates an exemplary process 200 for providing augmented fluoroscopy. The augmented fluoroscopy process 200 may include projecting a 3D lesion onto a 2D X-ray image as an overlay. The augmented fluoroscopy may display any number of overlays corresponding to multiple lesions or targets. The augmented fluoroscopy may display overlays for any desired features in addition to the lesion or target. The tomosynthesis imaging mode and the augmented fluoroscopy mode can be accessed at any stage during a surgical session (e.g., during navigation from the driving mode, during the execution of an operation at the target site, etc.).
[0034] Both process 100 and process 200 may begin with acquiring C-arm or O-arm video or imaging data using an imaging device, such as a C-arm imaging system 105 or 205, respectively. The C-arm or O-arm imaging system may include a radiation source (e.g., an X-ray source) and a detector (e.g., an X-ray detector or X-ray imaging device). The C-arm imaging system has one or more X-ray sources positioned on an arm 1340 having a "C" shape 1340 facing one or more X-ray detectors, and the C-arm may be rotated over a range of angles around the patient. An O-arm is similar to a C-arm, but consists of a complete non-destructive ring ("O") and can be rotated 360° around the patient. As used herein, the term O-arm may be used interchangeably with the term C-arm throughout this specification, unless the context indicates otherwise.
[0035] In some cases, a single C-arm source may provide video or imaging data for the two processes 100 and 200. In some cases, different C-arm sources may provide video or imaging data for the two processes 100 and 200. In some embodiments, raw video frames may be used for both tomosynthesis and fluoroscopy. However, while tomosynthesis may require unique frames from the C-arm, fluoroscopy views or extended fluoroscopy may operate using duplicate frames from the C-arm because they are live video. However, the methods herein may provide a unique frame check algorithm so that video frames for tomosynthesis are processed to ensure uniqueness. For example, as shown in process 160, upon receiving a new image frame, if the current mode is tomosynthesis, the image frame may be processed to determine whether it is a unique frame or a duplicate. The uniqueness check may be based on an image intensity comparison threshold. For example, duplicate frames may be identified by comparing the overall average intensity between two frames, by summing the absolute difference in intensity between the same pixels in the two frames across all pixels, or by summing over the square or other power of the difference in intensity between the same pixels in the two frames. For example, when the intensity difference relative to the previous frame is below a predetermined threshold, the frame may be identified as a duplicate and removed from use for tomosynthesis reconstruction. In some cases, unique or duplicate frames may be identified based on other factors. For example, a uniqueness check may be based on changes in stochastic noise within the image, even if the image has the same average image intensity. As an example, frames may be identified as duplicates based on the same average image intensity, but if a pixel-by-pixel comparison shows differences between the images, the frames may still be determined to be unique. If the current mode is fluoroscopy, the image frames may not be processed to check for uniqueness.
[0036] As shown, two processes 100 and 200 may detect video or imaging frames from a C-arm source at 110 and 210, respectively. In some cases, the video or imaging frames may be normalized. In some cases, normalization may be applied to the image frames to change the range of pixel intensity values within the video or imaging frames. Generally, normalization generates an n-dimensional grayscale image I: with intensity values within the range (Min, Max):
[0037]
number
[0038]
number
[0039] Accurate camera pose and camera parameters are important for both tomosynthesis image reconstruction and extended fluoroscopy overlay. Marker tracking accuracy can affect pose estimation accuracy or performance. The present disclosure provides an improved method for tracking markers in a sequence of video frames. The method can enable tomosynthesis reconstruction with an improved success rate, enable a larger sweep angle for tomosynthesis imaging, eliminate ghosting in 3D reconstructed tomosynthesis images (due to incorrect pose estimation from frame marker mistracking), improve reconstruction quality by using all images and more uniform angular sampling, and accelerate the tomosynthesis reconstruction process.
[0040] The present disclosure may provide an improved, robust marker tracking method with improved success rates and higher speeds. As shown in two processes 100 and 200, the same marker detection at 115 and 215, respectively, may be shared by both processes. As discussed in further detail in FIGS. 6A-6C , which show an example tomosynthesis board, an X-ray projection of a marker on a tomosynthesis board may be the marker in an X-ray image (e.g., obtained via a C-arm). Markers may be detected at 115 and 215 using any suitable image processing technique. For example, OpenCV's blob detection algorithm may be used to detect markers that are blob-shaped. In some cases, the detected marker (e.g., a blob) may be detected as having specific characteristics, such as position, shape, size, color, darkness / brightness, opacity, or other suitable characteristics of the marker.
[0041] As shown, two processes 100 and 200 may match markers to a board pattern at 120 and 220, respectively. The markers detected at operations 115 and 215 may be matched to a tomosynthesis board (e.g., the tomosynthesis board described with respect to FIG. 6 ). As described above, the markers may exhibit any number of various physical characteristics (e.g., position, shape, size, color, darkness / brightness, opacity, etc.) that may be detected at 115 and 215 and used to match the markers to the board pattern at 120 and 220. For example, the tomosynthesis board may have different types of markers, such as large blobs and small blobs. In some cases, the large blobs and small blobs may generate a pattern that may be used to match a marker pattern in a video or image frame to a pattern on the tomosynthesis board. In some cases, after operations 120 and 220, processes 100 and 200 may branch.
[0042] As shown, after matching the markers to the board pattern 120, the process 100 may find the best marker match across all video or image frames at 125. The initial marker match may be a match between the markers in the frame and the tomosynthesis board. In some cases, the patterns of matched markers may be compared on the tomosynthesis board to find the best match using Hamming distance. For each frame, a match with a pattern match score (e.g., the number of matched markers divided by the total number of detected markers) may be obtained. The best match may be determined at 125 as the match with the highest pattern match score among all frames. In some cases, one or more image frames with the top pattern match scores may be identified.
[0043] Process 100 may perform inter-frame tracking 130. At a high level, inter-frame tracking 130 may include propagating marker matches from the best match determined in 125 to the rest of the image frames by robust tomosynthesis marker tracking. In some cases, (i) markers in pairs of consecutive frames may be first matched, (ii) each marker in the first frame may then be matched to k nearest markers in the second frame, (iii) for each matching pair of markers, the motion displacement between the two frames may be calculated, (iv) all markers in the first frame may be transferred to the second frame along with their motion displacement, (v) if the motion displacement between a given transferred point from the first frame and a given point location in the second frame is less than a threshold and the two given marker types are the same, the match may be an inlier, and (vi) the best match may be the motion with the most inliers. From the calculated tomosynthesis marker tracking 130, the existing marker alignment in the current frame is transferred to the marker alignment in the next frame. This process is repeated for all frames at 135 until marker alignment for all frames can be found, and the markers in all frames are aligned to the tomosynthesis board.
[0044] In the augmented fluoroscopy process 200, after aligning markers in a video or image frame to a tomosynthesis board at 220, it can be determined at 225 whether the pattern match is unique. Camera pose estimation using markers for augmented fluoroscopy can be more difficult than that for tomosynthesis reconstruction because (i) only a single video or image frame may be available for augmented fluoroscopy, and (ii) motion information may not be available to disambiguate the pose estimation. The augmented fluoroscopy algorithm can provide a criterion for measuring the uniqueness of the match across the entire tomosynthesis board. In some cases, the marker pattern on the tomosynthesis board can be designed to ensure that the pattern in each subarea is unique. In some cases, the tomosynthesis board pattern may be optimized to maximize the Hamming distance between patches (e.g., any 5x5 patch). In some cases, if the board is rotated by 180 degrees, either physically or via a C-arm setting, the in-plane 180-degree rotation may be taken into account when optimizing the best pattern so that co-registration is minimized. Details regarding the patch / marker matching algorithm and unique marker design are provided later in this specification.
[0045] If the match is unique, according to criteria for measuring uniqueness 225, the camera pose may be correctly estimated and process 200 may proceed to pose estimation operation 230. Otherwise, at 225, the enhanced fluoroscopy overlay is not displayed and process 200 proceeds to operation 250, which may indicate that the enhanced fluoroscopy overlay is available.
[0046] Turning to imaging device pose estimation, processes 100, 200 may recover rotation and translation by minimizing reprojection error from 3D-2D point correspondences to perform pose estimation 140, 230, respectively. In some cases, a perspective n-point (PnP) pose calculation can be used to recover camera pose from n-pair point correspondences. The minimal form of the PnP problem can be P3P, which can be solved with three-point correspondences. For each tomosynthesis frame, there may be multiple marker alignments, and estimation methods such as the Random Sampling with Consensus (RANSAC) variant of the PnP solver can be used for pose estimation. In some cases, pose estimation 140, 230 can be further refined by using a nonlinear minimization method to minimize reprojection error, starting from an initial pose estimate by the PnP solver.
[0047] In tomosynthesis reconstruction 145, process 100 may perform tomosynthesis reconstruction based on pose estimation 140. In some cases, tomosynthesis reconstruction operation 145 may be implemented as a model in Python (or other suitable programming language) using the open-source ASTRA (MATLAB and Python Toolbox of High-Performance GPU Primitives for 2D and 3D Tomography) toolbox (or other suitable toolbox or package). In tomosynthesis reconstruction, the inputs to the model may be: (i) undistorted, inpainted (inpainting: a process of restoring damaged images) projection images; (ii) estimated projection matrices, such as the pose of each projection; and (iii) the size, resolution, and estimated position of the target tomosynthesis reconstruction volume. The output of the model is tomosynthesis reconstruction (e.g., a volume in NifTI format) 145. Thus, in operation 150, process 100 may optionally finish outputting a tomosynthesis reconstruction for a C-arm system, which may include a 3D volume with a combination of X-ray projection images acquired by the C-arm at various angles.
[0048] Act 235 may include projecting the lesion onto the video frame using the estimated pose from act 230 and the pre-calibrated camera parameters from act 245. By way of example, the lesion may be modeled as an ellipsoid projected from the video or image frame onto the 2D fluoroscopic image as an ellipse. Note that the lesion may be modeled using any suitable shape, color, transparency, or other graphical indicator. An augmented fluoroscopic overlay may be displayed over the live fluoroscopic view corresponding to the lesion projected onto the x-ray image 240. The lesion may be a 3D lesion, which is projected onto the 2D fluoroscopic image based at least in part on the camera matrix or pose estimate associated with each 2D fluoroscopic image. The information about the lesion may include 3D position information obtained from the tomosynthesis process. In some cases, the shape and size of the lesion may be based on a 3D model of the lesion (created from a pre-operative CT or any predetermined parameters). Further details regarding obtaining lesion information are provided elsewhere herein.
[0049] State Machine The tomosynthesis-augmented fluoroscopy overlay method described above can be utilized by a tracking system that provides the user with the real-time location of the lesion and the relative positions of the scope or needle and the lesion for correcting navigation. Figure 3 shows an exemplary system 300 of various state machines for implementing a tracking system based at least in part on tomosynthesis and live fluoroscopy with the real-time location of the lesion. At a high level, the state machines included in system 300 can read a set of inputs and change to different states based on those inputs. System 300 can include, in its states, a tracking subsystem 310, a vision subsystem 320, a localization subsystem 330, a system control subsystem 340, a media control subsystem 350, and a user input subsystem 360.
[0050] In some cases, the information about each state machine may include a functional description of key functions, system configuration parameters owned by the state machine, a state transition diagram, a table containing details of the state transitions, or a table presenting all input and output data for the state machine.
[0051] The tracking subsystem 310 may include two state machines, smTomoConfigManager 312 and smTomo 314, as well as helper classes that support interfaces between the tracking subsystem 310 and other subsystems, software, and hardware components. The tracking subsystem 310 may utilize RTI data contracts and implement those described with respect to smTomoConfigManager 312 and smTomo 314. smTomoConfigManager 312 may be responsible for loading tomosynthesis-related configuration parameters from a configuration file and sending the parameters to other state machines through data contracts. In some cases, configuration parameters have default values (e.g., previous values, recommended values, optimal values, etc.) that can be overridden by values specified in the configuration file. smTomo 314 may receive configuration parameters from smTomoConfigManager 312. smTomo 314 may retrieve and process fluoroscopic images from smFluoroFrameGrabber 322 of the vision subsystem 320. smTomo314 may receive user commands and call a tomosynthesis dynamic link library (DLL) module to process and generate intermediate files before tomosynthesis reconstruction. smTomo314 may also provide the captured unique fluoroscopic images to a treatment interface UI (e.g., as described in connection with Figures 19-24) for tip position selection for triangulation calculations to obtain the tip's 3D coordinates. Once reconstruction is complete, the reconstructed volume may be provided to the treatment interface UI for display so that the user can identify and select lesion position coordinates. The tip-to-lesion offset can be obtained and broadcast to the navigation unit for target driving updates. smTomo314 may be responsible for receiving normalized fluoroscopic images, passing them to algorithms, estimating the fluoroscopic image pose, generating intermediate files, and calling a reconstruction module (e.g., a 2D and 3D tomography toolbox with high-performance GPU speedup) to generate the reconstruction results.smTomo314 can perform triangulation calculations to obtain tip coordinates and tip-to-lesion vector calculations based on the EM sensor location and lesion location. The resulting reconstructions may be displayed in the Treatment UI for user lesion selection, and lesion information may be broadcast for enhanced fluoroscopy overlay through data subscriptions.
[0052] 4 illustrates an example configuration state machine, smTomoConfigManager 400, which may be a more detailed diagram of smTomoConfigManager 312 of FIG. 3. In some cases, smTomoConfigManager 400 may read tomosynthesis-related configuration parameters. If an entry is not found in the configuration file for a tomosynthesis-related configuration parameter, smTomoConfigManager 400 may instead retrieve a default value (e.g., a previous value, a recommended value, an optimal value, etc.). In some cases, smTomoConfigManager 400 may broadcast the tomosynthesis-related configuration parameters through an RTI data contract.
[0053] FIG. 5 illustrates an example configuration state machine, smTomo 500, which may be a more detailed diagram of smTomoConfigManager 312 of FIG. 3. smTomo 314 may receive configuration parameters from smTomoConfigManager (e.g., smTomoConfigManager 312 or smTomoConfigManager 400), for example, in UpdateConfig module 510. In some cases, smTomo 500 may receive normalized fluoroscopic image frames from smFluoroFrameGrabber (e.g., smFluoroFrameGrabber 322). In some cases, smTomo 500 may generate intermediate files for reconstruction (e.g., tomosynthesis reconstruction) via an algorithm module, for example, in GenerateReconstruction module 525. In some cases, smTomo 500 can calculate tip coordinates (e.g., via CalculatingTipLesionOffset module 545). In some cases, smTomo 500 may receive EM sensor data (e.g., from smRegistration 322). Using the EM sensor data, smTomo 500 may calculate the average EM coordinate and obtain the maximum deviation from the average EM coordinate. In some cases, smTomo 500 may be responsible for pose estimation and generation of intermediate images for tomosynthesis reconstruction. If a configuration parameter is not found in the configuration file, a default value (e.g., general, average, typical, etc.) may be used.
[0054] Marker board (tomosynthesis board) In some embodiments, the systems herein may provide a marker board (tomosynthesis board) with a unique marker design that aids in pose estimation with improved efficiency and accuracy. The unique marker design may advantageously allow for a large sweep angle. A large sweep angle can beneficially improve reconstruction quality (e.g., improve the axial field of view). FIGS. 6A-6C show an example of a tomosynthesis board 600A with the layering shown in marker design layout 600B and layout 600C. The marker board described with respect to FIGS. 6A-6C may be applied to one or more of the tomosynthesis or enhanced fluoroscopy techniques also described herein.
[0055] The tomosynthesis board 600A may include a physical pattern specific to the translation or rotation. The physical pattern may be formed of markers of various sizes in a predetermined code pattern. For example, as shown, the tomosynthesis board 600A may include dots of different sizes forming the code pattern. In some embodiments, the code pattern may be 3D. In some cases, the dots may be large and small blobs (e.g., beads) arranged on two layers in a grid pattern according to the marker design layout 600B (with an offset in the z-direction of the board, as shown in layout 600C). In some cases, the offset in the two planes may be sufficient (e.g., the offset is at least 20 mm, 30 mm, 40 mm, 50 mm, etc.) so that the 3D pattern of markers may enable calibration or pose estimation of the imaging device using a single 2D image of the marker. In some cases, the 3D pattern of markers may enable calibration or pose estimation with improved accuracy by utilizing multiple 2D images from the projections. In such cases, the offset between the two planes may be small (e.g., 10 mm or less, 20 mm or less, 30 mm or less, etc.). In some embodiments, the markerboard may have a 2D pattern. For example, dots of various sizes may be arranged on the same plane.
[0056] The blobs may be made from a material that is visible on an X-ray image, such as metal. The two-layer marker design shown in layout 600C, a side view of marker design layout 600B, improves the accuracy of pose estimation using tomosynthesis board 600A.
[0057] The marker design layout 600B may have a code pattern of a predetermined size. In some cases, the marker design layout 600B may be a size-coded pattern such that the pattern within each subarea 610 is unique ("1" represents a large bead and "0" represents a small bead). The subareas 610 may be any shape or size, and the pattern within the subarea is unique. The marker design layout 600B may be optimized to maximize the edit distance (e.g., a metric for determining dissimilarity between patterns, strings, etc.) between patches of the tomosynthesis board 600A. In some cases, the edit distance may be measured using the Hamming distance between patches. The patches may be square or rectangular, or some other shape. The patches may be small (e.g., a 3x2 patch, a 4x6 patch, a 5x5 patch, etc.). The patches may be large (e.g., a 5x7 patch, a 2x9 patch, a 9x9 patch, etc.). In some cases, the unique pattern within each subarea may be designed such that the distance between patches having a particular size (e.g., 3x2 patches, 4x6 patches, 5x5 patches, etc.) may be maximized. Details regarding marker matching algorithms are provided later in this specification.
[0058] If the tomosynthesis board 600A is rotated either physically or by rotation in a C-arm setup, in-plane rotation (e.g., 90 degrees, 180 degrees, 270 degrees, etc.) may be considered when designing the marker design layout 600B to minimize co-registration. In some cases, vertical or horizontal flips may be considered in the marker design layout 600B. Multiple rows of marker blobs, such as those shown in the side view of layout 600C, may be interlaced in layers (e.g., 2 layers, 3 layers, 5 layers, 10 layers, etc.) on the tomosynthesis board 600A.
[0059] Pattern Matching Technology 7A-7C show example images used for pattern matching for blob detection. As with Figures 8 and 9, the images and techniques described with respect to Figures 7A-7C may also be applied to one or more of the tomosynthesis or enhanced fluoroscopy techniques described herein.
[0060] FIG. 7A shows an example of blob detection of markers on an image 700A of a tomosynthesis board. Image 700A includes an X-ray projection of a blob on the tomosynthesis board (e.g., as described with respect to FIGS. 6A-6C). The blob is shown as a marker in image 700A. Blobs may be detected using any number of image processing techniques, machine learning (e.g., computer vision) techniques, masking techniques, or statistical techniques. For example, blobs may be detected using any suitable blob detection algorithm. Each detected blob may be marked with various characteristics, such as a center location and a radius, as shown in image 700A. Blobs may be classified as large or small markers according to their size (e.g., thresholded by the median size of all markers). The large and small markers may generate a pattern that is used to match the blob pattern on the tomosynthesis board. Although image 700A shows the markers as blobs, many different patterns, shapes, non-patterns, shading, colors, etc. may be used (e.g., various arrays of polygons, lines, grid patterns, writing, symbols, etc.). Generally, the markers may be implemented in a variety of different ways, provided that in some cases the markers may be useful for aligning the tomosynthesis images to a spatial location (e.g., with respect to the machine, with respect to the patient, etc.).
[0061] FIG. 7B shows an example of candidate points on a tomosynthesis board image 700B. Candidate points on the tomosynthesis board grid can be selected for initial markers on the grid alignment. In some cases, a homography model can be used to remove outliers in the initial markers for grid alignment. A homography can be calculated based on candidate points between points in an X-ray image (e.g., image 700A or 700B) and the tomosynthesis board. For example, an estimation technique such as RANSAC can calculate a homography based on candidate points between points in the X-ray image and the tomosynthesis board. Estimation techniques such as RANSAC, PROSAC (Progressive Sample Concensus), and NAPSAC (N Adjacent Points Sample Concensus) can estimate parameters of a mathematical model from a set of observations contaminated by outliers. The estimation technique can repeatedly sample observations, reject outlier samples that do not fit the model, and retain inlier samples that fit the model.
[0062] The estimation technique may implement a model that can be refined using inlier data obtained through various optimization methods. In some cases, once the homography of one layer of the tomosynthesis board is calculated, the remainder of the marker on that layer can be extracted if the projection of the blob is close enough to the marker. The markers remaining on the image can be fitted to other layers (e.g., the second layer) of the tomosynthesis board.
[0063] In some cases, the initial marker match is a match between the markers in image 700B and the tomosynthesis board grid. The initial marker match may be calculated across one or more frames of image 700B. In some cases, the initial match may be the best-matched frame (e.g., the frame with the highest match score among all frames tested, which may in some cases be all frames in image 700B). The initial match with the best-matched frame may serve as a starting point for propagating the marker match to the remaining frames of image 700B. Thus, once the initial marker match is established, in some cases, the pattern of matched markers may be "slid" across image 700B of the tomosynthesis board to find the remaining best match (e.g., using Hamming distance). For each frame, a pattern match score (e.g., the number of matched markers divided by the total number of detected markers) may be obtained; for example, FIG. 7C shows an example of marker extraction on image 700C illustrating pattern matching and calculating a pattern match score. The best match (eg, the highest match score among all frames) may be selected as the starting point for pattern matching for all frames in the tomosynthesis sweep.
[0064] As shown in FIG. 8 , which illustrates a process 800 for robust tomosynthesis marker matching, once a best-matched frame is obtained (e.g., via the techniques described with respect to FIGS. 7A-7C ), the best frame match can be propagated to all other frames in the tomosynthesis image. In some cases, process 800 may begin with obtaining a pair of consecutive frames having a first marker and a second marker, respectively, at 805 and 815. Process 800 may further include detecting markers included in the pair of consecutive frames (e.g., via computer vision techniques), at 810 and 820, respectively. Process 800 may further include matching the markers included in the pair of consecutive frames obtained at 805 and 815 (e.g., via k-nearest neighbors). Process 800 may further include, for each pair of matching markers, calculating a motion displacement between the pair of consecutive frames obtained at 805 and 815. Process 800 may further include, for each first marker in the first frame acquired at 805, transferring (e.g., mapping) the first marker to a second frame acquired at 815. The transfer of the first marker to a second marker in the second frame is illustrated in FIG. 9, which shows an example result of marker tracking across a tomosynthesis frame sequence (of two consecutive frames) on an image of a tomosynthesis board.
[0065] Referring again to FIG. 8 , for each first marker transferred to the second frame, if the distance between the transferred first marker and the corresponding second marker meets a threshold (e.g., a certain distance or less), then at 825, the match between the first marker and the second marker is an inlier. An initial match may be generated based on distance (e.g., all points within the distance are a match). In some cases, the best match is the match with the most inliers. Process 800 may be iterative or iterative, transferring existing marker matches in the current frame to the next (e.g., successive) frame and repeating for all frames at 830 until markers in all frames are matched to blobs (e.g., beads) on the tomosynthesis board at 835. In some cases, operation 830 may include taking all of the above matches, finding the motion with the highest number of matched marker points, and calling these matched point pairs inliers.
[0066] Pose estimation technique using markers in images FIG. 10 shows an example diagram 1000 of camera pose estimation. Reconstructing accurate camera pose and camera parameters can be an important aspect of both tomosynthesis image reconstruction and extended fluoroscopy overlay. As discussed above with respect to FIGS. 3 and 5, smTomo 314 and 500, respectively, can be involved in estimating the pose of fluoroscopy images (e.g., via triangulation). The described pose estimation systems, methods, and techniques may also be applied to one or more of the tomosynthesis or extended fluoroscopy techniques described herein.
[0067] In FIG. 1000, a pinhole camera model is shown. The pinhole camera model in FIG. 1000 can be used to describe the geometry of an X-ray projection. As shown in FIG. 1000, pose estimation may involve recovering the camera rotation and translation (camera pose) by minimizing the reprojection error from the 3D-2D point correspondence. In some cases, an optimization algorithm can be used to refine the camera calibration parameters by minimizing the reprojection error. The optimization algorithm may be a least-squares algorithm, such as global Levenberg-Marquardt optimization.
[0068] Recovering the camera pose may further include estimating the pose of a calibrated camera given a set of n 3D points in the world and their corresponding 2D projections in the image. The camera pose may include six degrees of freedom involving the rotation (e.g., roll, pitch, yaw) and 3D translation of the camera relative to the world. Point-by-Point (PnP) pose calculation may be used to recover the camera pose from n pairs of point correspondences. Thus, in some cases, n = 3, and thus the minimal form of the PnP problem is P3P, which can be solved with three-point correspondences. For each tomosynthesis frame, there may be multiple marker alignments, and RANSAC or other variants of the PnP solver may be used for camera pose estimation. Once estimated, the pose may be further refined by using a nonlinear minimization method to minimize the reprojection error, starting from an initial pose estimate by the PnP solver.
[0069] Performing camera pose estimation for tomosynthesis reconstruction may include acquiring an undistorted image (e.g., from a robotic bronchoscopy system). The undistorted image may undergo some pre-processing (e.g., image inpainting, etc.). The undistorted image may be normalized using a normalization algorithm. For example, the undistorted image may conform to Beer's law:
[0070]
number
[0071] Camera Pose-Based Tomosynthesis Reconstruction The estimated camera pose or directly measured camera pose may be utilized in the reconstruction of a 3D volumetric image, i.e., tomosynthesis reconstruction. In some cases, a projection matrix (e.g., an estimated camera pose matrix) may be obtained. In addition, in some cases, physical parameters of the tomosynthesis reconstruction (e.g., size, resolution, position, volume, geometry, etc.) may be obtained. One or more inputs of the normalized image, the projection matrix, or the physical parameters may enable the generation of a reconstructed volume for tomosynthesis reconstruction. To generate a reconstructed volume from the input, an algorithm (e.g., the PM2 vector algorithm) may convert the camera-format projection matrix into vector variables (e.g., in the ASTRA toolbox). Another algorithm may be the same as or similar to the ASTRA FDK Recon algorithm, which may invoke the FDK (Feldkamp, Davis, and Kress) reconstruction module, in which the normalized projection image may be cosine-weighted, ramp-filtered, and then back-projected into a volume according to cone-beam geometry. Finally, in some cases, additional algorithms may convert the reconstructed volume (e.g., as output from the ASTRA FDK Recon algorithm) into an appropriate format. For example, the NifTI processing algorithm may save the reconstructed volume as a NifTI image with an affine matrix.
[0072] Augmented Fluoroscopy with Camera Pose Estimation Performing camera pose estimation for augmented fluoroscopy may enable the goal of projecting a lesion onto an X-ray image to be achieved. This disclosure provides a method for accurately projecting a 3D lesion onto a 2D fluoroscopy image using accurate camera pose and camera parameters. The camera calibration and pose estimation approach for generating an augmented layer or overlay of a lesion on a 2D image may be similar to that described for augmented fluoroscopy. However, camera pose estimation for augmented fluoroscopy may be more challenging than pose estimation for tomosynthesis reconstruction because, in some cases, only a single frame is available for augmented fluoroscopy and motion information may not be available (e.g., to remove ambiguity in the pose estimation). One or more criteria may be implemented to measure the uniqueness of a match to the tomosynthesis board. If the match meets the criteria, the match may be determined to be unique. Furthermore, if the match is unique, the camera pose may be determined to be correctly estimated. The estimated camera pose and pre-calibrated camera parameters may be used to project a 3D lesion onto a fluoroscopy video frame (2D image). If the match is not unique and the camera pose is not estimated correctly, the augmented fluoroscopy overlay may not be displayed.
[0073] In some embodiments, an enhancement layer or target / lesion overlay is displayed on a live fluoroscopic view or a 2D fluoroscopic image in fluoroscopy mode. The target / lesion (e.g., one or more lesions) overlay may be modeled as a 3D shape (e.g., an ellipsoid, prism, sphere, etc.) whose projection on the fluoroscopic image is a 2D shape (e.g., an ellipse, polygon, circle, etc.). In some cases, the shape, size, or appearance of the one or more lesion overlay may be based at least in part on the projection of a lesion 3D model (e.g., a 3D mesh model) onto the 2D fluoroscopic image.
[0074] FIG. 11 shows an example of an augmented fluoroscopy projection 1100 with a 3D lesion model projected onto a 2D plane (e.g., an image plane), consistent with examples described herein. As shown in this example, the lesion may be modeled as a 3D mesh object with multiple corner points. The 3D mesh model may be generated from a preoperative CT scan or during planning. In the illustrated example, the corner points are projected onto the 2D fluoroscopy image, and the corner points form a projected polyline contour (from the outermost point). Alternatively, the shape or appearance of the overlay for the lesion may be predetermined (e.g., a circle, marker, etc.), which may not be based on the 3D mesh model from imaging.
[0075] In some cases, the location of the overlay may be determined based at least in part on a target / lesion location determined from tomosynthesis or reconstructed 3D tomosynthesis images and a pose estimate associated with the 2D fluoroscopic image.
[0076] Marker-free posture estimation technology The relative camera pose at which images are acquired is an input required for tomosynthesis reconstruction and augmented fluoroscopy of a 3D volume. Methods and systems for accurately determining the relative camera pose at which images are acquired can be utilized to provide the pose input required for tomosynthesis and augmented fluoroscopy. In some embodiments, the camera pose may be acquired without markers. In some cases, the methods herein may acquire camera poses without utilizing markers, which beneficially allows higher quality images to be achieved because markers can partially obscure the image. For example, in tomosynthesis mode, a region of the image around each marker is typically excised from the image before performing tomosynthesis, reducing the overall amount of information available to generate a 3D reconstruction of the volume, so a higher quality 3D reconstruction of the volume may be achieved without the presence of markers in the image.
[0077] As shown in FIG. 13 , the attitude or motion of the fluoroscopy (tomosynthesis) imaging system can be measured directly using any suitable motion / position sensor 1310 located on the fluoroscopy (tomosynthesis) imaging system. The motion / position sensor can include, for example, an inertial measurement unit (IMU), one or more gyroscopes, velocity sensors, accelerometers, magnetometers, position sensors (e.g., Global Positioning System (GPS) sensors), visual sensors (e.g., imaging devices such as cameras capable of detecting visible, infrared, or ultraviolet light), proximity or range sensors (e.g., ultrasonic sensors, lidar, time-of-flight, or depth cameras), altitude sensors, attitude sensors (e.g., compasses), or magnetic field sensors (e.g., magnetometers, electromagnetic sensors, wireless sensors). In some cases, the fluoroscopy system may include rotary or linear encoders or similar means for measuring the rotational motion of the arm relative to a structure that supports and holds the arm in place. Encoders may also be used to provide the attitude of the imaging device. In some cases, one or more sensors for tracking the movement and position of the fluoroscopy (tomosynthesis) imaging station may be located on the imaging station or remotely from the imaging station, such as a wall-mounted camera 1320. Various poses may be captured by one or more sensors, as described above.
[0078] In some cases, when the relative pose of the light source and detector is known from a motion / position sensor, markers (e.g., a pattern of blobs or beads in a frame for estimating pose) may not be required to estimate the camera pose. In some cases, when pose information is available from multiple sources, such as from both direct pose measurements (e.g., motion / position sensors) and pose estimation (e.g., image analysis of features in a frame), pose information from multiple sources (e.g., direct measurements and estimated pose) can be combined to provide a more accurate pose estimate. For example, direct pose measurements and computer vision-based estimated poses may be averaged (or weighted) to generate a final pose for the imaging system.
[0079] In some cases, the C-arm imaging system may undergo only rotation around the axis of rotation, without global translation, for tomosynthesis reconstruction or extended fluoroscopy, and the orientation information required for each image may include only the relative angle between images. The relative angle between images can be measured by many of the methods described above. For example, a 3D accelerometer can be attached to the C-arm, and the direction of acceleration due to Earth's gravity can be used to determine the relative change in camera angle as the C-arm is rotated. If the C-arm can both rotate and translate, the camera's full six degrees of freedom (6DOF) may need to be known as input to the tomography or extended fluoroscopy. In this case, for example, a binocular optical "localizer" system 1320, together with localizer fiducial markers 1350 attached to the C-arm 1340, may provide full 6DOF information regarding the (x, y, z) position and (Rx, Ry, Rz) orientation of the fiducial markers in the frame. A (single) camera calibration process can be performed to determine the translational and rotational transformations from the localizer fiducial marker frame to the camera frame. After calibration, the 6DOF pose of the camera may be known as each image is acquired based on captured data from the localizer.
[0080] Robotic Bronchoscopy System FIG. 12 illustrates examples of robotic bronchoscopy systems 1200, 1230 according to some examples. The robotic bronchoscopy systems can implement the methods, subsystems, and functional modules described above. As shown in FIG. 12 , the robotic bronchoscopy system 1200 can include a steerable catheter assembly 1220 and a robotic support system 1210 for supporting or carrying the steerable catheter assembly. The steerable catheter assembly can be a bronchoscope. In some embodiments, the steerable catheter assembly can be a disposable robotic bronchoscope. In some embodiments, the robotic bronchoscopy system 1200 can include an instrument drive mechanism 1213 attached to an arm of the robotic support system. The instrument drive mechanism can be provided by any suitable controller device (e.g., a handheld controller), which may or may not include a robotic system. The instrument drive mechanism can provide a mechanical and electrical interface to the steerable catheter assembly 1220. The mechanical interface can allow the steerable catheter assembly 1220 to be releasably coupled to the instrument drive mechanism. For example, the handle portion of the steerable catheter assembly may be attached to the instrument drive mechanism via quick installation / release means such as a magnet, a spring-loaded level, etc. In some cases, the steerable catheter assembly may be manually coupled to or released from the instrument drive mechanism without the use of tools.
[0081] The steerable catheter assembly 1220 can include a handle portion 1223 that can include components configured to process image data, provide power, or establish communication with other external devices. For example, the handle portion 1223 can include circuitry and communication elements that enable electrical communication between the steerable catheter assembly 1220, the instrument drive mechanism 1213, and any other external system or device. In another example, the handle portion 1223 can include circuit elements such as a power supply for powering the endoscope's electronics (e.g., camera and LED light). In some cases, the handle portion can be in electrical communication with the instrument drive mechanism 1213 via an electrical interface (e.g., a printed circuit board) so that image / video data or sensor data can be received by the instrument drive mechanism's communication module and transmitted to other external devices / systems. Alternatively, or in addition, the instrument drive mechanism 1213 can provide only a mechanical interface. The handle portion can be in electrical communication with a modular wireless communication device or any other user device (e.g., a portable / handheld device or controller) to transmit sensor data or receive control signals. Further details regarding the handle portion are provided later in this specification.
[0082] The steerable catheter assembly 1220 can include a flexible elongate member 1211 coupled to a handle portion. In some embodiments, the flexible elongate member can include a shaft, a steerable tip, and a steerable section. The steerable catheter assembly can be a single-use robotic bronchoscope. In some cases, only the elongate member can be disposable. In some cases, at least a portion of the elongate member (e.g., shaft, steerable tip, etc.) can be disposable. In some cases, the entire steerable catheter assembly 1220, including the handle portion and elongate member, can be disposable. The flexible elongate member and handle portion are designed so that the entire steerable catheter assembly can be disposed of at low cost. More details regarding flexible elongate members and steerable catheter assemblies are provided later in this specification.
[0083] In some embodiments, the provided bronchoscope system may also include a user interface. As shown in exemplary system 1230, the bronchoscope system may include a therapy interface module 1231 (user console side) or a therapy control module 1233 (patient and robot side). The therapy interface module may allow an operator or user to interact with the bronchoscope during a surgical procedure. In some embodiments, the therapy control module 1233 may be a handheld controller. The therapy control module may optionally include proprietary user input devices and one or more add-on elements removably coupled to existing user devices to improve the user input experience. For example, a physical trackball or roller may replace or complement at least one function of a virtual graphical element displayed on a graphical user interface (GUI) (e.g., a navigation arrow displayed on a touchpad) by providing similar functionality to the graphical element it replaces. Examples of user devices may include, but are not limited to, a mobile device, a smartphone / cell phone, a tablet, a personal digital assistant (PDA), a laptop or notebook computer, a desktop computer, a media content player, etc. More details regarding user interface devices and user consoles are provided later in this specification.
[0084] The user console 1231 can be attached to the robotic support system 1210. Alternatively, or in addition, the user console or a portion of the user console (e.g., the treatment interface module) can be mounted on a separate mobile cart.
[0085] The present disclosure provides a robotic endoluminal platform that integrates tool-in-lesion tomosynthesis technology. In some cases, the robotic endoluminal platform may be a bronchoscopy platform. The platform may be configured to perform one or more operations consistent with the methods described herein. FIG. 13 illustrates an example of a robotic endoluminal platform and its components or subsystems according to some embodiments of the present invention. In some embodiments, the platform may include a robotic bronchoscopy system and one or more subsystems that can be used in combination with the robotic bronchoscopy system of the present disclosure.
[0086] In some embodiments, one or more subsystems may include an imaging system, such as a fluoroscopy imaging system, to provide real-time imaging of the target site (e.g., including the lesion). To better visualize and provide 3D coordinates of anatomical structures, multiple 2D fluoroscopy images can be used to create a tomosynthesis or cone-beam CT (CBCT) reconstruction. FIG. 13 shows an example of a fluoroscopy (tomosynthesis) imaging system 1300. For example, the fluoroscopy (tomosynthesis) imaging system can perform precise lesion position tracking or intralesion tool confirmation before or during a surgical procedure, such as those described above. In some cases, the location of the lesion can be tracked based on position data related to the fluoroscopy (tomosynthesis) imaging system / station (e.g., a C-arm) and image data captured by the fluoroscopy (tomosynthesis) imaging system. The location of the lesion can be registered with the coordinate frame of the robotic bronchoscopy system.
[0087] In some cases, the position, pose, or motion of the fluoroscopic imaging system may be measured / estimated to register the coordinate frame of the image to the robotic bronchoscopy system or to construct a 3D model / image. In some cases, the pose of the imaging system may be estimated using pose estimation methods as described elsewhere herein. For example, a unique markerboard-based pose estimation method may be used to obtain the pose of the imaging device associated with each 2D image.
[0088] Alternatively, the attitude or motion of the fluoroscopy (tomosynthesis) imaging system can be measured directly using any suitable motion / position sensor 1310 located on the fluoroscopy (tomosynthesis) imaging system. The motion / position sensor can include, for example, an inertial measurement unit (IMU), one or more gyroscopes, velocity sensors, accelerometers, magnetometers, position sensors (e.g., Global Positioning System (GPS) sensors), visual sensors (e.g., imaging devices such as cameras capable of detecting visible, infrared, or ultraviolet light), proximity or range sensors (e.g., ultrasonic sensors, lidar, time-of-flight, or depth cameras), altitude sensors, attitude sensors (e.g., compasses), or magnetic field sensors (e.g., magnetometers, electromagnetic sensors, radio sensors). In some cases, the fluoroscopy system can include rotary or linear encoders or similar means for measuring the rotational motion of the arm relative to a structure that supports and holds the arm in place. Encoders can also be used to provide the attitude of the imaging device. In some cases, one or more sensors for tracking the motion and position of the fluoroscopy (tomosynthesis) imaging station may be located on the imaging station, such as a wall-mounted camera 1320, or may be located remotely from the imaging station. Various poses may be captured by one or more sensors, as described above. If the relative pose of the light source and detector is known from a motion / position sensor, it is not necessary to use a pattern of blobs or beads in the frame to estimate the pose. In some cases, when pose information is available from multiple sources, such as from both direct pose measurements (e.g., motion / position sensors) and pose estimation (e.g., image analysis of features in the frame), the pose information from the multiple sources (e.g., direct measurements and estimated pose) may be combined to provide a more accurate pose estimation. For example, direct pose measurements and estimated poses based on computer vision may be averaged (or weighted) to generate a final pose for the imaging system.
[0089] In some embodiments, the location of the lesion can be segmented in the image data captured by the fluoroscopy (tomosynthesis) imaging system with the aid of the signal processing unit 1330. One or more processors of the signal processing unit can be configured to further overlay the treatment location (e.g., the lesion) on the real-time fluoroscopy image / video. For example, the processing unit can be configured to generate an augmentation layer containing augmentation information such as the location of the treatment location or target site. In some cases, the augmentation layer can also include a graphical marker indicating a path to the target site. The augmentation layer can be a substantially transparent image layer containing one or more graphic elements (e.g., boxes, arrows, etc.). The augmentation layer can be overlaid on an optical view of the optical image or video stream captured by the fluoroscopy (tomosynthesis) imaging system or displayed on a display device. The transparency of the augmentation layer allows the optical image to be viewed by the user along with the graphical elements overlaid on the optical image. In some cases, both the segmented lesion image and an optimal path for navigation of the elongated member to reach the lesion can be overlaid on the real-time tomosynthesis image. This may allow the operator or user to visualize the exact location of the lesion as well as the planned path of bronchoscope movement. In some cases, segmented and reconstructed images (e.g., CT images described elsewhere) provided prior to operation of the systems described herein may be overlaid on the real-time image.
[0090] In some embodiments, one or more subsystems of the platform may comprise one or more treatment subsystems, such as a manual or robotic instrument (e.g., a biopsy needle, a biopsy forceps, a biopsy brush) or a manual or robotic treatment instrument (e.g., an RF ablation instrument, a Cryo instrument, a microwave instrument, etc.).
[0091] In some embodiments, one or more subsystems of the platform may include a navigation and localization subsystem. The navigation and localization subsystem may be configured to construct a virtual airway model based on preoperative images (e.g., preoperative CT images or tomosynthesis). The navigation and localization subsystem may be configured to identify a segmented lesion location in the 3D-rendered airway model, and based on the location of the lesion, the navigation and localization subsystem may generate an optimal path from the main bronchus to the lesion with a recommended approach angle toward the lesion for performing a surgical procedure (e.g., biopsy).
[0092] In some cases, to aid in reaching the target tissue location, the position and movement of the medical instrument may be registered with an intraoperative 3D image of the patient's anatomy. In some cases, this may be achieved by determining a transformation or other navigation solution from the frame of reference of the 3D image to the frame of reference of the EM field, allowing the location of the lesion within the 3D model of the patient's anatomy to be updated based on data from the intraoperative 3D image of the patient's anatomy. The transformation (co-registration) between the frame of reference of the 3D image and the frame of reference of the EM or other navigation system may include three rotations between frames and three translations between frames.
[0093] The present disclosure may provide a co-registration method for co-registering a reference frame of a 3D image with a navigation reference frame (e.g., a reference coordinate system of an EM field). In some cases, the co-registration method may utilize markers visible in the image dataset to establish a reference frame for the 3D image. The marker reference frame and the EM reference frame (or other navigation reference frame) may have a known transformation relationship (e.g., the rotation and translation between the marker reference frame and the EM reference frame are known during equipment setup or from mechanical constraints of the device). The position of the patient's anatomical structures is found in the 3D image reference frame, and from a mechanical configuration or setting, a transformation from the 3D image reference frame to the navigation reference frame is obtained, and the position of the patient's anatomical structures in the navigation reference frame may be updated based on the measured position of the patient in the 3D image.
[0094] In some cases, instead of obtaining the rotation and translation between the marker reference frame and the navigation reference frame (e.g., the EM reference frame) from the equipment setup or device mechanical constraints, only the rotation of the marker reference frame relative to the navigation reference frame (e.g., both the marker frame and the EM generator are fixed to the bed and the frame's (x, y, z) axes are parallel to the bed's major axes) is obtained from the equipment setup or device mechanical constraints. Meanwhile, the translation between the frames is obtained based on real-time measurements. For example, the translational relationship can be obtained by measuring the (x, y, z) positions of features / structures (e.g., the tip, any fiducial markers, parts of the endoscope, etc.) in both the navigation system and the imaging system. For example, in EM navigation, the (x, y, z) position of the endoscope tip is measured within the frame of the EM navigation system. The (x, y, z) position of the endoscope tip in the 3D image reference frame is measured by positioning the tip within a 3D dataset that includes the tip simultaneously with the EM measurements. It should be noted that any structure / feature (e.g., endoscope tip, tool, marker on tool, etc.) whose position can be measured in both the navigation system and the imaging system can be utilized to determine the translational relationship between the two frames.
[0095] In some cases, both the rotational and translational relationship between the two frames can be obtained by using a structure or feature that can be located (x, y, z) by both the marker reference frame and the navigation reference frame. For example, in EM navigation, both the (x, y, z) position and (Rx, Ry, Rz) angular orientation of the endoscope tip within the frame of the EM navigation system are measured. The structure or feature can be built into the endoscope, which is opaque to x-rays, allowing the position and angular orientation of the structure to be determined by 3D reconstruction. Based on the endoscope tip (x, y, z) and (Rx, Ry, Rz) determined in both the 3D image frame and the EM frame, the two reference frames can be co-registered.
[0096] In some cases, radiopaque markers fixed to the EM (or other navigation system) can be utilized to obtain the transformation matrix. For example, by construction or calibration, the markers fixed to the EM system may have a known translation and orientation relative to the EM frame. The markers on the EM frame may be visible in the 3D image, which can be used to determine the translation and orientation of the markers fixed to the EM system relative to the markers in the 3D image. The method can then combine the transformations to determine the location of the physiology (e.g., lesion) in the EM frame as EM Frame_T_Lesion = EM Frame_T_EMMarker * EMMarker_T_3D Frame * 3D Frame_T_Lesion, where each Frame 2_T_Frame 1 label refers to a 4x4 rotational and translational transformation matrix that provides the (x, y, z) location of a point in the Frame 2 coordinate system given the (x, y, z) location of the same point in the Frame 1 coordinate system.
[0097] In some cases, the co-registration method may include independently measuring the relative position and orientation of the c-arm camera with respect to the navigational reference frame, for example, using one set of 3D localization tools 1310 or 1350 physically connected to the camera and fixed to a structure 1340 with a known orientation relative to the camera, and a second set of 3D localization tools fixed to a structure physically connected to the EM frame with a known orientation relative to the EM frame.
[0098] In some cases, the co-registration method may include structures or features relative to the C-arm that allow the EM navigation (or other navigation system) to measure the camera's translation and angle relative to the EM frame. For example, a 6DOF EM sensor may be fixed to the C-arm so that the sensor's position and orientation can be measured by the EM navigation system within the EM frame. The EM sensor's position and orientation relative to the camera may be known from construction or calibration, and measuring the EM sensor's position and orientation thus provides the camera's position and orientation within the EM frame. A 3D image may be reconstructed from multiple 2D projections using individual camera poses, as described elsewhere herein. Because the camera poses are already within the EM frame, the reconstructed 3D image frames are automatically co-registered to the EM reference frame (they are the same frame). This method advantageously eliminates the need to use computer vision or features within the image for co-registration.
[0099] With image-guided instruments registered to the image, the instruments can navigate natural or surgically created passageways in anatomical systems such as the lungs, colon, intestines, kidneys, heart, circulatory system, etc. In some cases, after a medical instrument (e.g., needle, endoscope) reaches the target location or after a surgical procedure is completed, 3D imaging can be performed to confirm that the instrument or operation is at the target location.
[0100] In a registration step prior to driving the bronchoscope to the target site, the system may register a rendered virtual view of the airway to the patient's airway. Image registration may consist of a single registration step or a combination of a single registration step and real-time sensory updates to the registration information. The registration process may include finding a transformation that aligns objects (e.g., airway model, anatomical site) between different coordinate systems (e.g., EM sensor coordinates and patient 3D model coordinates based on preoperative CT imaging). Registration is described in more detail below.
[0101] Once registered, all airways can be aligned with the preoperatively rendered airways. While driving the robotic bronchoscope toward the target site, the position of the bronchoscope within the airways can be tracked and displayed. In some cases, the position of the bronchoscope relative to the airways can be tracked using a position sensor. Other types of sensors (e.g., cameras) can also be used in place of or in conjunction with positioning sensors using sensor fusion techniques. A positioning sensor, such as an electromagnetic (EM) sensor, can be embedded in the distal tip of the catheter, and an EM field generator can be placed next to the patient's torso during the procedure. The EM field generator can localize the EM sensor position in 3D space or can localize the EM sensor position and orientation with five or six degrees of freedom (5DOF or 6DOF), consisting of three spatial coordinates and two or three orientation angles. This can provide a visual guide to the operator when driving the bronchoscope toward the target site.
[0102] In real-time EM tracking, an EM sensor, comprising one or more sensor coils embedded at one or more locations and orientations within a medical instrument (e.g., the tip of an endoscopic tool), measures variations in an EM field generated by one or more static EM field generators positioned near the patient. The position information detected by the EM sensor is stored as EM data. The EM field generator (or transmitter) can be placed near the patient to generate a low-intensity, low-frequency alternating magnetic field that can be detected by the embedded sensor. The alternating magnetic field induces small currents in the sensor coils of the EM sensor, which can be analyzed to determine the distance and angle between the EM sensor and the EM field generator. These distances and orientations can be registered intraoperatively with respect to the patient's anatomy (e.g., a 3D model) to determine a registration transformation that aligns a single location in a coordinate system with a location in a preoperative model of the patient's anatomy.
[0103] In some embodiments, the platform herein can utilize a fluoroscopic imaging system to determine the position and orientation of medical instruments and patient anatomical structures within the coordinate system of the surgical environment. In particular, the systems and methods herein can employ mobile C-arm fluoroscopy as a low-cost, mobile, real-time qualitative assessment tool. Fluoroscopy is an imaging modality that acquires real-time video of patient anatomical structures and medical instruments. The fluoroscopy system can include a C-arm system that provides positional flexibility and is capable of orbital, horizontal, or vertical movement via manual or automated control. Fluoroscopic image data from multiple perspectives in the surgical environment (i.e., the fluoroscopic imaging device is moved between multiple positions) can be compiled to generate two-dimensional or three-dimensional tomographic images. When using a fluoroscopic imager system including a digital detector (e.g., a flat-panel detector), the generated and compiled fluoroscopic image data can enable the segmentation of planar images in parallel planes according to tomosynthesis imaging techniques. The C-arm imaging system can include a source (e.g., an X-ray source) and a detector (e.g., an X-ray detector or X-ray imaging device). The X-ray detector can generate an image representing the intensity of the received X-rays. The imaging system can reconstruct a 3D image based on multiple 2D images acquired from a wide range of angles. In some cases, the rotation angle range can be at least 120 degrees, 130 degrees, 140 degrees, 150 degrees, 160 degrees, 170 degrees, 180 degrees, or more. In some cases, the 3D image can be generated based on the orientation of the X-ray imaging device.
[0104] The bronchoscope or catheter may be disposable. FIG. 14 illustrates an example of a flexible endoscope 1400 according to some embodiments of the present disclosure. As shown in FIG. 14, the flexible endoscope 1400 may include a handle / proximal portion 1409 and a flexible elongate member that is inserted into a subject. The flexible elongate member may be the same as described above. In some embodiments, the flexible elongate member may include a proximal shaft (e.g., insertion shaft 1401), a steerable tip (e.g., tip 1405), and a steerable segment (active bending section 1403). The active bending section and proximal shaft section may be the same as those described elsewhere herein. The endoscope 1400 may also be referred to as a steerable catheter assembly, as described elsewhere herein. In some cases, the endoscope 1400 may be a disposable robotic endoscope. In some cases, the entire catheter assembly may be disposable. In some cases, at least a portion of the catheter assembly may be disposable. In some cases, the entire endoscope can be released from the instrument drive mechanism and disposed of. In some embodiments, the endoscope can contain varying levels of stiffness along the shaft to improve functional operation.
[0105] The endoscope or steerable catheter assembly 1400 may include a handle portion 1409, which may include one or more components configured to process image data, provide power, or establish communication with other external devices. For example, the handle portion may include circuitry and communication elements that enable electrical communication between the steerable catheter assembly 1400, an instrument drive mechanism (not shown), and any other external systems or devices. In another example, the handle portion 1409 may include circuit elements such as a power supply for powering the endoscope's electronics (e.g., a camera, electromagnetic sensors, and LED lights).
[0106] One or more components located in the handle may be optimized so that expensive and complex components may be allocated to the robotic support system, handheld controller, or instrument drive mechanism, thereby reducing costs and simplifying the design of the single-use endoscope. The handle or proximal portion may provide an electrical and mechanical interface that allows electrical and mechanical communication with the instrument drive mechanism. The instrument drive mechanism may include a set of motors that are actuated to rotationally drive a set of pull wires of the catheter. The handle portion of the catheter assembly may be mounted on the instrument drive mechanism such that its pulley / capstan assembly is driven by the set of motors. The number of pulleys may vary based on the pull wire configuration. In some cases, one, two, three, four, or more pull wires may be utilized to articulate the flexible endoscope or catheter.
[0107] The handle portion can be designed to enable the robotic bronchoscope to be low-cost and disposable. For example, classic manual and robotic bronchoscopes may have cables at the proximal end of the bronchoscope handle. The cables often include illumination fibers, camera video cables, and other sensor fibers or cables, such as electromagnetic (EM) sensors or shape-sensing fibers. Such complex cables can be expensive, adding to the cost of the bronchoscope. The provided robotic bronchoscope may have an optimized design so that simplified structures and components can be employed while maintaining mechanical and electrical functionality. In some cases, the handle portion of the robotic bronchoscope may employ a cable-free design while providing a mechanical / electrical interface to the catheter.
[0108] An electrical interface (e.g., a pulley-mounted circuit board) may allow image / video data or sensor data to be received by the instrument drive's communications module and transmitted to other external devices / systems. In some cases, the electrical interface may establish electrical communication without cables or wires. For example, the interface may include pins soldered onto an electronics board, such as a pulley-mounted circuit board (PCB). For example, a receptacle connector (e.g., a female connector) may be provided on the instrument drive as a mating interface. This may advantageously allow an endoscope to be quickly plugged into the instrument drive or robotic support without utilizing an extra cable. Such a type of electrical interface may also serve as a mechanical interface, such that when the handle portion is plugged into the instrument drive, both a mechanical and electrical connection are established. Alternatively, or in addition, the instrument drive may provide only the mechanical interface. The handle portion may electrically communicate with the modular wireless communication device or any other user device (e.g., a portable / handheld device or controller) to transmit sensor data or receive control signals.
[0109] In some cases, the handle portion 1409 may include one or more mechanical control modules, such as a luer 1411 for interfacing with an irrigation / aspiration system. In some cases, the handle portion may include a lever / knob for articulation control. Alternatively, the articulation control may be located on a separate controller attached to the handle portion via the instrument drive mechanism.
[0110] The endoscope may be attached to a robotic support system or a handheld controller via an instrument drive mechanism. The instrument drive mechanism may be provided by any suitable controller device (e.g., a handheld controller), which may or may not include a robotic system. The instrument drive mechanism may provide a mechanical and electrical interface to the steerable catheter assembly 1400. The mechanical interface may allow the steerable catheter assembly 1400 to be releasably coupled to the instrument drive mechanism. For example, a handle portion of the steerable catheter assembly may be attached to the instrument drive mechanism via quick installation / release means, such as a magnet, a spring-loaded level, or the like. In some cases, the steerable catheter assembly may be manually coupled to or released from the instrument drive mechanism without the use of tools.
[0111] In the illustrated example, the distal tip of the catheter or endoscope shaft is configured to articulate / bend in two or more degrees of freedom to provide a desired camera field of view or control the direction of the endoscope. As shown in this example, an imaging device (e.g., a camera) and a position sensor (e.g., an electromagnetic sensor) 1407 are disposed at the tip of the catheter or endoscope shaft 1405. For example, the camera's line of sight can be controlled by controlling the articulation of the active bending section 1403. In some cases, the angle of the camera can be adjustable so that the line of sight can be adjusted without, or in addition to, articulating the distal tip of the catheter or endoscope shaft. For example, the camera can be oriented at an angle (e.g., tilted) relative to the axial direction of the endoscope tip using optical components.
[0112] The distal tip 1405 may be a rigid component that allows for positioning of sensors, such as electromagnetic (EM) sensors, imaging devices (e.g., cameras), and other electronic components (e.g., LED light sources), embedded in the distal tip.
[0113] In real-time EM tracking, an EM sensor, comprising one or more sensor coils embedded at one or more positions and orientations within a medical instrument (e.g., the tip of an endoscopic tool), measures fluctuations in an EM field generated by one or more EM field generators positioned near the patient. The position information detected by the EM sensor is stored as EM data. The EM field generator (or transmitter) can be placed near the patient to generate a low-intensity alternating magnetic field that can be detected by the embedded sensor. The alternating magnetic field induces a small current in the sensor coil of the EM sensor, which can be analyzed to determine the distance and angle between the EM sensor and the EM field generator. For example, the EM field generator may be positioned close to the patient's torso during a procedure to locate the EM sensor position in 3D space, or to locate the EM sensor position and orientation within 5DOF or 6DOF. This can provide a visual guide for the operator when maneuvering the bronchoscope toward the target site.
[0114] The endoscope may have a unique design for the elongate member. In some cases, the active bending section 1403 and the proximal shaft of the endoscope may be comprised of a single tube incorporating a series of cuts (e.g., reliefs, slits, etc.) along its length to allow for improved flexibility, desirable stiffness, and anti-extradition features (e.g., features that define a minimum bend radius).
[0115] As described above, the active bending segment 1403 can be designed to allow bending (e.g., articulation) in two or more degrees of freedom. Greater degrees of bending, such as 180 degrees and 270 degrees (or other articulation parameters for clinical indications), can be achieved by the unique structure of the active bending segment. In some cases, a variable minimum bend radius along the axis of the elongate member can be provided such that the active bending segment can include two or more different minimum bend radii.
[0116] Articulation of the endoscope can be controlled by applying force to the distal tip of the endoscope via one or more pull wires. One or more pull wires can be attached to the distal end of the endoscope. In the case of multiple pull wires, pulling one wire at a time can change the orientation of the distal tip up and down, left and right, or in any direction needed. In some cases, the pull wires can be tethered to the distal tip of the endoscope, passed through a bending section, and entered a handle where they are connected to a drive component (e.g., a pulley). This handle pulley can interact with an output shaft from the robotic system.
[0117] In some embodiments, the proximal end or portion of one or more pull wires may be operably coupled to various mechanisms (e.g., gears, pulleys, capstans, etc.) within the handle portion of the catheter assembly. The pull wires may be metal wires, cables, or threads, or may be polymeric wires, cables, or threads. The pull wires may also be made from natural or organic materials or fibers. The pull wires may be any type of suitable wire, cable, or thread capable of supporting various types of loads without significant deformation or breakage. The distal end / portion of one or more pull wires may be anchored or integrated into the distal portion of the catheter such that operation of the pull wires by the control unit may steer or articulate at least the distal portion (e.g., flexible section) of the catheter (e.g., up, down, pitch, yaw, or any direction therebetween) or apply force or tension to the distal portion.
[0118] The pull wires may be made from any suitable material, such as stainless steel (e.g., SS316), metals, alloys, polymers, nylon, or biocompatible materials. The pull wires may be wires, cables, or threads. In some embodiments, different pull wires may be made from different materials to vary the load-bearing capabilities of the pull wires. In some embodiments, different sections of the pull wire may be made from different materials to vary stiffness or load-bearing capacity along the pull. In some embodiments, the pull wires may be utilized for the transmission of electrical signals.
[0119] The proximal design can improve device reliability without introducing extra costs, enabling a low-cost, single-use endoscope. In another aspect of the present invention, a single-use robotic endoscope is provided. The robotic endoscope can be a bronchoscope or the same as the steerable catheter assembly described elsewhere herein. Conventional endoscopes can be complex in design and are typically designed to be reused after a procedure, requiring extensive cleaning, disinfection, or sterilization after each procedure. Existing endoscopes are often designed with complex structures to ensure they can withstand cleaning, disinfection, and sterilization processes. The provided robotic bronchoscope can be a single-use endoscope, which can beneficially reduce cross-contamination between patients and infections. In some cases, the robotic bronchoscope can be delivered to healthcare professionals in a pre-sterilized package and intended to be discarded after a single use.
[0120] As shown in FIG. 15 , the robotic bronchoscope 1510 can include a handle portion 1513 and a flexible elongate member 1511. In some embodiments, the flexible elongate member 1511 can include a shaft, a steerable tip, and a steerable / active bending section. The robotic bronchoscope 1510 can be the same as the steerable catheter assembly described in FIG. 14 . The robotic bronchoscope can be a disposable robotic endoscope. In some cases, only the catheter can be disposable. In some cases, at least a portion of the catheter can be disposable. In some cases, the entire robotic bronchoscope can be disengaged from the instrument drive mechanism and discarded. In some cases, the bronchoscope can contain varying levels of stiffness along its shaft to improve functional operation. In some cases, the minimum bend radius along the shaft can vary.
[0121] The robotic bronchoscope can be releasably coupled to an instrument drive mechanism 1520. The instrument drive mechanism 1520 can be mounted to an arm of a robotic support system or to any actuation support system as described elsewhere herein. The instrument drive mechanism can provide a mechanical and electrical interface to the robotic bronchoscope 1510. The mechanical interface can allow the robotic bronchoscope 1510 to be releasably coupled to the instrument drive mechanism. For example, a handle portion of the robotic bronchoscope can be attached to the instrument drive mechanism via quick installation / release means such as a magnet and spring-loaded level. In some cases, the robotic bronchoscope can be manually coupled to or released from the instrument drive mechanism without the use of tools.
[0122] FIG. 16 shows an example of an instrument drive mechanism 1600B that provides a mechanical interface to a handle portion 1613 of a robotic bronchoscope. As shown in the example, the instrument drive mechanism 1600B may include a set of motors that are actuated to rotationally drive a set of pull wires of a flexible endoscope or catheter. The handle portion 1613 of the catheter assembly may be mounted on the instrument drive mechanism such that its pulley assembly or capstan is driven by the set of motors. The number of pulleys may vary based on the pull wire configuration. In some cases, one, two, three, four, or more pull wires may be utilized to articulate the flexible endoscope or catheter.
[0123] The handle portion can be designed to make the robotic bronchoscope low-cost and disposable. For example, classic manual and robotic bronchoscopes may have cables at the proximal end of the bronchoscope handle. The cables often include illumination fibers, camera video cables, and other sensor fibers or cables, such as electromagnetic (EM) sensors or shape-sensing fibers. Such complex cables can be expensive and increase the cost of the bronchoscope. The provided robotic bronchoscope may have an optimized design so that simplified structures and components can be employed while maintaining mechanical and electrical functionality. In some cases, the handle portion of the robotic bronchoscope may adopt a cable-free design while providing a mechanical / electrical interface to the catheter.
[0124] FIG. 17 shows an example of a distal tip 1700 of an endoscope. In some cases, the distal portion or tip of the catheter 1700 can be substantially flexible so that it can be steered in one or more directions (e.g., pitch, yaw). The catheter can include a tip section, a bending section, and an insertion shaft. In some embodiments, the catheter can have variable bending stiffness along the longitudinal axis. For example, the catheter can include multiple sections with different bending stiffnesses (e.g., flexible, semi-rigid, and rigid). The bending stiffness can be varied by selecting materials with different stiffness / rigidity, varying the structure (e.g., cut, pattern) in different segments, adding additional support components, or any combination of the above. In some embodiments, the catheter can have a variable minimum bend radius along the longitudinal axis. Selecting different minimum bend radii at different locations along the catheter's length can beneficially provide anti-prolapse capabilities while still allowing the catheter to reach difficult-to-reach areas. In some cases, the proximal end of the catheter does not need to be highly flexible, and therefore the proximal portion of the catheter may be reinforced with additional mechanical structures (e.g., additional layers of material) to achieve greater flexural stiffness. Such a design may provide support and stability to the catheter. In some cases, variable flexural stiffness may be achieved by using different materials during the extrusion of the catheter. This may advantageously allow for different stiffness levels along the catheter shaft in the extrusion manufacturing process without additional fastening or assembly of different materials.
[0125] The distal portion of the catheter can be steered by one or more pull wires 1705. The distal portion of the catheter can be made of any suitable material, such as a copolymer, polymer, metal, or alloy, so that it can be bent by the pull wires. In some embodiments, the proximal or terminal ends of the one or more pull wires 1705 can be coupled to a drive mechanism (e.g., a gear, pulley, capstan, etc.) via an anchoring mechanism as described above.
[0126] The pull wires 1705 may be metal wires, cables, or threads, or may be polymer wires, cables, or threads. The pull wires 1705 may also be made from natural or organic materials or fibers. The pull wires 1705 may be any type of suitable wire, cable, or thread capable of supporting various types of loads without significant deformation or breakage. The distal ends or portions of one or more pull wires 1705 may be tethered or integrated into the distal portion of the catheter such that operation of the pull wires by the control unit may steer or articulate (e.g., up, down, pitch, yaw, or any direction in between) at least the distal portion (e.g., flexible section) of the catheter or apply a force or tension to the distal portion.
[0127] The catheter may have dimensions such that one or more electronic components can be integrated into the catheter. For example, the outer diameter of the distal tip may be approximately 4 to 4.4 millimeters (mm), and the diameter of the working channel may be approximately 2 mm, allowing one or more electronic components to be embedded in the catheter wall. However, it should be noted that, depending on different applications, the outer diameter may be within any range, less than 4 mm or greater than 4.4 mm, and the diameter of the working channel may be within any range, depending on the tool dimensions or specific application.
[0128] The one or more electronic components may include an imaging device, an illumination device, or a sensor. In some embodiments, the imaging device may be a video camera 1713. The imaging device may include optical elements and an image sensor for capturing image data. The image sensor may be configured to generate image data in response to a wide range of light wavelengths or a specific wavelength of light. Various image sensors for capturing image data, such as complementary metal oxide semiconductor (CMOS) or charge coupled device (CCD), may be used. The imaging device may be a low-cost camera. In some cases, the image sensor may be provided on a circuit board. The circuit board may be an imaging printed circuit board (PCB). The PCB may include multiple electronic elements for processing the image signal. For example, the circuit for a CCD sensor may include an analog-to-digital converter and an amplifier for amplifying and converting the analog signal provided by the CCD sensor. Optionally, the image sensor may be integrated with an amplifier and a converter for converting the analog signal to a digital signal, so that a circuit board is not required. In some cases, the output of the image sensor or circuit board may be image data (digital signals) that may be further processed by the camera's camera circuitry or processor. In some cases, the image sensor may comprise an array of optical sensors.
[0129] The illumination device may include one or more light sources 1711 positioned at the distal tip. The light sources may be light emitting diodes (LEDs), organic LEDs (OLEDs), quantum dots (QDs), arrays or combinations of multiple LEDs, OLEDs, or QDs, or any other suitable light source. In some cases, the light sources may include miniature LEDs for compact designs or dual-tone flash LED illumination.
[0130] The imaging and illumination devices may be integrated into the catheter. For example, the distal portion of the catheter may include suitable structure matching at least the dimensions of the imaging and illumination devices. The imaging and illumination devices may be embedded within the catheter. FIG. 18 shows an exemplary distal portion of a catheter with integrated imaging and illumination devices. The camera may be located in the distal portion. The distal tip may have structure to receive the camera, illumination device, or position sensor. For example, the camera may be embedded within a cavity 1810 at the distal tip of the catheter. The cavity 1810 may be integrally formed with the distal portion of the cavity and may have dimensions matching the length / width of the camera so that the camera cannot move relative to the catheter. The camera may be adjacent to the working channel 1820 of the catheter and provide a near-field view of the tissue or organ. In some cases, the attitude or orientation of the imaging device may be controlled by controlling the rotational motion (e.g., roll) of the catheter.
[0131] Power to the camera may be provided by a wired cable. In some cases, the cable wires may be in a wire bundle, providing power to the camera as well as lighting elements or other circuitry at the distal tip of the catheter. The camera or light source may be powered from a power source located in the handle portion via wire, copper wire, or any other suitable means running through the length of the catheter. In some cases, real-time images or video of the tissue or organ may be transmitted wirelessly to an external user interface or display. The wireless communication may be WiFi, Bluetooth, RF communication, or other forms of communication. In some cases, images or video captured by the camera may be broadcast to multiple devices or systems. In some cases, image or video data from the camera may be transmitted to a processor located in the handle portion along the length of the catheter via wire, copper wire, or any other suitable means. Image or video data may be transmitted to an external device / system via wireless communication components in the handle portion. In some cases, the system may be designed so that the wires are not visible or exposed to the operator.
[0132] In traditional endoscopy, illumination may be provided by a fiber optic cable that transmits light from a light source located at the proximal end of the endoscope to the distal end of the robotic endoscope. In some embodiments of the present disclosure, to reduce design complexity, a small LED light may be employed and embedded in the distal section of the catheter. In some cases, the distal section may include a structure 1430 with dimensions that match the dimensions of the small LED light source. As shown in the illustrated example, two cavities 1430 may be integrally formed with the catheter to accommodate two LED light sources. For example, the outer diameter of the distal tip may be approximately 4 to 4.4 millimeters (mm), and the diameter of the catheter's working channel may be approximately 2 mm, allowing two LED light sources to be embedded in the distal end. The outer diameter may be any range smaller than 4 mm or larger than 4.4 mm, and the diameter of the working channel may be any range depending on the dimensions of the tool or the specific application. Any number of light sources may be included. The internal structure of the distal section may be designed to accommodate any number of light sources.
[0133] In some cases, each of the LEDs may be connected to a power wire that may lead to the proximal handle. In some embodiments, the LEDs may be soldered to separate power wires that are later bundled together to form a single strand. In some embodiments, the LEDs may be soldered to pulling wires that provide power. In other embodiments, the LEDs may be crimped or connected directly to a single pair of power wires. In some cases, a protective layer, such as a thin layer of biocompatible adhesive, may be applied to the front surface of the LEDs to provide protection while still allowing light to be emitted. In some cases, an additional cover 1831 may be placed on the advancing end face of the distal tip to allow for precise positioning of the LEDs as well as sufficient room for the adhesive. The cover 1831 may be constructed of a transparent material with a refractive index similar to that of the adhesive so that the illumination light is not obstructed.
[0134] User Interface Example The systems, methods, and techniques described herein may be implemented, at least in part, through the use of a user interface that may be presented on a graphical user interface (e.g., UI 2740 of FIG. 27). FIGS. 19-26 illustrate exemplary user interfaces. At a high level, the user interface may be used to perform and interpret tomosynthesis and augmented fluoroscopy.
[0135] A graphical user interface (GUI) may allow a user to switch between multiple modes in a guidance workflow. In some cases, a user interface for tomosynthesis may be accessible from a user interface for driving or navigation. For example, once a user drives the endoscope via driving or navigation interface 2500, as shown in FIG. 25, the user may select to enter tomosynthesis mode by clicking icon 2501. For example, upon clicking icon 2501 to switch to tomosynthesis mode, the tomosynthesis mode GUI (e.g., 1900 in FIG. 19) is displayed. The tomosynthesis mode GUI may allow the user to return to driving mode at any time, such as by clicking an icon in header 1901.
[0136] From the driving screen 2500, the user can continue to view the camera feed 2505 from the bronchoscope and use the controller to drive through the lungs. The user can choose to configure the driving GUI 2500 and add or remove additional views. For example, the driving screen can be configured to display a virtual endoluminal view 2507 and a virtual lung 2509, which is a computer-generated 3D model of the lung. The user can be permitted to add, remove, or swap out one or more of the other views, such as axial, coronal, and sagittal CT.
[0137] The virtual endoluminal view 2507 provides the user with a computer-generated view of the camera feed, along with graphical elements (e.g., ribbons) showing the path to the currently selected target. In some cases, the path is also represented in a virtual lung 2509. The user can switch to tomosynthesis mode at any given time. For example, once the endoscope tip is within the target biopsy range, the user can enable tomosynthesis mode to help verify the relative distance to the lesion by clicking icon 2501. More details regarding tomosynthesis operations and GUI are provided later in this specification. After the tomosynthesis process is complete, the user can return to the driving screen 2500.
[0138] In some cases, upon completion of tomosynthesis, the virtual endoluminal view may display a floating target based on the results of the tomographic scan. FIG. 26 shows an example of a virtual endoluminal view 2600 displaying a target 2601 along with a graphical element 2603 (e.g., a ribbon) indicating the path to the target. The angle of the target 2615 is displayed as seen from the perspective of the working channel as the tool (e.g., a needle instrument) will exit the bronchoscope. In some cases, the exit axis of the working channel may not be aligned with the axis of the endoscope distal tip (the example in FIG. 17 shows the exit axis 1721 of the working channel 1703). The perspective of the working channel may be based on the known dimensions, structure, or configuration of the distal tip (e.g., the exit axis 1721 of the working channel relative to the endoscope tip is the imaging device 1713) and / or the real-time orientation and position of the distal tip. The angle of the target 2615 relative to the exit axis of the working channel can be determined based at least in part on the layout of the working channel within the distal tip, the real-time position and orientation of the distal tip, and the location of the target. The target and angle arrows 2615 can be useful to assist the user in aligning the tool with the lesion before performing the biopsy. The user can also choose to repeat the tomosynthesis process while the tool is expected to be within the lesion to increase confidence in the biopsy.
[0139] The virtual endoluminal panel displays a rendered view of the internal airway 2600. In some cases, the virtual endoluminal panel may allow the user to enter targeting mode 2610. In some cases, when the user switches to targeting mode 2610, the rendered internal airway may disappear and the target 2611 may be displayed in free space (e.g., depicted as a filled oval shape) when the target is within a predetermined proximity range from the tip. The predetermined proximity range may be determined by the system or may be configurable by the user. In some cases, graphical elements (e.g., crosshairs 2613 and arrows 2615) may appear in the center of the panel with a triangular shape indicator around its edge to show the target's location relative to the direction the scope is pointing.
[0140] In some cases, after tomosynthesis is complete, the automated guidance workflow may allow the user to adjust the position of the lesion (target) based at least in part on the tomosynthesis calculations. In some cases, the tomosynthesis calculations may include a relationship between the position of the lesion and the position of the scope tip. For example, the position of the lesion may be automatically updated based on the relationship between the scope tip and the lesion according to the tomosynthesis calculations. In some cases, the user may toggle tomosynthesis calculation adjustments for the target via a graphical icon 2503 shown in the operating screen 2500 of FIG. 25. When the toggle is on, the position of the target in the virtual lung and virtual endoluminal panels may be adjusted or updated to reflect the calculations made by the tomosynthesis process based on the user selection and the lesion. When the toggle is off, such calculations may be ignored, and the position of the scope tip may rely solely on the EM data, and the position of the target may rely solely on the planned target on the CT scan. In alternative cases, the user may choose to adjust the position of the scope instead of or in addition to adjusting the position of the lesion / target.
[0141] In some cases, enhanced fluoroscopy may be available in fluoroscopy mode after tomosynthesis is complete. A user can enable enhanced fluoroscopy mode, such as via a toggle 2401 displayed in the fluoroscopy panel user interface 2400, to turn on enhanced fluoroscopy mode. Fluoroscopy view mode may be accessed from drive mode during the entire navigation process. For example, a user may switch from drive mode to fluoroscopy view mode via a drive screen. The fluoroscopy view may provide real-time fluoroscopy images / videos. The fluoroscopy panel user interface 2400 may display an enhanced fluoroscopy feature that allows a user to enable / disable enhancements to the fluoroscopy view. For example, if enhanced fluoroscopy is switched on (“enabled”), an overlay of the target / lesion 2403 may be displayed on the fluoroscopy view. In some cases, the option to turn enhanced fluoroscopy on / off may be available even though tomosynthesis is complete. If enhanced fluoroscopy is turned on ("enabled") before completion of tomosynthesis (when target locations are not available), there may be no target / lesion overlay display. Availability of target / lesion information from tomosynthesis can be obtained as described above. For example, lesion information can be broadcast for enhanced fluoroscopy overlay through data contracts between state machines as described above.
[0142] Existing endoscopic systems utilizing tomosynthesis technology may not be compatible with all types of imaging devices (e.g., C-arm systems). For example, current endoscopic systems may be compatible with selected C-arm systems or may require tedious setup for each C-arm system. The endoscopic systems herein employ improved tomosynthesis algorithms as described above that may be compatible with any type of C-arm with minimal or reduced information about the C-arm system. For example, the systems herein may provide a user interface that allows for easy and convenient setup of the C-arm system.
[0143] FIG. 19 shows an example user interface 1900 for a tomosynthesis process dashboard. As shown, the user interface 1900 includes a header, a camera panel, a step indicator, instructions, visual guidance, an end tomography function, and a progress button. The user interface 1900 for tomosynthesis may be accessible from a user interface for driving or navigation. Each of the headers may remain present, and the camera panel may remain visible throughout the tomosynthesis process. A user of the user interface 1900 may be guided through tomosynthesis by screens that may be broken down into a series of steps, along with instructions to the user of where they are in the tomosynthesis process (see step indicators). Instructions and visual guidance in the form of images or videos may be displayed within each screen. At any point during the tomosynthesis process, the user may be able to exit the tomosynthesis screen of the user interface 1900 and return to the driving user interface. The progress button may also allow the user to navigate through the steps of the tomosynthesis process as needed.
[0144] FIG. 20 shows an example user interface 2000 of a C-arm settings dashboard. As shown, the user interface 2000 includes a C-arm dropdown and a C-arm setting. In the user interface 2000, a user can select a connected and compatible C-arm from the dropdown. Once a C-arm is selected, possible settings for that model of C-arm may be displayed. The displayed settings may be default settings, previous settings, recommended settings, optimal settings, etc. Once a C-arm setting is selected (e.g., by a user), the user may be instructed to adjust the C-arm to the selected setting.
[0145] Once the imaging device is configured in the GUI (e.g., user interface 2000), the user may be guided to capture a fluoroscopic image using the imaging device and may be guided to select a scope position via the GUI. FIG. 21 shows an exemplary user interface 2100 for a scope selection dashboard. As shown, user interface 2100 includes a fluoroscopic image and angle control. A fluoroscopic image may be displayed in user interface 2200. The fluoroscopic image may be a 2D image without extension. The user may be able to use a slider shown in user interface 2100 to scroll through different angles of the scope captured from the C-arm and select one or more fluoroscopic images to select the scope position. For example, the user may click on a fluoroscopic image showing the position of the scope tip. FIG. 22 shows an exemplary user interface 2200 for a selection crosshair panel. User interface 2200 may show a more detailed view of the fluoroscopic image in user interface 2100. User interface 2200 may include a selection crosshair. The user interface 2200 can display selection across hairs upon scope selection on a fluoroscopic image displayed in the user interface 2100 showing the location of the scope.
[0146] In some cases, after selecting the scope position, the user may be guided to select the location of the target (e.g., lesion). FIG. 23 shows an exemplary user interface 2300 for a lesion selection (target selection) dashboard. As shown, the user interface 2300 may display a reconstructed tomography 2310, a CT panel 2320, a selection crosshair 2315, a scroll bar 2313, a reset button, a depth indicator 2311, instructions, brightness and contrast controls, and a view angle indicator. In the user interface 2300, once a tomosynthesis scan is acquired, the user may be presented with reconstructed tomography images 2310 from pairs and multiple orientations of the CT 2320. As shown on the user interface 2300, within each set of images, the tomosynthesis images 2310 and the CT scan 2320 may be displayed with their corresponding view angles (e.g., the view angle is shown in the upper left corner). A diagonal line 2315 may be displayed in the user interface 2300 across all scans for the user to mark lesion selection. The layers of each scan can be analyzed via a scroll bar 2313 superimposed on the tomosynthesis image, and the depth of the displayed view 2311 is indicated within the image. The view can be reset to default by clicking the reset button. Instructions and brightness and contrast controls can be provided below the scan to guide the user through the process and allow them to adjust the image view as needed.
[0147] FIG. 24 shows an exemplary user interface 2400 of the extended fluoroscopy panel. As shown, the user interface 2400 includes a user-selected lesion location indicator and an extended fluoroscopy toggle. In some cases, after the tomosynthesis process is completed, a target location overlay is available (e.g., based on the target location determined from tomosynthesis and projected onto the 2D fluoroscopy image as described in FIG. 11 and elsewhere herein), which may enable the extended fluoroscopy feature 2401. The location of the user-selected lesion is shown on the fluoroscopy panel as an overlay on the user interface 2400. The overlay can be toggled (e.g., by the user) via the extended fluoroscopy toggle. However, if the extended fluoroscopy toggle is enabled and the overlay is not available (e.g., the camera pose could not be reconstructed), no changes are displayed on the fluoroscopy view.
[0148] Computer Systems The present disclosure provides computer systems programmed to implement the methods of the present disclosure. Figure 27 illustrates a computer system 2701 programmed or otherwise configured to operate any of the methods, systems, processes, or techniques described herein (including systems or methods for generating tomosynthesis reconstructions or augmented fluoroscopy views described herein). For example, user interface 2740 can present one or more of the user interfaces described with respect to Figures 19-26.
[0149] The computer system 2701 can coordinate various aspects of the present disclosure, such as techniques for tomosynthesis (e.g., tomosynthesis reconstruction) or fluoroscopy (e.g., enhanced fluoroscopy). The computer system 2701 can be a user's electronic device or a computer system located remotely relative to the electronic device. The electronic device can be a mobile electronic device.
[0150] The computer system 2701 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 2705, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 2701 also includes memory or memory locations 2710 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 2715 (e.g., a hard disk), a communication interface 2720 (e.g., a network adapter) for communicating with one or more other systems, and peripheral devices 2725, such as cache, other memory, data storage, or an electronic display adapter. The memory 2710, storage unit 2715, interface 2720, and peripheral devices 2725 communicate with the CPU 2705 via a communication bus (solid lines), such as a motherboard. The storage unit 2715 may be a data storage unit (or data repository) for storing data. The computer system 2701 can be operably coupled to a computer network ("network") 2730 using the communication interface 2720. Network 2730 may be the Internet, an Internet or extranet, or an intranet or extranet in communication with the Internet. Network 2730 may, in some cases, be a telecommunications or data network. Network 2730 may include one or more computer servers, which may enable distributed computing, such as cloud computing. Network 2730 may, in some cases, implement a peer-to-peer network, which may enable devices coupled to computer system 2701 to act as clients or servers, with the aid of computer system 2701.
[0151] The CPU 2705 can execute instructions on a computer-readable medium, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 2710. The instructions may be directed to the CPU 2705, which can then program or otherwise configure the CPU 2705 to implement the methods of the present disclosure. Examples of operations performed by the CPU 2705 may include fetch, decode, execute, and writeback.
[0152] The CPU 2705 may be part of a circuit, such as an integrated circuit. One or more other components of the system 2701 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0153] The storage unit 2715 can store files such as drivers, libraries, and saved programs. The storage unit 2715 can store user data, such as user preferences and user programs. The computer system 2701 can optionally include one or more additional data storage units external to the computer system 2701, such as located on a remote server that communicates with the computer system 2701 via an intranet or the Internet.
[0154] Computer system 2701 can communicate with one or more remote computer systems via network 2730. For example, computer system 2701 can communicate with a remote computer system of a user (e.g., a medical device operator). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 2701 via network 2730.
[0155] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 2701, such as, for example, on memory 2710 or electronic storage unit 2715. The instructions can be code stored on a computer-readable medium and can be provided in the form of software. During use, the code can be executed by the processor 2705. In some cases, the code can be retrieved from the storage unit 2715 and stored in memory 2710 for easy access by the processor 2705. In some situations, the electronic storage unit 2715 can be eliminated, and the machine-executable instructions are stored in memory 2710.
[0156] The code may be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be supplied in a programming language that may be selected to allow the code to be executed in a pre-compiled or as-compiled manner.
[0157] Aspects of the systems and methods provided herein, such as computer system 2701, can be embodied in programming. Various aspects of the present technology may be thought of as “products” or “articles of manufacture,” typically in the form of computer-readable media that store instructions as code or associated data carried on or embodied in some type of computer-readable medium. The machine-executable code can be stored in an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of a computer, processor, or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or portions of the software may, from time to time, be communicated via the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, for example, from an administrative server or host computer to an application server computer platform. Thus, another type of medium that may carry software elements includes optical, electrical, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical landline networks, and over various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered software-bearing media. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable media" refer to any medium or media that participate in providing instructions to a processor for execution.
[0158] Thus, a computer-readable medium such as a computer-executable code may take many forms, including, but not limited to, a tangible storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices in any computer, such as may be used to implement the databases, etc., shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, and optical fiber, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may be in the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, a floppy disk, a flexible disk, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, a DVD or DVD-ROM, any other optical medium, punch cards, paper tape, any other physical storage medium with a pattern of holes, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave carrying data or instructions, a cable or link carrying such a carrier wave, or any other medium from which a computer can read programming code or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0159] The computer system 2701 can include or communicate with an electronic display 2735 with a user interface (UI) 2740 for providing, for example, tomosynthesis (e.g., tomosynthesis reconstruction) or fluoroscopy (e.g., augmented fluoroscopy) data, such as text, video, images, etc. Examples of a UI include, but are not limited to, a graphical user interface (GUI), a web-based user interface, or an application programming interface (API). The UI 2740 can also be used for input, in some cases, via touchscreen functionality.
[0160] The methods, systems, instructions, and techniques of the present disclosure may be implemented by one or more algorithms, which may be implemented by software when executed by the central processing unit 2705. The algorithm may include the steps of: (a) providing a first graphical user interface (GUI) for a tomosynthesis mode and a second GUI for a fluoroscopy view mode for viewing a portion of an endoscopic device and a target within an object; (b) receiving a sequence of fluoroscopy image frames including a portion of an endoscopic device, a marker, and a target, the sequence of fluoroscopy image frames corresponding to various orientations of an imaging system acquiring the sequence of fluoroscopy image frames; (c) upon switching to the tomosynthesis mode, i) performing a uniqueness check on the sequence of fluoroscopy image frames and ii) generating a reconstructed 3D tomosynthesis image based at least in part on the orientation of the imaging system estimated using the markers; and (d) upon switching to the fluoroscopy view mode, i) generating an estimated orientation of the imaging system associated with the fluoroscopy image frames from the sequence of fluoroscopy image frames based at least in part on the markers included in the fluoroscopy image frames, and ii) generating an overlay of the target to be displayed on the fluoroscopy image frames based at least in part on the estimated orientation. In some embodiments, fluoroscopic images for tomosynthesis and extended fluoroscopic models may be acquired using cone beam CT (CBCT).
[0161] In some embodiments, the algorithm may perform operations including: (a) navigating an endoscopic device toward a target within an object in a navigation mode of a graphical user interface (GUI), the GUI displaying a virtual view having visual elements to guide navigation of the endoscopic device; (b) upon switching to a tomosynthesis mode of the GUI, i) receiving a sequence of fluoroscopic image frames corresponding to various orientations of an imaging system that includes a portion of the endoscopic device and the target and that acquires the sequence of fluoroscopic image frames, ii) generating a reconstructed 3D tomosynthesis image based at least in part on the orientation of the imaging system, and iii) determining a position of the target based at least in part on the reconstructed 3D tomosynthesis image; and (c) upon switching to a fluoroscopic view mode of the GUI, i) obtaining an imaging system orientation associated with the fluoroscopic image frames acquired in the fluoroscopic view mode, and ii) generating an overlay of the target to be displayed on the fluoroscopic image frame based at least in part on the orientation of the imaging system and the position of the target determined in (b).
[0162] In some embodiments, the virtual view in the navigation mode includes rendering a graphical representation of the target and an indicator showing the angle of the target relative to the exit axis of the working channel of the endoscopic device when the distal tip of the endoscopic device is determined to be within a predetermined proximity of the target. In some embodiments, the position of the target displayed in the navigation mode is updated based on the position of the target determined in (b). In some embodiments, the pose of the imaging system in the tomosynthesis mode is estimated using markers included in the sequence of fluoroscopic image frames. In some embodiments, the pose of the imaging system in the tomosynthesis mode is measured by one or more sensors.
[0163] In some embodiments, the pose of the imaging system relative to the fluoroscopic image frames in a fluoroscopic view mode is estimated using markers included in the fluoroscopic image frames. In some cases, the markers have a 3D pattern. In some examples, the markers include multiple features arranged on at least two different planes. In some cases, the markers include multiple features of different sizes arranged in a coding pattern. In some examples, the coding pattern includes multiple subareas, each with a unique pattern. In some examples, the pose of the imaging system is estimated by aligning patches of multiple features in the fluoroscopic image frames to the coding pattern.
[0164] In some embodiments, a pose of the imaging system relative to the fluoroscopic image frames in a fluoroscopic view mode is measured by one or more sensors. In some embodiments, in a tomosynthesis mode, the sequence of fluoroscopic image frames is processed by performing a uniqueness check on the sequence of fluoroscopic image frames. In some embodiments, the uniqueness check further includes determining whether a fluoroscopic image frame from the sequence of fluoroscopic image frames is unique based at least in part on an intensity comparison.
[0165] Exemplary Methods 28 shows an exemplary method 2800 for presenting one or both of a tomosynthesis reconstructed image or an enhanced fluoroscopic image in a guidance workflow. The method may include navigating an endoscopic device toward a target via a drive UI 2801, receiving an instruction to switch to tomosynthesis mode from a drive UI 2803, generating a target position and alignment angle for aligning a tool to the target within a tomosynthesis mode UI 2805, receiving a command to switch to a fluoroscopic viewing mode and displaying enhanced fluoroscopic features on a fluoroscopy panel 2807, and, upon enabling the enhanced fluoroscopic features, displaying an overlay indicating the target position on the fluoroscopic view based at least in part on the target position determined during tomosynthesis 2809.
[0166] For example, a user may navigate 2801 the endoscopic device toward a target via a first UI, such as the drive UI described above, and upon receiving a command to switch to tomosynthesis imaging mode, provide a second UI 2803 that displays a tomosynthesis reconstruction, the tomosynthesis reconstruction being generated by: (i) acquiring one or more fluoroscopic images or 2D scans across a region of interest of the patient, where at least a portion of the fluoroscopic images across the region of interest include first image data corresponding to a plurality of markers, and the reconstructed tomosynthesis image includes a plurality of tomosynthesis slices; and displaying an indicator indicating a target position and an angle indicator for aligning a tool with the target, where the target position and angle are determined at least in part based on user input received via the second UI. The tomosynthesis image is reconstructed based on the fluoroscopic images and the plurality of markers.
[0167] The method may include receiving a user input to switch to a fluoroscopy mode. The fluoroscopy mode may provide a third UI displaying an enhanced fluoroscopy function that allows the user to enable / disable an enhanced overlay displayed on the fluoroscopy view. The enhanced fluoroscopy overlay is generated based at least in part on a target location identified in the tomosynthesis imaging. The third UI may be accessed from the first UI. In some embodiments, the fluoroscopy images for the tomosynthesis and the fluoroscopy images for the fluoroscopy views may be acquired using cone-beam CT (CBCT).
[0168] In some cases, the navigation mode UI or driving UI may automatically update once tomosynthesis is complete. For example, the virtual endoluminal view of the driving UI may display a floating target based on the results of a tomographic scan. The virtual endoluminal view may be similar to that illustrated in FIG. 26, in which the target is displayed along with a graphical element (e.g., a ribbon) indicating the path to the target. The angle of the target is also displayed as seen from the perspective of the working channel as the tool (e.g., a needle instrument) exits the bronchoscope. The angle of the target relative to the exit axis of the working channel may be determined based at least in part on the layout of the working channel within the distal tip, the real-time location and orientation of the distal tip, and the target location obtained from the tomosynthesis results. The target and angle arrows may help assist the user in aligning the tool with the lesion before performing the biopsy. The user may also choose to repeat the tomosynthesis process while the tool is expected to be within the lesion to increase confidence in the biopsy.
[0169] The navigation mode UI or driving UI may also provide the user with a targeting mode, as described in FIG. 26 . The user may switch to targeting mode, in which the rendered internal airway may disappear and the target may be displayed in free space (e.g., depicted as a filled oval shape) when the target is within a predetermined proximity range from the tip. The predetermined proximity range may be determined by the system or may be configurable by the user. In some cases, a graphical element (e.g., a crosshair and an arrow) may appear in the center of the panel with a triangular-shaped indicator around its edge to indicate the target's location relative to the direction the scope is pointing. As described above, visual indicators such as the location of the crosshair, arrow, etc. may be determined at least in part based on tomosynthesis results.
[0170] The method 2800 may implement one or more of the systems, methods, computer-readable media, techniques, processes, acts, etc. described herein.
[0171] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. The present invention is not intended to be limited by the specific examples provided herein. While the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. Accordingly, it is contemplated that the present invention also encompasses any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
[0172] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It is understood that various alternatives to the embodiments of the invention described herein may be employed.
[0173] Whenever the terms "at least," "greater than," or "greater than or equal to" precede the first number in a series of two or more numbers, the term "at least," "greater than," or "greater than or equal to" applies to each and every number in the series. For example, 1, 2, or 3 or more is equivalent to 1 or more, 2 or more, or 3 or more.
[0174] Whenever the term "no more than," "less than," or "less than or equal to" precedes the first number in a series of two or more numbers, the term "no more than," "less than," or "less than or equal to" applies to each and every number in the series. For example, 3, 2, or 1 or less is equivalent to 3 or less, 2 or less, or 1 or less.
[0175] Any reference herein to the term "or" is intended to mean what is also known as "inclusive or" or "logic OR," and when used as a logical statement, the phrase "A or B" is understood to be true when either A or B is true, or when both A and B are true, when used as a list of elements. The phrase "A, B, or C" is intended to include all combinations of the elements recited in the phrase, e.g., any element selected from the group consisting of A, B, C, (A,B), (A,C), (B,C), and (A,B,C), and so on, if additional elements are listed. Furthermore, it is also understood that the indefinite article "a" or "an" and the corresponding related definite article "the" or "said," respectively, are intended to mean one or more, unless otherwise stated, implied, or physically impossible. Furthermore, it will be understood that the phrases "at least one of A and B, etc.", "at least one of A or B, etc.", "selected from A and B, etc.", and "selected from A or B, etc." are each intended to mean, for example, any listed element individually or any combination of two or more elements, e.g., any element from the group consisting of "A," "B," and "A and B together," etc.
[0176] Certain inventive embodiments herein contemplate numerical ranges. When a range exists, it includes the endpoints of the range. Furthermore, all subranges and values within the ranges exist as if explicitly written out. The terms "about" or "approximately" can mean within an acceptable error range of a value, which depends in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, in accordance with the practice in the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. When values are described in this application and claims, unless otherwise specified, the term "about" can be assumed to mean within an acceptable error range for the particular value.
[0177] It should be noted that the various example or suggested ranges described herein are specific to those example embodiments and are not intended to limit the scope or reach of the disclosed technology, but again merely provide example ranges of frequencies, amplitudes, etc. associated with those respective embodiments or use cases.
[0178] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. The present invention is not intended to be limited by the specific examples provided herein. While the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. Therefore, it is contemplated that the present invention also encompasses any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. 1. A computer-implemented method for an endoscopic device, comprising: (a) providing a first graphical user interface (GUI) for a tomosynthesis mode and a second GUI for a fluoroscopy view mode for viewing a portion of the endoscopic device and a target within a subject; (b) receiving a sequence of fluoroscopic image frames including a portion of the endoscopic device, a marker, and the target, the sequence of fluoroscopic image frames corresponding to various orientations of an imaging system acquiring the sequence of fluoroscopic image frames; (c) upon switching to the tomosynthesis mode, i) performing a uniqueness check on the sequence of fluoroscopic image frames, and ii) generating a reconstructed 3D tomosynthesis image based at least in part on the pose of the imaging system estimated using the markers; (d) upon switching to the fluoroscopic view mode, i) generating an estimated pose of the imaging system associated with a fluoroscopic image frame from the sequence of fluoroscopic image frames based at least in part on the markers included in the fluoroscopic image frame, and ii) generating an overlay of the target to be displayed on the fluoroscopic image frame based at least in part on the estimated pose. A computer-implemented method comprising:
2. The computer-implemented method of claim 1 , wherein the uniqueness check is not performed in the fluoroscopic view mode.
3. 2. The computer-implemented method of claim 1, wherein the uniqueness check comprises determining whether a fluoroscopic image frame from the sequence of fluoroscopic image frames is unique based at least in part on an intensity comparison.
4. The computer-implemented method of claim 1 , wherein the marker has a 3D pattern.
5. The computer-implemented method of claim 4 , wherein the marker comprises a plurality of features disposed on at least two different planes.
6. The computer-implemented method of claim 1 , wherein the marker has a plurality of features of different sizes arranged in a coded pattern.
7. The computer-implemented method of claim 6 , wherein the coding pattern includes a plurality of sub-areas, each sub-area having a unique pattern.
8. 7. The computer-implemented method of claim 6, wherein in the tomosynthesis mode, the pose of the imaging system is estimated by matching patches of the plurality of features in the sequence of fluoroscopic image frames to the coding pattern.
9. The computer-implemented method of claim 8 , further comprising identifying one or more fluoroscopic image frames having a high pattern match score.
10. 7. The computer-implemented method of claim 6, wherein in the fluoroscopic view mode, the estimated pose of the imaging system is generated by matching patches of the plurality of features in the fluoroscopic image frame to the coding pattern.
11. The computer-implemented method of claim 1 , wherein the first GUI is configured to display the reconstructed 3D tomosynthesis image and receive user input on the reconstructed 3D tomosynthesis image indicating the location of the target.
12. 12. The computer-implemented method of claim 11, wherein the second GUI displays the fluoroscopic image frame with the overlay of the target, and the position of the target displayed on the fluoroscopic image frame is based at least in part on the position of the target.
13. The computer-implemented method of claim 1 , wherein the shape of the overlay is based at least in part on a 3D model of the target projected onto the fluoroscopic image frame based on the estimated pose.
14. The computer-implemented method of claim 13 , wherein the 3D model is generated based on computed tomography images.
15. The computer-implemented method of claim 1 , wherein the second GUI provides a graphical element for enabling or disabling the display of the overlay.
16. 1. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to: (a) providing a first graphical user interface (GUI) for a tomosynthesis mode and a second GUI for a fluoroscopy view mode for viewing a portion of an endoscopic device and a target within a subject; (b) receiving a sequence of fluoroscopic image frames including a portion of the endoscopic device, a marker, and the target, the sequence of fluoroscopic image frames corresponding to various orientations of an imaging system acquiring the sequence of fluoroscopic image frames; (c) upon switching to the tomosynthesis mode, i) performing a uniqueness check on the sequence of fluoroscopic image frames, and ii) generating a reconstructed 3D tomosynthesis image based at least in part on the pose of the imaging system estimated using the markers; (d) upon switching to the fluoroscopic view mode, i) generating an estimated pose of the imaging system associated with a fluoroscopic image frame from the sequence of fluoroscopic image frames based at least in part on the markers included in the fluoroscopic image frame; and ii) generating an overlay of the target to be displayed on the fluoroscopic image frame based at least in part on the estimated pose. A non-transitory computer-readable medium for causing operations to be performed, including:
17. 17. The non-transitory computer-readable medium of claim 16, wherein the uniqueness check is not performed in the fluoroscopic viewing mode.
18. 17. The non-transitory computer-readable medium of claim 16, wherein the uniqueness check comprises determining whether a fluoroscopic image frame from the sequence of fluoroscopic image frames is unique based at least in part on an intensity comparison.
19. The non-transitory computer-readable medium of claim 16 , wherein the marker has a 3D pattern.
20. 20. The non-transitory computer-readable medium of claim 19, wherein the marker comprises a plurality of features disposed on at least two different planes.
21. 17. The non-transitory computer-readable medium of claim 16, wherein the marker has a plurality of features of different sizes arranged in a coded pattern.
22. 22. The non-transitory computer-readable medium of claim 21, wherein the coding pattern includes multiple sub-areas, each sub-area having a unique pattern.
23. 22. The non-transitory computer-readable medium of claim 21 , wherein in the tomosynthesis mode, the pose of the imaging system is estimated by matching patches of the plurality of features in the sequence of fluoroscopic image frames to the coding pattern.
24. 24. The non-transitory computer-readable medium of claim 23, wherein the operations further comprise identifying one or more fluoroscopic image frames having a high pattern match score.
25. 22. The non-transitory computer-readable medium of claim 21 , wherein in the fluoroscopic view mode, the estimated pose of the imaging system is generated by matching patches of the plurality of features in the fluoroscopic image frames to the coding pattern.
26. 17. The non-transitory computer-readable medium of claim 16, wherein the first GUI is configured to display the reconstructed 3D tomosynthesis image and receive user input on the reconstructed 3D tomosynthesis image indicating the location of the target.
27. 27. The non-transitory computer-readable medium of claim 26, wherein the second GUI displays the fluoroscopic image frame with the overlay of the target, and a position of the target displayed on the fluoroscopic image frame is based at least in part on a position of the target.
28. 17. The non-transitory computer-readable medium of claim 16, wherein the shape of the overlay is based at least in part on a 3D model of the target projected onto the fluoroscopic image frame based on the estimated pose.
29. 30. The non-transitory computer-readable medium of claim 28, wherein the 3D model is generated based on computed tomography images.
30. 17. The non-transitory computer-readable medium of claim 16, wherein the second GUI provides a graphical element for enabling or disabling the display of the overlay.
31. 1. A computer-implemented method for an endoscopic device, comprising: (a) navigating the endoscopic device toward a target within an object in a navigation mode of a graphical user interface (GUI), the GUI displaying a virtual view having visual elements to guide navigation of the endoscopic device; (b) upon switching the GUI to a tomosynthesis mode, i) receiving a sequence of fluoroscopic image frames corresponding to various orientations of an imaging system including a portion of the endoscopic device and the target and acquiring the sequence of fluoroscopic image frames, ii) generating a reconstructed 3D tomosynthesis image based at least in part on the orientations of the imaging system, and iii) determining a position of the target based at least in part on the reconstructed 3D tomosynthesis image; (c) upon switching to a fluoroscopic view mode of the GUI, i) obtaining a pose of the imaging system associated with fluoroscopic image frames acquired in the fluoroscopic view mode, and ii) generating an overlay of the target to be displayed on the fluoroscopic image frames based at least in part on the pose of the imaging system and the position of the target determined in (b); A computer-implemented method comprising:
32. 32. The computer-implemented method of claim 31, wherein the virtual view in the navigation mode includes rendering a graphical representation of the target and an indicator showing the angle of the target relative to an exit axis of a working channel of the endoscopic device when the virtual view in the navigation mode determines that the distal tip of the endoscopic device is within a predetermined proximity range of the target.
33. 32. The computer-implemented method of claim 31, wherein the position of the target displayed in the navigation mode is updated based on the position of the target determined in (b).
34. 32. The computer-implemented method of claim 31, wherein the pose of the imaging system in the tomosynthesis mode is estimated using markers included in the sequence of fluoroscopic image frames.
35. 32. The computer-implemented method of claim 31, wherein the orientation of the imaging system in the tomosynthesis mode is measured by one or more sensors.
36. 32. The computer-implemented method of claim 31, wherein the pose of the imaging system relative to the fluoroscopic image frames in the fluoroscopic view mode is estimated using markers included in the fluoroscopic image frames.
37. 37. The computer-implemented method of claim 36, wherein the marker has a 3D pattern.
38. 38. The computer-implemented method of claim 37, wherein the marker comprises a plurality of features disposed on at least two different planes.
39. 38. The computer-implemented method of claim 37, wherein the marker has a plurality of features of different sizes arranged in a coded pattern.
40. 40. The computer-implemented method of claim 39, wherein the coding pattern includes a plurality of sub-areas, each sub-area having a unique pattern.
41. 40. The computer-implemented method of claim 39, wherein the pose of the imaging system is estimated by matching patches of the plurality of features in the fluoroscopic image frames to the coding pattern.
42. 32. The computer-implemented method of claim 31, wherein the pose of the imaging system relative to the fluoroscopic image frame in the fluoroscopic view mode is measured by one or more sensors.
43. 32. The computer-implemented method of claim 31, wherein in the tomosynthesis mode, the sequence of fluoroscopic image frames is processed by performing a uniqueness check on the sequence of fluoroscopic image frames.
44. 44. The computer-implemented method of claim 43, further comprising determining whether a fluoroscopic image frame from the sequence of fluoroscopic image frames is unique based at least in part on an intensity comparison.
45. 1. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to: (a) navigating the endoscopic device toward a target within an object in a navigation mode of a graphical user interface (GUI), the GUI displaying a virtual view having visual elements to guide navigation of the endoscopic device; (b) upon switching the GUI to a tomosynthesis mode, i) receiving a sequence of fluoroscopic image frames corresponding to various orientations of an imaging system including a portion of the endoscopic device and the target and acquiring the sequence of fluoroscopic image frames, ii) generating a reconstructed 3D tomosynthesis image based at least in part on the orientations of the imaging system, and iii) determining a position of the target based at least in part on the reconstructed 3D tomosynthesis image; (c) upon switching to a fluoroscopic view mode of the GUI, i) obtaining a pose of the imaging system associated with a fluoroscopic image frame acquired in the fluoroscopic view mode; and ii) generating an overlay of the target to be displayed on the fluoroscopic image frame based at least in part on the pose of the imaging system and the position of the target determined in (b). A non-transitory computer-readable medium for causing operations to be performed, including:
46. 46. The non-transitory computer-readable medium of claim 45, wherein the virtual view in the navigation mode includes rendering a graphical representation of the target and an indicator showing the angle of the target relative to an exit axis of a working channel of the endoscopic device when the distal tip of the endoscopic device is determined to be within a predetermined proximity range of the target.
47. 46. The non-transitory computer-readable medium of claim 45, wherein the position of the target displayed in the navigation mode is updated based on the position of the target determined in (b).
48. 46. The non-transitory computer-readable medium of claim 45, wherein the pose of the imaging system in the tomosynthesis mode is estimated using markers included in the sequence of fluoroscopic image frames.
49. 46. The non-transitory computer-readable medium of claim 45, wherein the orientation of the imaging system in the tomosynthesis mode is measured by one or more sensors.
50. 46. The non-transitory computer-readable medium of claim 45, wherein the pose of the imaging system relative to the fluoroscopic image frames in the fluoroscopic view mode is estimated using markers included in the fluoroscopic image frames.
51. 51. The non-transitory computer-readable medium of claim 50, wherein the marker has a 3D pattern.
52. 52. The non-transitory computer-readable medium of claim 51, wherein the marker comprises a plurality of features disposed on at least two different planes.
53. 52. The non-transitory computer-readable medium of claim 51, wherein the marker has a plurality of features of different sizes arranged in a coded pattern.
54. 54. The non-transitory computer-readable medium of claim 53, wherein the coding pattern includes multiple sub-areas, each sub-area having a unique pattern.
55. 54. The non-transitory computer readable medium of claim 53, wherein the pose of the imaging system is estimated by matching patches of the plurality of features in the fluoroscopic image frames to the coding pattern.
56. 46. The non-transitory computer readable medium of claim 45, wherein the pose of the imaging system relative to the fluoroscopic image frame in the fluoroscopic view mode is measured by one or more sensors.
57. 46. The non-transitory computer-readable medium of claim 45, wherein in the tomosynthesis mode, the sequence of fluoroscopic image frames is processed by performing a uniqueness check on the sequence of fluoroscopic image frames.
58. 46. The non-transitory computer-readable medium of claim 45, wherein one or more of the operations further comprise determining whether a fluoroscopic image frame from the sequence of fluoroscopic image frames is unique based at least in part on an intensity comparison.