System and method for navigating a catheter through a lumen
The system uses a steerable catheter with an imaging device and machine learning algorithms to generate precise navigation trajectories, addressing the challenge of navigating bronchoscopes to lung lesions by integrating pre-operative and real-time imaging data for improved biopsy accuracy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE BRIGHAM & WOMEN S HOSPITAL INC
- Filing Date
- 2025-10-31
- Publication Date
- 2026-05-07
AI Technical Summary
Surgeons face challenges in accurately navigating a bronchoscope to lung lesions due to similar-appearing anatomical locations in lung airways, necessitating precise registration of pre-operative CT images with real-time camera views during lung biopsies.
A system and method utilizing a steerable catheter with an imaging device, a navigation module, and a computer system to generate predicted trajectories based on pre-operative data and consecutive images, incorporating machine learning algorithms for pose and depth estimation to guide catheter navigation through lumens.
Enhances the accuracy of catheter navigation by providing real-time recommended trajectories, reducing the risk of deviation from the intended path and improving the precision of lung biopsy procedures.
Smart Images

Figure US2025053640_07052026_PF_FP_ABST
Abstract
Description
BWH 2024-548-03Quarles 129319.01124SYSTEM AND METHOD FOR NAVIGATING A CATHETER THROUGH A LUMENCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application is based on and claims priority from U.S. Patent Application Ser. No. 63 / 714,676, filed on October 31, 2024, and U.S. Patent Application Ser. No. 63 / 760,904, filed on February 20, 2025, the entire disclosures of which are each incorporated herein by reference.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] N / ABACKGROUND
[0003] Lung cancer is the most common cause of cancer-related deaths in the United States and is one of the most frequently diagnosed malignancies. Early diagnosis is crucial for improving survival rates. Typically, after an initial computer tomography (CT) screening for a suspicious lesion, a lung biopsy is conducted to confirm the presence of cancer. In a standard lung biopsy procedure, a patient’s pre-operative CT scan is taken and used by the surgeon to plan the path to the suspected lesion. During surgery, the surgeon navigates the bronchoscope to the lesion site following the pre-operative planned path. The surgeon uses the real-time camera image to navigate the bronchoscope, which is challenging due to similar-appearing anatomical locations in lung airways through camera images. Therefore, surgeons refer to the pre-operative CT image to verify the correct path. Therefore, accurate registration of the patient’s CT model to real-time camera view is essential for surgical guidance.SUMMARY
[0004] Accordingly, the present disclosure provides systems, apparatus, and methods for tissue assessment and navigation, and more particularly to exemplary embodiments of a flexible optical imaging probe, a navigation module, and methods for using the same.1QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0005] In one embodiment, a system for generating predicted trajectories for navigating a catheter through a lumen, including: a steerable catheter including a proximal end, a distal end, a catheter tip, and a tool channel extending from the proximal end to the distal end; an imaging device disposed within the steerable catheter; a console including a system controller, a display controller, and a display; and a navigation module configured to: receive preoperative data and multiple consecutive images from the imaging device from the system controller, and output information based on the received preoperative data and the multiple consecutive images, the information including at least one of: a location, a direction, or a recommended trajectory for navigating the catheter through the lumen.
[0006] In another embodiment, a system for generating predicted trajectories for navigating a catheter through a lumen, including: a steerable catheter including a proximal end, a distal end, a catheter tip, and a tool channel extending from the proximal end to the distal end; an imaging device disposed within the tool channel; a display; and a computer system operatively connected to the imaging device and the display, the computer system configured to: receive preoperative data and multiple consecutive image frames acquired by the imaging device, estimate pose and depth of the catheter based on the preoperative data and the multiple consecutive image frames, generate a recommended trajectory for the catheter to follow to navigate the lumen based on the estimated pose and depth, and output the recommended trajectory to the display.
[0007] In yet another embodiment, a method for tracking catheter location for navigation of a catheter through a lumen, including: determining a predicted pose iteration including: receiving, by a computer system, preoperative data and multiple consecutive images acquired by a catheter including an imaging tool; and outputting, by the computer system, information based on the received preoperative data and the multiple consecutive images, the information including at least one of: a location, a direction, or a recommended trajectory for navigating the catheter through the lumen.
[0008] In still another embodiment, a method for generating predicted trajectories for navigating a catheter through a lumen, including: providing a steerable catheter including a proximal end, a distal end, a catheter tip, and a tool channel extending from the proximal end to2QB\99263808.1BWH 2024-548-03Quarles 129319.01124 the distal end, wherein an imaging device is disposed within the tool channel; receiving, from a computer system operatively connected to the imaging device and a display, preoperative data and multiple consecutive image frames acquired by the imaging device; estimating, using the computer system, pose and depth of the catheter based on the preoperative data and the multiple consecutive image frames; generating, using the computer system, a recommended trajectory for the catheter to follow to navigate the lumen based on the estimated pose and depth; and outputting, using the computer system, the recommended trajectory to the display.
[0009] In yet another embodiment, a method of training a machine learning algorithm, including: initializing, using a computer system, a network architecture configured with a plurality of encoders, decoders, and layers; accessing, using the computer system, training data including preoperative data and imaging data; and developing, using the computer system, weights and biases of the plurality of encoders, decoders, and layers based on the training data, wherein developing the weights and biases of the network architecture is completed based on minimizing a loss function.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Various objects, features, and advantages of the disclosed subject matter can be more fully appreciated with reference to the following detailed description of the disclosed subject matter when considered in connection with the following drawings, in which like reference numerals identify like elements.
[0011] FIG. 1 shows a schematic representation of a robotic catheter system being used in a medical environment.
[0012] FIG. 2 shows a schematic representation of robotic catheter system including system controller and display controller.
[0013] FIGS. 3A-3C show schematic diagrams of bendable sections of a robotic catheter system. FIG. 3A shows an exemplary construction of a steerable catheter with a ring-shaped component. FIG. 3B and FIG. 3C show exemplary catheter tip manipulations by actuating one or more bending segments of the steerable catheter. FIG. 3B shows manipulation of the distal3QB\99263808.1BWH 2024-548-03Quarles 129319.01124 segment of the steerable section. FIG. 3C shows manipulation of the middle segment of the steerable section.
[0014] FIGS. 4A and 4B show schematic diagrams of embodiments of a controller.
[0015] FIG. 5 shows a schematic diagram of a system controller and / or display controller.
[0016] FIG. 6 shows a schematic diagram of a navigation module.
[0017] FIG. 7 shows a flow chart of a method to generate a recommended trajectory for a catheter navigating a lumen.
[0018] FIG. 8 shows a schematic of a computer system.
[0019] FIGS. 9A-9B show examples of the effect of incorrect pixel intensities on depth estimation. FIG. 9A shows an original image with the incorrect camera intensities on the image periphery. FIG. 9B shows an incorrect depth estimated at the image periphery using the selfsupervised pose and depth estimate method.
[0020] FIG. 10 shows an exemplary self-supervised camera artifact estimation (CAM) network.
[0021] FIG. 11 shows an exemplary SS-PDE network architecture and training pipeline.
[0022] FIGS. 12A-12B show an example of a method of centerline correction for translation and rotational components. FIG. 12A shows translation correction, and FIG. 12B shows rotational correction.
[0023] FIG. 13 shows trajectories generated from predicted poses with centerline correction (green, left) and without centerline correction (yellow, right). The trajectory generation from predicted poses without centerline correction (yellow, right) drifts out of the airway.
[0024] FIG. 14 shows airway labeling in Case 2 (female, 88 years old, EM-navigated transbronchial biopsy). The color-coded spheres represent the bronchoscope tip location, while the colored lines denote airway centerlines, with each color corresponding to a specific airway segment. SS-DPE accurately identifies the bronchoscope’s location within the airways as it4QB\99263808.1BWH 2024-548-03Quarles 129319.01124 advances from the right main bronchus (brown, right-hand portion of inset) to subsequent bronchial segments (green to blue, center to left-hand portion of inset).
[0025] FIG. 15 shows a comparison of 3cGAN and SS-DPE methods on localization error. Results are statistically significant, *p<0.05, by t-test.
[0026] FIG. 16 shows a comparison of 3cGAN and SS-DPE methods on correct Airway Label Percentage.
[0027] FIGS. 17A-17D show a comparison of 3cGAN and SSDPE methods on localization error in different lobes of the lung. FIG. 17A. shows right upper lobe (RUL). FIG. 17B shows left upper lobe (LUL). FIG. 17C shows right lower lobe (RLL). FIG. 17D shows left lower lobe (LLL). SS-DPE has significantly less localization error in all lung lobes except for in the right upper lobe, where results are comparable.
[0028] FIGS. 18A-18D show trajectories generated by the 3cGAN methods (left image; trajectory shown in green) and SS-DPE method (right image; trajectory shown in green) against the reference electromagnetic tracking (EM) data (left and right images; trajectory shown in red). FIG. 18A shows the right upper lobe. FIG. 18B shows the left upper lobe. FIG. 18C shows the right lower lobe. FIG. 18D shows the left lower lobe. Overall, the SS-DPE trajectories remain closer to the reference EM data than those produced by 3cGAN.
[0029] FIGS. 19A-19D show an illustration of SS-DPE method integrated into 3D Slicer software for navigational bronchoscopy using a phantom derived from Case 6, a 67-year-old male who underwent robotically guided transbronchial biopsy. FIG. 19A shows the lung airway model generated from the CT scan. The green line represents the ground truth (GT) trajectory obtained from EM tracking, while the yellow dot indicates the predicted camera location in CT coordinate space. FIG. 19B shows an image of the camera view. FIG. 19C shows the depth map predicted by DepthNet within SS-DPE. FIG. 19D shows the virtual lung model generated from the CT scan.5QB\99263808.1BWH 2024-548-03Quarles 129319.01124DETAILED DESCRIPTION
[0030] In accordance with some embodiments of the disclosed subject matter, mechanisms (which can include, for example, systems and methods) for generating predicted trajectories for navigating a catheter through a lumen are provided.
[0031] Robotic Catheter System
[0032] An embodiment of a robotic catheter system 100 is described in reference to FIG. 1 through FIG. 5. FIG. 1 illustrates a simplified representation of a medical environment, such as an operating room, where a robotic catheter system 100 can be used. FIG. 2 illustrates a functional block diagram of the robotic catheter system 100. FIGS. 3A-C represents the catheter and bending. FIGS. 4A-B illustrates a logical block diagram of the robotic catheter system 100. In this example, the system 100 includes a system console 102 (computer cart) operatively connected to a steerable catheter 104 via a robotic platform 106. The robotic platform 106 includes one or more than one robotic arm 108 and a linear translation stage 110.
[0033] In FIG. 1, a user 112 (e.g., a physician) controls the robotic catheter system 100 via a user interface unit (operation unit) to perform an intraluminal procedure on a patient 114 positioned on an operating table 116. The user interface may include at least one of a main display 118 (a first user interface unit), a secondary display 120 (a second user interface unit), and / or a handheld controller 124 (a third user interface unit). The main display 118 may include, for example, a large display screen attached to the system console 102 or mounted on a wall of the operating room and may be, for example, designed as part of the robotic catheter system 100 or be part of the operating room equipment. Optionally, there is a secondary display 120 that is a compact (portable) display device configured to be removably attached to the robotic platform 106. Examples of the secondary display 120 include a portable tablet computer or a mobile communication device (a cellphone).
[0034] The steerable catheter 104 is actuated via an actuator unit 122. The actuator unit 103 is removably attached to the linear translation stage 110 of the robotic platform 106. The handheld controller 124 may include a gamepad-like controller with a joystick having shift levers and / or push buttons. It may be a one-handed controller or a two-handed controller. In one6QB\99263808.1BWH 2024-548-03Quarles 129319.01124 embodiment, the actuator unit 122 is enclosed in a housing having a shape of a catheter handle. One or more access ports 126 are provided in or around the catheter handle. The access port 126 is used for inserting and / or withdrawing end effector tools and / or fluids when performing an interventional procedure of the patient 114.
[0035] The system console 102 includes a system controller 128, a display controller 130, and the main display 118. The main display 118 may include a conventional display device such as a liquid crystal display (LCD), an OLED display, a QLED display or the like. The main display 118 provides a graphic interface unit (GUI) configured to display one or more views. These views include live view image 132, an intraoperative image 134, and a preoperative image 136, and other procedural information 138. Other views that may be displayed include a model view, a navigational information view, and / or a composite view. The live image view 132 may be an image from a camera at the tip of the catheter. This view may also include, for example, information about the perception and navigation of the catheter 104. This view may provide information of a recommended trajectory to navigate a lumen. The preoperative image 136 may include pre-acquired 3D or 2D medical images of the patient acquired by conventional imaging modalities such as computer tomography (CT), magnetic resonance imaging (MRI), or ultrasound imaging. The intraoperative image 134 may include images used for image guided procedure such images may be acquired by fluoroscopy or CT imaging modalities. Intraoperative image 134 may be augmented, combined, or correlated with information obtained from a sensor, camera image, or catheter data.
[0036] In the various embodiments where a catheter tip tracking sensor 140 is used, the sensor may be located at the distal end of the catheter. The catheter tip tracking sensor 140 may be, for example, an electromagnetic (EM) sensor. If an EM sensor is used, a catheter tip position detector 142 is included in the robotic catheter system 100; this catheter tip position detector would include an EM field generator operatively connected to the system controller 128. Suitable electromagnetic sensors for use with a steerable catheter are well-known and described, for example, in U.S. Pat. No.: 6,201,387 and international publication WO 2020 / 194212A1, each of which is incorporated herein by reference.7QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0037] In various embodiments, a catheter tip with a tracking sensor such as an EM sensor is used to generate ground truth data for final testing to verify how well the method works. In cases such as these, the performance of a catheter system that includes a trained machine learning algorithm is compared with the performance of a catheter system with an EM sensor.
[0038] Similar to FIG. 1, the diagram of FIG. 2 illustrates the robotic catheter system 100 includes the system controller 128 operatively connected to the display controller 130, which is connected to the display unit 118, and to the hand held control 124. The system controller 128 is also connected to the actuator unit 122 via the robotic platform 106, which includes the linear translation stage 110. The actuator unit 122 includes a plurality of motors 144 that control the plurality of drive wires 160. These drive wires travel through the steerable catheter 104. One or more access ports 126 may be located on the catheter. The catheter includes a proximal section 148 located between the actuator and the proximal bending section 152 where they actuate the proximal bending section. Three of the six drive wires 160 continue through the distal bending section 1 6 where they actuate this section and allow for a range of movement. This figure is shown with two bendable sections (152 and 156). Other embodiments as described herein can have three bendable sections (see FIG. 3). In some embodiments, a single bending section may be provided, or alternatively, four or more bendable sections may be present in the catheter.
[0039] FIG. 3 A shows an exemplary embodiment of a steerable catheter 104. The steerable catheter 104 includes a non-steerable proximal section 148, a steerable distal section 150, and a catheter tip 158. The proximal section 148 and distal bendable section 150 (including 152, 154 and 156) are joined to each other by a plurality of drive wires 160 arranged along the wall of the catheter. The proximal section 148 is configured with thru-holes or grooves or conduits to pass drive wires 160 from the distal section 150 to the actuator unit 122. The distal section 150 is comprised of a plurality of bending segments including at least a distal segment 156, a middle segment 154, and a proximal segment 152. Each bending segment is bent by actuation of at least some of the plurality of drive wires 160 (driving members). The posture of the catheter may be supported by non-illustrated supporting wires (support members) also arranged along the wall of the catheter (see U.S. Pat. Pub. US2021 / 0308423, incorporated herein by reference). The proximal ends of drive wires 160 are connected to individual actuators or8QB\99263808.1BWH 2024-548-03Quarles 129319.01124 motors 144 of the actuator unit 122, while the distal ends of the drive wires 160 are selectively anchored to anchor members in the different bending segments of the distal bendable section 150.
[0040] Each bending segment is formed by a plurality of ring-shaped components (rings) with thru-holes, grooves, or conduits along the wall of the rings. The ring-shaped components are defined as wire-guiding members 162 or anchor members 164 depending on their function within the catheter. Anchor members 164 are ring-shaped components onto which the distal end of one or more drive wires 160 are attached. Wire-guiding members 162 are ring-shaped components through which some drive wires 160 slide through (without being attached thereto).
[0041] Detail “A” in FIG. 3 A illustrates an exemplary embodiment of a ring-shaped component (a wire-guiding member 162 or an anchor member 164). Each ring-shaped component includes a central opening which forms the tool channel 168, and plural conduits 166 (grooves, sub-channels, or thru-holes) arranged lengthwise equidistant from the central opening along the annular wall of each ring-shaped component. Inside the ring-shaped component, an inner cover such as is described in US 2021 / 0369085 and US 2022 / 0126060 (each of which is incorporated herein by reference), may be included to provide a smooth inner channel and provide protection. The non-steerable proximal section 148 is a flexible tubular shaft and can be made of extruded polymer material. The tubular shaft of the proximal section 148 also has a central opening or tool channel 168 and plural conduits 166 along the wall of the shaft surrounding the tool channel 168. An outer sheath may cover the tubular shaft and the steerable section 150. In this manner, at least one tool channel 168 formed inside the steerable catheter 104 provides passage for an imaging device and / or end effector tools from the insertion port 126 to the distal end of the steerable catheter 104. In some embodiments, an imaging device is disposed within the steerable catheter 104. In some embodiments, an imaging device is disposed in a tool channel 168.
[0042] The actuator unit 122 includes one or more servo motors or piezoelectric actuators. The actuator unit 122 bends one or more of the bending segments of the catheter by applying a pushing and / or pulling force to the drive wires 160. As shown in FIG. 3 A, each of the three bendable segments of the steerable catheter 104 has a plurality of drive wires 160. If each9QB\99263808.1BWH 2024-548-03Quarles 129319.01124 bendable segment is actuated by three drive wires 160, the steerable catheter 104 has nine driving wires arranged along the wall of the catheter. Each bendable segment of the catheter is bent by the actuator unit 122 by pushing or pulling at least one of these nine drive wires 160. Force is applied to each individual drive wire in order to manipulate / steer the catheter to a desired pose. The actuator unit 122 assembled with steerable catheter 104 is mounted on the linear translation stage 110. Linear translation stage 110 includes a slider and a linear motor. In other words, the linear translation stage 110 is motorized, and can be controlled by the system controller 128 to insert and remove the steerable catheter 104 to / from the patient’s bodily lumen.
[0043] An imaging device 170 that can be inserted through the tool channel 168 includes an endoscope camera (videoscope) along with illumination optics (e.g., optical fibers or LEDs). The illumination optics provides light to irradiate the lumen and / or a lesion target which is a region of interest within the patient. End effector tools refer endoscopic surgical tools including clamps, graspers, scissors, staplers, ablation or biopsy needles, and other similar tools, which serve to manipulate body parts (organs or tumorous tissue) during examination or surgery. The imaging device 170 may be what is commonly known as a chip-on-tip camera and may be color or black-and-white.
[0044] In some embodiments, a tracking sensor 140 (e.g., an EM tracking sensor) is attached to the catheter tip 158. In this embodiment, steerable catheter 104 and the tracking sensor 140 can be tracked by the tip position detector 142. Specifically, the tip position detector 142 detects a position of the tracking sensor 140, and outputs the detected positional information to the system controller 100. The system controller 128, receives the positional information from the tip position detector 142, and continuously records and displays the position of the steerable catheter 104 with respect to the patient’s coordinate system. The system controller 128 controls the actuator unit 122 and the linear translation stage 110 in accordance with the manipulation commands input by the user 112 via one or more of the user interface units (the handheld controller 124, a GUI at the main display 118 or touchscreen buttons at the secondary display 120). A tracking sensor 140 may be used for training a machine learning algorithm. A catheter with an EM tracking sensor may be used to generate training data to train a navigation module.10QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0045] FIG. 3B and FIG. 3C show exemplary catheter tip manipulations by actuating one or more bending segments of the steerable catheter 104. As illustrated in FIG. 3B, manipulating only the most distal segment 156 of the steerable section changes the position and orientation of the catheter tip 158. On the other hand, manipulating one or more bending segments (152 or 154) other than the most distal segment affects only the position of catheter tip 158, but does not affect the orientation of the catheter tip. In FIG. 3B, actuation of distal segment 155 changes the catheter tip from a position Pl having orientation 01, to a position P2 having orientation 02, to position P3 having orientation 03, to position P4 having orientation 04, etc. In FIG. 3C, actuation of the middle segment 154 changes the position of catheter tip 158 from a position Pl having orientation 01 to a position P2 and position P3 having the same orientation 01. Here, it should be appreciated by those skilled in the art that exemplary catheter tip manipulations shown in FIG. 3B and FIG. 3C can be performed during catheter navigation (e.g., while inserting the catheter through tortuous anatomies). In the present disclosure, the exemplary catheter tip manipulations shown in FIG. 3B and FIG. 3C apply namely to the targeting mode applied after the catheter tip has been navigated to a predetermined distance (a targeting distance) from the target.
[0046] FIG. 4A illustrates the system controller 128 executes software programs and controls the display controller 130 to display a navigation screen (e.g., a live view image 132) on the main display 118 and / or the secondary display 120. The display controller 130 may include a graphics processing unit (GPU) or a video display controller (VDC). FIG. 4B provides another embodiment of a system controller 128. While a tracking sensor 140 may be provided, it is not necessary for the catheter navigation as provided herein, and is not a necessary component of the system. The robotic catheter (also described as a steerable catheter) 104 is controlled through an actuator unit 122. The system controller 128 is shown in FIG. 4B as containing the SS-DPE 129. This embodiment additionally provides an optional voice detector 125 for embodiments where voice control is used to support navigation of the robotic catheter. (See WO application Ser. No. PCT / US24 / 37930, herein incorporated by reference).
[0047] FIG. 5 illustrates components of the system controller 128 and / or the display controller 130. The system controller 128 and the display controller 130 can be configured separately. Alternatively, the system controller 128 and the display controller 102 can be11QB\99263808.1BWH 2024-548-03Quarles 129319.01124 configured as one device. In either case, the system controller 128 and the display controller 130 comprise substantially the same components. Specifically, the system controller 128 and display controller 130 may include a central processing unit (CPU 182) comprised of one or more processors (microprocessors), a random-access memory (RAM 184) module, an input / output (I / O 186) interface, a read only memory (ROM 180), and data storage memory (e.g., a hard disk drive (HDD 188) or solid-state drive (SSD)).
[0048] The ROM 180 and / or HDD 188 store the operating system (OS) software, and software programs necessary for executing the functions of the robotic catheter system 100 as a whole. The RAM 184 is used as a workspace memory. The CPU 182 executes the software programs developed in the RAM 184. The I / O 186 inputs, for example, positional information to the display controller 130, and outputs information for displaying the navigation screen to the one or more displays (main display 118 and / or secondary display 120). In the embodiments descried below, the navigation screen is a graphical user interface (GUI) generated by a software program but, it may also be generated by firmware, or a combination of software and firmware.
[0049] The system controller 128 may control the steerable catheter 104 based on any known kinematic algorithms applicable to continuum or snake-like catheter robots. For example, the system controller controls the steerable catheter 104 based on an algorithm known as follow the leader (FTL) algorithm. By applying the FTL algorithm, the most distal segment 156 of the steerable section 150 is actively controlled with forward kinematic values, while the middle segment 154 and the proximal segment 152 (following sections) of the steerable catheter 104 move at a first position in the same way as the distal section moved at the first position or a second position near the first position.
[0050] Additionally or alternatively, the system controller 128 may control the steerable catheter 104 based on a navigation module, such as a SS-DPE method. The SS-DPE method may determine predicted trajectories for the catheter system to follow.
[0051] The display controller 130 acquires position information of the steerable catheter 104 from system controller 102. Alternatively, the display controller 130 may acquire the position information directly from the tip position detector 142. The steerable catheter 104 may12QB\99263808.1BWH 2024-548-03Quarles 129319.01124 be a single-use or limited-use catheter device. In other words, the steerable catheter 104 can be attachable to, and detachable from, the actuator unit 122 to be disposable.
[0052] During a procedure, the display controller 130 can generate and outputs a live- view image or other view(s) or a navigation screen to the main display 118 and / or the secondary display 120. This view can optionally be registered with a 3D model of a patient’s anatomy (a branching structure) and the position information of at least a portion of the catheter (e.g., position of the catheter tip 158) by executing pre-programmed software routines. Upon completing navigation to a desired target, one or more end effector tools can be inserted through the access port 126 at the proximal end of the catheter, and such tools can be guided through the tool channel 168 of the catheter body to perform an intraluminal procedure from the distal end of the catheter.
[0053] The tool may be a medical tool such as an endoscope camera, forceps, a needle or other biopsy or ablation tools. In one embodiment, the tool may be described as an operation tool or working tool. The working tool is inserted or removed through the working tool access port 126. In the embodiments below, an embodiment of using a steerable catheter to guide a tool to a target is explained. The tool may include an endoscope camera or an end effector tool, which can be guided through a steerable catheter under the same principles. In a procedure there is usually a planning procedure, a registration procedure, a targeting procedure, and an operation procedure.
[0054] Navigation System
[0055] Described herein is a navigation system 200 used to generate recommended trajectories to navigate a lumen (see FIG. 6). Navigation system 200 is also referred to as a navigation module. Navigation module 200 is connected with system console 102 as well as displays including main display 118 and / or second display 120. In some cases, the navigation module is referred to as SS-DPE module or SS-DPE method.
[0056] Navigation module 200 receives image frames from imaging device 170. These images may be live view images 130 and / or intraoperative images 132. Navigation module 200 receives multiple consecutive images from a time series or video. Optionally, navigation module 200 receives preoperative image 136. Preoperative image 136 may be a CT image. Preoperative image 136 may be a CT image of a lumen a user wishes to navigate using a catheter system.13QB\99263808.1BWH 2024-548-03Quarles 129319.01124Based on the image frames from imaging device 170 and preoperative image 136, navigation module 200 generates trajectories to navigate a lumen. Navigation module 200 may optionally receive other procedural information 138 to generate trajectories.
[0057] The image frames may be made up of a first and second image. Navigation module 200 can predict an incremental pose based on the image frames. An incremental pose is a pose calculated for a specific timepoint corresponding to image frames at a given time. The incremental pose may be calculated using live image data. For instance, let the first frame and second frame be source and target frames, and It. Navigation module 200 can predict an incremental pose for Ii and I2, and update the predicted incremental pose for I2 and I3, then I3 and I4, etc. The incremental pose is calculated once each iteration. An iteration can be consecutive frames. Alternatively, the iteration can be non-consecutive frames. An iteration can be every frame, every other frame, every 3 frames, every 4 frames, every 5 frames, every 6 frames, every 7 frames, every 8 frames, every nine frames, or every ten frames. In some embodiments, the iteration is three frames when the collection speed is ten frames per second.
[0058] The first image, or source image, may be a combination of a plurality of image frames (e.g., a composite image). The plurality of images may include 2 frames, 3 frames, 4 frames, 5 frames, 6 frames, 7 frames, 8 frames, 9 frames, 10 frames, 15 frames, or 20 frames. The frames may be averaged together. In some embodiments, the first image and second image may be consecutive frames. In some embodiments, an image may be excluded if the quality of the image is too low. The quality of an image can be determined automatically (e.g., quantified and compared to a threshold) or manually (e.g., decided on by a skilled user). In these embodiments, the second image can be one later frame (e.g., the incremental pose may be calculated for I3 and Ii, while frame I2 may not be used). Skipping frames helps capture objects moving at different speeds. Skipping frames is equivalent to capturing the higher motion or higher speed in a video.
[0059] Navigation module 200 may include a self supervised depth and pose estimation module (SS-DPE module) 202. Navigation module 200 and / or SS-DPE module 202 may further include a trained model, which may include at least one of a pose consistent loss (PC-Loss) module 204, a long term loss (LT-Loss) module 206, a centerline correction module 208, or a14QB\99263808.1BWH 2024-548-03Quarles 129319.01124 camera artifact masking (CAM) module 210. CAM module 210 module may include trained algorithms or networks. As such, the CAM module 204 may be interchangeably referred to as “CAM network.” Similarly, SS-DPE 202 module may be referred to as “SS-DPE network.”
[0060] Self-supervised Depth Pose Estimation Module
[0061] SS-DPE module 202 is used to estimate the relative pose and depth between camera frames (e.g., the incremental pose and incremental depth). SS-DPE module 202 may comprise a neural network. In some embodiments, SS-DPE comprises two neural networks: PoseNet, which is used to predict camera pose, and DepthNet, which is used to predict camera depth.
[0062] Let the two consecutive image frames be source and target frames, Is and It, respectively. ‘Msis the camera pose from the target to source 3D coordinate frame. It comprises of 3D rotation, R and translation, T, between the target and the source frame as given in Eq. 2.(Eq. 2)sMt= [YJ]
[0063] Let Dt be the depth image of the target frame and K be the camera intrinsic matrix, which is pre-computed and is assumed to be fixed for the whole image dataset. Then, we can transform a 2D image pixel, Xt in the target frame to a 3D point, Xt by Eq. 3 :
[0064] The 3D backprojected point At can be re-projected in the source image Is^tas shown in Eq. 4:(Eq. 4) xs= [K\0]Mt^sXt
[0065] Here, the pose estimation network, PoseNet estimates MHS, parameterized as the three translation and 3 Euler angle components. The depth estimation network, DepthNet estimates the scene depth, Dt.
[0066] In some embodiments, SS-DPE module 202 is trained to minimize more than one loss metric. In some embodiments, SS-DPE module 202 is trained to minimize three loss metrics. The three loss metrics may be photometric loss, forward-backward pose consistency loss (PC -loss, may include PC-Loss module 204) and long-term pose consistency loss (LT-Loss, may15QB\99263808.1BWH 2024-548-03Quarles 129319.01124 include LT-Loss module 206). The SS-DPE module may output pose estimates based on inputs from CAM module 210.
[0067] The SS-DPE network provides relative pose estimates. Using these pose estimates, a trajectory may be obtained by consecutively multiplying the relative pose estimates. The estimated trajectory was scaled by multiplying it with a scale factor. The scale factor was calculated by taking the ratio of the predicted trajectory length to the ground truth (GT) trajectory length.
[0068] Centerline Correction Module
[0069] Centerline correction module 208 may be used to refine the trajectories generated by the SS-DPE module. Centerline correction further improves the predicted relative pose estimates (e.g., the incremental poses) by using the prior information of the centerlines of branch lung airways in the CT model. The centerline correction module can register the imaging data (e.g., the first and second image) to the preoperative data, and adjust the incremental pose to a position towards a closest point on the centerline.
[0070] The trajectory generated from the predicted relative pose estimates is drifts away from the target trajectory, as the error from the previous pose estimates accumulate over the subsequent camera poses. To correct for the drift, the translation and rotational components are adjusted to be closer towards the nearest centerline. For translation, the closest points on the centerline, Ca and Cb corresponding to two consecutive estimated camera positions a and b are located. Then the new camera position, b, was t distance from the original estimated camera position b on the perpendicular between the original camera position and the projected point as shown in FIG. 9A. If the distance between the a and b did not match a and b, we added the residual distance to obtain the final camera position as bnew, as shown in FIG. 12A. For rotational correction, the rotation angle between the vectors defined by camera positions Ca and Cb and a and b is calculated, and a rotation matrix with the scaled rotational angle, r is obtained. The rotational matrix thus obtained was used to transform the obtained estimated camera position, as shown in FIG. 12B.
[0071] In some embodiments, the centerline is calculated prior to a procedure. Throughout the procedure of navigating a lumen, the predicted position of the steerable catheter16QB\99263808.1BWH 2024-548-03Quarles 129319.01124 and / or imaging device is translated towards the closest position on the centerline. The movement may be capped by a maximum translation value for the movement per frame. A maximum translation value may range between 0.5 mm to 1.5 mm per frame, when images are collected at ten frames per second.
[0072] The centerline correction module may adjust the incremental pose towards the point on the centerline at every iteration. Alternatively, the incremental pose can be adjusted towards the centerline periodically at a frequency less than a frequency of the predicted incremental pose iteration. For example, if the frequency of the predicted incremental pose iteration is every frame, the frequency of the period may be every other frame. Periodically can refer to any length of time equal to or greater than the time used for an iteration. The period can be every second iteration, every third iteration, every fourth iteration, every fifth iteration, every seventh iteration, every eight iteration, every ninth iteration, or every tenth iteration.
[0073] CAM module
[0074] CAM module 210 generates a mask to remove the corrupt pixel intensities from the SS-DPE module. CAM module may or may not be used for training the SS-DPE network. In some embodiments, CAM module 210 masks at least a portion of the first image, and at least a portion of the second image. In some embodiments, the CAM module applies the same mask to the first and second image. In some embodiments, the CAM module applies a different mask to the first image and the second image. The SS-DPE method gives the 6 degrees of freedom (6DOF) pose between two consecutive image frames by learning the inter-frame and intra-frame intensity differences occurring because of the motion and the 3D structure of the scene, respectively. Therefore, for the SS-DPE method, inter-frame intensity change is a strong cue for learning the relative 3D pose. However, camera images often have corrupt intensities at the periphery because of the positioning of the camera lens relative to the imaging plane. In order to mitigate the effect of these intensities on network training, a practical self- supervised method, referred to as CAM, may be used. CAM module 210 assumes that the underlying artifact is a combination of multiple Gaussian masks and it is possible to decompose the original image into an artifact-free image and a combination of 2D Gaussians. In some embodiments, the image artifacts are modeled as a combination of four 2D Gaussian masks. In some embodiments, the17QB\99263808.1BWH 2024-548-03Quarles 129319.01124 image artifacts are modeled using Gaussian masks with at least four parameters. In some embodiments, the mask is generated in order to remove artifacts at a peripheral portion of one or more of the images, particularly where the camera images have corrupt intensities at the periphery due to the positioning of the camera lens relative to the imaging plane.
[0075] CAM module 204 may include multiple neural networks. In one embodiment, CAM module 210 includes two neural networks: a Gauss-Network trained to estimate image artifacts, and a Decomposition Network trained to generate an artifact-free image. In some embodiments, the Decomposition Network has a U-Net architecture.
[0076] In some embodiments, adjusting the incremental pose is an adjustment relative to an estimated radius of a lumen. The estimated radius of the lumen may be obtained from a preoperative image, such as a CT image. In some embodiments, the radius of the lumen is obtained from the difficulty index as described in U.S. Pat. Pub. 2023 / 0277245, herein incorporated by reference in its entirety.
[0077] In some embodiments, the navigation system, including the CAM module, SS- DPE module, and Centerline correction module, are collectively referred to as “pose estimation module” or “SS-DPE method.” This is not meant to be limiting to the specific SS-DPE module, but the overall method comprising the three modules.
[0078] In some embodiments, the navigation system only generates an incremental pose estimate. Generating an incremental depth estimate, or generating a depth map, are optional. In addition, the preoperative data is only used for the centerline correction module. In some embodiments, the centerline correction module may not be used; in these embodiments, preoperative data is not used.
[0079] Training a Machine Learning Model
[0080] Machine learning models are used in various aspects of the navigation algorithm. Each model (also referred to as a network or algorithm) is trained based on training data to complete a specific task. In general, the machine learning model can be trained by optimizing model parameters based on minimizing a loss function. As one non-limiting example, the loss function may be a mean squared error loss function.18QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0081] The method includes accessing training data with a computer system. Accessing the training data may include retrieving such data from a memory or other suitable data storage device or medium. Alternatively, accessing the training data may include acquiring such data and transferring or otherwise communicating the data to the computer system.
[0082] Appropriate feature selection can be implemented to reduce the risk of overfitting when the input variables are high-dimensional. As a non-limiting example, a forward stepwise selection can be used, which starts with an empty feature set and adds one feature at each step that maximally improves a pre-defined criterion until no more improvement can be achieved. To avoid overfitting, the accuracy computed on a validation set can be used as an evaluation criterion; when the sample size is limited, cross-validation accuracy can be adopted.
[0083] Training a machine learning model may include initializing the model, such as by computing, estimating, or otherwise selecting initial model parameters. Training data can then be input to the initialized machine learning model, generating outputs and / or uncertainty data that indicate an uncertainty in the outputs. The quality of the output data can then be evaluated, such as by passing the output data to the loss function to compute an error. The current machine learning model can then be updated based on the calculated error (e.g., using backpropagation methods based on the calculated error).
[0084] The current machine learning model can be updated by updating the model parameters in order to minimize the loss according to the loss function. When the error has been minimized (e.g., by determining whether an error threshold or other stopping criterion has been satisfied), the current model and its associated model parameters represent the trained machine learning model.
[0085] The one or more trained neural networks are then stored for later use. Storing the neural network(s) may include storing network parameters (e.g., weights, biases, or both), which have been computed or otherwise estimated by training the neural network(s) on the training data. Storing the trained neural network(s) may also include storing the particular neural network architecture to be implemented. For instance, data pertaining to the layers in the neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be stored.19QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0086] Exemplary training data, network architectures, and loss functions for different trained algorithms is provided below. The training data and loss functions are not meant to be limiting.
[0087] CAM Decomposition Net: The CAM Decomposition Net can follow a U-Net architecture including an encoder and decoder, each having a number of consecutive blocks. The encoder blocks may include convolution layers of a specific size, and a set stride. The encoder may further include batch normalization and an activation layer with a specific activation function such as ReLU activation. The output of the encoder (e.g., the bottleneck layer) may be used as input to the CAM Gauss-Net. The decoder may include upsampling layers and convolutional layers. The convolutional layers may have a specific size and a given stride. The decoder may further include batch normalization and a ReLU activation layer. Training data may include ground truth original images and reconstructed images. A combination of SSIM and LI losses may be used.
[0088] CAM Gauss Net: The CAM Gauss Net can include multiple blocks. The first consecutive blocks may include a convolutional layer, an instance normalization layer, an activation layer, and a dropout layer. The output of these first consecutive blocks may be flattened before providing to the second group of blocks. The second group of blocks may include fully connected layers with an activation function. The output layer may be followed by a sigmoid layer to scale the values into a range between 0 and 1. A specific optimizer, learning rate, epoch number, and batch size may be used. Training data may include ground truth original images and reconstructed images. A combination of SSIM and LI losses may be used.
[0089] SS-DPE Net (including SS-DPE Depth Net and SS-DPE Pose Net): A selfsupervised framework may be used for the SS-DPE net. Monodepth2 may be used as a backbone for the SS-DPE net. The encoders of PoseNet and DepthNet may be initialized with specific encoder’s weights. A specific optimizer, learning rate, epochs, and batch size may be selected to train the SS-DPE net. The input of the SS-DPE Depth net may be the Source Frame, processed by the CAM module. The input of the SS-DPE Pose Net may be the Target Frame, processed by the CAM module. For overall self-supervised training, a photometric loss may be minimized. The photometric loss may be a combination of structure similarity and L2 loss. For PoseNet20QB\99263808.1BWH 2024-548-03Quarles 129319.01124 training, forward-backward pose consistency loss (PC-loss) may be used. This enforces the inverse property of rigid motion on the relative pose outputs of PoseNet. AL2 loss may be calculated on the Euler angles corresponding to the translation and rotation of consecutive image pairs. A long-term pose consistency loss (LT-Loss) may be imposed to minimize the accumulation of errors in downstream poses. An L2 loss may be used on the calculated Euler angles and translation predictions of consecutive image pairs.
[0090] Illustrative Computer System
[0091] FIG. 7 shows an example process 700 to track catheter location for navigating a catheter through a lumen. At step 702, preoperative data, a first image, and a second image may be received. The preoperative data and first and second images may be received with a computer system. The preoperative data may be a CT scan of the lumen to navigate. The first and second images may be acquired by an imaging device attached to the catheter. At step 704, process 700 may estimate an incremental pose based on the first image and second image. At step 706, process 700 may generate catheter pose information registered to preoperative data based on the incremental pose. At 708, process 700 may display the catheter pose information to a user. The user may be a medical professional, e.g., to a doctor in a medical environment.
[0092] In FIG. 8, an example 800 of a system (e.g., a data processing system) for generating pose information in accordance with some embodiments of the disclosed subject matter is shown. In some embodiments, computing device 804 and / or server 816 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable computer, a server computer, a virtual machine being executed by a physical computing device, etc. As described herein, system 800 can present catheter pose information to a user (e.g., a researcher and / or a physician).
[0093] In some embodiments, communication network 802 can be any suitable communication network or combination of communication networks. In some embodiments, communication network 802 can be any suitable communication network or combination of communication networks. For example, communication network 802 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to- peer network (e.g., a Bluetooth network), a cellular network (e.g., a 4G network, a 5G network,21QB\99263808.1BWH 2024-548-03Quarles 129319.01124 etc., complying with any suitable standard, such as CDMA, GSM, LTE, LTE Advanced, WiMAX, etc ), a wired network, etc. In some embodiments, communication network 802 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semiprivate network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Communications links shown in FIG. 8 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, etc.
[0094] FIG. 8 additionally shows an example of hardware that can be used to implement computing device 804 and server 816 in accordance with some embodiments of the disclosed subject matter. In some embodiments, computing device 804 can be used to execute one or more set of instructions to generate catheter pose information.
[0095] As shown in FIG. 8, computing device 804 can include one or more hardware processor 806, one or more displays 808, one or more inputs 810, one or more communications 812, and / or memory 814. In some embodiments, processor 806 can be any suitable hardware processor or combination of processors, such as central processing unit, a graphics processing unit, etc. In some embodiments, display 808 can include any suitable display devices, such as a computer monitor, a touchscreen, a television, etc. In some embodiments, inputs 810 can include any suitable input device and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, etc.
[0096] In some embodiments, communication systems 812 can include any suitable hardware, firmware, and / or software for communicating information over communication network 802 and / or any other suitable communication networks. For example, communications systems 812 can include one or more transceivers, one or more communication chips and / or chip sets, etc. In a more particular example, communications systems 812 can include hardware, firmware and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, etc.
[0097] In some embodiments, memory 814 can include any suitable storage device or devices that can be used to store instructions, values, etc., that can be used, for example, by22QB\99263808.1BWH 2024-548-03Quarles 129319.01124 processor 806 to present content using display 808, to communicate with server 816 via communications system(s) 812, etc.
[0098] Memory 814 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 814 can include RAM, ROM, EEPROM, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, etc. In some embodiments, memory 814 can have encoded thereon a computer program for controlling operation of computing device 804. In such embodiments, processor 806 can execute at least a portion of the computer program to present content (e.g., images, user interfaces, graphics, tables, etc.), receive content from server 816, transmit information to server 816, etc.
[0099] In some embodiments, server 816 can include a processor 818, a display 820, one or more inputs 822, one or more communications systems 824, and / or memory 826. In some embodiments, processor 818 can be any suitable hardware processor or combination of processors, such as a central processing unit, a graphics processing unit, etc. In some embodiments, display 820 can include any suitable display devices, such as a computer monitor, a touchscreen, a television, etc. In some embodiments, inputs 822 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, etc.
[0100] In some embodiments, communications systems 824 can include any suitable hardware, firmware, and / or software for communicating information over communication network 802 and / or any other suitable communication networks. For example, communications systems 824 can include one or more transceivers, one or more communication chips and / or chip sets, etc. In a more particular example, communications systems 824 can include hardware, firmware and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, etc.
[0101] In some embodiments, memory 826 can include any suitable storage device or devices that can be used to store instructions, values, etc., that can be used, for example, by processor 818 to present content using display 820, to communicate with one or more computing devices 804, etc. Memory 826 can include any suitable volatile memory, non-volatile memory,23QB\99263808.1BWH 2024-548-03Quarles 129319.01124 storage, or any suitable combination thereof. For example, memory 826 can include RAM, ROM, EEPROM, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, etc. In some embodiments, memory 826 can have encoded thereon a server program for controlling operation of server 816. In such embodiments, processor 818 can execute at least a portion of the server program to transmit information and / or content (e.g., results of a tissue identification and / or classification, a user interface, etc.) to one or more computing devices 804, receive information and / or content from one or more computing devices 804, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone, etc.), etc.
[0102] In some embodiments, any suitable computer readable media can be used for storing instructions for performing the functions and / or processes described herein. For example, in some embodiments, computer readable media can be transitory or non-transitory. For example, non-transitory computer readable media can include media such as magnetic media (such as hard disks, floppy disks, etc.), optical media (such as compact discs, digital video discs, Blu-ray discs, etc ), semiconductor media (such as RAM, Flash memory, electrically programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), etc.), any suitable media that is not fleeting or devoid of any semblance of permanence during transmission, and / or any suitable tangible media. As another example, transitory computer readable media can include signals on networks, in wires, conductors, optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.
[0103] Exemplary Advantages
[0104] The methods described herein offer many advantages. Anon-exhaustive list of advantages is provided here. The navigation module described herein is unsupervised / self- supervised. The navigation module only requires two video images to produce a relative pose position with respect to the previous camera image frame. The centerline correction module minimizes accumulated error in a predicted trajectory. Other methods require determining depth maps and using the depth maps to generate trajectories. In contrast, the only time a depth map is necessarily used in the navigation module described herein is for an initial registration between24QB\99263808.1BWH 2024-548-03Quarles 129319.01124 the CT image and camera image. Other methods require reconstructing surface based on depth maps. In contrast, the navigation module described herein provides pose and trajectory estimation in CT coordinate space. This can be particularly useful, for example, when the clinician has created a surgical plan within the CT coordinate space. Thus, the pose and trajectory estimates can be easily compared to the surgical plan. The pose and trajectory estimates are also more resistant to CT-body divergence.
[0105] Examples
[0106] The examples below are meant to be illustrative, and are not limiting. Network architectures, training data, and methods of calculating loss may differ from those described below.
[0107] Electromagnetic Tracking (EMT) is an essential technology for navigated bronchoscopy, a promising tool that increases the effectiveness of transbronchial biopsies of lung cancer while reducing pneumothorax and other adverse events. EMT specifically tracks the tip of the bronchoscope within a pre-procedural CT scan, guiding clinicians through the airways to CT- visible lesions and facilitating tissue sample collection for cancer diagnosis. Despite its potential benefits, a challenge that hinders the broader clinical adoption of navigated bronchoscopy using EMT is CT-to-body divergence. This divergence occurs when the initial registration between EMT and the CT scan becomes invalid partway through the biopsy, leading to inaccurate guidance within the airway and towards the lesion. CT-to-body divergence contributes to significant navigation inaccuracies, with literature indicating a 32.8% decrease in diagnostic yield due to these errors.
[0108] In contrast, Vision-Based Tracking (VBT) was developed to address the limitations of EMT by utilizing the bronchoscopic video images that are inherently available to track the bronchoscope’s tip within a pre-procedural CT scan. Unlike EMT, VBT continuously tracks the bronchoscope’s view relative to the CT throughout the entire transbronchial biopsy procedure, making VBT less vulnerable to CT-to-body divergence. In more recent studies, VBT creates a depth map from the bronchoscopic view itself, which serves as a surrogate to estimate the position and orientation of the bronchoscope relative to the CT scan. This method is referred to as Depth-Based Pose Estimation (DBPE) in the rest of the manuscript. By registering the point25QB\99263808.1BWH 2024-548-03Quarles 129319.01124 cloud generated from the depth map of the bronchoscopic view to the model generated from the CT image, modern VBT techniques can localize the bronchoscope within the CT image.
[0109] As researchers continue to explore DBPE for VBT that uses depth maps as surrogates to localize the bronchoscope to the pre-procedural CT image, they have begun to report limitations in pose estimation based solely on depth maps. This challenge arises because different regions of the lung airways share very similar geometries and structures. As a result, the same point cloud can be registered to multiple areas within the lung on the CT scan, leading to several possible registration transformations. Consequently, the estimated camera poses within the virtual model created from the CT image often lack an accurate temporal relationship and can become unreliable in regions with homogeneous geometry and depth, such as the main bronchus of the lung.
[0110] Direct Pose Estimation (DPE), which does not rely on a depth map, can directly estimate the relative camera pose between two consecutive frames in an unsupervised manner. This DPE approach is particularly well-suited for clinical applications, as it addresses the challenges of acquiring labeled clinical data and eliminates the need for continuous point cloud- to-CT model registration to estimate camera pose. It simultaneously estimates the relative camera pose between two frames and the depth for the first image, thereby reconstructing the first image from the second. Following its success in general computer vision, DPE - especially selfsupervised methods - has gained interest in colonoscopy. A self-supervised approach leverages unlabeled data, which is often more readily available than the labeled datasets required for supervised learning. This allows models to learn useful representations and features without the high costs and time associated with manual annotation, enhancing scalability and flexibility across various applications.[OHl] Given the success of self-supervised DPE in colonoscopy, there is potential for applying self-supervised DPE in other endoluminal modalities, such as bronchoscopy. However, the application of self-supervised DPE in bronchoscopy is not straightforward, as the algorithm must correctly choose between multiple possible airways to reach the lesion of interest - a much more challenging task than simply moving forward in a single lumen. Thus, it remains uncertain26QB\99263808.1BWH 2024-548-03Quarles 129319.01124 whether self-supervised DPE can effectively track the camera in CT coordinate space during bronchoscopy.
[0112] Therefore, the objective of this study is to develop and analyze self- supervised DPE in navigated bronchoscopy for transbronchial biopsy. We hypothesize that self-supervised DPE can be effectively applied to navigated bronchoscopy under clinically realistic conditions, assisting bronchoscopists in completing transbronchial biopsies in clinically relevant phantoms. To optimize the performance of self-supervised DPE specifically for bronchoscopic applications, we introduce novel approaches not seen in colonoscopy, including an unsupervised method to mitigate intensity artifacts in images, two new loss functions, and proposed methods to address common issues in self-supervised pose estimation, such as drift. Drift occurs due to the accumulation of errors from successive relative pose estimates. Our method corrects for this drift by adjusting the pose estimates toward the closest point on the centerline of the lung airway in the CT model. Finally, we tested our approach using patient-derived phantoms and clinically oriented performance metrics to assess the usefulness and feasibility of navigated bronchoscopy, comparing our results to those of conventional VBT approaches.
[0113] Materials and Methods
[0114] Study Design
[0115] The primary hypothesis we tested is that self-supervised DPE can be effectively applied to navigated bronchoscopy under clinically realistic conditions, assisting bronchoscopists in completing transbronchial biopsies in clinically relevant phantoms. We have two working hypotheses. First, the DPE can provide superior navigation when compared to DBPE, in CT coordinate frame for transbronchial biopsy. Second, our proposed correction to the estimated poses, using centerline of the airway model, can lead to more robust navigation when compared to navigation using DPE alone. We tested our hypotheses using the dependent variables, average Trajectory Error (TE) and Percentage of Correct Labels (PCL).
[0116] To test our first working hypothesis, we develop a neural network-based DPE method to predict relative pose between two consecutive camera image frames and calculate the localization error and segment identification rate of our proposed DPE and DBPE in the 3D Slicer software.27QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0117] Navigating Bronchoscopy Using Self-Supervised DPE
[0118] Our pose-estimation pipeline consists of three components: a Camera Artifact Masking (CAM) module, a self- supervised depth and pose estimation network (SS-DPE), and a centerline correction module. The CAM module masks out corrupt image pixels, and the output of the CAM module is used to train the SS-DPE. The SS-DPE estimated relative pose between two camera image frames. The centerline correction module finetunes the pose predictions based on the centerline of the lung airway.
[0119] CAM Architecture and Training Details
[0120] CAM is used to remove corrupt pixel intensities, which can make the network estimates incorrect at the periphery. FIG. 9A and FIG. 9B show an example of these corrupt pixel intensities.
[0121] The CAM module assumes that the underlying artifact is a combination of multiple Gaussian masks and it is possible to decompose the original image into an artifact-free image and a combination of 2D Gaussians, which is modeled as a combination of four 2D Gaussian masks. For simplicity, we used a combination of 4 Gaussians with zero covariance in the two dimensions.
[0122] The neural network architecture of the CAM module is shown in FIG. 10. The module consists of two sub networks: the Decomposition Net and Gauss-Net. The aim of the Decomposition Net is to model an artifact-free image, while the Gauss-Net aims to estimate the image artifact. The Gauss-Net estimates the mean and standard deviation of the four masks in the x and y dimensions. Therefore, each mask has four parameters. These parameters are then used to generate four Gaussian masks, which are applied to the output of the Decomposition Net. The final loss function minimizes the reconstruction loss between the original image and the final recomposed image, as shown in FIGS. 9A-9B.
[0123] SS-DPE Network Architecture and Training Details
[0124] The SS-DPE network architecture is shown in FIG. 11. The Decomposition Net follows a U-Net architecture, consisting of an encoder and a decoder, each with four consecutive blocks. Each block of the encoder consists of two consecutive combinations of a 3x328QB\99263808.1BWH 2024-548-03Quarles 129319.01124 convolution layer with a stride of 1, followed by batch normalization and a ReLU activation layer. The output of each block is passed as input to the next block, with the spatial dimension halved and the channel dimension doubled after each block. Similarly, each decoder block consists of upsampling layers and two consecutive combinations of a 3*3 convolution layer with a stride of 1, followed by a batch normalization and a ReLU activation layer, similar to the encoder blocks. After each decoder block, the input is upsampled by a factor of two in the spatial dimension and decreased by a factor of two in the channel dimension. At each level, the input of each decoder block is concatenated with the output of the corresponding encoder block as a skip connection. We used a base filter size of 32, and the image input to the network is of size 192x 192x3.
[0125] The Gauss-Net consists of five blocks. The first three consecutive blocks consist of a convolution layer with a kernel size of 3 x3, an instance normalization layer, a ReLU activation layer, and a dropout layer with a probability of 0.2. Before applying the next blocks, the output is flattened. Then, two blocks of fully connected layers with ReLU activation units are applied. In the last block, a fully connected layer maps the activations to give an output of size lx16, followed by a sigmoid layer to scale the values between 0 and 1. The input to the Gauss- Net is the bottleneck layer of the Decomposition Net. For training we used Adam optimizer and a learning rate of 104We trained the network for 20 epochs with a batch size of 4.
[0126] Training Loss
[0127] For training, we used a self-supervised reconstruction loss that consists of a combination as SSIM and LI losses between the original image, Ion, and reconstructed image, Irecon, as this combination has shown good performance in image reconstruction tasks. 0 is a hyperparamater with a value of 0.55. Eq. 1 shows how reconstructed loss is calculated.
[0128] Self-Supervised Depth and Pose Estimation Module (SS-PDE)
[0129] Network Architecture and Training Details
[0130] We used the Monodepth2 network as the backbone for our SS-PDE pipeline. The encoders of both PoseNet and DepthNet were initialized with ResNetl8 encoder’s weights. The29QB\99263808.1BWH 2024-548-03Quarles 129319.01124Adam optimizer was used with Pi = 0.5 and P2 = 0.99. The learning rate was set to 104The SS- PDE network was trained for 25 epochs with batch size of 16.
[0131] During training, two consecutive images were first input into the pre-trained CAM network, described in the previous subsection, to generate corresponding Gaussian masks. These masks had values between 0 and 1 and were thresholded at 0.6. To mask out the peripheral corrupt pixel intensities, the thresholded masks were then multiplied with the original images, and the resulting images were fed into the SS-PDE network. The full training pipeline is illustrated in FIG. 8.
[0132] Training Losses
[0133] We used Automasking and photometric loss but omitted the use of a smoothing term from photometric loss.
[0134] 1 . Photometric Loss: For the overall self-supervised training we minimized the photometric loss, which is a combination of structure similarity and L2 loss, a is a hyperparameter with value of 0.65.
[0135] 2 Forward-Backward Pose Consistency Loss (PC-Loss): We enforced the inverse property of rigid motions on the relative pose outputs of PoseNet network. Following the inverse homomorphic property of rigid motion (Wang et al. (2019)), the relative rigid body transformation matrixsMt, from frame It to L should be inverse of the transformation matrixlMstransforming from frame Is to It. To calculate the PC-loss, we input the concatenated image pairs {It to Is} and {Is to It} to the SS-DPE network in two passes to obtain the corresponding Euler angles and translations: (0S, as, fl, t x, tsy, tsz} and {ft, at, Pt, ttx, tiy, ttz} . Then, the Euler angles corresponding tosMt~Jare {-ps, -as, -0s} and the translation {t'sx, t'sy, t'sz}, tst =JRxIts, where 7L is obtained from {~ps, -as, -0S} . Finally, the PC-loss is the L2 loss between {-(3s, -as, -0 s, t sx, t SV, t'sz} and {ft, at, Pt, ttx, tty, ttz} .
[0136] 3. Long-term Pose Consistency Loss: Sequential multiplication of pose estimates lead to the accumulation of errors in the downstream poses, which leads to the drift in the estimated trajectory when compared to the ground truth trajectory. Problem of drift is well30QB\99263808.1BWH 2024-548-03Quarles 129319.01124 known in self-supervised ego-motion estimates. Therefore, to mitigate the error, we enforced a long-term pose constraint. Let2MI,3M2, and3MI be the relative transformations between the camera frames 1 and 2, 2 and 3, and 1 and 3, respectively. Then, the rigid body transformations2MI,3M2, and3MI should follow3MI =2MI *3M2. To calculate the LT-loss, we first input the image pairs { 1, 2} and {2, 3 } and { 1, 3 } to the SS-DPE network in separate passes to obtain the corresponding Euler angles and translations. The {03, as, ps, tsx, tsy, tsz} and {0’s, a’s, P's, t'sx, t'sy, t'sz} are the Euler angles and translation predictions of input frames { 1, 3} and corresponding to3MI. respectively. LT-loss is the L2 loss between {0s, as, Ps, tsx, tsy, t3z} and {0's, a' 3, P's, t'sx, t'sy, t Sz}.
[0137] Centerline Corrections: Pose correction with prior airway branching
[0138] To perform the centerline correction, the trajectory generated from the predicted poses is registered to the CT coordinate frame by estimating an initial registration transformation between the camera and CT coordinate system. To assess the initial Camera-to-CT transformation, we placed the depth map obtained from the DPE in a predetermined position within the CT coordinate space in the 3D Slicer software (Pieper et al. (2004)) and manually rotated it to roughly align with the carina in the lung CT model. The final transformation was then obtained using the ICP algorithm. This transformation was applied to the first estimated pose in the trajectory to transform the entire trajectory into the CT coordinate space. Then the centerline correction was applied on the transformed trajectory in the CT coordinate frame. FIG. 13 shows trajectories generated from predicted poses with centerline correction (green, left) and without centerline correction (yellow, right).
[0139] Phantoms Derived from Human Lung CT
[0140] We enrolled seven subjects to create patient-derived phantoms based on their CT scans. Six subjects underwent transbronchial lung biopsies within a predefined period before the study began, while the phantom for the seventh subject was derived from CT data of healthy individuals using 3D Slicer software. This resulted in a demographic profile of two females and four males, averaging 72.8 years (ranging from 62 to 87 years). Including a healthy subject enhanced the diversity of our training and validation datasets. To further ensure variety, we31QB\99263808.1BWH 2024-548-03Quarles 129319.01124 prepared a phantom with varying rigidity to represent the diverse mechanical properties of the lung. Note that sex was not considered in the study’s design.
[0141] First, from the CT, the lung airways were segmented from the patient’s CT scans using the method of Nardelli et al. (2015), implemented in 3D Slicer (Pieper et al. (2004)). To create phantom, a mold was created from the AB S plastic, which was broken down after the silicone was cured. The semi-deformable phantoms were made by pouring silicone into a box containing the solid ABS plastic model. Once the silicone dried, the solid model was cut out, leaving a hollow airway in the box-shaped phantom. The rigid phantom was 3D printed using ABS-R plastic material with water-soluble RapidRinse support material, which was washed away to reveal the hollow airway. We developed two deformable, four semi-deformable, and one rigid phantom. We only used the right lobe of the lung for the semi-deformable phantoms, while for the deformable and rigid phantoms, we used the entire lung.
[0142] Training Data
[0143] We randomly chose five phantoms to train the CAM and SS-DPE networks. Data was collected using a continuum robotic bronchoscope (Masaki et al. (2021)) that had an OmniVision, CA, USA, camera, with a 5-DOF EM sensor mounted on the tip of the robotically controlled catheter, which was inside the bronchoscope’s channel. We also mounted a custom- made fixture to keep the bronchoscope camera stationary concerning the catheter. In all phantoms, we maneuvered the bronchoscope to the far ends of the upper and lower lung lobes twice. In one phantom, we only reached the upper and lower lobes once. For training, this resulted in a total of 20 trajectories.
[0144] We also mounted a custom-made fixture to keep the bronchoscope camera stationary with respect to the catheter. In all phantoms, except the rigid phantom, we maneuvered the bronchoscope to the far ends of the upper and lower lung lobes twice. In the rigid phantom, we reached the upper and lower lobes only once. For training, this resulted in a total of 20 trajectories. For validation, we obtained pose estimates and corresponding depth maps using deformable phantom 2, semi-deformable phantoms 1 and 2, and the rigid phantom. Similar to the training set, we reached the far ends of the upper and lower lobes twice, except in the rigid phantom, where we reached them only once.32QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0145] Validation Procedure
[0146] Validation Metric: To test our hypotheses, we used two phantoms that were not included in the training set, along with two of the five phantoms used for training, in the validation study. Notably, different trajectories were utilized for the phantom employed in training and validation. For a total of 20 trajectories to four lobes of the phantoms, we conducted a mock transbronchial biopsy using a continuum robotic bronchoscope (Masaki et al. (2021)) to reach the furthest airway we could reach while maintaining a continuous sequence of images. In total, we had 20 trajectories in four phantoms for validation.
[0147] We used Correct Airway Label percentage and Mean Localization Error as validation metrics. The Airway Label percentage is the percentage of camera frames assigned correct airway labels, where labels are the airway segments (Tschirren et al. (2005); Keuth et al. (2024)). Airway label is a clinically relevant metric as reaching and path planning to the lesion of interest involves consecutively selecting the correct airway segment (Dolina et al. (2008)). Mean Localization Error is the mean of the Euclidean distance error between the GT position EM camera location and the predicted camera location. The mean Localization Error quantifies the navigational error.
[0148] For the GT airway labels, we first used The Vascular Modeling Toolkit (VMTK) (Izzo et al. (2018)) to generate a model of centerlines of the airways in the CT model. Then, the centerline model was separated into multiple line segments starting and terminating at each branching point. The GT airway label corresponding to each camera frame was the closest airway segment to a specific GT camera position in the CT coordinate space.
[0149] Ground Truth EM-to-CT Registration: To calculate mean Localization Error, we used EM tracker data, registered to the CT coordinate frame as GT. To register the EM tracker to the CT coordinate system, we mounted 5-12 fiducial markers on each lung phantom, visible in the CT scans. A CT scan of the phantoms was performed to locate the fiducials in the CT coordinate space. We used an EM tracker (Aurora, NDI, Waterloo, Canada) for EM tracking. We manually touched the EM probe to the fiducial markers and calculated the rigid registration between the CT and EM coordinate spaces. All rigid registrations had an error of less than 2mm. The EM tracker’s position in the CT coordinate space was used as GT to test our hypotheses.33QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0150] Depth-Based Pose Estimation as a Reference Model: We compared our method with Depth-Based Pose Estimation (DBPE) to evaluate the performance of our proposed SS- DPE. The DBPE used as reference data was generated by first obtaining a depth map for each camera image frame using the 3cGAN method (Banach et al. (2021)). We then employed the Iterative Closest Point (ICP) algorithm to register each depth map to the lung airway model, thereby determining the bronchoscope camera pose. We calculate and report the mean Localization Error and Airway Label percentage for both methods. Throughout the remainder of the paper, we will refer to the reference DBPE method as 3cGAN.
[0151] VBT Using Depth Maps: VBT using depth maps also used to collect reference data to investigate the performance of VBT by DPE. The VBT by depth map was possible by first obtaining depth maps using the cGAN method. We then performed an ICP algorithm to register each depth map to the lung airway model to produce pose from the bronchoscope. We applied the same centerline correction to the estimated trajectories described in the previous section. For both hypotheses, we calculated the validation metrics as outlined in the next section.
[0152] Data Analysis
[0153] Box plots and summary statistics were generated to compare Localization Error and Correct Airway Label Percentage between the 3cGAN and SS-DPE methods. A two-sample t-test was conducted to assess the statistical significance of performance differences between the two methods. In all tests, p-values of 0.05 or less were considered statistically significant.
[0154] Second, box plots were generated for Localization Error within each lobe, including the right lower lobe (RLL), left upper lobe (LUL), left lower lobe (LLL), and right upper lobe (RUL). Similarly, the Correct Airway Label Percentage was calculated by determining the percentage of correctly labeled localization data within each lobe.
[0155] Lastly, we applied a multivariate linear regression with Localization Error and Correct Airway Label Percentage as the dependent variables to identify the factors contributing to SS-DPE performance. The independent variables included the addition of Artifact Removal, Pose Consistency Loss, Long- Term Loss as a loss function, and Centerline Correction as a procedural step. Each coefficient and p-value were calculated and tabulated.34QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0156] Results
[0157] All validation studies were completed successfully. The comparison of localization error and correct airway label percentage between the 3cGAN and SS-DPE methods is show in FIG. 15 and FIG. 16.
[0158] The SS-DPE showed a significantly less (therefore, improved) Localization Error than 3cGAN (p=0.014) did, with the mean of 20.9 ± 7.5 mm and 27.7 ± 12.6 mm for SS-DPE and 3cGAN, respectively, as shown in FIG. 15.
[0159] The 3cGAN achieved a success rate of 48.7 ±26.3% in identifying the bronchoscope’s location within pre-labeled airway segments. SS-DPE had a seemingly higher success rate of 55.1 ± 20.0% with no statistically significant different from 3cGAN, as shown in FIG. 16. FIG. 14 presents an example of airway label identification.
[0160] In RLL, LUL, and LL, SS-DPE seemingly demonstrated a reduction in mean Localization Error compared to 3cGAN but not in RUL (FIGS. 17A-17D). Specifically, SS-DPE produced mean Localization Errors of 12.7 ± 3.3 mm in RLL, 19.8 ± 3.0 mm in LUL, and 14.2 ± 1.7 mm in LLL, and in 12.6 ± 5.4 in RUL, all below 20 mm. FIGS. 18A-18D show trajectories generated by the 3cGAN method and SS-DPE method for each lobe of the lung. In each lobe of the lung, the SS-DPE trajectory was closer to the ground-truth trajectory than the 3cGAN trajectory.
[0161] Pose Consistency Loss significantly improve Localization Error (p=0.003), emphasizing the importance of letting the model maintain pose consistency when using SS-DPE. Centerline Correction (CC) had only a marginal effect, reducing mean Localization Error in SS- DPE from 21.9 ± 8.8 mm without CC to 20.9 ± 7.5 mm with CC1. Similarly, CC showed a slight improvement in Correct Airway Label Percentage in SS-DPE, increasing it from 50.1 ± 18.6% without CC to 55.1 ± 20.0% with CC. The p-value for the impact of CC on Correct Airway Label Percentage is close to 0.05, which diminished the effectiveness of the Centerline Correction in statistical analysis.
[0162] FIG. 19A-19D shows the SS-DPE method integrated into the 3D Slicer software for navigational bronchoscopy. The trajectory predicted by SS-DPE closely matched the ground35QB\99263808.1BWH 2024-548-03Quarles 129319.01124 truth trajectory (see FIG. 19 A). FIG. 19B shows the camera view and FIG. 19C shows the depth map predicted by Depth Net within SS-DPE. FIG. 19D shows the virtual lung model generated from the CT scan. Note that the camera view is registered to the CT coordinate space.coefficient p- value coefficient p-va-lueCamera Artefact Masking 1.8 0.124 -5.4 0.053Pose Consistency Loss -3.4 0.003* 1.9 0.482Long Term Loss -1.7 0.139 4.4 0.114Centerline Correction -0.90.070
[0163] Table 1 : Effect of factors contributing to SS-DPE performance on Localization Error and Correct Airway Label Percentage.
[0164] Discussion and Conclusion
[0165] SS-DPE for navigated bronchoscopy was introduced and validated using clinically derived phantoms constructed from CT scans of patients who underwent transbronchial biopsy. We compared SS-DPE to state-of-the-art depth-based pose estimation method (3cGAN), finding that SS-DPE significantly reduced mean localization error (20.9 mm vs. 27.7 mm) and achieved a comparable Correct Airway Label Percentage (55.1% vs. 48.7%). Our results indicated that SS-DPE outperformed 3cGAN in most lobes, except the right upper lobe, which is particularly challenging to navigate due to difficulties in maintaining a smooth motion of the bronchoscope.
[0166] Although the focus of the past studies has been primarily on depth estimation and the metric used for 6DoF pose estimation is not directly comparable, our results are in concordance with the relevant existing studies. For example, similar to others our direct pose estimation method performs significantly better than the depth based pose estimation, that can cause abrupt tracking failure. Furthermore, our mean localization error 16.14 mm, estimated on all lobes of the lung, is close to the tracking error of 14.5 mm reported by others, although unlike ours, other groups have not revealed the anatomical lobe location of the data collection. Moreover, similar to certain groups we demonstrate that the inclusion of Pose Consistency loss significantly improves the Localization Error. It is because the Pose Consistency Loss36QB\99263808.1BWH 2024-548-03Quarles 129319.01124 encourages the SS-DPE network to give more robust pose estimates by constraining the network to predict poses that are mathematically consistent with the inverse property of rigid motion. However, we also note that the long-term consistency loss has less impact on the SS-DPE performance, which was introduced to alleviate the problem of drift by constraining the long range motion. This could be because of our proposed centerline correction method that uses the prior airway centerline geometry to correct for drift, thus lessening the impact of long-term loss on the localization error.
[0167] Our method builds upon the successful application of self-supervised depth and pose estimation network in colonoscopy. Therefore, we adopted a similar architecture and reconstruction loss as a backbone to train our network. However, while previous studies have primarily focused on depth estimation, with limited attention to 6DOF pose estimation performance, we introduced several key modifications to the existing network to achieve robust pose estimates in transbronchial biopsy with CT. Additionally, unlike pose estimation in a single lumen, transbronchial navigation is much more sensitive to drift, as the algorithm must move forward and choose the correct airway. To address this, we introduced an inexpensive centerline correction method that uses a centerline model of the airway obtained from the CT.
[0168] To conclude, we presented a self-supervised direct pose estimation method that tracks the tip of the bronchoscope within a pre-procedural CT scan using only the native camera images from the bronchoscope. To achieve this, we extended existing SS-DPE studies in endoluminal applications for transbronchial biopsy and thoroughly validated clinically relevant, human-derived phantoms. We successfully demonstrated the clinical feasibility of SS-DPE and significantly improved performance compared to state-of-the-art depth-based pose estimation methods.
[0169] Tables 4-5 summarize results of different phantoms, comparing different models, losses, and trajectory error.37QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0170] Table 1: Full phantom results (deformable).
[0171] Table 2: Rigid phantom results.
[0172] Table 3: Box Phantom 1 Results.38QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0173] Table 4: Box Phantom 2 Results.
[0174] Table 5: Trajectory Error for Phantom Models A and D.
[0175] Autonomous Driving
[0176] The system and methods for generating predicted trajectories for navigating a catheter through a lumen may be used with an autonomous driving system or method of navigating the catheter. For example, the systems and methods as described in WO 2025 / 019378, WO 2024 / 238180, and / or WO 2025 / 019377, each fded on 7 / 12 / 2024, and herein incorporated by reference in their entirety may be combined with the systems and methods for generating predicted trajectories as described herein.
[0177] An autonomous method and an autonomous navigation robot system can include one or more actuators to steer and move the catheter where the controller or an additional controller is configured to perform three steps: 1) perception step, 2) planning step), and 3) control step. After the navigation module outputs a recommended trajectory to navigate a lumen based on the incremental pose, the controller may use this trajectory.
[0178] Error checking39QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0179] In some embodiments, the vision-based tracking algorithm can be used to check for errors, such as if the catheter is damaged and / or the patient shows abnormality that would affect the treatment.
[0180] When a gamepad or other controller controls the robotic catheter, the input value to move it and time points at the beginning and end of motion are stored in the controller. Using the two camera image frames captured at the two time points, SS-DPE estimates a relative camera pose and stores the relative camera pose in the controller. As shown in FIG. 4B, the SS- DPE 129 can be a component of the system controller 128 to facilitate communication between the SS-DEP and system controller. The controller calculates the difference between the input value from the gamepad and the estimated relative camera pose stored in the controller. The calculated difference is stored in the controller with a timestamp. If the difference between the input value from the gamepad and the estimated relative camera pose exceeds the predetermined threshold, the controller indicates this information, for example by showing an alert on the monitor. This difference may be the instantaneous difference or the sum of the calculated differences over a period of time. This allows the operator to check if the robotic catheter system is damaged and / or the patient shows abnormality.
[0181] The mean localization error of SS-DPE can be evaluated at the time of shipment of the robotic catheter system and this (or some multiple of this localization error) can be set as the default threshold. If needed, the operator can freely set the predetermined threshold.
[0182] When the autonomous navigation mode is applied to control the robotic catheter system, the input value from the autonomous navigation mode may be stored in the controller instead of the input value from the gamepad. In some embodiments, if the difference between the input value from the autonomous navigation and the estimated relative camera pose exceeds the predetermined threshold, an error state is entered. This error state may include, for example, stopping the autonomous navigation and showing an alert on the monitor. Accordingly, in some embodiments the system may further include an error module configured to carry out the abovedescribed procedures. The error module may compare a relative cameral pose (e.g., incremental pose) to a navigation input value (e.g., obtained from a gamepad or other controller) and generate a difference. The error module may then compare the difference to a particular (e.g.,40QB\99263808.1BWH 2024-548-03Quarles 129319.01124 predetermined) threshold and output, based on the difference exceeding the particular threshold, a status that an error state has been entered. In some embodiments, outputting the status that an error state has been entered may be at least one of output on a display and / or pause operation of the system.
[0183] References - Each of the following documents is incorporated by reference in its entirety:
[0184] Banach, A., King, F., Masaki, F., Tsukada, H., Hata, N., 2021. Visually navigated bronchoscopy using three cycle-consistent generative adversarial network for depth estimation. Medical image analysis 73, 102164.
[0185] Borrego-Carazo, J., Sanchez, C., Castell s-Rufas, D., Carrabina, J., Gil, D., 2023. Bronchopose: an analysis of data and model configuration for vision- based bronchoscopy pose estimation. Computer Methods and Programs in Biomedicine 228, 107241.
[0186] Chen, A., Pastis, N., Furukawa, B., Silvestri, G.A., 2015. The effect of respiratory motion on pulmonary nodule location during electromagnetic navigation bronchoscopy. Chest 147, 1275-1281.
[0187] Dolina, M.Y., Cornish, D.C., Merritt, S.A., Rai, L., Mahraj, R., Higgins, W.E., Bascom, R., 2008. Interbronchoscopist variability in endobronchial path selection: a simulation study. Chest 133, 897-905.
[0188] Folch, E.E., Pritchett, M.A., Nead, M.A., Bowling, M.R., Murgu, S.D., Krimsky, W.S., Murillo, B.A., LeMense, G.P, Minnich, D.J., Bansal, S., et al., 2019. Electromagnetic navigation bronchoscopy for peripheral pulmonary lesions: one-year results of the prospective, multicenter navigate study. Journal of Thoracic Oncology 14, 445-458.
[0189] Godard, C., Mac Aodha, O., Firman, M., Brostow, G.J., 2019. Digging into selfsupervised monocular depth estimation, in: Proceedings of the IEEE / CVF international conference on computer vision, pp. 3828-3838.
[0190] Guo, L., Nahm, W., 2024. A cgan-based network for depth estimation from bronchoscopic images. International Journal of Computer Assisted Radiology and Surgery 19, 33—36.41QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0191] He, Q., Feng, G., Bano, S., Stoyanov, D., Zuo, S., 2024. Monolot: Selfsupervised monocular depth estimation in low-texture scenes for automatic robotic endoscopy. IEEE Journal of Biomedical and Health Informatics.
[0192] Izzo, R., Steinman, D., Manini, S., Antiga, L., 2018. The vascular modeling toolkit: a python library for the analysis of tubular structures in medical images. Journal of Open Source Software 3, 745.
[0193] Karaoglu, M.A., Brasch, N., Stollenga, M., Wein, W., Navab, N., Tombari, F., Ladikos, A., 2021. Adversarial domain feature adaptation for bronchoscopic depth estimation, in: Medical Image Computing and Computer Assisted Intervention-MICCAI 2021 : 24th International Conference, Strasbourg, France, September 27-October 1, 2021, Proceedings, Part IV 24, Springer, pp. 300-310.
[0194] Keuth, R., Heinrich, M., Eichenlaub, M., Himstedt, M., 2024. Airway label prediction in video bronchoscopy: capturing temporal dependencies utilizing anatomical knowledge. International Journal of Computer Assisted Radiology and Surgery 19, 713-721.
[0195] Li, Y, 2023. Endodepthl: Lightweight endoscopic monocular depth estimation with cnn-transformer, in: 2023 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), IEEE. pp. 4344-4351.
[0196] Makris, D., Scherpereel, A., Leroy, S., Bouchindhomme, B., Faivre, J.B., Remy, J., Ramon, P, Marquette, C.H., 2007. Electromagnetic navigation diagnostic bronchoscopy for small peripheral lung lesions. European Respiratory Journal 29, 1187-1192.
[0197] Masaki, F., King, F., Kato, T., Tsukada, H., Colson, Y, Hata, N., 2021. Technical validation of multi-section robotic bronchoscope with first person view control for transbronchial biopsies of peripheral lung. IEEE Transactions on Biomedical Engineering 68, 3534-3542. doi: 10.1109 / TBME.2021.3077356.
[0198] Nadig, T.R., Thomas, N., Nietert, P.J., Lozier, J., Tanner, N.T., Memoli, J.S.W., Pastis, N.J., Silvestri, G.A., 2023. Guided bronchoscopy for the evaluation of pulmonary lesions: an updated meta-analysis. Chest 163, 1589-1598.42QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0199] Nardelli, P., Khan, K.A., Corvo, A., Moore, N., Murphy, M.J., Twomey, M., O’Connor, O.J., Kennedy, M.P., Estepar, R.S.J., Maher, M.M., et al., 2015. Optimizing parameters of an open-source airway segmentation algorithm using different CT images.Biomedical engineering online 14, 1-24.
[0200] Olson, C.F., Matthies, L.H., Schoppers, H., Maimone, M.W., 2000. Robust stereo ego-motion for long distance navigation, in: Proceedings IEEE Conference on Computer Vision and Pattern Recognition. CVPR 2000 (Cat. No. PR00662), IEEE. pp. 453-458.
[0201] Parisotto, E., Singh Chaplot, D., Zhang, J., Salakhutdinov, R., 2018. Global pose estimation with an attention-based recurrent network, in: Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp. 237-246.
[0202] Pickering, E.M., Kalchiem-Dekel, O., Sachdeva, A., 2018. Electromagnetic navigation bronchoscopy: a comprehensive review. AME Medical Journal 3.
[0203] Pieper, S., Halle, M., Kikinis, R., 2004. 3d slicer, in: 20042nd IEEE international symposium on biomedical imaging: nano to macro (IEEE Cat No. 04EX821), IEEE. pp. 632- 635.
[0204] Rau, A., Bano, S., Jin, Y, Azagra, P, Morlana, J., Kader, R., Sanderson, E., Matuszewski, B.J., Lee, J.Y., Lee, D.J., et al., 2024. Simcol3d — 3d reconstruction during colonoscopy challenge. Medical Image Analysis 96, 103195.
[0205] Recasens, D., Lamarca, J., Fa'cil, J.M., Montiel, J., Civera, J., 2021. Endo- depth-and-motion: Reconstruction and tracking in endoscopic videos using depth networks and photometric constraints. IEEE Robotics and Automation Letters 6, 7225-7232.
[0206] Sganga, J., Eng, D., Graetzel, C., Camarillo, D., 2019. Offsetnet: Deep learning for localization in the lung using rendered images, in: 2019 international conference on robotics and automation (ICRA), IEEE. pp. 5046- 5052.
[0207] Shao, S., Pei, Z„ Chen, W., Zhu, W., Wu, X., Sun, D., Zhang, B , 2022. Selfsupervised monocular depth and ego-motion estimation in endoscopy: Appearance flow to the rescue. Medical image analysis 77, 102338.43QB\99263808.1BWH 2024-548-03Quarles 129319.01124
[0208] Shen, M., Gu, Y, Liu, N., Yang, G.Z., 2019. Context-aware depth and pose estimation for bronchoscopic navigation. IEEE Robotics and Automation Letters 4, 732-739.
[0209] Tschirren, J., McLennan, G., Palagyi, K., Hoffman, E.A., Sonka, M., 2005. Matching and anatomical labeling of human airway tree. IEEE transactions on medical imaging 24, 1540-1547.
[0210] Wang, X., Maturana, D., Yang, S., Wang, W, Chen, Q., Scherer, S., 2019.Improving learning-based ego-motion estimation with homomorphism- based losses and drift correction, in: 2019 IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE. pp. 970-976.
[0211] Xiong, M., Zhang, Z., Zhong, W., Ji, J., Liu, J., Xiong, H., 2021. Self- supervised monocular depth and visual odometry learning with scale- consistent geometric constraints, in: Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence, pp. 963-969.
[0212] Yang, Z., Pan, J., Dai, J., Sun, Z., Xiao, Y, 2024. Self-supervised lightweight depth estimation in endoscopy combining CNN and transformer. IEEE Transactions on Medical Imaging.
[0213] Zhao, C., Shen, M., Sun, L., Yang, G.Z., 2019. Generative localization with uncertainty estimation through video-CT data for bronchoscopic biopsy. IEEE Robotics and Automation Letters 5, 258-265.
[0214] Zhao, H., Gallo, O., Frosio, I., Kautz, J., 2016. Loss functions for image restoration with neural networks. IEEE Transactions on computational imaging 3, 47-57.
[0215] WO 2024 / 098240
[0216] U.S. Pat. 12,087,007
[0217] CN Pat. Appl. 117689706
[0218] Thus, while the invention has been described above in connection with particular embodiments and examples, the invention is not necessarily so limited, and that numerous other44QB\99263808.1BWH 2024-548-03Quarles 129319.01124 embodiments, examples, uses, modifications and departures from the embodiments, examples and uses are intended to be encompassed by the claims attached hereto.45QB\99263808.1
Claims
BWH 2024-548-03Quarles 129319.01124CLAIMSWhat is claimed is:
1. A system for generating predicted trajectories for navigating a catheter through a lumen, comprising: a steerable catheter comprising a proximal end, a distal end, a catheter tip, and a tool channel extending from the proximal end to the distal end; an imaging device disposed within the steerable catheter; a console comprising a system controller, a display controller, and a display; and a navigation module configured to: receive preoperative data and multiple consecutive images from the imaging device from the system controller, and output information based on the received preoperative data and the multiple consecutive images, the information comprising at least one of: a location, a direction, or a recommended trajectory for navigating the catheter through the lumen.
2. The system of claim 1, wherein the information output by the navigation module does not include information from an electromagnetic (EM) sensor.
3. The system of claim 2, wherein the navigation module comprises a trained network comprising a camera artifact masking (CAM) module.46QB\99263808.1BWH 2024-548-03Quarles 129319.011244. The system of claim 1, wherein the navigation module comprises a trained network and comprises at least one of a pose consistency loss (PC-Loss) module, a long term loss (LT-Loss) module, a centerline correction module, or a camera artifact masking (CAM) module.
5. The system of claim 4, wherein the trained network of the navigation module comprises a pose consistency loss (PC-Loss) module.
6. The system of claim 5, wherein the multiple consecutive images comprises a first image and a second image consecutive with the first image, and wherein the PC-Loss module is configured to determine a pose consistency loss by: determining first translational and Euler angle components corresponding to a transformation of the first image to the second image, determining second translational and Euler angle components corresponding to a transformation of the second image to the first image, and determining the pose consistency loss based on a loss between the second translational and Euler angle components and an inverse of the first translational and Euler angle components.
7. The system of any one of claims 4-6, wherein the trained network of the navigation module comprises a Long Term Loss (LT-Loss) module.
8. The system of any one of claims 2-7, wherein the trained network of the navigation module comprises a centerline correction module.
9. The system of any one of claims 2-8, wherein the trained network of the navigation module comprises a camera artifact masking (CAM) module.47QB\99263808.1BWH 2024-548-03Quarles 129319.0112410. The system of any one of claims 2-9, wherein the CAM module is configured to mask at least a portion of each of the multiple consecutive images image prior to outputting the information based on the received preoperative data and the multiple consecutive images.
11. The system of any one of claims 2-10, wherein the CAM module is configured to remove artifacts by generating a mask to remove at least a portion of each of the multiple consecutive images.
12. The system of claim 11, wherein the CAM module is configured to generate the mask to remove a peripheral portion of each of the multiple consecutive images.
13. The system of any one of claims 2-12, wherein the CAM module models the artifacts as a combination of multiple Gaussian masks.
14. The system of any one of claims 2-13, wherein the CAM module comprises more than one trained network.
15. The system of any one of claims 1-14, wherein the navigation module comprises more than one trained network.
16. The system of any one of claims 1-15, wherein the multiple consecutive images comprises a first image and a second image consecutive with the first image, and wherein the navigation module is used to output an incremental pose of the second image relative to the first image.48QB\99263808.1BWH 2024-548-03Quarles 129319.0112417. The system of claim 16, wherein the navigation module characterizes the incremental pose using translational and Euler angle components of a translation matrix used to describe the first image and the second image.
18. The system of claim 17, wherein the navigation module comprises a trained network and a centerline correction module, and wherein the centerline correction module is further configured to: create a centerline of the lumen using the preoperative data; register the first image and the second image with the preoperative data; and adjust the incremental pose to a position towards a closest point on the centerline.
19. The system of claim 18, wherein a depth map is used to register the first image and the second image with the preoperative data.
20. The system of any one of claims 18 or 19, wherein the incremental pose is adjusted towards the closest centerline point each time the navigation module determines a predicted pose iteration.
21. The system of any one of claims 18-20, wherein the incremental pose is periodically adjusted towards the closest centerline point at a frequency less than a frequency at which the navigation module determines a predicted pose iteration.
22. The system of claim 21, wherein the adjustment of the incremental pose is an adjustment relative to an estimated radius of the lumen.49QB\99263808.1BWH 2024-548-03Quarles 129319.0112423. The system of any one of claims 1-22, wherein the system is further configured to output a depth map of the lumen.
24. The system of any one of claims 1-23, wherein the multiple consecutive images comprises a first image and a second image consecutive with the first image, and wherein the navigation module is further configured to: output an incremental pose of the second image relative to the first image, obtain an input value for navigation; and compare the input value for navigation to the incremental pose of the second image relative to the first image to generate a difference, wherein, when the difference exceeds a predetermined threshold, an error state is entered.
25. The system of claim 24, wherein the predetermined threshold is set as a mean localization error of the navigation module or a multiple thereof, wherein the mean localization error was evaluated during manufacture.
26. The system of claim 24, wherein obtaining an input value for navigation comprises estimating an input-based pose based on incremental values inputted to the system.
27. The system of any one of claims 24-26, wherein the navigation comprises autonomous navigation.
28. The system of any one of claims 1-27, wherein the navigation module is further configured to: display the error state on the display.50QB\99263808.1BWH 2024-548-03Quarles 129319.0112429. The system of any one of claims 1-28, wherein the navigation module is further configured to: send a command to stop the navigation when the error state is entered.
30. The system of any one of claims 1-29, wherein the distal end of the catheter is bendable.
31. The system of any one of claims 1-30, further comprising an actuator configured to bend the distal end of the catheter to steer the catheter based on the recommended trajectory.
32. The system of any one of claims 1-31, wherein the display comprises a plurality of displays which present, on separate displays of the plurality of displays: a live image from the imaging device, images from the preoperative data, and the recommended trajectory to navigate the lumen.
33. The system of any one of claims 1-32, wherein the lumen comprises a lung airway.
34. The system of any one of claims 1-33, wherein the system further comprises an error module configured to: compare an incremental pose to a navigation input value and generate a difference, and output, based on the difference exceeding a particular threshold, a status that an error state has been entered.51QB\99263808.1BWH 2024-548-03Quarles 129319.0112435. The system of claim 34, wherein the system, when outputting the status that the error state has been entered, is further configured to at least one of output the status that the error state has been entered on the display, or pause operation of the system.
36. A system for generating predicted trajectories for navigating a catheter through a lumen, comprising: a steerable catheter comprising a proximal end, a distal end, a catheter tip, and a tool channel extending from the proximal end to the distal end; an imaging device disposed within the tool channel; a display; and a computer system operatively connected to the imaging device and the display, the computer system configured to: receive preoperative data and multiple consecutive image frames acquired by the imaging device, estimate pose and depth of the catheter based on the preoperative data and the multiple consecutive image frames, generate a recommended trajectory for the catheter to follow to navigate the lumen based on the estimated pose and depth, and output the recommended trajectory to the display.
37. A method for tracking catheter location for navigation of a catheter through a lumen, comprising: determining a predicted pose iteration comprising: receiving, by a computer system, preoperative data and multiple consecutive images acquired by a catheter comprising an imaging tool; and52QB\99263808.1BWH 2024-548-03Quarles 129319.01124 outputting, by the computer system, information based on the received preoperative data and the multiple consecutive images, the information comprising at least one of a location, a direction, or a recommended trajectory for navigating the catheter through the lumen.
38. The method of claim 37, wherein the information output based on the received preoperative data and the multiple consecutive images does not include information from an electromagnetic (EM) sensor.
39. The method of claim 38, wherein determining the predicted pose iteration further comprises using a trained network comprising a camera artifact masking (CAM) module.
40. The method of claim 37, wherein determining the predicted pose iteration further comprises using a navigation module comprising a trained network, wherein the trained network comprises at least one of a pose consistency loss (PC-Loss), a long term loss module (LT-Loss), a centerline correction module, or a camera artifact masking (CAM) module.
41. The method of claim 40, wherein the trained network of the navigation module comprises a pose consistency loss (PC-Loss) module.
42. The method of claim 41, wherein the multiple consecutive images comprises a first image and a second image consecutive with the first image, and wherein the method further comprises determining a pose consistency loss by:53QB\99263808.1BWH 2024-548-03Quarles 129319.01124 determining, using the PC-Loss module, first translational and Euler angle components corresponding to a transformation of the first image to the second image, determining, using the PC-Loss module, second translational and Euler angle components corresponding to a transformation of the second image to the first image, and determining, using the PC-Loss module, the pose consistency loss based on a loss between the second translational and Euler angle components and an inverse of the first translational and Euler angle components.
43. The method of any one of claims 40-42, wherein the trained network of the navigation module comprises a long term loss (LT-Loss) module.
44. The method of any one of claims 40-43, wherein the trained network of the navigation module comprises a centerline correction module.
45. The method of any one of claims 40-44, wherein the trained network of the navigation module comprises a camera artifact masking (CAM) module.
46. The method of any one of claims 40-45, wherein the CAM module is configured to mask at least a portion of each of the multiple consecutive images prior to outputting the information.
47. The method of any one of claims 40-46, wherein the CAM module is configured to remove artifacts by generating a mask to remove at least a portion of each of the multiple consecutive images.54QB\99263808.1BWH 2024-548-03Quarles 129319.0112448. The method of claim 47, wherein the CAM module is configured to generate the mask to remove a peripheral portion of each of the multiple consecutive images.
49. The method of any one of claims 40-48, wherein the CAM module models the artifacts as a combination of multiple Gaussian masks.
50. The method of any one of claims 40-49, wherein the CAM module comprises more than one trained network.
51. The method of any one of claims 40-50, wherein the computer system comprises more than one trained network.
52. The method of claim 51, wherein the multiple consecutive images comprises a first image and a second image consecutive with the first image, and wherein the computer system is used to output an incremental pose of the second image relative to the first image.
53. The method of claim 52, wherein the computer system characterizes the incremental pose using translational and Euler angle components of a translation matrix used to describe the first image and the second image.
54. The method of any one of claims 40-53, wherein the computer system comprises a trained network and a centerline correction module, and wherein the method further comprises: creating a centerline of the lumen using the preoperative data using the centerline correction module;55QB\99263808.1BWH 2024-548-03Quarles 129319.01124 registering the first image and second image with the preoperative data; and adjusting the incremental pose to a position towards a closest point on the centerline.
55. The method of claim 54, wherein a depth map is used to register the first image and second image with the preoperative data.
56. The method of any one of claims 54 or 55, wherein the incremental pose is adjusted towards the closest centerline point each time the navigation module determines the predicted pose iteration.
57. The method any one of claims 54-56, wherein the incremental pose is periodically adjusted towards the closest centerline point at a frequency less than a frequency at which the navigation module determines the predicted pose iteration.
58. The method of claim 57, wherein the adjustment of the incremental pose is an adjustment relative to an estimated radius of the lumen.
59. The method of any one of claims 37-58, wherein the preoperative data comprises a CT scan of the lumen.
60. The method of any one of claims 37-59, wherein the multiple consecutive images are acquired by the catheter while navigating the catheter through the lumen.
61. The method of any one of claims 37-60, wherein the method further comprises outputting a depth map of the lumen.56QB\99263808.1BWH 2024-548-03Quarles 129319.0112462. The method of any one of claims 37-61, wherein the multiple consecutive images comprises a first image and a second image consecutive with the first image, and wherein the method further comprises: outputting an incremental pose of the second image relative to the first image; obtaining an input value for navigation; and comparing the input value for navigation to the incremental pose of the second image relative to the first image to generate a difference, wherein, when the difference exceeds a predetermined threshold, an error state is entered.
63. The method of claim 62, wherein the predetermined threshold is set as a mean localization error of the computer system or a multiple thereof, wherein the mean localization error was evaluated during manufacture.
64. The method of claim 62, wherein obtaining an input value for navigation comprises estimating an input-based pose based on incremental values inputted to the system.
65. The method of any one of claims 62-64, wherein the navigation comprises autonomous navigation.
66. The method of any one of claims 37-65, wherein the computer system is further configured to: display the error state on a display.
67. The method of any one of claims 62-66, wherein the computer system is further configured to:57QB\99263808.1BWH 2024-548-03Quarles 129319.01124 send a command to stop the navigation when the error state is entered.
68. The method of any one of claims 37-67, further comprising displaying a live image from the imaging tool, the preoperative data, and the recommended trajectory to navigate the lumen.
69. The method of any one of claims 37-68, further comprising navigating the lumen.
70. The method of any one of claims 37-69, further comprising: comparing, using an error module, an incremental pose to a navigation input value and generate a difference, and outputting, using the error module and based on the difference exceeding a particular threshold, a status that an error state has been entered.
71. The method of claim 70, wherein outputting the status that the error state has been entered further comprises at least one of: outputting the status that the error state has been entered on a display, or pausing operation of the system.
72. A method for generating predicted trajectories for navigating a catheter through a lumen, comprising: providing a steerable catheter comprising a proximal end, a distal end, a catheter tip, and a tool channel extending from the proximal end to the distal end, wherein an imaging device is disposed within the tool channel;58QB\99263808.1BWH 2024-548-03Quarles 129319.01124 receiving, from a computer system operatively connected to the imaging device and a display, preoperative data and multiple consecutive image frames acquired by the imaging device; estimating, using the computer system, pose and depth of the catheter based on the preoperative data and the multiple consecutive image frames; generating, using the computer system, a recommended trajectory for the catheter to follow to navigate the lumen based on the estimated pose and depth; and outputting, using the computer system, the recommended trajectory to the display.
73. A method of training a machine learning algorithm, comprising: initializing, using a computer system, a network architecture configured with a plurality of encoders, decoders, and layers; accessing, using the computer system, training data comprising preoperative data and imaging data; and developing, using the computer system, weights and biases of the plurality of encoders, decoders, and layers based on the training data, wherein developing the weights and biases of the network architecture is completed based on minimizing a loss function.
74. The method of claim 73, wherein the network architecture comprises a SS-DPE module.
75. The method of any one of claims 73-74, wherein the SS-DPE module is trained using multiple SS-DPE loss functions.59QB\99263808.1BWH 2024-548-03Quarles 129319.0112476. The method of claim 75, wherein the multiple SS-DPE loss functions comprise at least one of photometric loss, forward-backward pose consistency loss (PC-Loss), and long-term poses consistency loss (LT-Loss).
77. The method of any one of claims 75-76, wherein the SS-DPE module is selfsupervised.60QB\99263808.1
Citation Information
Patent Citations
Autonomous Robotic Catheter for Minimally Invasive Interventions
US20210236773A1
Artificial intelligence coregistration and marker detection, including machine learning and using results thereof
US20220346885A1
Axial Insertion and Movement Along a Partially Constrained Path for Robotic Catheters and Other Uses
US20230098497A1