Fully automatic real-time intelligence interpretation of intraoperative x-ray and CT for catheter localization in vascular networks
A deep learning-based system for real-time catheter localization in vascular networks addresses the challenges of tracking and visualization in neurointervention, enhancing precision and reducing physician workload.
Patent Information
- Application Number
- PCT/US2025/031346
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-29
- Filing Date
- 2025-05-29
- Publication Date
- 2025-12-04
AI Technical Summary
Existing catheter tracking systems in neurointerventional procedures, such as X-ray fluoroscopy and endovascular robotics, face challenges with poor image contrast, noise, device movement, and inability to accurately visualize and track catheters, leading to increased radiation exposure and labor intensity for physicians.
A system utilizing a trained AI model for real-time catheter localization in vascular networks, employing deep learning methods to segment and track catheters and guidewires, enabling precise visualization and navigation through vascular networks.
Enhances precision and autonomy in neurointerventional procedures by accurately tracking catheters and guidewires, reducing labor intensity and improving patient outcomes through automated navigation.
Smart Images

Figure US2025031346_04122025_PF_FP_ABST
Abstract
Description
Attorney Docket No.10063-094WO1 OTT202334 FULLY AUTOMATIC REAL-TIME INTELLIGENCE INTERPRETATION OF INTRAOPERATIVE X-RAY AND CT FOR CATHETER LOCALIZATION IN VASCULAR NETWORKS RELATED APPLICATION
[0001] This PCT application claims priority to, and the benefit of, U.S. Provisional Patent Application No. 63 / 652,816, filed May 29, 2024, entitled “FULLY AUTOMATIC REAL-TIME INTELLIGENCE INTERPRETATION OF INTRAOPERATIVE X-RAY AND CT FOR CATHETER LOCALIZATION IN VASCULAR NETWORKS,” which is incorporated by reference herein in its entirety. BACKGROUND
[0002] Catheters and guidewires, as well as endovascular robotics, are employed in various interventional procedures, e.g., X-ray-guided fluoroscopy, X-ray-guided, or CT-guided interventional procedures, among others.
[0003] Intraoperative X-ray fluoroscopy, as an example, while beneficial, can provide poor contrast and be subject to image noise and unexpected device movements that can compromise procedure efficacy and pose serious risks to patients, physicians, and staff, including increased radiation exposure, eyestrain, and orthopedic injuries. Endovascular robotics offers improved safety, consistency, and patient access in cerebrovascular interventions, with remote intervention already validated for cardiovascular and peripheral vascular indications. However, such robots cannot see or track catheters based on fluoroscopy data.
[0004] There is a need for improved catheter tracking and visualization. SUMMARY
[0005] An exemplary system and method are disclosed to intraoperative X-ray or CT data for catheter localization in vascular networks and body lumen, e.g., for real-time visualization and controls. The exemplary system employs a trained AI model that can improve upon localization in the vascular networks and other body lumen by estimating portions of the vasculatures or lumen as separate masks that can be defined as separate layers or datasets to which a catheter, guide wires, intravascular instrument, or other instruments, can be tracked for localization and / or navigation, e.g., of the catheter or intravascular instrument. The exemplary system and method may be employed with fully automated instruments or manual systems that provide visualization to a surgeon or clinician.Attorney Docket No.10063-094WO1 OTT202334
[0006] A study was conducted to develop and evaluate deep learning methods for catheter segmentation and tip position tracking in cerebral angiography. By enabling computers to "see" catheters and track their position, the study provides a first step towards a new paradigm of automation in neurointervention in which I-augmented systems empowered with computer vision could “watch” expert neurointerventional physicians and learn to autonomously navigate, analogous to how self-driving cars learn from expert drivers. This could make neurointerventions less labor-intensive and more efficient, reducing stress on neuro-interventional physicians and improving patient outcomes.
[0007] Telerobotic networks could disseminate life-saving stroke care, and systems augmented with artificial intelligence (AI) could enhance precision and autonomy by observing expert interventionalists, similar to how self-driving cars can learn from expert drivers.
[0008] Visual perception of catheters and guidewires on X-ray fluoroscopy is essential for neurointervention. Endovascular robots with teleoperation capabilities are being developed, but they cannot “see” intravascular devices, which precludes artificial intelligence (AI) augmentation that could improve precision and autonomy. Deep learning has not been explored for neurointervention, and prior works in cardiovascular scenarios are inadequate as they only segment device tips, while neurointervention and other interventions may require segmentation of the entire structure due to co-axial devices. Disclosed herein is an automatic and accurate image-based catheter segmentation method in cerebral angiography using deep learning.
[0009] In an aspect, provided is a system including: a processor; and a memory having instructions stored thereon, wherein execution of the instructions by the processor, can cause the processor to: receive real-time fluoroscopy data (e.g., X-ray fluoroscopy data) (e.g., bi-plane DSA images in DICOM format or NIFTI format) having vasculatures; receive real-time positional data (measured or estimated) for an intravascular instrument having a tip, including positional data for the tip or an instrument landmark of the instrument; determine, via a trained AI model (e.g., nnUNet, Unet, TransUNet, Tag-DL, etc.) configured to receive in its input the received fluoroscopy data, a plurality of masks each having the fluoroscopy image with an estimated portions of the vasculatures, including a first 2D mask having an estimated first blood vessel in the fluoroscopy data, a second 2D mask having an estimated second blood vessel in the fluoroscopy data, and a third 2D mask having an estimated third vessel in the fluoroscopy data, wherein the trained AI model can have been trained to identify the first blood vessel, the second blood vessel, and the third blood vessel (e.g., common carotid artery, external carotid artery, internal carotid artery, among others); determine a plurality of positional data (e.g., centerline data)Attorney Docket No.10063-094WO1 OTT202334 for the estimated first blood vessel, the estimated second blood vessel, and the estimated third blood vessel, including a first 2D positional data, a second 2D positional data, and a third 2D positional data; and combine the plurality of 2D positional data to generate a combined positional data having a plurality of 2D position data each associated with a blood vessel; determine the positional data for the tip or instrument landmark in relation to the plurality of 2D position data; and direct output of the combined positional data and the positional data for the tip or instrument landmark for subsequent use, via a graphical user interface, in tracking, monitoring, or guiding the tip or instrument landmark through the vasculatures.
[0010] In another aspect, provided is a method for operating any of the disclosed systems, the method including: receiving real-time fluoroscopy data (e.g., X-ray fluoroscopy data) (e.g., bi- plane DSA images in DICOM format or NIFTI format) having vasculatures; receiving real-time positional data (measured or estimated) for an intravascular instrument having a tip, including positional data for the tip or an instrument landmark of the instrument; determining, via a trained AI model (e.g., nnUNet, Unet, TransUNet, Tag-DL, etc.) configured to receive in its input the received fluoroscopy data, a plurality of masks each having the fluoroscopy image with an estimated portions of the vasculatures, including a first 2D mask having an estimated first blood vessel in the fluoroscopy data, a second 2D mask having an estimated second blood vessel in the fluoroscopy data, and a third 2D mask having an estimated third vessel in the fluoroscopy data, wherein the trained AI model can have been trained to identify the first blood vessel, the second blood vessel, and the third blood vessel (e.g., common carotid artery, external carotid artery, internal carotid artery, among others); determining a plurality of positional data (e.g., centerline data) for the estimated first blood vessel, the estimated second blood vessel, and the estimated third blood vessel, including a first 2D positional data, a second 2D positional data, and a third 2D positional data; and combining the plurality of 2D positional data to generate a combined positional data having a plurality of 2D position data each associated with a blood vessel; determining the positional data for the tip or instrument landmark in relation to the plurality of 2D position data; and directing output of the combined positional data and the positional data for the tip or instrument landmark for subsequent use, via a graphical user interface, in tracking, monitoring, or guiding the tip or instrument landmark through the vasculatures.
[0011] In yet another aspect, provided is a non-transitory computer-readable medium having instructions stored thereon, wherein execution of the instructions by a processor can cause the processor to perform any of the disclosed methods or operate any of the disclosed systems.Attorney Docket No.10063-094WO1 OTT202334
[0012] Other systems, methods, features, and / or advantages will be or may become apparent to one with skill in the art upon examination of the following drawings and detailed descriptions. It is intended that all such additional systems, methods, features, and / or advantages be included within this description and be protected by the accompanying claims. BRIEF DESCRIPTION OF DRAWINGS
[0013] Fig.1A shows an example system for catheter localization and visualization in vascular networks and body lumen, e.g., for real-time visualization and controls in accordance with an illustrative embodiment.
[0014] Figs. 1B and 1C show the system of Fig. 1A being utilized for catheter navigation in accordance with an illustrative embodiment.
[0015] Fig. 1D shows an example method for the system of Figs. 1A and 1B in accordance with an illustrative embodiment.
[0016] Figs.2A, 2B, and 2C show an example deep learning model and operations employed in the system of Figs.1A and 1B in accordance with an illustrative embodiment.
[0017] Fig.3A shows an example operational flow for the exemplary system configured with deep neural networks and rule-based algorithms.
[0018] Figs.3B – 3K shows example operations / modules of the flow of Fig.3A in accordance with an illustrative embodiment.
[0019] Figs. 4A – 4C show metrics employed in the evaluation of the exemplary system and methods.
[0020] Figs.5A – 5F show experimental results in a study conducted to develop and evaluate catheter segmentation and tip position tracking operation employed in the exemplary system and methods.
[0021] Figs. 6A – 6B show an example output of the deep learning operation in the catheter segmentation and tip position tracking operation conducted in a part of the study.
[0022] Figs.7A – 7E show performance results of deep learning operation and the tip position tracking operation and vessel segmentation and navigation operations of the exemplary system and method.
[0023] Figs.8A – 8N show clinical results of the exemplary system and method employed in the study.Attorney Docket No.10063-094WO1 OTT202334 DETAILED DESCRIPTION
[0024] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate aspects, can also be provided in combination with a single aspect. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single aspect, can also be provided separately or in any suitable subcombination. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure. DEFINITIONS
[0025] In this specification and in the claims that follow, reference will be made to a number of terms, which shall be defined to have the following meanings:
[0026] As used herein, “comprising” is to be interpreted as specifying the presence of the stated features, integers, steps, or components as referred to, but does not preclude the presence or addition of one or more features, integers, steps, or components, or groups thereof. Moreover, each of the terms “by”, “comprising,” “comprises”, “comprised of,” “including,” “includes,” “included,” “involving,” “involves,” “involved,” and “such as” are used in their open, non-limiting sense and may be used interchangeably. Further, the term “comprising” is intended to include examples and aspects encompassed by the terms “consisting essentially of” and “consisting of.” Similarly, the term “consisting essentially of” is intended to include examples encompassed by the term “consisting of.
[0027] As used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a compound”, “a composition”, or “a cancer”, includes, but is not limited to, two or more such compounds, compositions, or cancers, and the like.
[0028] It should be noted that ratios, concentrations, amounts, and other numerical data can be expressed herein in a range format. It can be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as “about” that particular value in addition to the value itself. For example, if the value “10” is disclosed, then “about 10” is also disclosed. Ranges can be expressed herein as from “about” one particular value, and / or to “about” another particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it can beAttorney Docket No.10063-094WO1 OTT202334 understood that the particular value forms a further aspect. For example, if the value “about 10” is disclosed, then “10” is also disclosed.
[0029] When a range is expressed, a further aspect includes from the one particular value and / or to the other particular value. For example, where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure, e.g. the phrase “x to y” includes the range from ‘x’ to ‘y’ as well as the range greater than ‘x’ and less than ‘y’. The range can also be expressed as an upper limit, e.g. ‘about x, y, z, or less’ and should be interpreted to include the specific ranges of ‘about x’, ‘about y’, and ‘about z’ as well as the ranges of ‘less than x’, less than y’, and ‘less than z’. Likewise, the phrase ‘about x, y, z, or greater’ should be interpreted to include the specific ranges of ‘about x’, ‘about y’, and ‘about z’ as well as the ranges of ‘greater than x’, greater than y’, and ‘greater than z’. In addition, the phrase “about ‘x’ to ‘y’”, where ‘x’ and ‘y’ are numerical values, includes “about ‘x’ to about ‘y’”.
[0030] It is to be understood that such a range format is used for convenience and brevity, and thus, should be interpreted in a flexible manner to include not only the numerical values explicitly recited as the limits of the range, but also to include all the individual numerical values or sub- ranges encompassed within that range as if each numerical value and sub-range is explicitly recited. To illustrate, a numerical range of “about 0.1% to 5%” should be interpreted to include not only the explicitly recited values of about 0.1% to about 5%, but also include individual values (e.g., about 1%, about 2%, about 3%, and about 4%) and the sub-ranges (e.g., about 0.5% to about 1.1%; about 5% to about 2.4%; about 0.5% to about 3.2%, and about 0.5% to about 4.4%, and other possible sub-ranges) within the indicated range.
[0031] As used herein, the terms “about,” “approximate,” “at or about,” and “substantially” mean that the amount or value in question can be the exact value or a value that provides equivalent results or effects as recited in the claims or taught herein. That is, it is understood that amounts, sizes, formulations, parameters, and other quantities and characteristics are not and need not be exact, but may be approximate and / or larger or smaller, as desired, reflecting tolerances, conversion factors, rounding off, measurement error and the like, and other factors known to those of skill in the art such that equivalent results or effects are obtained. In some circumstances, the value that provides equivalent results or effects cannot be reasonably determined. In such cases, it is generally understood, as used herein, that “about” and “at or about” mean the nominal value indicated ±10% variation unless otherwise indicated or inferred. In general, an amount, size, formulation, parameter or other quantity or characteristic is “about,” “approximate,” or “at orAttorney Docket No.10063-094WO1 OTT202334 about” whether or not expressly stated to be such. It is understood that where “about,” “approximate,” or “at or about” is used before a quantitative value, the parameter also includes the specific quantitative value itself, unless specifically stated otherwise.
[0032] As used herein, the terms “optional” or “optionally” means that the subsequently described event or circumstance can or cannot occur, and that the description includes instances where said event or circumstance occurs and instances where it does not.
[0033] EXAMPLE SYSTEMS AND METHODS
[0034] Example System
[0035] Fig. 1A depicts an example system for catheter localization and visualization in vascular networks and body lumen, e.g., for real-time visualization and controls in accordance with an illustrative embodiment. The system can take natural intra-operative X-ray fluoroscopy data and extract the catheter’s location within the arteries. The method is two internal processes. The first component extracts the catheter shape and tip coordinate (red dot) from the raw fluoroscopy frame. The second component extracts and labels the vessels from the raw digital subtraction angiography (DSA) mask frames. If pre-operative 3D-CT angiography is available for the specific patient, it can use this data to fill in missing vessels proximally. If 3D-CT angiography is not available, then the 2D information is only used for downstream location inference. This information is then combined to determine the intra-arterial location of the catheter.
[0036] Fig. 1B depicts an example system configured for arterial visualization. While prior works have segmented arteries, combined multi-class labeled segmentations and leveraging of pre- operative 3D data have yet to be demonstrated. This process will further augment guidance by filling in missing sections of arteries and providing global context that may be missing from the 2-dimensional image alone. (DSA = digital subtraction angiography).
[0037] Fig. 1C shows the system of Figs. 1A or 1B being utilized for catheter navigation in accordance with an illustrative embodiment.
[0038] In Fig.1A, system 100 (e.g., for instrument localization analysis) includes a processor; and a memory having instructions stored thereon, wherein execution of the instructions by the processor, can cause the processor to: receive real-time fluoroscopy data 102 (e.g., X-ray fluoroscopy data) (e.g., bi-plane DSA images in DICOM format or NIFTI format) having vasculatures and real-time positional data 104 (measured or estimated) for an intravascular instrument 106 (e.g., intravascular robot) having a tip 108, including positional data 104 for the tip or an instrument landmark 108 of the instrument 106. The system 100 is configured to determine(via a trained AI model 110 (e.g., nnUNet, Unet, TransUNet, Tag-DL, etc.) that toAttorney Docket No.10063-094WO1 OTT202334 receive in its input the received fluoroscopy data 102) a plurality of masks each having the fluoroscopy image with estimated portions of the vasculatures, including a first 2D mask having an estimated first blood vessel 114a in the fluoroscopy data 102, a second 2D mask having an estimated second blood vessel 114b in the fluoroscopy data 102, and a third 2D mask having an estimated third vessel 114c in the fluoroscopy data 102. The trained AI model 110 can have been trained to identify the first blood vessel 114a, the second blood vessel 114b, and the third blood vessel 114c (e.g., common carotid artery, external carotid artery, internal carotid artery, among others). The system 100 then determines a plurality of positional data (e.g., centerline data) for the estimated first blood vessel 114a, the estimated second blood vessel 114b, and the estimated third blood vessel 114c, including a first 2D positional data, a second 2D positional data, and a third 2D positional data; and combine the plurality of 2D positional data to generate a combined positional data 112 having a plurality of 2D position data each associated with a blood vessel. The system 100 can determine o receive the positional data 104 for the tip or instrument landmark 108 in relation to the plurality of 2D position data. In the example shown in Fig. 1A, the coordinates 104 of the tip or instrument landmark 108 is determined via a train AI model 120. The system 100 then directs output 116 (e.g., visualization output) of the combined positional data 112 and the positional data 104 for the tip or instrument landmark 108 for subsequent use, via a graphical user interface 118, in tracking, monitoring, or guiding the tip or instrument landmark 108 through the vasculatures.
[0039] In some aspects, the execution of the instructions by the processor can further cause the processor to modify the plurality of positional data to smooth and / or remove discontinuity artifacts among the first 2D positional data, the second 2D positional data, and the third 2D positional data.
[0040] In some aspects, the execution of the instructions by the processor can further cause the processor to identify (i) a first origin point for the second 2D positional data in the first 2D positional data and (ii) a second origin point for the third 2D positional data in the first 2D positional data, wherein the first origin point and second origin point can be employed in the instructions to determine the positional data 104 for the tip or instrument landmark 108 in relation to the plurality of 2D position data.
[0041] In some aspects, the execution of the instructions by the processor can further cause the processor to identify (i) a bifurcation point in the first 2D positional data, wherein the bifurcation point can be employed in the instructions to determine the positional data 104 for the tip or instrument landmark 108 in relation to the plurality of 2D position data.Attorney Docket No.10063-094WO1 OTT202334
[0042] In some aspects, the execution of the instructions by the processor can further cause the processor to generate visualization output 116 for the graphical user interface 118, the visualization output 116 including a first visualization object of the tip or instrument landmark 108 and a second visualization object of the instrument 106, wherein the first visualization object and the second visualization object can be combined with the real-time fluoroscopy data 102.
[0043] In some aspects, the execution of the instructions by the processor can further cause the processor to: generate visualization output 116 for the graphical user interface 118, the visualization output 116 including a third visualization object for an indication of the tip or instrument landmark 108 in the first blood vessel 114a, the second blood vessel 114b, or the third blood vessel 114c, wherein the third visualization object can be combined with the real-time fluoroscopy data 102.
[0044] In some aspects, the trained AI model 110 can have been trained using digital subtraction angiography (DSA) images.
[0045] In some aspects, the system can be employed for robotic neurointervention or robotic endovascular intervention.
[0046] In some aspects, the system can be employed for cerebrovascular intervention, endovascular intervention (e.g., diagnostic, therapeutic, or both), diagnostic cerebral angiography (DCA), cerebrovascular evaluation (e.g., ischemic stroke, aneurysms, and arteriovenous malformations (AVMs)), mechanical thrombectomy, or embolization (among others).
[0047] In some aspects, the real-time positional data 104 (measured or estimated) for the intravascular instrument 106 can be determined via a second trained AI model 120 that can generate the real-time positional data 104 from the real-time fluoroscopy data 102.
[0048] Example Computing System. The exemplary system and method may operate with a computing system to perform a sequence of computer-implemented acts or program modules running on a computing system comprising a processing unit. The processing unit may be a standard programmable processor that performs arithmetic and logic operations necessary for the operation of the computing device. As used herein, processing unit and processor refers to a physical hardware device that executes encoded instructions for performing functions on inputs and creating outputs, including, for example, but not limited to, microprocessors (MCUs), microcontrollers, graphical processing units (GPUs), and application-specific circuits (ASICs). Thus, while instructions may be discussed as executed by a processor, the instructions may be executed simultaneously, serially, or otherwise executed by one or multiple processors. TheAttorney Docket No.10063-094WO1 OTT202334 computing device may also include a bus or other communication mechanism for communicating information among various components of the computing device.
[0049] Example Methods
[0050] Example Method #1. Fig. 1D shows an example method of operating the system of Figs. 1A or 1B in accordance with an illustrative embodiment. In Fig. 1D, the method 200 for operating the disclosed systems includes: receiving 202 real-time fluoroscopy data (e.g., X-ray fluoroscopy data) (e.g., bi-plane DSA images in DICOM format or NIFTI format) having vasculatures; receiving 204 real-time positional data (measured or estimated) for an intravascular instrument having a tip, including positional data for the tip or an instrument landmark of the instrument; determining 206, via a trained AI model (e.g., nnUNet, Unet, TransUNet, Tag-DL, etc.) configured to receive in its input the received fluoroscopy data, a plurality of masks each having the fluoroscopy image with an estimated portions of the vasculatures, including a first 2D mask having an estimated first blood vessel in the fluoroscopy data, a second 2D mask having an estimated second blood vessel in the fluoroscopy data, and a third 2D mask having an estimated third vessel in the fluoroscopy data, wherein the trained AI model can have been trained to identify the first blood vessel, the second blood vessel, and the third blood vessel (e.g., common carotid artery, external carotid artery, internal carotid artery, among others); determining 208 a plurality of positional data (e.g., centerline data) for the estimated first blood vessel, the estimated second blood vessel, and the estimated third blood vessel, including a first 2D positional data, a second 2D positional data, and a third 2D positional data; and combining 210 the plurality of 2D positional data to generate a combined positional data having a plurality of 2D position data each associated with a blood vessel; determining 212 the positional data for the tip or instrument landmark in relation to the plurality of 2D position data; and directing 214 output of the combined positional data and the positional data for the tip or instrument landmark for subsequent use, via a graphical user interface, in tracking, monitoring, or guiding the tip or instrument landmark through the vasculatures.
[0051] In some aspects, the method can further include modifying the plurality of positional data to smooth and / or remove discontinuity artifacts among the first 2D positional data, the second 2D positional data, and the third 2D positional data.
[0052] In some aspects, the method can further include identifying (i) a first origin point for the second 2D positional data in the first 2D positional data and (ii) a second origin point for the third 2D positional data in the first 2D positional data, wherein the first origin point and secondAttorney Docket No.10063-094WO1 OTT202334 origin point can be employed in determining the positional data for the tip or instrument landmark in relation to the plurality of 2D position data.
[0053] In some aspects, the method can further include identifying (i) a bifurcation point in the first 2D positional data, wherein the bifurcation point can be employed in determining the positional data for the tip or instrument landmark in relation to the plurality of 2D position data.
[0054] In some aspects, the method can further include generating visualization output for the graphical user interface, the visualization output including a first visualization object of the tip or instrument landmark and a second visualization object of the instrument, wherein the first visualization object and the second visualization object can be combined with the real-time fluoroscopy data.
[0055] In some aspects, the method can further include generating visualization output for the graphical user interface, the visualization output including a third visualization object for an indication of the tip or instrument landmark in the first blood vessel, the second blood vessel, or the third blood vessel, wherein the third visualization object can be combined with the real-time fluoroscopy data.
[0056] In some aspects, the trained AI model can be trained using digital subtraction angiography (DSA) images.
[0057] In some aspects, the method can be employed for robotic neurointervention or robotic endovascular intervention.
[0058] In some aspects, the method can be employed for cerebrovascular intervention, endovascular intervention (e.g., diagnostic, therapeutic, or both), diagnostic cerebral angiography (DCA), cerebrovascular evaluation (e.g., ischemic stroke, aneurysms, and arteriovenous malformations (AVMs)), mechanical thrombectomy, or embolization (among others).
[0059] In some aspects, the real-time positional data (measured or estimated) for the intravascular instrument can be determined via a second trained AI model that generates the real- time positional data from the real-time fluoroscopy data.
[0060] In some aspects, the trained AI model can have been trained to employ manual segmentation of the internal carotid artery (ICA), external carotid artery (ECA), and common carotid artery on DSA images as three separate masks.
[0061] In yet another aspect, provided is a non-transitory computer-readable medium having instructions stored thereon, wherein execution of the instructions by a processor can cause the processor to perform any of the disclosed methods or operate any of the disclosed systems.Attorney Docket No.10063-094WO1 OTT202334
[0062] In some aspects, the execution of the instructions by the processor can further cause the processor to modify the plurality of positional data to smooth and / or remove discontinuity artifacts among the first 2D positional data, the second 2D positional data, and the third 2D positional data.
[0063] In some aspects, the execution of the instructions by the processor can further cause the processor to identify (i) a first origin point for the second 2D positional data in the first 2D positional data and (ii) a second origin point for the third 2D positional data in the first 2D positional data, wherein the first origin point and second origin point can be employed in the instructions to determine the positional data for the tip or instrument landmark in relation to the plurality of 2D position data.
[0064] In some aspects, the execution of the instructions by the processor can further cause the processor to identify (i) a bifurcation point in the first 2D positional data, wherein the bifurcation point can be employed in the instructions to determine the positional data for the tip or instrument landmark in relation to the plurality of 2D position data.
[0065] In some aspects, the execution of the instructions by the processor can further cause the processor to generate visualization output for the graphical user interface, the visualization output including a first visualization object of the tip or instrument landmark and a second visualization object of the instrument, wherein the first visualization object and the second visualization object can be combined with the real-time fluoroscopy data.
[0066] In some aspects, the execution of the instructions by the processor can further cause the processor to generate visualization output for the graphical user interface; the visualization output includes a third visualization object for an indication of the tip or instrument landmark in the first blood vessel, the second blood vessel, or the third blood vessel, wherein the third visualization object can be combined with the real-time fluoroscopy data.
[0067] Example Method #2. In some embodiments, the exemplary system can employ a tip- detection operation and calculate a tip-distance error as a hybrid operation.
[0068] Fig.3A shows an example operational flow for the exemplary system configured with deep neural networks and rule-based algorithms. Fig. 3B shows an example of the tip-detection operation that may be used in the operational flow shown in FIG.3A.
[0069] In the example shown in Fig.3A, the operation flow in Fig.3A, subpanel (a) shows the operation for the entire algorithm. Fig.3A, subpanel (b) shows the main data elements determined in the operation flow in Fig.3A.Attorney Docket No.10063-094WO1 OTT202334
[0001] The operation flow Fig.3A includes a catheter segmentation operation 205 and multi- vessel segmentation operation 207 that each implements deep learning operations that operate on scanned images acquired from a patient. The operational flow further implements a tip coordinate determination operation 209 to determine and localize the coordinates of the tip of the catheter. The operational flow further implements vessel segmentation and localization operations to determine the vessels as a geometric object to which navigation and localization can be performed. Table 1A shows example operations performed in Fig.3A and corresponding example algorithmic implementation for them. Final outputs of the operation flow, including catheter tip coordinates and location (determined by 209), vessel landmarks (determined by 215), path deviation (determined by 211), and path plan (determined by 217), can be displayed as overlay images on a monitor using algorithm 200c. Table 1A Tip detection operation 209 See Figs. 3B, 3C, 3D, 3E, and 3F. In Fig. 3B, the repair ti b i l t d i th l ith i Fi 3E he m he m
[0070] In Fig.3A, the exemplary system first receives a plurality of images as inputs, including (i) a raw frame acquired from a continuous image acquisition process 201 (shown as 201’) and (ii) a pre-contrast (i.e., non-contrast, FNC) and post-contrast (i.e., contrast, FC) frames acquired from a single acquisition process 203 (shown as 203’) from the raw frame.
[0071] The exemplary system then (i) performs catheter segmentation 205 (shown having a correspondence to 205’) on the raw frame (shown as current frame 201) and the pre-contrast frameAttorney Docket No.10063-094WO1 OTT202334 203 to generate a catheter segmentation masks 302, (ii) determines tip coordinates 209 (shown as 209’) using the tip-detection operation in Fig. 3B, (iii) determines artery location 213 (shown as 213’) using algorithm 200k in Fig. 3H, subpanels (a)-(b), and (iv) determines path deviation 211 (shown as 211’), including path deviation and distance of the tip to the endpoint, using algorithm 200m in Fig.3J and 200n in Fig.3K, respectively. Concurrently, the exemplary system performs (i) multi-vessel segmentation 207 (shown as 207’) on the pre-contrast and post-contrast frames using algorithm 200g in Fig.3D, algorithm 200h in Fig.3E, and algorithm 200j in Fig.3G and (ii) landmark detection 215 of vessels and path generation 217 (shown as 217’) for the catheter tip to move through the vessels using algorithm 200l in Fig. 3I, subpanels (a)-(b). Fig. 3D shows an algorithmic implementation for generating a topological representation of the carotid bifurcation from the multi-channel vessel segmentation mask. Fig.3E shows an algorithmic implementation for removing spurs generated from skeletonizing the combination of two vessel segmentation masks. Fig. 3B implements the skeletonization operation first before a repair operation (Fig.3E) to achieve the goal to repair the mask. The skeletonization function works both any tubular segmented masks, including vessel or device. Fig. 3G shows an algorithmic implementation for deriving the origin points of the internal and external carotid arteries in the carotid bifurcation. Fig. 3I shows an algorithmic implementation for creating a path that terminates at a specific distance into the target vessel. As shown, the algorithm selects the vessel of interest, traverses a certain distance into the final vessel, and then joins that path with the proximal centerlines. The exemplary system then displays the final outputs (e.g., catheter tip coordinates and location, vessel landmarks, path deviation, and path plan) as overlay images using algorithm 200c.
[0072] Tip-detection operation. Fig.3B shows a tip localization and detection operation. The detection is performed using a Zhang-Suen skeletonization method 223 configured with a logic to exclude endpoints that are not catheter tips by differential weighting of edges using a skeletonized mask 225.
[0073] Tip-distance Error Calculation. Fig.3C shows an algorithmic implementation (shown as algorithm 200f) for calculating a tip-distance error after detecting the tip coordinates 227 (e.g., tx, ty). In some embodiments, the exemplary method can calculate the bifurcation point, internal carotid arteries origin, external carotid arteries origin, and errors between the ground truth and prediction for each vessel mask.
[0074] Segmentation mask generation. Fig.3D shows an algorithmic implementation (shown as algorithm 200g) for generating a topological representation of the carotid bifurcation from the multi-channel vessel segmentation mask. As shown, the proximal vessel can be the commonAttorney Docket No.10063-094WO1 OTT202334 carotid artery, and the distal vessels can be the external carotid artery (ECA) and the internal carotid artery (ICA) (line 220). Skeletons of proximal and distal vessel pairs can be assembled and repaired (lines 222) using a custom pruning method, as shown in Fig. 3E. Then, the complete skeletons can be assembled through a logical OR operation on the set of pairwise vessel skeletons (line 224).
[0075] Artifact removal. Fig. 3E shows an algorithmic implementation (shown as algorithm 200h) for removing spurs generated from skeletonizing the combination of two vessel segmentation masks (in Fig. 3D). In Fig. 3E, the input (line 226) to algorithm 200h can be a skeleton, or path, created by joining two vessel segmentation masks, assuming that the path should be smooth and not have any occurring branches. To find spurs, the algorithm 200h uses a branch detection method (line 228), which creates a list of locations of spurs for removal (line 230). The detected branch can be clipped, removing a residual spur piece using a threshold of 20 pixels (lines 232). The remaining endpoints (line 234), if under a maximum distance, can be rejoined by a straight line to repair the path (line 236). The hyperparameters for dilation size and maximum endpoint distance can be chosen to be 4 pixels and 50 pixels (line 238), respectively.
[0076] Coordinate determination. Fig. 3F shows an algorithmic implementation (shown as algorithm 200i) for deriving the bifurcation point of a bifurcation in 2-dimensional imaging. The algorithm 200i can take a 2-dimensional skeletonized vessel representation and a threshold of 12 as inputs (lines 24) and return the coordinates of the carotid bifurcation (line 246). The algorithm 200i can achieve the carotid bifurcation through thresholding convolution of the skeleton with a branch point detection kernel (lines 242) and then selecting the correct branch point from the list of detected points (lines 244). For DSA roadmaps, the carotid bifurcation can be the first bifurcation encountered when traversing from the bottom of the frame. The threshold of 12 can be selected as input (lines 240).
[0077] Artery origin determination. Fig. 3G shows an algorithmic implementation (shown as algorithm 200j) for deriving the origin points of the internal and external carotid arteries in the carotid bifurcation. The algorithm 200j can take the multi-channel segmentation mask as input (line 250) and return a list of coordinates corresponding to the vessel origins. In the algorithm 200j, the proximal vessel can be the common carotid artery, and the distal vessels can be the external (ECA) and internal carotid arteries (ICA).
[0078] Location inference. Fig.3H shows an algorithmic implementation (shown as algorithm 200k) for the location inference that incorporates a correction for vessel deformation caused by catheter navigation. Subpanel (a) shows implementation 254 for deriving the arterial location fromAttorney Docket No.10063-094WO1 OTT202334 the catheter and vessel masks without accounting for vessel deformation. Subpanel (b) shows implementation 256 for deriving the arterial location from the catheter and vessel masks, accounting for vessel deformation.
[0079] Path traversal. Fig.3I shows an algorithmic implementation (shown as algorithm 200l) for creating a path that terminates at a specific distance into the target vessel. As shown, the algorithm selects the vessel of interest (lines 258), traverses a certain distance into the final vessel (lines 260), and then joins that path with the proximal centerlines (lines 262).
[0080] In-path deviation analysis. Fig. 3J shows an algorithmic implementation (shown as algorithm 200m) for calculating the in-plane deviation of the catheter tip from the path calculated in FIG.3I.
[0081] Endpoint distance calculation. Fig. 3K shows an algorithmic implementation (shown as algorithm 200n) for calculating the distance to the destination point of the pre-specified path using the deviation and closest point on the path calculated in FIG.3H. EXAMPLES Example 1: Automated Catheter Segmentation and Tip Detection in Cerebral Angiography with Topology-Aware Geometric Deep Learning
[0082] A study was conducted to develop and evaluate deep learning methods for catheter segmentation and tip position tracking in cerebral angiography. By enabling computers to "see" catheters and track their position, the study provides a first step towards a new paradigm of automation in neurointervention in which I-augmented systems empowered with computer vision could “watch” expert neurointerventional physicians and learn to autonomously navigate, analogous to how self-driving cars learn from expert drivers. This could make neurointerventions less labor-intensive and more efficient, reducing stress on neuro-interventional physicians and improving patient outcomes.
[0083] The dataset used for the study is one of the largest complete catheter and guidewire annotations in the literature and the first in neurointervention. Additionally, the study demonstrated strategies for improving performance that leverage intraoperative data and exploit the catheter's geometric properties, which are translatable to other scenarios. The generalizability of the exemplary methods to unseen patients and views suggests that it can be fine-tuned for other settings, such as animal models for pre-clinical development and evaluation. Ongoing studies evaluating these methods’ performance in various fluoroscopy systems, including those integrated with robotic systems and haptic feedback, are an important step towards realizing the full potential of artificial intelligence in neurointervention.Attorney Docket No.10063-094WO1 OTT202334
[0084] Discussion: Accurate and automatic frame-by-frame segmentation of the catheter and guidewire is the critical first step toward translating this technology to clinical use. In the context of this task, the combined catheter / guidewire structure is referred to as “catheter” for simplicity. Catheter segmentation in fluoroscopy is challenging due to the low signal-to-noise ratio (SNR) in fluoroscopy images, confounding by other wire-like structures, and sparsity of catheter pixels (<2%) compared to background pixels [8]. Some groups have used external signals to sense catheter position, such as optical, piezoelectric, and magnetic trackers [9]. While effective in theory, these approaches do not generalize well to the existing workflow of interventional suites that simultaneously use wires and catheters from multiple vendors without aftermarket modification. Consequently, detecting the catheter shape and position directly from fluoroscopy images is desirable
[0010] .
[0085] Recent work has utilized deep learning (DL) methods, particularly convolutional neural networks (CNNs) and vision transformer (ViT) architectures, that have demonstrated superiority over conventional image processing methods [8], [9],
[0011] -
[0014] . Through an iterative process supervised by human-annotated data, DL methods learn to optimally transform raw input data to a desired output by modulating millions of numerical weights in multilayered, interconnected “neurons”
[0015] . However, DL methods have not been developed for catheter segmentation in neurointervention. Most prior work in cardiovascular, peripheral vascular, or phantom data does not visualize the complete catheter structure, which is critical in neurointerventions where devices are co-axially passed over each other [9],
[0011] -
[0013] . Manual annotation of the complete catheter structure is considerably more laborious than annotating the tip alone, which has been a significant barrier to generating suitable datasets. Additionally, potential strategies to improve segmentation performance remain unexplored, including incorporating data from digital subtraction angiography (DSA) road-mapping and exploiting geometric symmetries of the catheter. Furthermore, the centerline Dice is a topology-aware metric that rewards segmentation results that preserve connectivity, which is more suitable for catheter segmentation than the conventional Dice score used in most prior work, as illustrated in FIG.2C and TABLE 1B
[0016] . TABLE 1B. Comparison of our work against similar studies describing catheter and / or guidewire segmentation with deep learning on X-ray fluoroscopy. (TDE = Tip Distance Error, CDE = Centerline distance error, EDM = Error Distance Measure, PCI = percutaneous coronary. Fully Procedures # Patients Method Object Se mented Performance )Attorney Docket No.10063-094WO1 OTT202334 CDE: 0.2 Ambrosini et Liver 28 Yes Full catheter mm l 2 17 12 th t i ti 4 )e e o e, s s u y es ga es eep ea g e o s o ca e e seg e a on in fluoroscopy data from cerebrovascular interventions and evaluates the impact of application- specific and geometric considerations on performance. The study develops a topology-aware geometric deep learning (TAG-DL) method and compares this against state-of-the-art segmentation methods, UNet, nnUNet, and TransUNet, a hybrid CNN / ViT architecture, to determine the optimal model architectures, inputs, and performance metrics for catheter segmentation. Currently, in the neurointerventional realm, robotics has largely been “robotic- assisted” rather than a true automated procedure [4,17,18]. The technology presented below will be an important step towards complete automation, which is the ultimate goal in the field of robotics.
[0087] Materials and Methods
[0088] Patient Selection and Data Collection: Due to the lack of publicly available data, this study developed an annotated fluoroscopy dataset to support the experiment. The fluoroscopy videos in this dataset were collected prospectively from procedures performed in the neuro- interventional radiology department. The study defined a fluoroscopy sequence as a set of s 2D images, S =(I1, I2,…, Is), where frames are separated by 0.1 seconds in time, corresponding to videos acquired at 10 frames per second
[0012] . The study included fluoroscopy sequences based on the following criteria: (1) use of DSA roadmapping (2) containing only 0.035” guidewire and / orAttorney Docket No.10063-094WO1 OTT202334 5-Fr diagnostic catheter in-frame, and (3) catheter or guidewire tip at or proximal to the intracranial internal carotid artery. This reflects the conditions of diagnostic cerebral angiography, which may be done as either a standalone procedure or the initial component of a therapeutic procedure
[0019] .
[0089] Manual Annotation of Fluoroscopy Video Sequences: The combined structure of the catheter and guidewire were manually segmented as a single mask in frames of fluoroscopy sequences using MRIcron (NITRC, Washington, USA) by an expert MD / PhD student (6 years of experience) under supervision of an experienced imaging scientist (15 years of experience). The catheter and guidewire are two distinct devices, with the guidewire being thinner than the hollow catheter it travels through. Since either the catheter or guidewire may be the most distal intravascular device used for navigation, the study treated them as a single segmentation mask, referred to as “catheter,” where the catheter pixels have a value of 1 and background pixels have a value of 0.
[0090] Deep Learning Model Architecture and Input: The TAG-DL deep learning model was built on a rotation-reflection equivariant U-Net [20-22]. This network structure guarantees identical image output due to input image changes corresponding to four rotations, each having two reflections
[0021] . In addition to the equivalent U-Net without grouped convolutions, the study implemented two other state-of-the-art models for medical image segmentation, nnUNet and TransUNet [23-25]. The nnUNet is a self-configuring UNet pipeline that determines optimal model architecture and hyperparameters based on the characteristics of the input data
[0024] . The TransUNet model incorporates vision transformers into the structure of UNet that are able to model global context
[0023] . Prior work has used variants of UNet and TransUNet for catheter segmentation, but TAG-DL and nnUNet have not been explored for this task.
[0091] Figs.2A, 2B, and 2C show an example deep learning model and operations employed in the system of Fig. 1A in accordance with an illustrative embodiment. Specifically, Fig. 2A shows a visual representation of the implementation of TAG-DL for catheter segmentation. The framework takes the current frame and digital subtraction angiography (DSA) mask as inputs and outputs a binary segmentation mask, all of which are 2D frames (Z2). The TAG-DL network architecture is a convolutional neural network (CNN) based on UNet with grouped convolutions that enforce rotation-reflection equivariance for the dihedral 4 symmetry group (D4).
[0092] Fig. 2B depicts an example of digital subtraction angiography (DSA) roadmap mask frames. The top row shows the current frame, the non-contrast mask frame, and the result of subtracting this mask from the current frame. The bottom row shows the analogous process with the contrast DSA mask frame. Fig. 2C is an illustration of centerline-Dice metric superiority toAttorney Docket No.10063-094WO1 OTT202334 Dice score in differentiating two segmentation results with equivalent Dice scores, where Label 2 is missing the tip and a section of the catheter body (right), while the Label 1 is missing only edge pixels (middle). The centerline-Dice appropriately scores Label 1 higher with preserved topology higher than Label 2, which misses most of the distal catheter. This translated to more accurate tip localization (red).
[0093] Five different input schemes were tested to assess the impact of DSA mask frame on the model’s performance on unseen cases: (1) current frame only, (2) current frame with a non- contrast mask frame as second channel, (3) current frame with a non-contrast mask frame subtracted prior to input as one channel (4) current frames with a contrast mask frame as the second channel, and (5) current frame with a contrast mask frame subtracted prior to input. An example of the contrast and non-contrast mask frames, along with the result of subtraction with the current frame, is shown in Fig. 2B. Prior to input, all frames and masks were resized to (512,512) and intensities were scaled from 0 to 1.
[0094] Deep Learning Model Training: For training and validation, the study used 37 sequences (2,860 frames) from 26 patients. For UNet, TransUNet, and TAG-DL, 10 sequences (600 frames) from this set were used for model validation. In the nnUNet pipeline, 20% of the training data is automatically selected for validation
[0024] . All sequences in the training and validation are frontal-view, and there are no lateral-view sequences. The validation data does not directly train the network; however, in each pass through the training data (or epoch), the model’s performance is measured on the validation dataset. All models were trained on either NVIDIA Titan RTX or NVIDIA RTX A6000 graphical processing units (GPUs). Further details of network hyperparameters are in TABLE 2. TABLE 2. Hyperparameters used in the design and training of deep learning networks. *The nnUNet pipeline automatically allocates training data into 250 batches. This approximate is calculated as follows: (80% training split * 2,860 frames) ÷ 250 batches ≅ 9.2 frames per batch. Model Architecture Hyperparameter )Attorney Docket No.10063-094WO1 OTT202334 20% of total training Validation Dataset 10 sequences 10 sequences 10 sequences datasetned model was tested on 17 sequences from 14 unseen patients (971 frames). Eleven sequences are in frontal-view (655 frames), and six are in lateral-view (316 frames), which the models are not trained on. The models output a continuous value between 0 and 1 for each pixel, and a threshold of 0.5 was used to generate a binary segmentation mask from the raw predictions. Each predicted binary segmentation mask was evaluated against the human-annotated label with centerline Dice score, Dice score, and tip-distance error. The centerline Dice scores are given in Fig. 4A and Equations 1-3. Fig. 4B shows equations for Dice score, precision, and recall. (TP = true positive, FP = false positive, TN = true negative, FN = false negative). ^^^^^^^^^^^ ^^^^^^^^^ (^ |^^ ∩ ^^|^^^^) =
[0096] The pixel-level Dice score, precision, and recall can be measured for each frame using Equations 4-6, where TP = true positive, FP = false positive, TN = true negative, and FN = false negative. ^^^ = 2^^(2^^ + $^ + $%)(Eq.4) ^^^^^^^^^^^ ^^^ =Attorney Docket No.10063-094WO1 OTT202334 (Eq.6)
[0097] The superior quantification of topology by the centerline-Dice, compared to the conventional Dice score, is illustrated in Fig. 2C. The conventional Dice score assumes that all pixels segmented are of equal importance; however, this is not true for catheter segmentation, where sacrificing the edge pixels to preserve a connected centerline structure produces a more acceptable result, but could receive a similar score to a thicker but broken segmentation result that excludes multiple sections of the catheter. The tip-distance error per frame was calculated as the Euclidean distance between the coordinates of the ground truth tip and the predicted tip. The tip coordinate was derived from the predicted catheter segmentation mask through skeletonization, followed by the detection of all endpoints and selection of the endpoint furthest from the edge of the frame for frontal sequences and closest to the top edge for lateral sequences
[0012] . Predicted masks were dilated for 3 iterations prior to skeletonization to remove potential artifactual branches followed by removing all pixels not part of the longest connected skeleton.
[0098] Results. Catheters and guidewires were manually segmented on 3,831 fluoroscopy frames from 54 distinct fluoroscopy sequences collected prospectively from 40 patients undergoing diagnostic and therapeutic neuro interventions. The demographic characteristics of the patients included are displayed in TABLE 3. Catheter pixels represented 0.12% to 1.5% of total pixels in each frame. All guidewires in the dataset are 0.035” Glidewire® (Terumo, GR3508). Diagnostic catheters represented in the dataset are 5 Fr Angled Taper Glidecath® (Terumo, CG508), 5 Fr Beacon Tip DAV (Cook Medical, G08699), 5 Fr Simmons / Sidewinder 2 (Terumo CG511), and 5 Fr Beacon Tip Sim2 (Cook Medical, G08422). These devices have been used in both conventional and robotic cerebral angiography, although current robotic systems are limited to a 0.018” guidewire
[0026] . All fluoroscopy sequences are derived from videos taken at 10 frames per second. TABLE 3. Demographic characteristics of patients included in the dataset (n = 40 patients) of 54 distinct fluoroscopy sequences. Characteristic ValueAttorney Docket No.10063-094WO1 OTT202334 Caucasian 29 Black 6
[0099] Cag p g nseen patients in frontal-view sequences, the best overall performing network, nnUNet, achieved a mean centerline- Dice score of 0.98 ± 0.01 on fluoroscopy sequences (n = 11 patients) with a median tip-distance error of 0.48 mm (IQR 0.88 mm, n = 655 frames). This was better than TAG-DL method, which achieved a mean centerline-Dice score of 0.95 ± 0.05 (p < 0.05) and a median tip-distance error of 0.52 mm (IQR: 1.36 mm, p < 0.01) on the same dataset.
[0100] In the lateral view, TAG-DL achieved the best segmentation performance with a mean centerline-Dice score of 0.92 ± 0.03 on unseen fluoroscopy sequences (n = 6 patients). This was not significantly better than nnUNet, which achieved a mean centerline-Dice score of 0.91± 0.02 (p = 0.52). In the lateral view, nnUNet achieved the best median tip-distance error of 0.38 (IQR 1.88) mm (n = 316 frames), which was not significantly better than TAG-DL, which achieved a median tip-distance error of 0.59 mm (IQR: 3.13 mm, p = 0.97). A summary of comparative model performance is displayed in TABLE 4. Boxplots of the performance distribution for all models and frame-by-frame performance over time for nnUNet and TAG-DL is illustrated in Figs.5A-5F. Specifically, Fig. 5A depicts performance distributions on unseen frontal-view fluoroscopy sequences. Boxplots of centerline-Dice (cl-Dice) and tip-distance error are broken down by input scheme and model architecture. Fig. 5B depicts TAG-DL performance on unseen frontal-view fluoroscopy sequences. Performance metrics were plotted for each of the 655 frames from fluoroscopy sequences from 11 unseen patients. All frames depict a frontal view of the patient. Red lines separate distinct patients. Each sequence of fluoroscopy frames corresponds to a videoAttorney Docket No.10063-094WO1 OTT202334 with a frame rate of 10 frames per second. Fig.5C depicts nnUNet performance on unseen frontal- view fluoroscopy sequences. Performance metrics were plotted for each of the 655 frames from fluoroscopy sequences from 11 unseen patients. All frames depict a frontal view of the patient. Red lines separate distinct patients. Each sequence of fluoroscopy frames corresponds to a video with frame rate of 10 frames per second. Fig. 5D depicts performance distributions on unseen lateral-view fluoroscopy sequences. Boxplots of centerline-Dice (cl-Dice) and tip-distance error are broken down by input scheme and model architecture. Fig. 5E depicts TAG-DL performance on unseen lateral view fluoroscopy sequences. Performance metrics were plotted in time order for 316 frames from 6 distinct fluoroscopy sequences from 4 unseen patients. Red lines separate distinct patients. All frames depict a lateral view of the patient, which the AI is not trained on. Each sequence of fluoroscopy frames corresponds to a video taken originally at 10 frames per second. Fig. 5F depicts nnUNet performance on unseen lateral view fluoroscopy sequences. Performance metrics were plotted in time order for 316 frames from 6 distinct fluoroscopy sequences from 4 unseen patients. Red lines separate distinct patients. All frames depict a lateral view of the patient, which the AI is not trained on. Each sequence of fluoroscopy frames corresponds to a video taken originally at 10 frames per second. TABLE 4. Performance of deep learning networks on unseen data. Four network architectures (UNet, TAG-DL, TransUNet, and nnUNet) were trained with three different inputs and tested on unseen fluoroscopy sequences in both frontal and lateral views. Dice, centerline Dice (cl-Dice), are reported as means over the number of patients in the respective dataset. Precision (positive- predictive value), and recall (sensitivity) scores are reported as means over the entire set of frames in each dataset. Tip-distance error is reported as median (inter-quartile range) over the entire set of frames in each dataset. “Contrast” and “Non-Contrast” in the first and second column refer to the DSA Input masks with contrast and without contrast, respectively. (cl-Dice = centerline Dice, PPV = positive predictive value). Deep Learning Model Performance Metrics onAttorney Docket No.10063-094WO1 OTT202334 Contrast DSA TAG-DL 0.93 0.85 0.73 (2.53) Mask Frame nnUNet 0.98 0.91 0.54 (1.04)
[0101] A notable finding is that all networks, except the plain UNet, gained a substantial performance boost from the inclusion of a DSA roadmapping mask in the input compared to the current frame alone. This was true for both unseen frontal and lateral-view sequences, with the magnitude of improvement being greater for the unseen lateral-view sequences (p < 0.001). Additionally, the pure convolutional neural networks, TAG-DL and nnUNet, outperformed TransUNet (p < 0.001). The plain UNet, which has an equivalent structure to TAG-DL, but without rotation-reflection equivariant convolutions, was unable to learn to segment the catheter and labeled the entire frame as “catheter” in all scenarios.
[0102] A visual representation of model performance on tip tracking in three unseen patients, including one lateral view is shown in Fig. 6A. Fig.6A depicts TAG-DL tip position tracking on unseen fluoroscopy sequences. Each row shows a different unseen patient, with the third row showing a patient in the lateral view. On the left column, the last fluoroscopy frame from each sequence is displayed with a partially overlaid digital subtraction angiography (DSA) roadmap.Attorney Docket No.10063-094WO1 OTT202334 The center column shows the human-labeled catheter segmentation mask (ground truth) with a plot of all annotated tip positions from start (in red) to finish (in green) for the set of frames in the sequence. The right column shows the corresponding AI-predicted catheter and tip labels for the set of frames. Sequences featuring the catheter tip traveling through a branch or across a large percentage of the frame were selected for this figure to most effectively illustrate tip tracking.
[0103] On visual inspection, the most common failure mode was false negative detections breaking the connected catheter structure with occasional false positive detections of contrast and other wire-like structures, which is shown in four unseen patients in Fig.6B. Fig.6B depicts visual examples of TAG-DL failure. (1st row) The distal tip of the catheter is missed by the AI in segmentation, resulting in a falsely proximal estimation of tip position. (2nd row) Contrast leakage from the distal catheter dip was incorrectly detected as part of the catheter, causing a false distal estimation of tip position. (3rd row) The catheter structure is broken in multiple points, resulting in a falsely proximal detection of the tip position. (4th row). The AI falsely detected another line- like structure longer than the catheter, resulting in a grossly incorrect tip detection.
[0104] TABLE 5 shows the overall efficacy of TAG-DL. TABLE 5. Data efficiency study of TAG-DL. The study trained TAG-DL with approximately 25%, 50%, and 75% of the number of patients in the training dataset and compared the final test performance with the TAG-DL network trained with 100% of patients in the training dataset. All networks are tested on the same test dataset for both frontal and lateral views used throughout this manuscript. (SD = standard deviation, IQR = interquartile range). Number of Patients Used in TAG-DL Training (Percentage of Total %Attorney Docket No.10063-094WO1 OTT202334 Tip-Distance error 1.04 (5.01) 0.91 (4.84) 0.85 (5.92) 0.59 (3.13) (median (IQR)) (mm) onin Cerebrovascular Interventions using Deep Learning Model
[0105] The study also developed an exemplary hybrid system using deep learning methods for tracking the catheter’s intravascular location in real time during live fluoroscopy. The hybrid system can receive raw fluoroscopy and DSA data, use deep learning to create representations of the catheter and vessels and synthesize this information through interconnected rule-based algorithms using anatomical knowledge for catheter localization.
[0106] Fig.7A (also detailed in Fig.3A) shows a visual representation of the implementation of the exemplary system configured with (i) deep neural networks (e.g., TAG-DL, nnUNet, etc.) and (ii) rule-based algorithms. As shown, the hybrid system can generate both the arterial location of the catheter and richer location information. Two neural networks, e.g., catheter perception (shown as 205 in Fig.3A) and vessel perception (shown as 207 in Fig.3A) can extract the catheter and the distinct vessels in the frame as intermediate outputs, and then a series of rule-based algorithms (e.g., expert system with encoded anatomic knowledge) - comprising tip detection 209, path deviation 211, location inference 213, landmark detection 215, and path generation 217 in Fig.3A - generate the final outputs using the intermediate outputs.
[0107] Discussion: Deterministically localizing catheters using external markers or specialized magnetic tracking modalities are not scalable [7’]-[10’]. Methods of localizing catheters in natural intraoperative fluoroscopy cannot be performed using current state-of-the-art image processing systems due to image noise, sparsity of catheter pixels, and interference from other wire-like structures [8’], [11’], [12’]. However, deep learning models can be used to track catheter tips in X-ray fluoroscopy on a frame-by-frame basis through catheter segmentation. This is an essential first step but lacks the arterial context needed for precise navigation, like knowing a car’s position in space without knowing the street, position on the street, or the exact address. To achieve localization within the vascular network, human-level perception of relevant arteries should be replicated and integrated.
[0108] During catheter navigation, a comprehensive understanding of the vascular network, including vessel origins, terminations, and connections is required to necessitate multi-class semantic segmentation to model vessel topology, particularly in 2D projections with overlapping vessels. The system configured for vessel segmentation, discussed in Example 1, focused on binary segmentation for diagnostic purposes, predominantly in coronary angiography, and did not addressAttorney Docket No.10063-094WO1 OTT202334 overlapping vessels in cerebrovascular DSA (shown in Fig.7B) [19’], [20’]. Specifically, Fig.7B is an illustration of various segmentation paradigms for the carotid artery bifurcation. Binary segmentation could not capture the topology of the carotid bifurcation. A segmentation facilitating true multi-class and overlap between ICA and ECA was necessary to model the topology of the carotid bifurcation.
[0109] Deep learning methods are superior to traditional image processing methods in differentiating vessels from other structures, despite anatomical variation among patients posing challenges for rule-based algorithms and handcrafted features [21’], [22’]. Therefore, deep learning methods are needed to interpret imperfect DSA images, particularly when vessels overlap, to support precise catheter navigation.
[0110] TABLE 6 shows the features provided by the exemplary hybrid system. TABLE 6. Example features provided by the exemplary hybrid system. Feature Description Multi-Class Vessel Se mentation The h brid s stem can se ment multi le overla in vessels with el el nd h) ng on
[0112] Patient Selection and Data Collection: There were no public datasets for catheter segmentation or digital subtraction angiography roadmaps for cerebrovascular interventions. The study developed an annotated dataset from cerebrovascular interventions. TABLE 7 shows the inclusion criteria for the subsets created for catheter segmentation, vessel segmentation, and catheter location. The study obtained all data in DICOM format and converted the data into NIFTI format using XMedCon [25’]. The study collected demographic information for each patient, along with the type of procedure, date of procedure, and surgeon. TABLE 7. Inclusion criteria for catheter segmentation, vessel segmentation, and catheter location inference datasets in the study.Attorney Docket No.10063-094WO1 OTT202334 Inclusion criteria Description (A) Inclusion criteria for catheter segmentation 1. Guidewire and catheter in frame he ne ng es he he eres: Binary catheter segmentation masks were annotated in MRICron, while digital subtraction angiography (DSA) images were annotated in 3D Slicer [26’], [27’]. The internal carotid artery (ICA), external carotid artery (ECA), and common carotid artery (CCA) were segmented on DSA images as three separate masks on the origin and termination of relevant vessels. Overlap was allowed only between the internal and external carotid arteries.
[0114] Deep Learning Model Training: Figs. 7C-7D shows the training loop for vessel and catheter segmentation in the joint-task formulation in the exemplary hybrid system. Specifically, Figs. 7C-7D show a neural network training process for overlapping multi-class vessel segmentation with a compound topology-aware loss function incorporating a centerline-Dice loss function. All fluoroscopy and DSA images were resized to (512, 512) dimension from their original dimension, and the array values were scaled between 0 and 1 before input. A customized loss function was developed for this task with a weighted three-term combination of centerline Dice loss, conventional Dice loss, and vessel-wise binary cross entropy (BCE), as shown in Equation 7 and TABLE 8. +^,-. / ^^0 = 1 ∗ +234 + 5 ∗ +6 / ^^ + 7 ∗ +^80 / ^^(Eq.7) TABLE 8. Definition of custom loss function enhancing the composite binary.Attorney Docket No.10063-094WO1 OTT202334 Given the following definitions: • 9^^^0as the predicted binary catheter segmentation mask by the model. thnd edour different values of gamma (0, 0.25, 0.50, and 1.00) corresponding to 0%, 11%, 20%, and 33% representation of centerline-Dice loss in the overall loss, with equal representation between conventional dice and binary cross-entropy losses in the remaining proportion.
[0116] The joint-task learning loss function can be defined per Equation 8 and TABLE 9. +Q, / ^: = 1 ∗ +234 + 5 ∗ +6 / ^^ + 7 ∗ +^80 / ^^ + R+4S^^^T8:U(Eq.8) TABLE 9. Definition of custom loss function enhancing the composite binary cross entropy and Dice loss function used in nnUNet with the topological information from the centerline Dice loss.Attorney Docket No.10063-094WO1 OTT202334 The joint-task learning between catheter and vessel can be formulated as follows: • VW^^^^8as the binary vessel segmentation mask in the same FOV. The distinct identity of the an s Y l). inted four values of lambda (0, 0.25, 0.50, and 1.00) and two values of gamma (0 and 0.25) with equal representation between conventional dice and binary cross-entropy losses in the remaining proportion. The catheter and vessel segmentation nnUNet models were trained for 1000 epochs, and the best network was chosen through five-fold cross-validation [28’].
[0118] Segmentation Performance Assessment: For both catheter and vessel, segmentation performance was assessed through the Dice score and centerline-Dice score, which quantified topological accuracy as described in FIGS. 3A-3C and Equations 1-6, along with clinically significant landmarks [29’]. For catheter segmentation, segmentation scores and tip-distance error (in millimeters) for each frame can be calculated using the method shown in FIGS. 2E-2F [15’]. For vessel segmentation, segmentation scores were calculated for each DSA image for the overall combined structure and each distinct vessel. For each vessel mask, the exemplary method can calculate the bifurcation point, ICA origin, ECA origin, and errors between the ground truth and prediction, respectively, as shown in FIGS.3D-3G.
[0119] Discrete Arterial Location Inference: Location inference at the vessel label was assessed at a per-label classification problem with five options: “CCA,” “ICA,” “ECA,” “ICA or ECA,” and “Out of Range.” The method for deriving the arterial location from the catheter andAttorney Docket No.10063-094WO1 OTT202334 vessel masks, with and without accounting for deformation, is shown in FIG. 3H. Localization performance was assessed at both the frame level and the patient level. At the frame level, the performance of the discrete location inference method was assessed with a confusion matrix between the ground truth location labels and the predicted location levels [30’]. The centerline- Dice scores for both catheter and vessel segmentations were also calculated for each frame and compared with the location inference result [31’], [32’].
[0120] Continuous Location Inference: The continuous location inference comprises path generation, which calculates the deviation of the catheter tip from the path and the distance along the path to the destination point. The destination point for each test case was chosen through a visual assessment of the catheter location at the end of the sequence defined by two variables: the target vessel and the distance to the target in centimeters. The method for creating a path from a designated starting proximal vessel to the destination point is shown in Fig.3I. The paths created by the method in Fig. 3I were visually and quantitatively assessed by measuring the Euclidean distance between the start point and end point between the ground truth and the AI-generated paths, along with the differences in path lengths.
[0121] A path to the destination was generated for each DSA roadmap. For each frame in the hold-out test dataset, the in-plane deviation of the tip coordinates from the generated path and the distance to the endpoint was calculated using the methods shown in Figs.3J-3K.
[0122] Hardware Prototype of the Exemplary System. To develop a hardware prototype for real-time inference, the study encapsulated the system within a commercial tower workstation (Lenovo P620), augmented with a dedicated GPU (NVIDIA RTX A4500) and two video capture cards (StarTech) for live video input through a DVI connection, as shown in Fig.7E. Specifically, Fig. 7E is an illustration of a joint-task learning process where the vessel segmentation mask is used in training the catheter segmentation network by panelizing catheter segmentation predictions outside the vessel mask. Since the connection to fluoroscopy was unavailable, test fluoroscopy sequences were played on a separate computer and received by the system through a DVI connection, as in a fluoroscopy setting. Fig.3A, subpanel (b) shows an example operation flow of the prototype of the system. OpenCV functions were used to design a custom pipeline that queried the video capture card, processed the input, passed it through the system, and then displayed the inference results on an external monitor.
[0123] The total inference time of the pipeline was measured, along with its component steps. TABLE 10 shows the degradation in input quality and shift in distribution assessed by comparing the signal-to-noise ratio (SNR), peak signal-to-noise ratio (PSNR), and Earth Mover’s DistanceAttorney Docket No.10063-094WO1 OTT202334 (EMD) [33’]. A higher PSNR (30-50 dB) indicated better fidelity to the original image, while EMD assessed the shift in pixel intensity distributions, comparing how closely images captured through a frame grabber resembled the original NIFTI images [34’]. Finally, the study compared the vessel segmentation performance of the trained nnUNet deep learning model on frame-grabbed inputs versus original NIFTI images, quantifying the change in centerline-Dice scores. TABLE 10. Definitions for signal-to-noise ratio (SNR) and Earth Mover’s Distance (EMD) to quantify the noise generated by the frame grabber process relative to the original input image and the shift in distribution, respectively. ^ ^%^ = 10 ∗ log ^` / _^T8P]a^` / ^^s n.. 1 to February 2024 from diagnostic cerebral angiography (64%), arterial embolization (26%), mechanical thrombectomy (8%), and carotid artery stenting (2%) procedures. All data were captured from three Siemens Artis bi-plane wall-mounted fluoroscopy systems. The dataset for vessel segmentation included 241 DSA images from 103 patients, of which the network wasAttorney Docket No.10063-094WO1 OTT202334 trained on 154 DSA images from 68 patients and tested on an unseen group of 87 images from 35 patients. Additional characteristics at both the patient and image level are shown in TABLE 11. Of these characteristics, 118 (53.1%) were acquired in frontal view, and 104 (46.8%) were acquired in lateral view in the overall dataset. A subset of the catheter segmentation dataset comprising 2,670 frames from 32 patients was used to train and test the joint-task model and assess the impact of the varying influence of extravascular and topological loss.
[0125] Figs.8A – 8N show clinical results of the exemplary system and method employed in the study. TABLE 11. Demographic characteristics of patients in the dataset used for vessel segmentation at the patient and image levels. There were 241 images in the dataset from 103 patients since some patients had multiple images. Age is shown as mean ± standard deviation. Categorical variables are reported as numbers (percentages). Patients Images Characteristic Total Training Test (N = 241) (N = 103) (N = 68) (N = 35) Age 64.3 ± 14.2 64.3 ± 13.6 66.0 ± 15.3 63.9 ± 15.0 Sex Female 64 (62.1%) 43 (63.2%) 21 (60.0%) 150 (62.2%) Male 39 (37.9%) 25 (36.8%) 14 (40.0%) 91 (37.8%) Race Caucasian 75 (72.8%) 53 (77.9%) 22 (62.9%) 161 (72.5%) Black 17 (16.5%) 11 (16.3%) 6 (17.1%) 36 (16.2%) Asian 6 (5.8%) 1 (1.5%) 5 (14.3%) 9 (4.1%) Declined 4 (3.9%) 2 (2.9%) 2 (5.7%) 14 (6.3%) Native American 1 (1.0%) 1 (1.5%) 0 (0.0%) 2 (0.9%) Ethnicity Hispanic / Latino 27 (26.2%) 20 (29.4%) 7 (20.0%) 151 (68.0%) Not Hispanic / Latino 73 (70.9%) 47 (69.1%) 26 (74.3%) 63 (28.4%) Declined 3 (2.9%) 1 (1.5%) 2 (5.7%) 8 (3.6%) Procedure Diagnostic Cerebral 61 (59.2%) 39 (57.4%) 22 (62.9%) 142 (64.0%) Angiogram Arterial 26 (25.2%) 20 (29.4%) 6 (17.1%) 57 (25.7%) EmbolizationAttorney Docket No.10063-094WO1 OTT202334 Mechanical 13 (12.6%) 7 (10.3%) 6 (17.1%) 18 (8.1%) Thrombectomy Carotid Stenting 3 (2.9%) 2 (2.9%) 1 (2.9%) 5 (2.3%) Surgeon 1 (AH) 31 (30.1%) 21 (30.9%) 11 (31.4%) 88 (36.5%) 2 (KY) 31 (30.1%) 20 (29.4%) 18 (51.4%) 87 (36.1%) 3 (YJZ) 22 (21.4%) 14 (20.6%) 1 (2.9%) 39 (16.2%) 4 (GWB) 19 (18.4%) 13 (19.1%) 5 (14.3%) 27 (11.2%) Year Collected 2021 20 (19.4%) 19 (27.9%) 1 (2.9%) 35 (14.5%) 2022 11 (10.7%) 8 (11.8%) 3 (8.6%) 16 (6.6%) 2023 64 (62.1%) 41 (60.3%) 23 (65.7%) 173 (71.8%) 2024 8 (7.8%) 0 (0.0%) 8 (22.9%) 17 (7.1%)
[0126] Vessel Segmentation Results: Fig. 8A shows visual results of vessel segmentation and landmark detection in 6 unseen patients, performed by the exemplary system, in decreasing order of roadmap quality. Specifically, Fig. 8A depicts visual results of vessel segmentation and landmark detection in 6 unseen patients, performed by the exemplary system, in decreasing order of roadmap quality. Each row contains the subtracted digital subtraction angiography (DSA) roadmap, the ground truth segmentation mask, and the AI prediction at increasing levels of centerline-Dice loss represented by a weighting factor (γ). The bifurcation point and origins of the internal carotid artery and external carotid artery are also plotted for each patient. As shown, on unseen patients, the predicted vessel segmentation was anatomically accurate across DSA images of varying input quality.
[0127] Fig.8B shows the vessel segmentation scores across DSA images. As shown, the best- performing network achieved a centerline-Dice score of 0.94 ± 0.09. For the network with γ = 0, the distribution of segmentation performance, measured by the centerline-Dice score, was overall worse for the external carotid artery relative to the internal carotid and common carotid arteries (p < 0.001 for ECA vs. ICA and ECA vs. CCA).
[0128] Fig. 8C shows the landmark detection performance of the exemplary system. As shown, the three landmarks extracted were the origin of the internal carotid artery, the origin of the external carotid artery, and the point of bifurcation. The error was calculated as the Euclidean difference between the landmarks extracted from the ground truth segmentation and the landmarks extracted from the predicted segmentation. Among the predicted segmentations, there were 6Attorney Docket No.10063-094WO1 OTT202334 (6.9%) images with at least one missing landmark due to an error in segmentations, a miss of a vessel, or errors in the logic of the landmark detection algorithms.
[0129] Fig. 8D shows example failure modes (e.g., errors in segmentations, missing vessels, or errors in the logic of landmark detection algorithms) in landmark detection. Among those with preserved landmark detection, the median error for the detection of the bifurcation of the common carotid artery (CCA) to the internal carotid (ICA) and external carotid (ECA) arteries was 1.26 [0.51 – 1.85] mm. The origin of the internal carotid artery was localized with a median error of 1.26 [IQR: 0.58 – 1.91] mm. The ICA origin was missed in 2 of 68 images and the error was greater than 20 mm in 5 of 68 images. The origin of the external carotid artery was detected with a median error of 1.51 [1.01 – 2.74] mm. There was no difference between the localization of the origins of the internal and external carotid arteries (p = 0.45).
[0130] Discrete Location Inference: The location inference method was tested on a hold-out test set of 726 frames of video from 10 patients unseen to both the catheter and vessel segmentation models. The overall accuracy of the inference method was 91.6% across all frames. The accuracy on the common carotid artery (n = 93 frames) was 89.3%, the internal carotid artery (n=155 frames) was 94.8%, and the external carotid artery (n = 96 frames) was 92.7%.
[0131] Fig.8E shows confusion matrices demonstrating location inference performance with and without correction for vessel deformation. The confusion matrix is in the x-y plane, while the centerline-Dice score range is the z-axis. TABLE 12 shows the patient-level outcomes of the location inference in the exemplary system. The average localization accuracy was 91.3% (n = 10 patients), ranging from 77.6% to 98.5% with a standard deviation of 7.5%. TABLE 12. Patient-level outcomes for location inference. Mean cl-Dice scores are shown as mean ± standard deviation at the patient level and among all 10 patients. Catheter Segmentation Vessel Segmentation Sequence Location ID Mean cl- TDE Inference Dice Overall CCA ICA ECA GT Vessel Accuracy111_3 0.97 ± 0.02 0.30 (0.42) 0.98 0.97 1.00 0.98 0.98 ± 0.01 0.970 118_4 0.96 ± 0.03 0.50 (0.80) 0.99 0.99 1.00 0.97 0.98 ± 0.01 0.961 28_3 0.90 ± 0.09 5.90 (8.73) 0.95 0.98 1.00 0.88 0.99 ± 0.01 0.920 34_1 0.97 ± 0.01 0.38 (0.53) 0.98 0.99 0.99 0.97 0.99 ± 0.00 0.985 51_1 0.99 ± 0.01 0.26 (0.36) 0.99 0.99 1.00 0.99 0.99 ± 0.01 0.969 59_2 0.81 ± 0.21 0.54 (0.69) 0.79 0.76 0.96 0.65 0.96 ± 0.00 0.792Attorney Docket No.10063-094WO1 OTT202334 60_3 0.98 ± 0.03 0.51 (2.01) 0.96 0.94 0.99 0.96 0.97 ± 0.02 0.940 65_1 0.98 ± 0.02 0.36 (0.51) 0.73 0.97 0.90 0.34 0.95 ± 0.03 0.965 68_5 0.70 ± 0.16 1.13 (3.35) 0.96 0.95 0.98 0.96 0.97 ± 0.01 0.776 69_4 0.99 ± 0.03 0.21 (0.29) 0.94 0.87 0.99 0.97 0.96 ± 0.04 0.941 Overall 0.93 ± 0.10 0.44 (2.73) 0.93 ± 0.94 ± 0.98 ± 0.87 ± 0.09 0.07 0.03 0.22 0.98 ± 0.15 0.92 ± 0.08
[0132] The association of catheter and vessel perception was assessed using correct and incorrect location inference classifications. In the univariate comparison of frames by classification outcomes (shown in TABLE 13), catheter perception is worse in frames that are misclassified (0.95 vs. 0.89, p < 0.001). Overall, vessel segmentation was not worse in misclassified frames; however, segmentation of the vessel corresponding to the ground truth location was worse (0.98 vs. 0.95, p < 0.001). Catheter segmentation tip detection was a more significant source of error in location inference. This impact of catheter and vessel segmentation performance was also assessed by plotting the distribution of catheter and vessel segmentation performance (per cl-Dice score) against the confusion matrix of location classification labels [31’].
[0133] Fig. 8F shows the plot demonstrating the distribution of catheter and vessel segmentation against the confusion matrix of location classification labels. The confusion matrix is in the x-y plane, while the centerline-Dice score range is the z-axis. TABLE 13. Comparison of catheter and vessel perception between correctly and incorrectly classified frames. The “Ground Truth” vessel segmentation refers to the segmentation of the vessel corresponding to the ground truth vessel for each frame. All cl-Dice scores are shown as mean ± standard deviation. Tip-distance error is shown as median (IQR). Location Inference Result (n = 580 frames) Performance Measure p-value Correct Incorrect (n=534) (n=46) Catheter Perception Segmentation 0.95 ± 0.08 0.89 ± 0.22 < 0.001 (cl-Dice) Tip-Distance Error (mm) 0.36 (0.80) 0.78 (5.24) < 0.001 Vessel Segmentation (cl-Dice) Overall 0.94 ± 0.07 0.94 ± 0.05 0.77Attorney Docket No.10063-094WO1 OTT202334 CCA 0.94 ± 0.06 0.91 ± 0.07 < 0.001 ICA 0.98 ± 0.03 0.98 ± 0.02 0.56 ECA 0.89 ± 0.19 0.92 ± 0.13 0.29 Ground Truth 0.98 ± 0.03 0.95 ± 0.05 < 0.001
[0134] Continuous Location Inference: Table 14 shows the quantitative metrics of path generation by the exemplary system. As shown, the median errors in starting point, ending point, and path length were 2.76 mm, 1.32 mm, and 2.58 mm, respectively. Two cases had errors in either the starting or ending vessel. Figs. 8G-8H show the visual results of the path generation by the exemplary system. TABLE 14. Comparison of AI-generated path to ground truth for 10 unseen patients. Destination Path Distance Errors (mm) Sequence ID Vessel Distance Total Path Length Start Point End Point Path Length28_3 ICA 7.5 cm 12.63 2.04 10.64 34_1 ICA 1 cm 9.41 cm 1.13 2.52 2.63 51_1 ICA 4 cm 10.66 cm 3.08 0.26 2.81 59_2 ICA 1 cm 4.14 cm 2.43 2.77 0.76 60_3 ICA 5 cm 7.53 cm 1.55 1.41 2.53 65_1 ICA 1 cm 7.73 cm 2.44 18.40 16.34 68_5 ICA 5.5 cm 7.58 cm 3.56 0.36 2.49 69_4 ECA 2.5 cm 5.08 cm 9.51 0.46 10.05 111_3 ECA 4 cm 10.22 cm 3.59 1.23 1.49 118_4 ECA 2 cm 6.51 cm 1.47 1.07 1.07 Overall - - 7.75 cm 2.76 1.32 2.58
[0135] For each frame in the hold-out test dataset, the study assessed the deviation of the tip coordinate from the centerline and the distance to the endpoint. The system had a median error of 0.36 (IQR: 0.56) mm in detecting path deviation and an error of 1.09 (IQR: 2.25) mm in judging the path length from the catheter tip to the destination point. Additionally, the system can represent pull-back and tip advancement events, as shown in Fig.8I.Attorney Docket No.10063-094WO1 OTT202334
[0136] Fig.8I shows path deviation and distance from an endpoint of the vessel path calculated frame-by-frame for two sequences (n=53 and 103 frames) having ground truth (GT) deviation values and AI predicted deviation values.
[0137] Fig. 8J shows the navigation components (e.g., discrete and continuous inference outputs) overlaid on three frames of an unseen test sequence. She visualization of navigational outputs on three frames of a sequence (n = 55 frames) in the beginning (frame 0), middle (frame 33), and end (frame 55). The column represents the raw input fluoroscopy data with overlaid DSA roadmap, ground truth human annotation, and AI prediction. The catheter tip is assessed for its deviation from the path and distance from the destination.
[0138] The study also assessed trends in the temporal performance of the system in two example patients. In the first patient, the lapse in localization performance was related to poor catheter segmentation and subsequent tip detection, whereas in the second patient, the lapse in localization performance was attributable to poor vessel segmentation.
[0139] Real-time Performance Results of the System: Figs.8K-8L show a hardware prototype and visual output of the exemplary system. Fig.8K shows the schematic and setup of the hardware prototype of the exemplary system, wherein the live fluoroscopy video was acquired by the server housing the AI models through a DVI connection. Fig.8L shows the inference occurring in real- time on the live fluoroscopy sequence, and results were displayed on a separate monitor.
[0140] On the NVIDIA A4500 GPU of the prototype, the mean inference time on single frame two-channel with the native nnUNet inference pipeline was 15.24 seconds (n = 10 runs). After modifying the nnUNet inference model of the prototype, the study decreased the mean inference time to 56 milliseconds, with a component means of 8 milliseconds for preprocessing and 47 milliseconds for the forward pass through the network. Additional steps completed frame-by- frame were calculations between the catheter and vessel masks (tip detection, location inference, path deviation, and distance to endpoint) and rendering the frame with overlaid information, which took a mean time of 11 milliseconds and 44 milliseconds, respectively. The mean duration of the sum of all frame-wise operations was 121 milliseconds, ranging from 94 to 147 milliseconds over an entire sequence, resulting in a potential performance speed of 6.8 frames per second.
[0141] The images acquired through a video capture card in the prototype were degraded in quality compared to training images, as shown in TABLE 15. As shown, SNR decreased from 2.2% to 33.1% across all test sequences, and PSNR values increased from 22.1 to 28.1. In addition to the loss in image quality, shifts in input distribution between the original image and the same image after video capture were measured with Wasserstein distances ranging from 0.02 to 0.06.Attorney Docket No.10063-094WO1 OTT202334 The mean observed decrease in cl-Dice was -0.035 (n = 10 patients) when using video capture images as input to the network compared to the original NIFTI files, showing strong performance despite shifts in input image quality and distribution. TABLE 16 shows the breakdown by patients. TABLE 15. Comparative assessment of image quality degradation between the frames inputted into the network from video capture compared to the original NIFTI file. Signal-to-Noise Ratio Sequence ChannelFromFrom VideoPSNR EMDNIFTI Capture ∆ % 28_3 Non-contrast 1.538 1.524 0.014 0.88 20.641 0.041 28_3 Contrast 1.589 1.571 0.019 1.19 20.759 0.040 34_1 Non-contrast 2.898 2.208 0.690 23.82 24.129 0.055 34_1 Contrast 2.995 2.339 0.656 21.90 24.979 0.049 51_1 Non-contrast 2.498 1.871 0.627 25.10 21.812 0.049 51_1 Contrast 2.688 2.025 0.663 24.66 21.876 0.047 59_2 Non-contrast 1.311 1.349 -0.038 -2.89 22.960 0.024 59_2 Contrast 1.356 1.395 -0.039 -2.88 22.952 0.024 60_3 Non-contrast 3.741 1.817 1.924 51.43 17.482 0.108 60_3 Contrast 3.939 1.903 2.035 51.68 17.189 0.110 65_1 Non-contrast 1.642 1.536 0.106 6.47 22.044 0.045 65_1 Contrast 1.725 1.619 0.106 6.16 22.158 0.045 68_5 Non-contrast 2.028 1.650 0.378 18.65 22.308 0.047 68_5 Contrast 2.095 1.703 0.393 18.75 22.409 0.045 69_4 Non-contrast 2.726 1.938 0.789 28.92 22.572 0.061 69_4 Contrast 2.910 1.948 0.963 33.08 22.117 0.063 111_3 Non-contrast 0.807 0.783 0.024 3.03 27.929 0.023 111_3 Contrast 0.822 0.799 0.023 2.83 28.062 0.023 118_4 Non-contrast 2.296 1.868 0.429 18.67 21.380 0.043 118_4 Contrast 2.375 1.930 0.445 18.73 21.415 0.041 Average 2.199 1.689 0.510 17.51 22.360 0.049 TABLE 16. Assessment of difference in performance of the segmentation network for vessels using video capture images compared to the original NIFTI file. Sequence View Segmentation Performance (cl-Dice)Attorney Docket No.10063-094WO1 OTT202334 Original On On Video Resolution NIFTI Capture ∆ %∆ 28_3 512 Lateral 0.826 0.78 -0.046 -5.57 34_1 960 Frontal 0.931 0.870 -0.061 -6.55 51_1 512 Frontal 0.930 0.906 -0.024 -2.58 59_2 512 Lateral 0.731 0.743 0.012 1.64 60_3 512 Frontal 0.892 0.844 -0.048 -5.38 65_1 512 Frontal 0.600 0.556 -0.044 -7.33 68_5 960 Lateral 0.861 0.853 -0.008 -0.93 69_4 512 Lateral 0.938 0.878 -0.060 -6.40 111_3 1024 Frontal 0.951 0.913 -0.038 -4.00 118_4 960 Lateral 0.910 0.873 -0.037 -4.07 Average 0.857 0.822 -0.035 - 4.12
[0142] Joint-Task Learning Results between Catheter and Vessel: The deep learning model in the system was trained on 1,944 frames from 22 unique patients and tested on 726 frames from 10 unique patients. The joint-task architecture improved both the catheter segmentation (p < 0.01) and tip-distance error (p < 0.01) relative to the baseline model, which had only the current and non-contrast frame. Incorporating centerline-Dice loss with a weight of 0.25 worsened performance in models with extravascular loss weights of 0.25 and 0.50.
[0143] Fig. 8M shows the distributions of all models compared to the baseline catheter segmentation models, with numerical results shown in TABLE 17. Specifically, Fig. 8M depicts the boxplots of catheter segmentation metrics for the 7 different model configurations. In Fig.8M and TABLE 17, the reduction in dataset size reduced the performance from 93% to 88.71% for the baseline catheter segmentation model. TABLE 17. Segmentation and tip-tracking performance of five deep learning network architectures on unseen lateral-view fluoroscopy sequences. λ is the weight of the extravascular prediction penalty, while γ is the weight of centerline-Dice loss in the overall loss function. The baseline loss function in nnUNet is a 1:1 ratio of Dice loss and cross-entropy loss, which are maintained in each level. (TDE = tip-distance error in millimeters, C = current frame, NC = pre- contrast DSA mask, VM = vessel mask). Segmentation Perfo nels Joint Los rmance Input Chan s Loc. Weights (n = 726 frames) Acc.Attorney Docket No.10063-094WO1 OTT202334 1 2 3 s t cl-Dice Dice Prec. Recall TDE C NC - - - 0.91 ± 0.82 ± 0.84 ± 0.82 ± 0.46 0.16 0.20 0.20 0.22 (1.94) 88.71% VM 0.00 0.00 0.94 ± 0.87 ± 0.88 ± 0.87 ± 0.39 0.12 0.12 0.13 0.11 (0.99) 89.26% C NC VM 0.10 0.00 0.94 ± 0.86 ± 0.86 ± 0.88 ± 0.42 0.11 0.11 0.14 0.09 (1.17) 88.98% C NC VM 0.25 0.00 0.94 ± 0.86 ± 0.85 ± 0.89 ± 0.38 0.12 0.13 0.16 0.09 (1.09) 89.26% C NC VM 0.50 0.00 0.93 ± 0.86 ± 0.85 ± 0.90 ± 0.41 0.12 0.12 0.15 0.09 (1.45) 88.98% C NC VM 0.25 0.25 0.80 ± 0.79 ± 0.84 ± 0.77 ± 0.50 0.16 0.14 0.18(1.20) 86.64% C NC VM 0.50 0.25 0.76 ± 0.77 ± 0.82 ± 0.75 ± 0.51 0.16 0.16 0.20 0.14 (1.64) 86.91%
[0144] To assess the potential of the loss function terms to penalize extravascular segmentations, the study identified frames that had false positives or missing segmentations in the baseline model and compared the frames to the predicted segmentations of vessel-guided models with six configurations of λ and γ, representing different levels of influence on vessel and topological guided segmentation. Fig. 8N shows the visual results of four frames from four different patients. Specifically, Fig. 8N depicts the visual results of a joint-task learning process at varying weights of extravascular penalty and centerline-Dice loss on four frames from four different patients.
[0145] Discussion
[0146] Discussion #1: The instant study demonstrates that deep learning methods can generate highly accurate catheter segmentation results, which can subsequently be used to track the tip position during cerebral angiography. This work is significant as it opens a new paradigm of innovation in neuro intervention, enabling automation previously hindered by computers' inability to "see" catheters in real-world fluoroscopy. Automation, even in robotic-assisted procedures, has been demonstrated to make them safer and is critical to advancing the field
[0027] . The dataset of complete catheter and guidewire annotations is one of the largest reported in the literature and the first in neurointervention. The strategies demonstrated for improving performance leverage intraoperative data and exploit the catheters' geometric properties, which are translatable to other scenarios for catheter segmentation. Surprisingly, these methods perform well in the lateral view despite not having been trained on this view, highlighting their generalizability and suggestingAttorney Docket No.10063-094WO1 OTT202334 invariance to specific anatomy. This is crucial for the translation of AI-augmented robotic navigation, which will require further research and development in animal models
[0028] . Even for human operators, highlighting the catheter and its tip could help reduce eye strain and provide helpful feedback to trainees.
[0147] The integration of digital subtraction angiography (DSA) road mapping masks into deep learning models significantly boosted performance on unseen data, likely due to the impact of learning guided by subtraction. This follows intuition as the subtraction of bony features in routine DSA road mapping accentuates the catheter for human operators. In the context of deep learning, this reduces the space of features that could confound the learning process. Additionally, this provides some temporal information, which has been shown to be helpful in prior work, without propagating error over time, allowing more reliable performance that can recover from temporary lapses [12,13]. Furthermore, models without this information performed significantly worse in the unseen lateral view, suggesting that subtraction can be leveraged by these networks for strong out-of-domain performance that is agnostic to specific anatomy. Notably, the contrast DSA mask, which highlights arterial anatomy, did not improve performance relative to the non- contrast mask. Since the catheter is located intra-arterially, it was expected that the contrast-filled DSA mask would address the inherent lack of attention in CNNs, which can be a major limitation in segmentation, especially due to the sparsity of the catheter (0.12 – 1.5%) compared to the background. However, the lack of benefit may reflect the variations in contrast injection and final roadmap quality. Nevertheless, the lack of reliance on contrast is preferrable due to both variability and potential toxicity.
[0148] The TAG-DL method, which combines rotation-reflection equivariance and topology- aware cl-Dice loss, outperformed its corresponding plain UNet by a large margin at this dataset size by leveraging inherent symmetries in catheter movement. Catheters can rotate arbitrarily and enter the frame from different directions, which deep learning models require many parameters to learn, while humans can perceive the catheter equally despite its orientation
[0020] . Incorporating this property into neural networks is part of the larger paradigm of geometric deep learning, which allows for better generalizable learning on smaller datasets
[0015] . However, the base nnUNet performed best overall, suggesting that automated optimization of the UNet structure and hyperparameters can achieve similar performance. TransUNet outperformed the plain UNet, demonstrating some benefit of modeling long-range dependency introduced by the transformer architecture [14,23]. However, it performed significantly worse than TAG-DL and nnUNet in thisAttorney Docket No.10063-094WO1 OTT202334 study. This is likely due to the lack of inductive bias in transformers, which require significantly more data to meet or exceed the performance of CNNs
[0014] .
[0149] Incorporating the topology-aware centerline Dice score both improved and provided a more meaningful assessment of the segmentation outcome. The centerline Dice and its corresponding loss, which preferentially rewards topologically accurate segmentations, better aligns with the goals of catheter segmentation than the conventional Dice score, which is used in prior literature
[0016] . This minimizes post-processing repair with more error-prone morphological operations. Nevertheless, even the cl-Dice has limitations in this setting. It assumes that all regions of the centerline are equally important, while, in practice, the tip is more relevant than the rest of the catheter body. Future work that incorporates tip-distance error into the loss function could reward accurate tip segmentation and further improve tip-tracking accuracy.
[0150] Discussion #2: The instant study also demonstrated an embodiment of the exemplary system employing machine learning algorithms to process and transform raw fluoroscopy and DSA data into a symbolic representation of catheters and vessels, which can then be interpreted by a rule-based module to produce both discrete and continuous assessments of location and deviation from the intended path.
[0151] By employing deep learning for narrow abstract perception tasks and using rule-based modules to produce the final outputs, the exemplary system can provide system-level explainability and deliver superior performance despite a limited dataset size compared to non- medical applications. Explainability and data efficiency are important considerations for medical applications, further supporting the utility of this approach [35’]. The study further demonstrated that the system can be deployed under real-world conditions, including performance at multiple frames per second and the ability to handle inputs degraded by standard video communication interfaces without application-specific fine-tuning.
[0152] The exemplary system is the first example of both discrete and continuous catheter localization in the context of patient-specific vascular networks in natural patient fluoroscopy. While previous studies demonstrated catheter tracking in a frame, catheters did not freely move throughout the frame and were constrained within patients’ specific anatomy [15’], [16’]. Thus, previous studies did not define the location state of the catheter in relation to the current intravascular location and path to a target vessel, which can be necessary for navigation. Furthermore, the navigational perception related to path generation, deviation, and distance to destination was reliant on the multi-layered perception of vessels in 2-dimensional imaging. Previous studies in 2D angiographic images performed binary segmentation with the goal ofAttorney Docket No.10063-094WO1 OTT202334 morphological analysis, which cannot be used for continuous location inference in the context of selective catheterization due to all vessels being considered a single entity, which prevented the distinct perception of vessels over and underlying each other [22’], [36’].
[0153] The development of a multi-class vessel segmentation method to handle DSA images in the instant study is a new contribution to the literature and a key feature for the exemplary system that facilitates the system to go beyond catheter tracking methods to achieve true localization. Segmentation of the carotid arteries has been previously described in CT and MR angiography [37’-39’], which differed from the instant study. CT and MR angiograms in previous studies were 3-dimensional modalities with long acquisition times, providing unambiguous representation of the 3-dimensional vascular tree but precluding direct use for real-time intraoperative navigation [37’], [40’], [41’]. Compared to conventional angiograms optimized for vessel morphological analysis and can be re-acquired as needed, DSA roadmaps used for catheter navigation were acquired on the fly and have overlapping or kissing vessels.
[0154] The vessel segmentation method in the exemplary system can handle variations in roadmap quality and represent the topology of the bifurcation despite the overlap, a capability not demonstrated by previous studies. Additionally, the study introduced landmark detection methods that provided clinically relevant metrics for grading vessel segmentations. For instance, although ECA segmentation was worse than ICA segmentation, the detection of ECA and ICA vessel origins showed no difference, indicating that segmentation near the bifurcation was preserved, which cannot be captured by generic pixel-level segmentation metrics. In the study, the reported landmark detection errors under 2 millimeters were lower than the mean diameters of the CCA (9.6 mm), ICA (7.6 mm), and ECA (4.9 mm) [42’].
[0155] The study demonstrated the suitability of the exemplary system for real-world applications through the strong performance of its prototype in real-time on images acquired from standard video capture hardware, which may be interoperable among various fluoroscopy systems. The 23.80% decrease in mean SNR, PSNR values between 20-30 dB, and an average EMD of 0.045 between frame grabber images and original NIFTI images reflected the information loss and distribution shifts in the process of acquiring the imaging data in this manner; however, the prototype of the system can perform under these conditions, likely due to the study using raw contrast and non-contrast images as inputs, mitigating the impact of distribution shifts. Additionally, the conditions under which the prototype was tested can be a greater challenge to the system than in deployment, as a direct connection to the fluoroscopy machine may acquire higher- resolution images.Attorney Docket No.10063-094WO1 OTT202334
[0156] The study also demonstrated a method of catheter segmentation that incorporated vessel information. This leveraged the real-world assumption that the catheter and its tip were intravascular and distal to the starting point, which had not been utilized before. While prior loss functions did not prioritize the parts of the segmentation near the tip, which was an important part of the catheter for clinical and navigational purposes, the study demonstrated a method for integrating this information to improve catheter segmentation and tip detection.
[0157] The information on the vessel segmentation mask produced an increase in both segmentation performance and tip-distance error, driven by improvements in precision and recall. In a previous study, there was no difference between including the contrast-filled DSA mask or the subtracted DSA roadmap as inputs into the catheter segmentation network [15’]. This approach may have failed previously because DSA roadmap quality varied depending on the contrast amount, the mask selection timing, and movement artifacts. The instant study overcame this by using the exemplary vessel segmentation method to level the quality between DSA images, which may have unmasked the network’s ability to consider vessel information for learning.
[0158] The increase in the influence of extravascular penalty was balanced out by other terms, suggesting that the network may learn different ways to utilize the vessel segmentation mask to improve catheter segmentation. Incorporating the centerline-Dice loss function did not improve the centerline-Dice score, which may be due to the soft skeletonization being utilized in the loss function, as opposed to the hard non-differentiable Zhang-Suen skeletonization used in the calculation of the score. A more faithful representation of the skeletonization may improve its contribution to the loss function.
[0159] The exemplary system can have some limitations. Firstly, the system perceives the entire structure of the catheter and guidewire as a single binary segmentation mask. While this is sufficient for localizing the distal tip, which is relevant for navigation, distinguishing the catheter and guidewire as separate masks is necessary for capturing device-over-device movements, which is integral to the successful exchange of co-axial devices. Additionally, the detection of therapeutic devices, such as embolization coils and stents, is not addressed by this method, even though these steps are necessary for completing an intervention.
[0160] In vessel perception, the example is limited to carotid bifurcation. However, in navigation to both proximal and distal arteries, the carotid bifurcation occurs in the normal intervention course. A concerted effort to collect DSA roadmaps of such vessels can be necessary to expand the capabilities of the exemplary system. However, procedures involving these distal vessels are less common and may occur on an emergent basis, as in mechanical thrombectomy forAttorney Docket No.10063-094WO1 OTT202334 acute ischemic stroke. Furthermore, large caliber vessels (e.g., the aorta) may be challenging to capture with DSA road mapping and are better suited for 2D / 3D fusion to generate in-frame vessel segmentation masks. Despite these challenges, the independence between catheter and vessel segmentation facilitates the system to accept vessel segmentations generated by other means, such as 2D / 3D fusion roadmaps derived from CT angiography.
[0161] Conclusion
[0162] The construction and arrangement of the systems and methods, as shown in the various implementations, are illustrative only. Although only a few implementations have been described in detail in this disclosure, many modifications are possible (e.g., variations in sizes, dimensions, structures, shapes, proportions of the various elements, values of parameters, mounting arrangements, use of materials, colors, orientations, etc.). For example, the position of elements may be reversed or otherwise varied, and the nature or number of discrete elements or positions may be altered or varied. Accordingly, all such modifications are intended to be included within the scope of the present disclosure. The order or sequence of any process or method steps may be varied or re-sequenced according to alternative implementations. Other substitutions, modifications, changes, and omissions may be made in the design, operating conditions, and arrangement of the implementations without departing from the scope of the present disclosure.
[0163] The present disclosure contemplates methods, systems, and program products on any machine-readable media for accomplishing various operations. The implementations of the present disclosure may be implemented using existing computer processors, or by a special purpose computer processor for an appropriate system, incorporated for this or another purpose, or by a hardwired system. Implementations within the scope of the present disclosure include program products, including machine-readable media for carrying or having machine-executable instructions or data structures stored thereon. Such machine-readable media can be any available media that can be accessed by a general-purpose or special-purpose computer or other machine with a processor. By way of example, such machine-readable media can comprise RAM, ROM, EPROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to carry or store desired program code in the form of machine-executable instructions or data structures, and which can be accessed by a general purpose or special purpose computer or other machine with a processor.
[0164] When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a machine, the machine properly views the connection as a machine-readable medium;Attorney Docket No.10063-094WO1 OTT202334 thus, any such connection is properly termed a machine-readable medium. Combinations of the above are also included within the scope of machine-readable media. Machine-executable instructions include, for example, instructions and data that cause a general-purpose computer, special-purpose computer, or special-purpose processing machine to perform a certain function or group of functions.
[0165] Although the figures show a specific order of method steps, the order of the steps may differ from what is depicted. Also, two or more steps may be performed concurrently or with partial concurrence. Such variation will depend on the software and hardware systems chosen and on the designer's choice. All such variations are within the scope of the disclosure. Likewise, software implementations could be accomplished with programming techniques with rule-based logic and other logic to accomplish the various connection steps, processing steps, comparison steps, and decision steps.
[0166] Machine Learning. In addition to the machine learning features described above, the analysis system can be implemented using one or more artificial intelligence and machine learning operations. The term “artificial intelligence” can include any technique that enables one or more computing devices or computing systems (i.e., a machine) to mimic human intelligence. Artificial intelligence (AI) includes but is not limited to knowledge bases, machine learning, representation learning, and deep learning. The term “machine learning” is defined herein to be a subset of AI that enables a machine to acquire knowledge by extracting patterns from raw data. Machine learning techniques include, but are not limited to, logistic regression, support vector machines (SVMs), decision trees, Naïve Bayes classifiers, and artificial neural networks. The term “representation learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, or classification from raw data. Representation learning techniques include, but are not limited to, autoencoders and embeddings. The term “deep learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, classification, etc., using layers of processing. Deep learning techniques include but are not limited to artificial neural networks or multilayer perceptron (MLP).
[0167] An artificial neural network (ANN) is a computing system including a plurality of interconnected neurons (e.g., also referred to as “nodes”). This disclosure contemplates that the nodes can be implemented using a computing device (e.g., a processing unit and memory as described herein). The nodes can be arranged in a plurality of layers, such as an input layer, an output layer, and optionally one or more hidden layers with different activation functions. An ANNAttorney Docket No.10063-094WO1 OTT202334 having hidden layers can be referred to as a deep neural network or multilayer perceptron (MLP). Each node is connected to one or more other nodes in the ANN. For example, each layer is made of a plurality of nodes, where each node is connected to all nodes in the previous layer. The nodes in a given layer are not interconnected with one another, i.e., the nodes in a given layer function independently of one another. As used herein, nodes in the input layer receive data from outside of the ANN, nodes in the hidden layer(s) modify the data between the input and output layers, and nodes in the output layer provide the results. Each node is configured to receive an input, implement an activation function (e.g., binary step, linear, sigmoid, tanh, or rectified linear unit (ReLU) function), and provide an output in accordance with the activation function. Additionally, each node is associated with a respective weight. ANNs are trained with a dataset to maximize or minimize an objective function. In some implementations, the objective function is a cost function, which is a measure of the ANN’s performance (e.g., error such as L1 or L2 loss) during training, and the training algorithm tunes the node weights and / or bias to minimize the cost function. This disclosure contemplates that any algorithm that finds the maximum or minimum of the objective function can be used for training the ANN. Training algorithms for ANNs include but are not limited to backpropagation. It should be understood that an artificial neural network is provided only as an example machine learning model. This disclosure contemplates that the machine learning model can be any supervised learning model, semi-supervised learning model, or unsupervised learning model. Optionally, the machine learning model is a deep learning model. Machine learning models are known in the art and are therefore not described in further detail herein.
[0168] A convolutional neural network (CNN) is a type of deep neural network that has been applied, for example, to image analysis applications. Unlike traditional neural networks, each layer in a CNN has a plurality of nodes arranged in three dimensions (width, height, depth). CNNs can include different types of layers, e.g., convolutional, pooling, and fully-connected (also referred to herein as “dense”) layers. A convolutional layer includes a set of filters and performs the bulk of the computations. A pooling layer is optionally inserted between convolutional layers to reduce the computational power and / or control overfitting (e.g., by downsampling). A fully-connected layer includes neurons, where each neuron is connected to all of the neurons in the previous layer. The layers are stacked similarly to traditional neural networks. GCNNs are CNNs that have been adapted to work on structured datasets such as graphs.
[0169] Other Supervised Learning Models. A logistic regression (LR) classifier is a supervised classification model that uses the logistic function to predict the probability of a target, which canAttorney Docket No.10063-094WO1 OTT202334 be used for classification. LR classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize an objective function, for example, a measure of the LR classifier’s performance (e.g., an error such as L1 or L2 loss), during training. This disclosure contemplates that any algorithm that finds the minimum of the cost function can be used. LR classifiers are known in the art and are therefore not described in further detail herein.
[0170] A Naïve Bayes’ (NB) classifier is a supervised classification model that is based on Bayes’ Theorem, which assumes independence among features (i.e., the presence of one feature in a class is unrelated to the presence of any other features). NB classifiers are trained with a data set by computing the conditional probability distribution of each feature given a label and applying Bayes’ Theorem to compute the conditional probability distribution of a label given an observation. NB classifiers are known in the art and are therefore not described in further detail herein.
[0171] A k-NN classifier is an unsupervised classification model that classifies new data points based on similarity measures (e.g., distance functions). The k-NN classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize a measure of the k-NN classifier’s performance during training. This disclosure contemplates any algorithm that finds the maximum or minimum. The k-NN classifiers are known in the art and are therefore not described in further detail herein.
[0172] A majority voting ensemble is a meta-classifier that combines a plurality of machine learning classifiers for classification via majority voting. In other words, the majority voting ensemble’s final prediction (e.g., class label) is the one predicted most frequently by the member classification models. The majority voting ensembles are known in the art and are therefore not described in further detail herein.
[0173] It is to be understood that the methods and systems are not limited to specific synthetic methods, specific components, or to particular compositions. It is also to be understood that the terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting.
[0174] As used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another implementation includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms anotherAttorney Docket No.10063-094WO1 OTT202334 implementation. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint.
[0175] “Optional” or “optionally” means that the subsequently described event or circumstance may or may not occur and that the description includes instances where said event or circumstance occurs and instances where it does not.
[0176] Throughout the description and claims of this specification, the word “comprise” and variations of the word, such as “comprising” and “comprises,” means “including but not limited to,” and is not intended to exclude, for example, other additives, components, integers or steps. “Exemplary” means “an example of” and is not intended to convey an indication of a preferred or ideal implementation. “Such as” is not used in a restrictive sense but for explanatory purposes.
[0177] Disclosed are components that can be used to perform the disclosed methods and systems. These and other components are disclosed herein, and it is understood that when combinations, subsets, interactions, groups, etc. of these components are disclosed while specific reference of each various individual and collective combinations and permutation of these may not be explicitly disclosed, each is specifically contemplated and described herein, for all methods and systems. This applies to all aspects of this application, including, but not limited to, steps in disclosed methods. Thus, if there are a variety of additional steps that can be performed it is understood that each of these additional steps can be performed with any specific implementation or combination of implementations of the disclosed methods.
[0178] The following patents, applications, and publications, as listed below and throughout this document, are hereby incorporated by reference in their entirety herein. Reference List #1 [1] Beaman C, Kaneko N, Meyers P, Tateshima S. A Review of Robotic Interventional Neuroradiology. American Journal of Neuroradiology.2021;42(5):808-814. [2] Karstensen L, Behr T, Pusch TP, Mathis-Ullrich F, Stallkamp J. Autonomous guidewire navigation in a two dimensional vascular phantom. Current Directions in Biomedical Engineering.2020;6(1). [3] Crinnion W, Jackson B, Sood A, et al. Robotics in neurointerventional surgery: a systematic review of the literature. Journal of NeuroInterventional Surgery.2022;14(6):539-545. [4] Costa M, Tataryn Z, Alobaid A, et al. Robotically-assisted neuro-endovascular procedures: Single-Center Experience and a Review of the Literature. Interv Neuroradiol. 2022:15910199221082475.Attorney Docket No.10063-094WO1 OTT202334 [5] Rabinovich EP, Capek S, Kumar JS, Park MS. Tele-robotics and artificial-intelligence in stroke care. J Clin Neurosci.2020;79:129-132. [6] Panesar SS, Volpi JJ, Lumsden A, et al. Telerobotic stroke intervention: a novel solution to the care dissemination dilemma. Journal of Neurosurgery.2019;132(3):971-978. [7] Kweon J, Kim K, Lee C, et al. Deep reinforcement learning for guidewire navigation in coronary artery phantom. IEEE Access.2021;9:166409-166422. [8] Zhang G, Wong H-C, Wang C, Zhu J, Lu L, Teng G. A Temporary Transformer Network for Guide-Wire Segmentation. Paper presented at: 202114th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP- BMEI)2021. [9] Chen K, Qin W, Xie Y, Zhou S. Towards real time guide wire shape extraction in fluoroscopic sequences: A two phase deep learning scheme to extract sparse curvilinear structures. Computerized Medical Imaging and Graphics.2021:101989.
[0010] Ramadani A, Bui M, Wendler T, Schunkert H, Ewert P, Navab N. A survey of catheter tracking concepts and methodologies. Medical Image Analysis.2022:102584.
[0011] Zhang G, Wong H-C, Zhu J, An T, Wang C. Jigsaw training-based background reverse attention transformer network for guidewire segmentation. International Journal of Computer Assisted Radiology and Surgery.2022:1-9.
[0012] Ambrosini P, Ruijters D, Niessen WJ, Moelker A, van Walsum T. Fully Automatic and Real-Time Catheter Segmentation in X-Ray Fluoroscopy.2017; Cham.
[0013] Zhou Y-J, Xie X-L, Zhou X-H, Liu S-Q, Bian G-B, Hou Z-G. Pyramid attention recurrent networks for real-time guidewire segmentation and tracking in intraoperative X-ray fluoroscopy. Computerized Medical Imaging and Graphics.2020;83:101734.
[0014] Park N, Kim S. How Do Vision Transformers Work? arXiv preprint arXiv:220206709. 2022.
[0015] Pang S, Du A, Orgun MA, et al. Beyond CNNs: Exploiting further inherent symmetries in medical image segmentation. IEEE transactions on cybernetics.2022.
[0016] Shit S, Paetzold JC, Sekuboyina A, et al. clDice-a novel topology-preserving loss function for tubular structure segmentation. Paper presented at: Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition2021.
[0017] Desai VR, Lee JJ, Sample T, Kleiman NS, Lumsden A, Britz GW. First in Man Pilot Feasibility Study in Extracranial Carotid Robotic-Assisted Endovascular Intervention. Neurosurgery.2021;88(3):506-514.Attorney Docket No.10063-094WO1 OTT202334
[0018] Nogueira RG, Sachdeva R, Al-Bayati AR, Mohammaden MH, Frankel MR, Haussen DC. Robotic assisted carotid artery stenting for the treatment of symptomatic carotid disease: technical feasibility and preliminary results. Journal of NeuroInterventional Surgery. 2020;12(4):341-344.
[0019] Qureshi AI, Agunbiade S, Huang W, et al. Changes in Neuroendovascular Procedural Volume During the COVID‐19 Pandemic: An International Multicenter Study. Journal of Neuroimaging.2021;31(1):171-179.
[0020] Veeling BS, Linmans J, Winkens J, Cohen T, Welling M. Rotation equivariant CNNs for digital pathology. Paper presented at: International Conference on Medical image computing and computer-assisted intervention2018.
[0021] Wong KK, Cummock JS, Li G, et al. Automatic segmentation in acute ischemic stroke: Prognostic significance of topological stroke volumes on stroke outcome. Stroke. 2022;53(9):2896-2905.
[0022] Ronneberger O, Fischer P, Brox T. U-net: Convolutional networks for biomedical image segmentation.2015.
[0023] Chen J, Lu Y, Yu Q, et al. Transunet: Transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:210204306.2021.
[0024] Isensee F, Jaeger PF, Kohl SA, Petersen J, Maier-Hein KH. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods. 2021;18(2):203-211.
[0025] Ronneberger O, Fischer P, Brox T. U-net: Convolutional networks for biomedical image segmentation. Paper presented at: Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 182015.
[0026] Beaman C, Saber H, Tateshima S. A technical guide to robotic catheter angiography with the Corindus CorPath GRX system. Journal of NeuroInterventional Surgery. 2022;14(12):1284-1284.
[0027] Britz GW, Panesar SS, Falb P, Tomas J, Desai V, Lumsden A. Neuroendovascular-specific engineering modifications to the CorPath GRX Robotic System. J Neurosurg.2019:1-7.
[0028] Desai VR, Lee JJ, Tomas J, Lumsden A, Britz GW. Initial Experience in a Pig Model of Robotic-Assisted Intracranial Arteriovenous Malformation (AVM) Embolization. Oper Neurosurg (Hagerstown).2020;19(2):205-209.Attorney Docket No.10063-094WO1 OTT202334
[0029] Liu Z, Mao H, Wu C-Y, Feichtenhofer C, Darrell T, Xie S. A ConvNet for the 2020s. arXiv preprint arXiv:220103545.2022.
[0030] Kim BM, Baek JH, Heo JH, Kim DJ, Nam HS, Kim YD. Effect of Cumulative Case Volume on Procedural and Clinical Outcomes in Endovascular Thrombectomy. Stroke. 2019;50(5):1178-1183. Reference List #2 [1'] Snelling BM, Sur S, Shah SS, et al. Transradial cerebral angiography: techniques and outcomes. Journal of neurointerventional surgery.2018. [2'] Beaman C, Saber H, Tateshima S. A technical guide to robotic catheter angiography with the Corindus CorPath GRX system. Journal of NeuroInterventional Surgery. 2022;14(12):1284-1284. [3'] Behr T, Pusch TP, Siegfarth M, Hüsener D, Mörschel T, Karstensen L. Deep reinforcement learning for the navigation of neurovascular catheters. Current Directions in Biomedical Engineering.2019;5(1):5-8. [4'] Ritter J, Karstensen L, Langejürgen J, Hatzl J, Mathis-Ullrich F, Uhl C. Quality-dependent Deep Learning for Safe Autonomous Guidewire Navigation. Current Directions in Biomedical Engineering.2022;8(1):21-24. [5'] Meng F, Guo S, Zhou W, Chen Z. Evaluation of a Reinforcement Learning Algorithm for Vascular Intervention Surgery. Paper presented at: 2021 IEEE International Conference on Mechatronics and Automation (ICMA)2021. [6'] Kweon J, Kim K, Lee C, et al. Deep reinforcement learning for guidewire navigation in coronary artery phantom. IEEE Access.2021;9:166409-166422. [7'] Negoro M, Tanimoto M, Arai F, et al. An intelligent catheter system robotic controlled catheter system. Interv Neuroradiol.2001;7(Suppl 1):111-113. [8'] Ramadani A, Bui M, Wendler T, Schunkert H, Ewert P, Navab N. A survey of catheter tracking concepts and methodologies. Medical Image Analysis.2022:102584. [9'] Chi W, Liu J, Rafii-Tari H, Riga C, Bicknell C, Yang GZ. Learning-based endovascular navigation through the use of non-rigid registration for collaborative robotic catheterization. Int J Comput Assist Radiol Surg.2018;13(6):855-864. [10'] Kim Y, Genevriere E, Harker P, et al. Telerobotic neurovascular interventions with magnetic manipulation. Science Robotics.2022;7(65):eabg9907.Attorney Docket No.10063-094WO1 OTT202334 [11'] Chen K, Wang C, Xie Y, Zhou S. A GPU-Based Automatic Approach for Guide Wire Tracking in Fluoroscopic Sequences. International Journal of Pattern Recognition and Artificial Intelligence.2019;33(08):1954025. [12'] Chen K, Qin W, Xie Y, Zhou S. Towards real time guide wire shape extraction in fluoroscopic sequences: A two phase deep learning scheme to extract sparse curvilinear structures. Computerized Medical Imaging and Graphics.2021:101989. [13'] Liu Z, Mao H, Wu C-Y, Feichtenhofer C, Darrell T, Xie S. A ConvNet for the 2020s. arXiv preprint arXiv:220103545.2022. [14'] Park N, Kim S. How Do Vision Transformers Work? arXiv preprint arXiv:220206709. 2022. [15'] Ghosh R, Wong K, Zhang YJ, Britz GW, Wong STC. Automated catheter segmentation and tip detection in cerebral angiography with topology-aware geometric deep learning. Journal of NeuroInterventional Surgery.2023:jnis-2023-020300. [16'] Ambrosini P, Ruijters D, Niessen WJ, Moelker A, van Walsum T. Fully Automatic and Real-Time Catheter Segmentation in X-Ray Fluoroscopy.2017; Cham. [17'] Zhou Y-J, Xie X-L, Zhou X-H, Liu S-Q, Bian G-B, Hou Z-G. Pyramid attention recurrent networks for real-time guidewire segmentation and tracking in intraoperative X-ray fluoroscopy. Computerized Medical Imaging and Graphics.2020;83:101734. [18'] Zhou Y-J, Liu S-Q, Xie X-L, et al. A Real-Time Multi-Task Framework for Guidewire Segmentation and Endpoint Localization in Endovascular Interventions. Paper presented at: 2021 IEEE International Conference on Robotics and Automation (ICRA)2021. [19'] Zhao C, Bober R, Tang H, et al. Semantic Segmentation to Extract Coronary Arteries in Invasive Coronary Angiograms. Journal of Advances in Applied & Computational Mathematics.2022;9:76-85. [20'] Yang S, Kweon J, Roh J-H, et al. Deep learning segmentation of major vessels in X-ray coronary angiography. Scientific reports.2019;9(1):1-11. [21'] Nasr-Esfahani E, Karimi N, Jafari MH, et al. Segmentation of vessels in angiograms using convolutional neural networks. Biomedical Signal Processing and Control.2018;40:240- 251. [22'] Nasr-Esfahani E, Samavi S, Karimi N, et al. Vessel extraction in X-ray angiograms using deep learning. Paper presented at: 201638th Annual international conference of the IEEE engineering in medicine and biology society (EMBC)2016.Attorney Docket No.10063-094WO1 OTT202334 [23'] Chen J, Lu Y, Yu Q, et al. Transunet: Transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:210204306.2021. [24'] Seah J, Boeken T, Sapoval M, Goh GS. Prime time for artificial intelligence in interventional radiology. CardioVascular and Interventional Radiology.2022;45(3):283- 289. [25'] Nolf E, Voet T, Jacobs F, Dierckx R, Lemahieu I. An open-source medical image conversion toolkit. Eur J Nucl Med.2003;30(Suppl 2):S246. [26'] Kikinis R, Pieper SD, Vosburgh KG.3D Slicer: a platform for subject-specific image analysis, visualization, and clinical support. In: Intraoperative imaging and image-guided therapy. Springer; 2013:277-289. [27'] Rorden C, Brett M. Stereotaxic display of brain lesions. Behav Neurol.2000;12(4):191- 200. [28'] Isensee F, Jaeger PF, Kohl SA, Petersen J, Maier-Hein KH. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods. 2021;18(2):203-211. [29'] Shit S, Paetzold JC, Sekuboyina A, et al. clDice-a novel topology-preserving loss function for tubular structure segmentation. Paper presented at: Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition2021. [30'] Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: Machine learning in Python. the Journal of machine Learning research.2011;12:2825-2830. [31'] Hunter JD. Matplotlib: A 2D graphics environment. Computing in science & engineering. 2007;9(03):90-95. [32'] Pajankar A.3D Visualizations in Matplotlib. In: Hands-on Matplotlib: Learn Plotting and Visualizations with Python 3. Springer; 2021:143-159. [33'] Nadipally M. Chapter 2 - Optimization of Methods for Image-Texture Segmentation Using Ant Colony Optimization. In: Hemanth DJ, Gupta D, Emilia Balas V, eds. Intelligent Data Analysis for Biomedical Applications. Academic Press; 2019:21-47. [34'] Bull DR, Zhang F. Chapter 4 - Digital picture formats and representations. In: Bull DR, Zhang F, eds. Intelligent Image and Video Compression (Second Edition). Oxford: Academic Press; 2021:107-142. [35'] Zhang Y, Weng Y, Lund J. Applications of explainable artificial intelligence in diagnosis and surgery. Diagnostics.2022;12(2):237.Attorney Docket No.10063-094WO1 OTT202334 [36'] Iyer K, Najarian CP, Fattah AA, et al. Angionet: a convolutional neural network for vessel segmentation in X-ray angiography. Scientific Reports.2021;11(1):18066. [37'] Anić M, Đukić T. Improved Three-Dimensional Reconstruction of Patient-Specific Carotid Bifurcation Using Deep Learning Based Segmentation of Ultrasound Images. Paper presented at: Serbian International Conference on Applied Artificial Intelligence2022. [38'] Zhou T, Tan T, Pan X, Tang H, Li J. Fully automatic deep learning trained on limited data for carotid artery segmentation from large image volumes. Quantitative Imaging in Medicine and Surgery.2021;11(1):67. [39'] Ziegler M, Alfraeus J, Bustamante M, et al. Automated segmentation of the individual branches of the carotid arteries in contrast-enhanced MR angiography using DeepMedic. BMC medical imaging.2021;21:1-10. [40'] Fu K, Liu Y, Wang M. Global Registration of 3D Cerebral Vessels to Its 2D Projections by a New Branch-and-Bound Algorithm. IEEE Transactions on Medical Robotics and Bionics.2021;3(1):115-124. [41'] Liao H, Lin W-A, Zhang J, Zhang J, Luo J, Zhou SK. Multiview 2D / 3D rigid registration via a point-of-interest network for tracking and triangulation. Paper presented at: Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition2019. [42'] Cobiella R, Quinones S, Konschake M, et al. The carotid axis revisited. Sci Rep. 2021;11(1):13847.
Claims
Attorney Docket No.10063-094WO1 OTT202334 What is claimed is:
1. A system comprising: a processor; and a memory having instructions stored thereon, wherein execution of the instructions by the processor, causes the processor to: receive real-time fluoroscopy data having vasculatures; receive real-time positional data for an intravascular instrument having a tip, including positional data for the tip or an instrument landmark of the instrument; determine, via a trained AI model configured to receive in its input the received fluoroscopy data, a plurality of masks each having the fluoroscopy image with an estimated portions of the vasculatures, including a first 2D mask having an estimated first blood vessel in the fluoroscopy data, a second 2D mask having an estimated second blood vessel in the fluoroscopy data, and a third 2D mask having an estimated third vessel in the fluoroscopy data, wherein the trained AI model was trained to identify the first blood vessel, the second blood vessel, and the third blood vessel; determine a plurality of positional data for the estimated first blood vessel, the estimated second blood vessel, and the estimated third blood vessel, including a first 2D positional data, a second 2D positional data, and a third 2D positional data; and combine the plurality of 2D positional data to generate a combined positional data having a plurality of 2D position data each associated with a blood vessel; determine the positional data for the tip or instrument landmark in relation to the plurality of 2D position data; and direct output of the combined positional data and the positional data for the tip or instrument landmark for subsequent use, via a graphical user interface, in tracking, monitoring, or guiding the tip or instrument landmark through the vasculatures.
2. The system of claim 1, wherein the execution of the instructions by the processor further causes the processor to: modify the plurality of positional data to smooth and / or remove discontinuity artifacts among the first 2D positional data, the second 2D positional data, and the third 2D positional data.
3. The system of claim 1 or 2, wherein the execution of the instructions by the processor further causes the processor to:Attorney Docket No.10063-094WO1 OTT202334 identify (i) a first origin point for the second 2D positional data in the first 2D positional data and (ii) a second origin point for the third 2D positional data in the first 2D positional data, wherein the first origin point and second origin point are employed in the instructions to determine the positional data for the tip or instrument landmark in relation to the plurality of 2D position data.
4. The system of claim 1 or 2, wherein the execution of the instructions by the processor further causes the processor to: identify (i) a bifurcation point in the first 2D positional data, wherein the bifurcation point is employed in the instructions to determine the positional data for the tip or instrument landmark in relation to the plurality of 2D position data.
5. The system of any one of claims 1-4, wherein the execution of the instructions by the processor further causes the processor to: generate visualization output for the graphical user interface, the visualization output comprising a first visualization object of the tip or instrument landmark and a second visualization object of the instrument, wherein the first visualization object and the second visualization object are combined with the real-time fluoroscopy data.
6. The system of any one of claims 1-5, wherein the execution of the instructions by the processor further causes the processor to: generate visualization output for the graphical user interface, the visualization output comprising a third visualization object for an indication of the tip or instrument landmark in the first blood vessel, the second blood vessel, or the third blood vessel, wherein the third visualization object is combined with the real-time fluoroscopy data.
7. The system of any one of claims 1-6, wherein the trained AI model was trained using digital subtraction angiography (DSA) images.
8. The system of any one of claims 1- 7, wherein the system is employed for robotic neurointervention or robotic endovascular intervention.Attorney Docket No.10063-094WO1 OTT202334 9. The system of any one of claims 1-8, wherein the system is employed for cerebrovascular intervention, endovascular intervention, diagnostic cerebral angiography (DCA), cerebrovascular evaluation, mechanical thrombectomy, or embolization.
10. The system of any one of claims 1-9, wherein the real-time positional data (measured or estimated) for the intravascular instrument is determined via a second trained AI model that generates the real-time positional data from the real-time fluoroscopy data.
11. A method for operating the system of claims 1-10, the method comprising: receiving real-time fluoroscopy data having vasculatures; receiving real-time positional data for an intravascular instrument having a tip, including positional data for the tip or an instrument landmark of the instrument; determining, via a trained AI model configured to receive in its input the received fluoroscopy data, a plurality of masks each having the fluoroscopy image with an estimated portions of the vasculatures, including a first 2D mask having an estimated first blood vessel in the fluoroscopy data, a second 2D mask having an estimated second blood vessel in the fluoroscopy data, and a third 2D mask having an estimated third vessel in the fluoroscopy data, wherein the trained AI model was trained to identify the first blood vessel, the second blood vessel, and the third blood vessel; determining a plurality of positional data for the estimated first blood vessel, the estimated second blood vessel, and the estimated third blood vessel, including a first 2D positional data, a second 2D positional data, and a third 2D positional data; and combining the plurality of 2D positional data to generate a combined positional data having a plurality of 2D position data each associated with a blood vessel; determining the positional data for the tip or instrument landmark in relation to the plurality of 2D position data; and directing output of the combined positional data and the positional data for the tip or instrument landmark for subsequent use, via a graphical user interface, in tracking, monitoring, or guiding the tip or instrument landmark through the vasculatures.
12. The method of claim 11 further comprising:Attorney Docket No.10063-094WO1 OTT202334 modifying the plurality of positional data to smooth and / or remove discontinuity artifacts among the first 2D positional data, the second 2D positional data, and the third 2D positional data.
13. The method of claim 11 or 12 further comprising: identifying (i) a first origin point for the second 2D positional data in the first 2D positional data and (ii) a second origin point for the third 2D positional data in the first 2D positional data, wherein the first origin point and second origin point are employed in determining the positional data for the tip or instrument landmark in relation to the plurality of 2D position data.
14. The method of claim 11 or 12, further comprising identifying (i) a bifurcation point in the first 2D positional data, wherein the bifurcation point is employed in determining the positional data for the tip or instrument landmark in relation to the plurality of 2D position data.
15. The method of any one of claims 11-14, further comprising generating visualization output for the graphical user interface, the visualization output comprising a first visualization object of the tip or instrument landmark and a second visualization object of the instrument, wherein the first visualization object and the second visualization object are combined with the real-time fluoroscopy data.
16. The method of any one of claims 11-15 further comprising: generating visualization output for the graphical user interface, the visualization output comprising a third visualization object for an indication of the tip or instrument landmark in the first blood vessel, the second blood vessel, or the third blood vessel, wherein the third visualization object is combined with the real-time fluoroscopy data.
17. The method of any one of claims 11-16, wherein the trained AI model was trained using digital subtraction angiography (DSA) images.
18. The method of any one of claims 11-17, wherein the method is employed for robotic neurointervention or robotic endovascular intervention.Attorney Docket No.10063-094WO1 OTT202334 19. The method of any one of claims 11-18, wherein the method is employed for cerebrovascular intervention, endovascular intervention, diagnostic cerebral angiography (DCA), cerebrovascular evaluation, mechanical thrombectomy, or embolization.
20. The method of any one of claims 11-19, wherein the real-time positional data (measured or estimated) for the intravascular instrument is determined via a second trained AI model that generates the real-time positional data from the real-time fluoroscopy data.
21. The method of any one of claims 11-20, wherein the trained AI model was trained employing manual segmentation of internal carotid artery (ICA), external carotid artery (ECA), and common carotid artery on DSA images as three separate masks.
22. A non-transitory computer-readable medium having instructions stored thereon, wherein execution of the instructions by a processor causes the processor to perform the method of any one of claims 11-21 or operate the system of any one of claims 1-10.
23. The non-transitory computer-readable medium of claim 22, wherein the execution of the instructions by the processor further causes the processor to: modify the plurality of positional data to smooth and / or remove discontinuity artifacts among the first 2D positional data, the second 2D positional data, and the third 2D positional data.
24. The non-transitory computer-readable medium of claim 22 or 23, wherein the execution of the instructions by the processor further causes the processor to: identify (i) a first origin point for the second 2D positional data in the first 2D positional data and (ii) a second origin point for the third 2D positional data in the first 2D positional data, wherein the first origin point and second origin point are employed in the instructions to determine the positional data for the tip or instrument landmark in relation to the plurality of 2D position data.
25. The non-transitory computer-readable medium of claim 22 or 23, wherein the execution of the instructions by the processor further causes the processor to:Attorney Docket No.10063-094WO1 OTT202334 identify (i) a bifurcation point in the first 2D positional data, wherein the bifurcation point is employed in the instructions to determine the positional data for the tip or instrument landmark in relation to the plurality of 2D position data.
26. The non-transitory computer-readable medium of any one of claims 22-25, wherein the execution of the instructions by the processor further causes the processor to: generate visualization output for the graphical user interface, the visualization output comprising a first visualization object of the tip or instrument landmark and a second visualization object of the instrument, wherein the first visualization object and the second visualization object are combined with the real-time fluoroscopy data.
27. The non-transitory computer-readable medium of any one of claims 22-26, wherein the execution of the instructions by the processor further causes the processor to: generate visualization output for the graphical user interface, the visualization output comprising a third visualization object for an indication of the tip or instrument landmark in the first blood vessel, the second blood vessel, or the third blood vessel, wherein the third visualization object is combined with the real-time fluoroscopy data.
28. The system of any one of claims 1-9, wherein the trained AI model was trained jointly from both catheter and vessel data.
29. The method of any one of claims 11-20, wherein the trained AI model was trained jointly from both catheter and vessel data.
Citation Information
Patent Citations
Methods and apparatuses for image guided medical procedures
US20070167801A1
System and Methods for Guiding a Medical Instrument
US20200237255A1
Anatomical feature tracking
US20220296312A1
Endoscope navigation system with updating anatomy model
US20220354380A1
Systems and methods for medical image fusion
US20240090859A1