Surgical instrument kinematics processing, navigation, and feedback
By identifying and evaluating the field of vision of surgical instruments, the consistency and accuracy of instrument motion monitoring in surgical operations are solved, and the accuracy and efficiency of surgical operations are improved.
Patent Information
- Application Number
- CN202380070703.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-11
- Filing Date
- 2023-10-09
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to monitor device movement consistently and accurately in surgical procedures, making it difficult for surgeons to provide meaningful feedback and evaluation, and data between different surgical rooms are difficult to compare, affecting the accuracy and efficiency of surgical procedures.
Through computer systems and methods, appropriate surgical instrument vision fields are identified and instrument kinematic behavior is evaluated based on these fields of view, providing consistent kinematic results to help surgical teams navigate and oriente patients within the body.
Consistency and accuracy monitoring of instrument movements in surgical procedures are achieved, providing meaningful feedback and evaluation, and improving the accuracy and efficiency of surgical procedures.
Smart Images

Figure CN120283262A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 415,220, titled "SURGICAL INSTRUMENT REFERENCE KINEMATICS DETERMINATION AND DISPLAY", filed on October 11, 2022; U.S. Provisional Application No. 63 / 415,225, titled "INTRAOPERATIVE SENSOR VISUAL FIELD CHARACTERIZATION AND PROCESSING", filed on October 11, 2022; and U.S. Provisional Application No. 63 / 415,231, titled "COMPUTER STRUCTURES AND INTERFACES FOR INTRAOPERATIVE SURGICAL NAVIGATION", filed on October 11, 2022, each of which is incorporated herein by reference in its entirety for all purposes. Technical Field
[0003] Various disclosed embodiments relate to systems and methods for evaluating the kinematic behavior of surgical instruments (e.g., for navigation, analysis, feedback) or for improved modeling of internal anatomical structures. Background Art
[0004] Although machine learning, network connectivity, surgical robots, and various other new technologies have great potential to improve healthcare efficacy and efficiency, during surgical procedures, many of these technologies and their applications cannot reach their full potential without accurately and consistently monitoring instrument movement. Mechanical encoders and similar technologies can facilitate monitoring of surgical instrument position, but they are often inconsistent between surgical operating rooms and may not provide data that is easily comparable to systems with different or no such monitoring technology. Additionally, mechanical solutions may require specific hardware and maintenance, imposing additional costs and the reluctance to keep the system fully calibrated. The inability to effectively and economically acquire instrument kinematic data across different surgical operating room configurations also makes it difficult to provide meaningful and consistent feedback to surgeons, analyze surgeon performance, and provide comparisons between the performance of surgeons and that of other practitioners. Variations in the acquisition methods across surgical operating rooms can bias the evaluation of different surgeons and their surgical procedures, leading to inaccurate and potentially harmful conclusions.
[0005] However, when the data from the in - body sensors is inappropriate (e.g., in the presence of moisture on the sensor, when sensor movement obscures the field of view, when occlusion prevents proper data collection, etc.), even such systems can be adversely affected. The downstream processing that does not anticipate such corruption may itself be damaged by these inputs, resulting in "garbage out" due to "garbage in". Even more unfortunately, in many downstream processing systems (especially those that iteratively consider previously generated results in their inputs), such incorrect outputs may produce a cascade of incorrect results and even subsequently cause valid inputs to be considered incorrectly because the process attempts to reconcile new valid inputs with previously considered invalid inputs. While simply excluding infeasible images is valuable, consistently identifying when, where, and how often such inappropriate data instances are encountered can itself provide information about the surgeon's performance and the status of the surgical procedure.
[0006] The availability of such high - quality kinematic results can then enable various downstream applications. For example, even experienced surgeons may find it difficult to navigate and orient themselves within the patient. The stress and distractions in the surgical operating room, the wide variations in anatomy among patient populations, and the variations in surgical operating room instruments and configurations can all easily cause confusion and disorientation. For example, during a colonoscopy, the surgeon may have difficulty knowing how much of the intestine remains to be examined, how much of the intestine has already been examined, the thoroughness of the examination of various regions of the intestine, how the examination compares to past or related examinations, etc. As staff shortages and an aging population continue to force hospitals to do more with fewer resources, there is also a need for surgical team members with different levels of experience to perform surgical procedures quickly, effectively, and consistently.
[0007] Accordingly, there is a need for computer systems and methods that can first identify the surgical instrument fields of view suitable for downstream instrument kinematic processing, as well as systems and methods for evaluating the kinematic behavior of surgical instruments based on those suitable fields of view, and finally, there is a need for systems and methods to assist the surgical team in consistently orienting and navigating within the patient using such kinematic results. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The various embodiments described herein can be better understood by reference to the following detailed description in conjunction with the accompanying drawings, in which like reference numerals represent the same or functionally similar elements:
[0009] Figure 1A is a schematic diagram of various elements that may occur in a surgical operating room during a surgical procedure for some embodiments;
[0010] Figure 1Bis a schematic diagram of various components that may be present in a surgical operating room during a surgical procedure using a surgical robot;
[0011] Figure 2A is a schematic diagram of an organ, in this example the large intestine, the cross-sectional view of which reveals the progress of a colonoscope during a surgical examination that may occur in some embodiments;
[0012] Figure 2B is a schematic diagram of the distal tip of a colonoscope that may be used in some embodiments;
[0013] Figure 2C is a schematic diagram of a portion of the colon, the cross-sectional view of which reveals the position of the colonoscope relative to multiple pockets;
[0014] Figure 2D is from Figure 2C a schematic representation of a visual image and a corresponding depth frame acquired by a camera from the perspective of the camera of the colonoscope depicted therein;
[0015] Figure 2E is a pair of images acquired from a colonoscope camera with a straight view and a colonoscope camera with a fish-eye view, which depict a grid-like pattern of orthogonal rows and columns in perspective, each of which images may be used in some embodiments;
[0016] Figure 3A is a schematic diagram of a computer-generated three-dimensional model of the large intestine in a first perspective in some embodiments, with a portion of the model highlighted;
[0017] Figure 3B is Figure 3A a schematic diagram of the computer-generated three-dimensional model in a second perspective;
[0018] Figure 3C is Figure 3A a schematic diagram of the computer-generated three-dimensional model in a third perspective;
[0019] Figures 4A - 4C is a schematic two-dimensional cross-sectional representation of the continuous progression of a colonoscope through the large intestine over time that may occur in some embodiments;
[0020] Figures 4D - 4F is in some embodiments a two-dimensional schematic representation of a depth frame generated from the corresponding field of view depicted in Figures 4A - 4C ;
[0021] Figure 4G is in some embodiments Figures 4D - 4F a schematic two-dimensional representation of a fusion operation (for creating a merged representation) between depth frames of;
[0022] Figure 5 is a flow chart of various operations in an example process for generating at least a portion of a computer model of an internal body structure such as an organ that can be implemented in some embodiments;
[0023] Figure 6 is a flow chart of a preprocessing variant of a process that can be implemented in some embodiments Figure 5 of;
[0024] Figure 7 is an example processing pipeline for generating at least a portion of a three-dimensional model of the large intestine from colonoscopy data acquisition that can be implemented in some embodiments;
[0025] Figure 8 is an example processing pipeline depicting a preprocessing variant of a pipeline that can be implemented in some embodiments Figure 7 of;
[0026] Figure 9A is an example processing pipeline for determining a depth map and a rough local pose from colonoscopy images using two different neural networks that can be implemented in some embodiments;
[0027] Figure 9B is an example processing pipeline for determining a depth map and a rough local pose from colonoscopy images using a single neural network that can be implemented in some embodiments;
[0028] Figure 10A is a flow chart of various operations in a neural network training process that can be performed on a network in some embodiments Figure 9A and Figure 9B ;
[0029] Figure 10B is a bar plot depicting an exemplary set of training results of a process that can occur in conjunction with some embodiments Figure 10A of;
[0030] Figure 11A is a flow chart of various operations in a new fragment determination process that can be implemented in some embodiments;
[0031] Figure 11B is a schematic side view representation of consecutive fields of view of an endoscope related to frustum overlap determination that can occur in some embodiments;
[0032] Figure 11C is a schematic time series of cross-sectional views depicting a collision of a colonoscope with the side wall of the colon and the resulting change in the field of view of the colonoscope camera that can occur in conjunction with some embodiments;
[0033] Figure 11D is a schematic representation of a set of segments corresponding to Figure 11C collisions that can be generated in some embodiments;
[0034] Figure 11E is a schematic network diagram illustrating various key - frame relationships after a graphical network pose optimization operation that can occur in some embodiments;
[0035] Figure 11F is a schematic diagram of a segment having an associated truncated signed - distance function (TSDF) grid relative to a full - model TSDF grid that can be generated in some embodiments;
[0036] Figure 12A is a schematic cross - sectional view of a colonoscope and the resulting colonoscope camera field of view within a portion of the colon during advancing and withdrawing movements that can occur in some embodiments;
[0037] Figure 12B is a schematic three - dimensional model of a portion of the colon having a middle center - line axis reference geometry that can be used in some embodiments;
[0038] Figure 12C is a schematic side perspective view of a colonoscope camera in an advancing orientation relative to the center - line reference geometry that can occur in some embodiments;
[0039] Figure 12D is a schematic side perspective view of a colonoscope camera in a withdrawing orientation relative to the center - line reference geometry that can occur in some embodiments;
[0040] Figure 12E is a schematic side perspective view of a colonoscope camera in an off - center - line orientation relative to the center - line reference geometry that can occur in some embodiments;
[0041] Figure 12F is a schematic side perspective view of a colonoscope camera in a rotating orientation relative to the center - line reference geometry that can occur in some embodiments;
[0042] Figure 12G is a schematic side perspective view of a colonoscope camera in a rotating off - center - line orientation relative to a curved center - line reference geometry;
[0043] Figure 12H is a schematic side perspective view of a series of points along the center - line reference geometry;
[0044] Figure 12I is a schematic perspective view of an orientation on a center - line reference geometry having a radial spatial context that can occur in some embodiments;
[0045] Figure 12J Is a schematic cross-sectional view of a patient's pelvic region during a robotic surgical procedure that can occur in some embodiments;
[0046] Figure 12K Is a schematic depiction of a patient's internal cavity having a cylindrical reference geometry;
[0047] Figure 12L Is a schematic depiction of a patient's internal cavity having a spherical reference geometry;
[0048] Figure 13A Is a schematic three-dimensional model of the colon in which advancement and withdrawal paths are depicted;
[0049] Figure 13B Is a flowchart illustrating various operations in an example intermediate centerline estimation process that can be implemented in some embodiments;
[0050] Figure 13C Is a schematic three-dimensional model of the colon having an associated pre-existing global centerline and local centerlines for segments that can occur in some embodiments;
[0051] Figure 13D Is a flowchart illustrating various operations in an example process for estimating local centerline segments that can be implemented in some embodiments;
[0052] Figure 13E Is a flowchart illustrating various operations in an example process for extending a global centerline with local centerlines of segments that can be implemented in some embodiments;
[0053] Figure 14 Is a schematic operation pipeline depicting the individual steps in an example process for updating a global mid-axis centerline with local segment centerlines that can be implemented in some embodiments;
[0054] Figure 15A Is a flowchart illustrating various operations in an example process for updating instrument kinematic data relative to a reference geometry during a surgical procedure that can be implemented in some embodiments;
[0055] Figure 15B Is a flowchart illustrating various operations in an example process for evaluating kinematic data that can be implemented in some embodiments;
[0056] Figure 15C Is a schematic representation of a model having spatial and temporal context regions that can be used in some embodiments;
[0057] Figure 15D Is a collection of graphical user interface (GUI) elements that can be implemented in some embodiments;
[0058] Figure 16A A schematic diagram of a three-dimensional colon model with a path pattern that can be presented in some embodiments;
[0059] Figure 16B A schematic diagram of a three-dimensional colon model with a middle centerline pattern and a corresponding velocity plot that can be presented in some embodiments;
[0060] Figure 16C A schematic diagram of a path pattern in a three-dimensional cavity model that can be presented in some embodiments;
[0061] Figure 16D A schematic diagram of a three-dimensional cavity model with a spherical reference geometry that can be presented in some embodiments;
[0062] Figure 16E A schematic image from a centerline reference geometry kinematic icon animation that can be presented in some embodiments;
[0063] Figure 16F A schematic image from a spherical reference geometry kinematic icon animation that can be presented in some embodiments;
[0064] Figure 17A A set of schematic GUI elements that can be implemented in some embodiments;
[0065] Figure 17B A three-dimensional colon model GUI element and a velocity plot GUI element that can be presented to a user in some embodiments;
[0066] Figure 17C An enlarged view of a reference geometry kinematic pattern depicted in the GUI of 17A;
[0067] Figure 17D A set of schematic GUI elements that can be implemented and appear in a robotic surgical procedure interface in some embodiments;
[0068] Figure 17E A set of schematic GUI elements that can be implemented and appear in a robotic surgical procedure interface in some embodiments;
[0069] Figure 17F is in Figure 17E An enlarged view of a guidance pattern depicted in the GUI element of
[0070] Figure 18A A schematic set of various surgical camera states and corresponding fields of view that can occur in some embodiments;
[0071] Figure 18BIs a schematic cross-section of a patient's pelvic region during a laparoscopic procedure that can occur in some embodiments;
[0072] Figure 18C Is a schematic perspective view of a surgical tool that obscures a portion of the field of view of a surgical camera that can occur in some embodiments;
[0073] Figure 18D Is a schematic collection of various surgical fields with graphical interface overlays that can occur in some embodiments;
[0074] Figure 19A Is a flowchart depicting various operations in an advanced classification process that can be implemented in some embodiments;
[0075] Figure 19B Is a flowchart depicting various operations in an image preprocessing process that can be implemented in some embodiments;
[0076] Figure 19C Is a flowchart depicting various operations in image classification post-processing that can be implemented in some embodiments;
[0077] Figure 20A Is a schematic neural network architecture diagram of an example neural network structure that can be used in some embodiments;
[0078] Figure 20B Is for creating Figure 20A A partial code listing of an example implementation for the network topology depicted in;
[0079] Figure 20C Is for performing forward propagation on an example network implementation of Figure 20B A partial code listing;
[0080] Figure 21A Is a flowchart depicting various operations in a neural network training and validation process that can be implemented in some embodiments;
[0081] Figure 21B Is a flowchart depicting various operations in a single-classifier or multi-classifier classification process that can be implemented in some embodiments;
[0082] Figure 22 Is a flowchart depicting various operations in an example process for inferring surgical performance characteristics from classification results that can be implemented in some embodiments;
[0083] Figure 23A Is a collection of schematic graphical elements that can appear in a GUI in some embodiments, including a visual image frame with a graphical interface overlay indicating the most recent classification quality assessment;
[0084] Figure 23B are schematic GUI elements that can be implemented in some embodiments for providing intraoperative image quality feedback;
[0085] Figure 23C are a pair of consecutive GUI states with intraoperative image quality feedback indications during a surgical procedure that can be implemented in some embodiments;
[0086] Figure 23D is a flowchart illustrating various operations in an example process for providing user feedback regarding infeasible images that can be implemented in some embodiments;
[0087] Figure 23E is a schematic diagram of a surgical tool that occludes a portion of the field of view of a surgical camera and that can occur in some embodiments;
[0088] Figure 24A is a set of schematic GUI elements for viewing a surgical procedure that can be implemented in some embodiments;
[0089] Figure 24B is a flowchart illustrating various operations in an example process for responding to a user playback position selection that can be implemented in some embodiments;
[0090] Figure 24C is a flowchart illustrating various operations in an example process for responding to a user model position selection that can be implemented in some embodiments;
[0091] Figure 25 is a schematic sequence of states of the model, view, and projection mapping regions of the GUI during a coverage assessment process that can be implemented in some embodiments;
[0092] Figure 26A is Figure 25 a magnified perspective view of the model state at time 2500b in
[0093] Figure 26B is Figure 25 a magnified perspective view of the projection mapping state at time 2500b in
[0094] Figure 26C is a schematic representation of a pair of relatively rotated surgical camera orientations and their corresponding views;
[0095] Figure 27 is a flowchart illustrating various operations in an example process for performing a Figure 25 coverage assessment process;
[0096] Figure 28What can be implemented in some embodiments Figure 25 A schematic state sequence of a model, view, and projection mapping region, but with additional graphical guidance;
[0097] Figure 29A A set of pairs of schematic and projection mapping regions of a local compass range that can be implemented in some embodiments, and corresponding schematic perspective views of a colon model;
[0098] Figure 29B A projection mapping GUI element with level-of-detail magnification that can be implemented in some embodiments;
[0099] Figure 30A A schematic representation of a continuous navigation compass that can be implemented in some embodiments;
[0100] Figure 30B A schematic representation of a discontinuous navigation compass that can be implemented in some embodiments;
[0101] Figure 30C What can occur in some embodiments when determining the relative position of a highlighted display on a compass for, e.g., in Figure 30A or Figure 30B A schematic representation of a series of states;
[0102] Figure 30D A schematic representation of a projection image mapping region with a bar chart color reference that can be implemented in some embodiments;
[0103] Figure 30E A schematic perspective view of an augmented reality representation of navigation guidance as seen from the perspective of a second camera that can occur in some embodiments;
[0104] Figure 31 Diagrams the various operations in an example process for rendering Figure 28 Graphical guidance;
[0105] Figure 32 A schematic block diagram illustrating the various components and their relationships in an example processing pipeline for iterative internal structure representation and navigation that can be implemented in some embodiments;
[0106] Figure 33A A schematic block diagram illustrating the various operational relationships between the components of an example surface parameterization process that can be implemented in some embodiments;
[0107] Figure 33B A schematic block diagram illustrating the various operational relationships between the components of an example surface flattening image update process that can be implemented in some embodiments;
[0108] Figure 33C is a schematic block diagram of various operational relationships between components of an example navigation compass update process that can be implemented in some embodiments;
[0109] Figure 34A depicts a schematic representation of various GUI panels that can be presented to a viewer during a first time point in a surgical procedure in some embodiments;
[0110] Figure 34B depicts, in some embodiments, during a second time point in a surgical procedure, that which can be presented to a viewer Figure 34A a schematic representation of various GUI panels in;
[0111] Figure 34C is a perspective view of a cylindrical reference geometry grid and an example intestinal grid geometry that can be used in some embodiments;
[0112] Figure 34D is a perspective view of a spherical reference geometry grid and an example cavity grid geometry that can be used in some embodiments;
[0113] Figure 34E is a perspective view of a cumulative convex hull reference geometry grid and an example cavity grid geometry that can be used in some embodiments;
[0114] Figure 34F is a schematic side view of an example correspondence between a portion of a reference geometry grid and a portion of a mesh geometry derived in a surgical procedure that can be used in some embodiments;
[0115] Figure 35A is a schematic collection of GUI elements in an example colonoscopy that can be presented to a viewer in some embodiments;
[0116] Figure 35B is a schematic collection of GUI elements in an example surgical robot inspection that can be presented to a viewer in some embodiments;
[0117] Figure 36A is a schematic representation of an incomplete model, contour determination, centerline guidance path, and navigation compass that can be implemented in some embodiments;
[0118] Figure 36B is a collection of schematic perspective views of various orientation axes relative to a centerline during an example compass alignment process that can be implemented in some embodiments;
[0119] Figure 37 is a block diagram of an example computer system that can be used in conjunction with some embodiments.
[0120] For ease of understanding, specific examples shown in the accompanying drawings are selected. Accordingly, the disclosed embodiments should not be limited to the specific details in the drawings or the corresponding disclosure. For example, the drawings may not be drawn to scale, the dimensions of some elements in the drawings may have been adjusted for ease of understanding, and the operations of the embodiments associated with the flowcharts may cover more, alternative, or fewer operations than those described herein. Thus, some components and / or operations may be separated into different blocks or combined into a single block in a different manner than depicted. The embodiments are intended to cover all modifications, equivalents, and alternatives falling within the scope of the disclosed examples, rather than limiting the embodiments to the specific examples described or depicted. Detailed Description
[0121] Overview of an exemplary surgical operating room
[0122] Figure 1A is a schematic diagram of various elements that may occur in a surgical operating room 100a during a surgical procedure, as regarding some embodiments. In particular, Figure 1A depicts a non-robotic surgical operating room 100a in which a surgeon 105a on the patient side performs an operation on a patient 120 with the help of one or more assisting members 105b, which themselves may be surgeons, physician's assistants, nurses, technicians, etc. The surgeon 105a may perform the operation using various tools, such as visualization tools 110b, such as laparoscopic ultrasound, video image acquisition endoscopes, etc., and mechanical end effectors 110a, such as scissors, retractors, dissectors, etc.
[0123] Visualization tool 110b provides an internal view of patient 120 to surgeon 105a, e.g., by displaying a visualization output from a camera that is mechanically and electrically coupled to visualization tool 110b. The surgeon can view the visualization output, e.g., through an eyepiece coupled to visualization tool 110b or on a display 125 configured to receive the visualization output. For example, where visualization tool 110b is a video image acquisition endoscope, the visualization output can be a color or grayscale image. Display 125 can allow an assisting member 105b to monitor the progress of surgeon 105a during the surgical procedure. The visualization output from visualization tool 110b can be recorded and stored for future viewing, e.g., using hardware or software on visualization tool 110a itself, acquiring the visualization output in parallel as it is provided to display 125, or acquiring the output from display 125 once it appears on the screen, etc. While two-dimensional video acquisition using visualization tool 110b can be widely discussed herein, e.g., when visualization tool 110a is an endoscope, it should be understood that in some embodiments, visualization tool 110b can acquire depth data instead of or in addition to two-dimensional image data (e.g., using a laser rangefinder, a stereoscope, etc.). Thus, it should be understood that when such data is available, the various two-dimensional operations discussed herein can be applied, with necessary modifications, to such three-dimensional depth data.
[0124] A single surgical procedure can include the execution of several sets of actions, each set of actions forming a discrete unit referred to herein as a task. For example, locating a tumor can constitute a first task, excising the tumor is a second task, and closing the surgical site is a third task. Each task can include multiple actions, e.g., the tumor excision task can require several cutting actions and several cauterizing actions. While some surgical procedures require tasks to be performed in a particular order (e.g., excision occurs before closing), the order and presence of some tasks in some surgical procedures can allow variation (e.g., canceling a prophylactic task or reordering an excision task in cases where the order is not valid). Transitions between tasks can require surgeon 105a to remove a tool from the patient, replace the tool with a different tool, or introduce a new tool. Some tasks can require visualization tool 110b to be removed and repositioned relative to its position in a previous task. While some assisting members 105b can assist with tasks related to the surgical procedure, such as administering anesthesia 115 to patient 120, assisting member 105b can also assist with these task transitions, e.g., anticipating the need for a new tool 110c.
[0125] Advances in technology have enabled procedures such as Figure 1A depicted to also be performed with robotic systems and have enabled procedures that cannot be performed in a non-robotic surgical operating room 100a. Specifically, Figure 1Bis a schematic illustration of various components present in a surgical operating room 100b during a surgical operation employing a surgical robot, such as the da Vinci TM surgical system, as may occur with some embodiments. Here, a patient-side cart 130 having tools 140a, 140b, 140c, and 140d respectively attached to each of a plurality of arms 135a, 135b, 135c, and 135d may occupy the position of a patient-side surgeon 105a. As previously described, one or more of the tools 140a, 140b, 140c, and 140d may include a visualization tool (here, visualization tool 140d), such as a video image endoscope, laparoscopic ultrasound, etc. An operator 105c, who may be a surgeon, may view the output of the visualization tool 140d via a display 160a on a surgeon's console 155. By manipulating a handheld input mechanism 160b and a pedal 160c, the operator 105c may communicate remotely with the tools 140a - 140d on the patient-side cart 130 in order to perform a surgical procedure on a patient 120. In fact, since the communication between the surgeon's console 155 and the patient-side cart 130 may occur across a telecommunications network in some embodiments, the operator 105c may or may not be in the same physical location as the patient-side cart 130 and the patient 120. The electronic device / console 145 may also include a display 150 depicting the patient's vital signs and / or the output of the visualization tool 140d.
[0126] Similar to the task transitions of a non-robot surgical operating room 100a, the surgical operation in operating room 100b may require the removal or replacement of tools 140a - 140d, including the visualization tool 140d, for various tasks, as well as the introduction of new tools (e.g., new tool 165). As previously described, one or more assistant members 105d may now anticipate such changes and work with the operator 105c to make any necessary adjustments as the surgical procedure progresses.
[0127] Similarly, similar to the non-robotic surgical operating room 100a, the output from the visualization tool 140d can be recorded here, for example, at the patient-side cart 130, at the surgeon's console 155, recorded from the display 150, and so on. Although some of the tools 110a, 110b, 110c in the non-robotic surgical operating room 100a can record additional data such as temperature, motion, conductivity, energy level, etc., the presence of the surgeon's console 155 and the patient-side cart 130 in the operating room 100b can help record much more data than just the data output from the visualization tool 140d. For example, the manipulation of the handheld input mechanism 160b by the operator 105c, the activation of the pedal 160c, eye movements within the display 160a, etc. can all be recorded. Similarly, the patient-side cart 130 can record tool activations throughout the surgical procedure (e.g., the application of radiant energy, the closing of scissors, etc.), the movement of the end effector, etc. In some embodiments, an in-operating room recording device such as Intuitive Data Recorder TM (IDR) can be used to record data, which can collect and store sensor data in a local or networked location.
[0128] Overview of an exemplary organ data acquisition
[0129] Whether in the non-robotic surgical operating room 100a or within the robotic surgical operating room 100b, there can be situations where the surgeon 105a, the assistant member 105b, the operator 105c, the assistant member 105d, etc. attempt to examine the organs or other internal body structures of the patient 120 (e.g., using the visualization tools 110b or 140d). For example, as Figure 2AAs shown and through section 205b, colonoscope 205d can be used to examine large intestine 205a. Although this detailed description will use the large intestine and colonoscope as specific examples for ease of understanding by the reader, it should be readily understood that the disclosed embodiments are not necessarily limited to the large intestine and colonoscope, and in fact are specifically not contemplated to be so limited here. Instead, it should be understood that the disclosed embodiments can equally be applied in conjunction with other organs and internal structures (such as the lungs, heart, stomach, arteries, veins, urethra, areas between organs and tissues, etc.), as well as in conjunction with other instruments (such as laparoscopes, thoracoscopes, sensor-bearing catheters, bronchoscopes, ultrasound probes, micro-robots (e.g., swallowed sensor platforms), etc.). Many such organs and internal structures will include folds, protrusions, and other structures that can occlude portions of the organ or internal structure from one or more perspectives. For example, the large intestine 205a shown here includes a series of small pits known as haustra (including haustra 205f and haustra 205g). Regardless of these haustra and various other challenges (including possible limitations of the visualization tools themselves) contributing to occlusion of the field of view, thoroughly examining the large intestine can be very difficult for a surgeon or an automated system.
[0130] In the depicted example, as the operator or automated system slides the colonoscope 205d forward, the colonoscope 205d can be navigated through the large intestine by adjusting the bending section 205i. The bending section 205i can also be adjusted to orient the distal tip 205c in a desired orientation. As the colonoscope travels through the large intestine 205a, perhaps from the descending colon all the way to the transverse colon and then to the ascending colon, an actuator in the bending section 205i can be used to guide the distal tip 205c along the centerline 205h of the intestine. The centerline 205h is a path along points that are substantially equidistant from the inner surface of the large intestine along its length. Prioritizing the movement of the colonoscope 205d along the centerline 205h can reduce the risk of collision with the intestinal wall, which could harm the patient 120 or cause discomfort to the patient 120. Although the colonoscope 205d is shown here as entering via the rectum 205e, it should be understood that laparoscopic incisions and other paths can also be used to access the large intestine of the patient 120 as well as other organs and internal body structures.
[0131] Figure 2B A close-up view of the distal tip 205c of the colonoscope 205d is provided. This exemplary tip 205c includes a visual image camera 210a (which can acquire, for example, color or grayscale images), a light source 210c, a flushing outlet 210b, and an instrument bay 210d (which can accommodate, for example, cautery tools, scissors, forceps, etc.), although variations in the distal tip design will be readily understood. For clarity, and as indicated by the ellipsis 210i, it should be understood that the bending section 205i can extend a significant distance behind the distal tip 205c.
[0132] As previously described, when the colonoscope 205d is advanced and retracted through the intestine, joints or other flexible actuators within the flexible section 205i can facilitate movement of the distal tip 205c in various directions. For example, referring to arrows 210f, 210g, 210h, an operator or an automated system can generally advance the colonoscope tip 205c in the Z direction shown by arrow 210f. Actuators in the flexible portion 205i can allow the distal end 205c to rotate about the Y axis or the X axis (possibly simultaneously), represented by arrows 210g and 210h respectively (and thus analogous to yaw and pitch respectively). In this way, the field of view 210e of the camera 210a can be adjusted to facilitate examination of structures other than those directly in front of the direction of movement of the colonoscope, such as areas obscured by haustral folds.
[0133] Specifically, Figure 2C is a schematic illustration of a portion of the large intestine, the cross-sectional view of which reveals the position of the colonoscope tip 205c relative to a plurality of haustral annular ridges. Interstitial tissue forming annular ridges can exist between each of the haustra 215a, 215b, 215c, 215d. In this example, an annular ridge 215h is formed between haustra 215a, 215b, an annular ridge 215i is formed between haustra 215b, 215c, and an annular ridge 215j is formed between haustra 215c, 215d. Although an operator may wish for the colonoscope to generally travel along the centerline 205h of the colon to minimize patient discomfort, the operator may also wish for the flexible portion 205i to reorient the distal tip 205c such that the field of view 210e of the camera 210a can observe portions of the colon obscured by the annular ridges.
[0134] For camera 210a, regions that are farther from light source 210c may appear darker than regions that are closer to light source 210c. Accordingly, annular ridge 215j may appear brighter than opposing wall 215f in the camera's field of view, and aperture 215g may appear very dark or completely dark to camera 210a. In some embodiments, distal tip 205c may include a depth sensor, e.g., a depth sensor in instrument bay 210d. Such a sensor may use, e.g., time-of-flight photon reflection data, ultrasound, a stereo vision image camera pair (e.g., on an additional camera in addition to camera 210a), etc. to determine depth. However, the various embodiments disclosed herein contemplate estimating depth data based on a visual image from a single visual image camera 210a on distal tip 205c. For example, a neural network may be trained to identify distance values corresponding to an image from camera 210a (e.g., changes in surface structure and brightness resulting from reflected light from light 210c at varying distances may provide sufficient correlation with depth between successive images for depth prediction). Some embodiments may employ a six-degree-of-freedom guidance sensor (e.g., a 3D sensor) provided by Northern Digital, Inc. to replace or be combined with the pose estimation methods described herein such that the methods described herein and the six-degree-of-freedom sensor provide complementary confirmation of each other's results.
[0135] Thus, for clarity, Figure 2D a visual image of a depth frame and a corresponding schematic representation taken from the perspective of the camera of the colonoscope depicted in Figure 2C are depicted. Here, annular ridge 215j occludes a portion of annular ridge 215i, which itself occludes a portion of annular ridge 215h, which in turn occludes a portion of wall 215f. When aperture 215g is within the camera's field of view, the aperture is far enough from the light source such that it may appear completely dark.
[0136] With the aid of a depth sensor, or by performing image processing on image 220a (and possibly previous or subsequent images that follow the movement of the colonoscope) using the systems and methods discussed herein, a corresponding depth frame 220b can be generated that corresponds to the same field of view from which the visual image 220a was produced. As shown in this example, depth frame 220b assigns depth values to some or all of the pixel locations in image 220a (however, it should be understood that the visual image and the depth frame will not always have values that directly map pixels to depth values, e.g., in the case where the depth frame has a smaller size than the visual image). It should be understood that in some embodiments, the depth frame itself, which includes a series of depth values, can be presented as a grayscale image (e.g., the maximum depth value mapped to value 0, the minimum depth value mapped to 255, and the resulting mapped values presented as a grayscale image). Thus, annular ridge 215j can be associated with the closest set of depth values 220f, annular ridge 215i can be associated with another set of depth values 220g, annular ridge 215h can be associated with yet another set of depth values 220d, back wall 215f can be associated with a distant set of depth values 220c, and the orifice 215g can be outside the depth sensing range (or completely black, outside the range of the light source), thereby incurring the maximum depth value 220e (e.g., a value corresponding to infinite or unknown depth). Although a single pattern is shown for each annular ridge in this schematic for ease of understanding by the reader, it should be understood that annular ridges rarely present a flat surface in the X-Y plane (in the direction of arrows 210h and 210g) at the distal tip. Thus, for example, many of the depth values within set 220f are not likely to be exactly the same value.
[0137] Although the visual image camera 210a can acquire a rectilinear image, it should be understood that in some embodiments, lenses, post-processing, etc. can be applied such that the image acquired from camera 210a is not rectilinear. For example, Figure 2Eis a pair of images 225b, 225c, which are respectively acquired from a colonoscope camera with a straight view and a colonoscope camera with a fish-eye view, and depict a grid-like square pattern 225a of orthogonal rows and columns in a perspective view. Such a square pattern can help determine the intrinsic parameters of a given camera. It should be understood that once the intrinsic parameters of the camera are known (which can be useful, for example, to normalize / dimension the different sensor systems into a similar form recognized by a machine learning architecture), the straight view can be achieved by undistorting the fish-eye view. The fish-eye view can allow the user to easily perceive a wider field of view than the straight view angle. Since the focal point of the fish-eye lens and other details of the colonoscope (such as the brightness of the light 210c) can vary between devices and even over time on the same device, it may be necessary to recalibrate the various processing methods of the specific device under discussion (considering the "intrinsic" of the device, such as focal length, principal point, distortion coefficient, etc.), or at least anticipate device variations when training and configuring the system.
[0138] Exemplary computer - generated organ model
[0139] During or after examining an internal body structure (such as the large intestine 205a) with a camera system (such as camera 210a), it may be desirable to generate a corresponding three-dimensional model of the organ or the cavity being examined. For example, various disclosed embodiments can generate a Truncated Signed Distance Function (TSDF) volume model based on depth data acquired during the examination, such as the TSDF model 305 of the large intestine 205a. Although TSDF is provided here as an example for the convenience of the reader, it should be understood that many suitable three-dimensional data formats exist. For example, a model in TSDF format can be easily converted to a vertex mesh or other desired model formats, and thus the reference to "model" in this document can be understood to refer to any such format. Therefore, the model can be textured with the images acquired via camera 210a, or can be colored, for example, with a vertex shader. For example, in the case where the colonoscope travels within the large intestine, the model can include an inner surface and an outer surface, where the inner surface is rendered with the texture acquired during the examination, and the outer surface is colored with vertex coloring. In some embodiments, only the inner surface can be rendered, or only a part of the outer surface can be rendered, so that the viewer can easily examine the inside of the organ.
[0140] Such computer-generated models can be used for various purposes. For example, portions of the model can be differently textured, highlighted via a contour (e.g., from the viewer's perspective, the outline of the area is projected onto the texture of the billboard vertex mesh surface in front of the model), called out with a 3D marker, or otherwise identified, and these portions are associated with, for example, portions of an examination identified by an operator, portions of an organ determined by the various embodiments disclosed herein to be inadequately viewed, organ structures of interest (such as polyps, tumors, abscesses, etc.). For example, portions 310a and 310b of the model can be vertex-colored or contoured in a color different from the rest of model 305 or in some other way different from the rest of model 305 to draw the operator's attention to inadequate viewing, for example, in cases where the operator fails to obtain a complete image acquisition of an organ region, moves too quickly through the region, only obtains a blurred image of the region, views the region when it is obscured by smoke, etc. Although a complete model of the organ is shown in this example, it should be understood that an incomplete model can equally be generated, for example, generated in real-time during an examination, generated after an incomplete examination, etc. In some embodiments, the model can be a non-rigid 3D reconstruction (e.g., combined with a physical model to represent the behavior of tissue with varying stiffness).
[0141] For clarity,[ Figure 3A , Figure 3B , Figure 3C each of[ Figure 3B shows the three-dimensional model 305 from a different perspective. Specifically, a coordinate system 320 is provided for the reader's reference, which has X - Y - Z axes represented by arrows 315a, 315c, 315b, respectively. If the model is rendered around the coordinate system 320 at the center of the model, then Figure 3B shows the model 305 rotated approximately 40 degrees 330a around the Y axis (i.e., in the X - Z plane 325) with respect to the orientation of the model 305 in Figure 3A . Similarly, Figure 3C depicts the model 305 further rotated approximately an additional 40 degrees 330b to an orientation that is almost perpendicular to the orientation in Figure 3A . It should be understood that the model 305 can be rendered only from the inside of the organ (e.g., where the colonoscope appears), only from the outside, or from both the inside and the outside (e.g., using two complementary texture meshes). In cases where the only available data is for the inside of the organ, the external texture can be vertex-colored, textured with a synthetic texture approximating the texture of the actual organ, simply transparent, etc. In some embodiments, only vertex coloring is used to render the outside. As discussed herein, the viewer may be able to view in a manner similar to Figure 3A , Figure 3B , Figure 3CRotate the model in a manner, as well as translate, scale, etc., so as to, for example, more closely study the identified regions 310a, 310b to plan subsequent surgeries, evaluate the relationship between the organ and the envisioned implant (e.g., surgical mesh, fiducial markers, etc.), and so on.
[0142] Exemplary frame generation and merging operations
[0143] Since depth data can be incrementally acquired throughout the examination, the data can be combined to facilitate the creation of a corresponding three-dimensional model of all or part of the internal body structure (such as model 305). For example, Figures 4A - 4C Presents a temporally continuous schematic two-dimensional cross-sectional representation corresponding to the actual three-dimensional view of the colonoscope field of view as the colonoscope travels through the colon.
[0144] Specifically, Figure 4A Depicts a two-dimensional cross-sectional view of the interior of the colon represented by a top portion 425a and a bottom portion 425b. As discussed, the interior of the colon, like many internal body parts, can contain various irregular surfaces, for example, where the pouches join, where polyps form, etc. Thus, when the colonoscope 405 is in the Figure 4A position, the camera coupled to the distal tip 410 can have an initial field of view 420a. Since the irregular surfaces may obscure parts of the interior of the colon, only certain surfaces, particularly surfaces 430a, 430b, 430c, 430d, and 430e, may be visible to the camera (and / or depth sensor) from this position. Moreover, since this is a cross-sectional view similar to Figure 2C this, it should be understood that such surfaces can correspond to the annular ridge surfaces that appear in image 220a. That is, although the surfaces are represented by lines here, it should be understood that these surfaces can correspond to three-dimensional structures, such as the annular ridges between the pouches, such as annular ridges 215h, 215i, 215j. Due to the limited field of view, the surgeon may not yet have seen the obscured regions, such as region 425c outside the field of view 420a. It should be understood that such limitations to the field of view can exist regardless of whether the camera image is rectilinear, fish-eye, etc.
[0145] As Figure 4BAs shown, as the colonoscope 405 is further advanced into the colon (from right to left in this depiction), the camera's field of view 420b can now sense surfaces 440a, 440b, and 440c. Naturally, portions of these surfaces can coincide with portions of the previously viewed surfaces, as in the case of surfaces 430a and 440a. If the field of view of the colonoscope continues to advance linearly without adjustment (e.g., rotation of the distal tip via the flexible section 205i), portions of the occluded surface can remain invisible. Here, for example, although the colonoscope advances, region 425c still does not appear within the camera's field of view 420b. Similarly, when the colonoscope 405 is advanced to Figure 4C the position where, surfaces 450a and 450b can now be visible in the field of view 420c. However, unfortunately, the colonoscope will continue to pass through region 425c, which does not appear in the field of view.
[0146] It should be understood that throughout the progression of the colonoscope 405, depth values corresponding to the internal structure of the colonoscope prior to it can be generated in real-time during the examination or by post-processing the acquired data after the examination. For example, in the case where the distal tip 205c does not include a sensor specifically designed for depth data acquisition, the system can instead use images from the camera to infer depth values (an operation that can be performed in real-time or near real-time using the methods described herein). There are various methods for determining depth values from images, including, for example, using a neural network trained to convert visual image data into depth values. For example, it should be understood that self-supervised methods can be used to generate a network for inferring depth from monocular images, such as the method found in the paper "Digging Into self-Supervised Monocular Depth Estimation", which was published in the form of arXiv TM preprint arXiv TM :1806.01260v4 and written by Clément Godard, Oisin Mac Aodha, Michael Firman, and Gabriel Brostow, and implemented in the Monodepth2 self-supervised model described in that paper. However, such methods do not specifically anticipate the unique challenges present in this endoscopic scenario and can be modified as described herein. In the case where the distal tip 205c does include a depth sensor, or in the case where stereoscopic vision images are available, depth values from various sources can be corroborated by the values from the monocular image method.
[0147] Thus, multiple depth values can be generated for each position of the colonoscope at which data is acquired to produce a corresponding depth data "frame". Here, Figure 4A the data in can produceFigure 4D depth frame 470a, Figure 4B the data in can generate Figure 4E depth frame 470, and Figure 4C the data in can generate Figure 4F depth frame 470c. Thus, depth values 435a, 435b, 435c, 435d, and 435e can correspond to surfaces 430a, 430b, 430c, 430d, and 430e, respectively. Similarly, depth values 445a, 445b, and 445c can correspond to surfaces 440a, 440b, and 440c, respectively, and depth values 455a and 455b can correspond to surfaces 450a and 450b.
[0148] Note that each depth frame 470a, 470b, 470c is acquired from the perspective of the distal tip 410, and the distal tip 410 can be used as the geometric origin 415a, 415b, 415c for each corresponding frame. Thus, each of the frames 470a, 470b, 470c can be considered with respect to the pose of the distal tip at the time of data acquisition (e.g., position and orientation represented by a matrix or quaternion), and if the depth data in the resulting frames is to be combined, e.g., to form a three-dimensional representation of an organ as a whole (such as model 305), then each of the frames 470a, 470b, 470c can be globally reoriented. This process (referred to as stitching or fusion) is schematically shown in Figure 4G where depth frames 470a, 470b, 470c are combined 460a, 460b to form the merged frame 480 460c. Example methods for stitching frames together are described herein.
[0149] Exemplary data processing operations
[0150] Figure 5 is a flowchart of various operations in an example process 500 for generating a computer model of at least a portion of an internal body structure that can be implemented in some embodiments. At block 505, the system can initialize a counter N to 0 (it should be understood that the flowchart is merely exemplary and is chosen for ease of understanding by the reader, and thus, many embodiments may not employ such a counter or Figure 5(specific operations disclosed in). At block 510, the computer system can allocate storage for the initial fragment data structure. As explained in more detail herein, a fragment is a data structure that contributes to all or part of the creation of a model and includes data of one or more depth frames. In some embodiments, a fragment can contain data related to a sequence of consecutive frames depicting similar regions of internal body structures and can share a large overlapping region over that region. Thus, the fragment data structure can include memory allocated to receive RGB visual images, visual feature correspondences between visual images, depth frames, relative poses between frames within the fragment, timestamps, and the like. At blocks 515 and 520, the system can then iterate over each image in the acquired video, incrementing a counter accordingly, and then at block 525 retrieve the corresponding next consecutive visual image of the video.
[0151] As shown in this example, the visual image retrieved at block 525 can then be processed in parallel by two different sub - processes, namely, a feature - matching - based pose - estimation sub - process 530a and a depth - determination - based pose - estimation sub - process 530b. However, naturally, it should be understood that the sub - processes can instead be executed sequentially. Similarly, it should be understood that parallel processing does not necessarily mean two different processing systems, as a single system can be used for parallel processing using, for example, two different threads (e.g., when the same processing resources are shared between two threads), etc.
[0152] The feature - matching - based pose - estimation sub - process 530a uses correspondences between features of the image (such as scale - invariant feature transform (SIFT) features) and these features when they appeared in a previous image to determine the local pose from the image. For example, the method specified in the paper "BundleFusion: Real - time Globally Consistent 3DReconstruction", which is available as an arXiv TM arXiv preprint TMPublished in the form of 1604.01093v3, written by Angela Dai, Matthias Niessner, Michael Zollhofer, Shahram Izadi, and Christian Theobalt. In particular, the feature correspondences for global pose alignment described in Section 4.1 of this paper, where the Kabsch algorithm is used for alignment. However, it should be understood that the exact method specified need not be used in every embodiment disclosed herein (e.g., it should be understood that various alternative correspondence algorithms suitable for feature comparison can be used). Instead, at block 535, any image features can be generated from the visual image that are suitable for pose recognition relative to the features of the previously considered image. For this purpose, SIFT features (as described in the "BundleFusion" paper cited above), Speeded-Up Robust Features (SURF), features from Features from Accelerated Segment Test (FAST), Binary Robust Independent Elementary Features (BRIEF) descriptors (as used in Oriented FAST and Rotated BRIEF (ORB)), Binary Robust Invariant Scalable Keypoints (BRISK), etc. can be used. In some embodiments, instead of using these conventional features, a neural network can be used to generate features (e.g., values from the layers of a UNet network, using the method specified in the 2021 paper "LoFTR: Detector-Free Local Feature Matching with Transformers", which is available as arXiv TM Preprint arXiv TM : Obtained from 2104.00680v1 and written by Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. Using the method specified in the paper "SuperGlue: Learning Feature Matching with Graph Neural Networks", which is available as arXiv TM Preprint arXiv TM : Obtained from 1911.11763v2 and written by Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew Rabinovich, etc.). Such customized features can be useful when applied to a specific internal body context, a specific camera type, etc.
[0153] At block 540, features of the same type can be generated (or retrieved if previously generated) for the previously considered images. For example, if M is 1, only the previous image will be considered. In some embodiments, each previous image can be considered (e.g., M is N - 1), similar to the "BundleFusion" method of Dai et al. Then, the features generated at block 540 can be matched to those generated at block 535. Then, these matching correspondences determined at block 545 themselves can be used to determine a pose estimate for the Nth image at block 550, e.g., by finding the optimal set of rigid camera transforms that best align the features of the N to N - M images.
[0154] In contrast to the feature-matching-based pose estimation sub-process 530a, the depth-determined pose estimation process 530b employs one or more machine learning architectures to determine pose and depth estimates. For example, in some embodiments, the estimation process 530b considers image N and image N-1, submitting this combination to a machine learning architecture that is trained to determine both the pose and depth frames for the image, as indicated by block 555 (although not shown here for clarity, it should be understood that in the absence of any previous images, or when N = 1, the system can simply wait until a new image arrives for consideration; thus, block 505 could be changed to initialize N to M such that there are a sufficient number of previous images for analysis). It should be understood that many machine learning architectures can be trained to generate both pose and depth frame estimates for a given visual image in this manner. For example, similar to sub-process 530a, some machine learning architectures can determine depth and pose by considering not only the Nth image frame as input, but also by considering multiple previous image frames (e.g., the Nth and N-1th images, the Nth to N-Mth images, etc.). However, it should be understood that machine learning architectures that consider only the Nth image to produce depth and pose estimates also exist and can be used. For example, block 555 can apply a single-image machine learning architecture generated according to various methods described in the paper "Digging Into self-Supervised Monocular Depth Estimation" cited above. The Monodepth2 self-supervised model described in that paper can be trained on images depicting the endoscopic environment. In cases where sufficient real-world endoscopic data is not available for this purpose, synthetic data can be used. In fact, while the self-supervised method of Godard et al. using real-world data did not envision using precise pose and depth data to train machine learning architectures, synthetic data generation can easily facilitate the generation of such parameters (e.g., because a virtual camera can be advanced through a computer-generated organ model in known distance increments), and thus can facilitate a fully supervised training method rather than the self-supervised method in their paper (however, synthetic images can still be used for self-supervised methods, such as when the training data includes both synthetic and real-world data). Such supervised training can be useful, for example, given the unique variations between certain endoscopes, operating environments, etc., which may not be fully represented in self-supervised methods. Whether prepared via self-supervised training, fully supervised training, or via other training methods, the model of block 555 predicts the depth frame and pose for the visual image here.It should be understood that various methods are used to supplement unbalanced synthetic datasets and real-world datasets, including, for example: the method described in the 2018 paper "T2Net: Synthetic-to-Realistic Translation for Solving Single-Image Depth Estimation Tasks", which is available as an arXiv preprint. TM arXiv preprint TM obtained as :1808.01454v1 and written by Chuanxia Zheng, Tat-Jen Cham, and Jianfei Cai; the method described in the 2019 paper "Geometry-Aware Symmetric Domain Adaptation for Monocular Depth Estimation", which is available as an arXiv preprint. TM arXiv preprint TM obtained as :1904.01870v1 and written by Shanshan Zhao, Huan Fu, Mingming Gong, and Dacheng Tao; the method described in the paper "Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks", which is available as an arXiv preprint. TM arXiv preprint TM obtained as :1703.10593v7 and written by Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros, and any suitable neural style transfer methods, such as the method described in the paper "Deep Photo Style Transfer", which is available as an arXiv preprint. TM arXiv preprint TM obtained as :1703.07511v3 and written by Fujun Luan, Sylvain Paris, Eli Shechtman, and Kavita Bala (e.g., applicable to results suggesting photo-realistic images).
[0155] Accordingly, as processing continues to block 560, the system may have an available pose determined at block 550, a second pose determined at block 555, and a depth frame determined at block 555. The pose determined at block 555 may be different from the pose determined at block 550, assuming they use different methods. If block 550 successfully finds a pose (e.g., a sufficient number of feature matches), the process may continue with the pose from block 550 and the depth frame generated at block 555 in subsequent processing (e.g., transitioning to block 580).
[0156] However, in some cases, the pose determination at block 550 may fail. For example, in the case where features cannot be matched at block 545, the system may not be able to determine a pose at block 550. Although this failure can occur during the normal course of image acquisition given the vast diversity of the body interior and conditions, this failure can also be caused, for example, when the operator moves the camera too quickly, resulting in blurring of the Nth frame, making it difficult or impossible to generate features at block 535. Instrument occlusion, biological matter occlusion, smoke (e.g., from a cautery device), or other irregularities can similarly result in poor feature generation or poor feature matching. Naturally, if such an image is subsequently considered at block 545, it can again lead to a failure in pose recognition. In such a case, at block 560, the system may transition to block 565, prepare the pose determined at block 555 for use in place of the pose determined at block 550 (e.g., adjust for differences in scale, format, etc., although in some embodiments it may be sufficient to perform the replacement at block 575 without preparation), and perform the replacement at block 575. In some embodiments, during the first iteration starting from block 515, since there is no previous frame in process 530a with which to perform a match at block 540, the system may similarly rely on the pose from block 555 for the first iteration.
[0157] At block 580, the system can determine whether the pose (whether from block 550 or block 555) and depth frame correspond to an existing segment being generated or whether they should be associated with a new segment. Various methods can be used to determine when a new segment will be generated. In some embodiments, a new segment can simply be generated after a fixed number (e.g., 20) of frames have been considered. In other embodiments, the number of matching features at block 545 can be used as a proxy for regional similarity. In cases where many features in a frame match those in its immediately preceding frame, it can be reasonable to assign the corresponding depth frame to the same segment (e.g., transitioning to block 590). Conversely, in cases where there are few enough matches, it can be inferred that the endoscope has moved to a substantially different region and thus the system should start a new segment at block 585a. Additionally, as described herein, at block 585b, the system can also perform integration of previously considered segments and global pose network optimization (for clarity, it will be recognized that the "local" poses of blocks 550 and 555, also referred to as "coarse" poses, are relative to consecutive frames, while the "global" pose is relative to the coordinates of the entire model). The process 1100 regarding Figure 11A provides an example method for performing block 580.
[0158] Using the available depth frames and poses and their determined corresponding segments, at block 590, the system can use pose estimation to integrate the depth frames with the current segment. For example, Simultaneous Localization and Mapping (SLAM) can be used to determine the pose of the depth frame relative to other frames in the segment. Since organs are typically non-rigid, non-rigid methods such as those described in the paper "As-rigid-as-possible surface modeling" by Olga Sorkine and Marc Alexa in Symposium on Geometry processing. Vol. 4. 2007 can be used. Furthermore, it should be understood that the exact methods specified therein need not be used in every embodiment. Similarly, some embodiments can employ methods from the dynamic fusion approach specified in the paper "Dynamic Fusion: Reconstruction and tracking of non-rigid scenes in real-time" by Richard A. Newcombe, Dieter Fox, and Steven M. Seitz published in "Proceedings of the IEEE conference on computer vision and pattern recognition, 2015". Dynamic fusion can be appropriate because many of the papers cited herein did not anticipate the non-rigidity of body tissues nor the artifacts caused by breathing, patient movement, surgical instrument movement, etc. Thus, the canonical models cited in that paper will correspond to the key-frame depth frames described herein. In addition to integrating the depth frames with their peer frames in the segment, at block 595, the system can also attach the pose estimation to the set of poses associated with the frames of the segment for future consideration (e.g., as discussed with respect to block 570, the collective pose can be used to improve the global alignment with other segments).
[0159] Once all the desired images from the video have been processed at block 515, the system can transition to block 570 and begin generating a complete or intermediate model of the organ by merging one or more newly generated segments with the help of the optimized pose trajectory determined at block 595. In some embodiments, block 570 can be skipped because the global pose alignment at block 585b may already include model generation operations. However, as described in more detail herein, in some embodiments, not all segments can be integrated into the final mesh upon acquisition, and thus block 570 can include selecting segments from a network (e.g., similar to the network described herein with respect to Figure 11E ).
[0160] Exemplary data processing operation pre - processing variant
[0161] Figure 6 is a flowchart of a preprocessing variant of a process that can be implemented in some embodiments. In particular, while most processing operations can generally remain as described above, the example process 600 can also seek to exclude visual images that are not suitable for downstream processing, thereby greatly improving the efficiency and effectiveness of the system. The consideration of unsuitable visual images not only consumes valuable resources pointlessly, but can also cause the system to attempt to reconcile the results of downstream processing of an unsuitable image with the results of other potentially suitable images. Thus, the consideration of a single unsuitable image can impede the proper consideration of subsequent suitable images. Accordingly, at block 625b, the system can provide the image to a non-informative frame filter, such as a neural network as described herein, to evaluate whether the image is suitable for downstream processing. If the filter finds that the image is "non-informative", downstream processing of the image can be abandoned and, instead, the next visual image can be considered, as indicated. Figure 5
[0162] Conversely, if the filter finds that the visual image is suitable, the count N of informative visual images can be incremented at block 620, and the visual image frame is then processed in parallel by two different sub-processes (a feature-matching-based pose estimation sub-process 530a and a depth-determination-based pose estimation sub-process 530b), as previously discussed. For clarity, here, the counter value N refers to the total count of image frames that have been found to be informative, rather than simply being all image frames (thus, N - 1 refers to the previous image that has been found to be informative, and not necessarily simply the previous image in chronological order).
[0163] Exemplary end - to - end data processing pipeline
[0164] For greater clarity, Figure 7 is a processing pipeline 700 for generating at least a portion of a three-dimensional model of the large intestine from colonoscopy data acquisition that can be implemented in some embodiments. Again, although the large intestine is shown here for ease of understanding, it should be understood that embodiments contemplate other organs and internal structures of patient 120.
[0165] Here, as the colonoscope 710 progresses through the actual large intestine 705, a camera or depth sensor can bring a new region of the intestine 705 into view. At Figure 7 At the moment depicted, region 715 of the intestine 705 is within the field of view of the endoscopic camera, thereby generating a two-dimensional visual image 720 of region 715. The computer system can use image 720 to generate both extracted features 725 (corresponding to process 530a) and deep neural network features 730 (corresponding to process 530b). In this example, the extracted features 725 generate pose 735. In contrast, the deep neural network features 730 can include a depth frame 740a and a pose 740b (however, in embodiments where pose 735 is always used, the neural network that generates pose 740b may be unnecessary).
[0166] As discussed, the computer system can use pose 735 and depth frame 740a in matching and verification operations 745, where the suitability of the depth frame and pose is considered. At blocks 750 and 755, a new frame can be integrated with other frames of the segment by determining the correspondence between the new frame and other frames of the segment and performing local pose optimization. When segment 760 is complete, the system can align the segment with previously collected segments via global pose optimization 765 (e.g., corresponding to block 585b). The computer system can then perform global pose optimization 765 on segment 760 to orient segment 760 relative to an existing model. After creating the first segment, the computer system can also use this global pose to determine keyframe correspondences between segments 770 (e.g., generating a network similar to that described herein with respect to Figure 11E the network).
[0167] The performance of global pose optimization 765 can involve referencing and updating database 775. This database can contain records of previous poses 775a, camera calibration intrinsics 775b, records of frame segment indices 775c, frame features including corresponding UV texture map data (such as acquired camera images of organs) 775d, and records of keyframe-to-keyframe matches 775e (e.g., as in Figure 11E the network). The computer system can integrate 780 database data (e.g., corresponding to block 570) at the end of an examination or in real-time during an examination to update 785 or generate a computer-generated model of an organ, such as a TSDF representation 790. In this example, the system operates in real-time and updates a pre-existing portion of the TSDF model 790a with a new set of voxels (or, for example, corresponding vertices and textures where the model is a polygonal mesh) 790b corresponding to the new segment 760 generated for region 715.
[0168] Exemplary end - to - end data processing pipeline pre - processing variant
[0169] For clarity, Figure 8 is an example processing pipeline 800 depicting a preprocessing variant of the Figure 7 pipeline that can be implemented in some embodiments. Similar toFigure 6 Additional preprocessing. As the colonoscope 710 advances through the actual large intestine 705, a camera or depth sensor can bring a new area of the intestine 705 into view. As previously mentioned, at the moment depicted in Figure 7 , an area 715 of the intestine 705 is within the field of view of the endoscopic camera, resulting in a two-dimensional visual image 720 of the area 715. After confirming 895 that the image is suitable for localization and mapping (e.g., as discussed with respect to box 625b), rather than discarding the frame 890c, the computer system can use the image 720 to generate both the extracted features 725 (corresponding to process 530a) and the deep neural network features 730 (corresponding to process 530b), as previously discussed.
[0170] Exemplary end - to - end data processing pipeline - exemplary pose and depth pipeline
[0171] It should be understood that there are various methods for determining the rough relative pose 740b and the depth map 740a (e.g., at box 555). Naturally, in the case where the inspection device includes a depth sensor, the depth map 740a can be directly generated from the sensor (naturally, this may not yield the pose 740b). However, many depth sensors impose limitations, such as time-of-flight limitations, which may reduce the applicability of the sensor for data acquisition within the organ. Therefore, it may be desirable to infer the pose and depth data from the visual image, as most inspection tools will already have generated the visual data for the surgeon to view in any case.
[0172] Inferring pose and depth from visual images can be difficult, especially when only monocular rather than stereo image data is available. Similarly, it may be difficult to obtain sufficient such data with corresponding depth values (if required for training) to appropriately train machine learning architectures, such as neural networks. There are indeed some techniques for obtaining pose and depth data from monocular images, such as the methods described in the "Digging Into self-Supervised Monocular DepthEstimation" paper cited herein, but these methods are not directly applicable to the context of the body interior (the work of Godard et al. is targeted at the field of autonomous driving), and thus cannot address the various unique challenges of this data.
[0173] Figure 9ADepicts an example processing pipeline 900a for obtaining depth and pose data from monocular images in an in-body context. Here, the computer system considers two temporally consecutive image frames from an endoscopic camera, an initial image acquisition 905a and a subsequent acquisition 905b after the endoscope has been advanced forward through the intestine (however, as indicated by the ellipsis 960, it should be readily understood that variants with more than two consecutive images are employed and the input to the neural network can be adjusted accordingly; similarly, corresponding operations for retraction and other camera movements should be understood). In pipeline 900a, the computer system supplies 910a the initial image acquisition 905a to a first depth neural network 915a, which is configured to produce 920a a depth frame representation 925 (corresponding to depth data 740a). It should be understood that in the case of considering more than two images, the image acquisition 905a can be, for example, the first image in a time series. Similarly, the computer system supplies 910b, 910c both the image 905a and the image 905b to a second pose neural network 915b to produce 920b a rough pose estimate 930 (corresponding to rough relative pose 740b). Specifically, the network 915b can predict a transformation 940 that accounts for the viewpoint difference between the image 905a (captured from orientation 935a) and the image 905b (captured from orientation 935b). It should be understood that in embodiments considering more than two consecutive images, the transformation 940 can be performed temporally between the first and the last of the images. In the case of considering more than two input images, all the input images can be provided to the network 915b.
[0174] Thus, in some embodiments, the depth network 915a can be a UNet-like network (e.g., a network having substantially the same layers as UNet) configured to receive a single image input. For example, the DispNet network described in the paper “Unsupervised Monocular Depth Estimation with Left-Right Consistency” can be used for the depth determination network 915a, which is available as an arXiv TM arXiv preprint TM: 1609.03677v3 is obtained and written by Clément Godard, Oisin Mac Aodha, and Gabriel J. Brostow. As mentioned above, the method from "Digging Into self-Supervised Monocular Depth Estimation" described above can also be used for the depth determination network 915a. Thus, the depth determination network 915a can be, for example, a UNet with a ResNet(50) or ResNet(101) backbone and a DispNet decoder. Some embodiments can also employ a depth consistency loss and a mask between two frames during training, as in the paper "Unsupervised scale-consistent depth and ego-motion learning from monocular video" which is available as arXiv TM Preprint arXiv TM : 1908.10553v2 is obtained and written by Jia-Wang Bian, Zhichao Li, Naiyan Wang, Huangying Zhan, Chunhua Shen, Ming-Ming Cheng, and Ian Reid, and in the paper "Unsupervised scale-consistent depth and ego-motion learning from monocular video" as well as in the method described in the paper "Unsupervised Learning of Depth and Ego-Motion from Video" which appears as arXiv TM Preprint arXiv TM : 1704.07813v2 and is written by Tinghui Zhou, Matthew Brown, Noah Snavely, and David G. Lowe.
[0175] Similarly, the pose network 915b (e.g., when the pose is not determined in parallel with one of the above methods for network 915a) can be a ResNet "encoder"-type network (e.g., ResNet(18) encoder), whose input layer is modified to accept two images (e.g., a 6-channel input that receives image 905a and image 905b as a concatenated RGB input). Then, the bottleneck features of this pose network 915b can be spatially averaged and passed through a 1x1 convolutional layer to output 6 parameters of the relative camera pose (e.g., three for translation and three for rotation for a given three-dimensional space). In some embodiments, another 1x1 head can be used to extract two luminance correction parameters, e.g., as described in the paper "D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry", which is available as an arXiv TM preprint arXiv TM : 2003.01060v2, and is written by Nan Yang, Lukas von Stumberg, Rui Wang, and Daniel Cremers. In some embodiments, each output can be accompanied by an uncertainty value 955a or 955b (e.g., using the method described in the D3VO paper). However, it should be recognized that many embodiments only generate pose and depth data without an accompanying uncertainty estimate. In some embodiments, the pose network 915b can alternatively be PWC-Net, as described in the paper "PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume" by Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz, available as an arXiv TM preprint arXiv TM : 1709.02371v3, or as described in the paper "Towards Better Generalization: Joint Depth-Pose Learning without PoseNet" by Wang Zhao, Shaohui Liu, Yezhi Shu, and Yong-Jin Liu, available as an arXiv TM preprint arXiv TM : 2004.01314v2.
[0176] It should be understood that the pose network can be trained using supervised or self-supervised methods, but with different losses. In supervised training, direct supervision of pose values (rotation, translation) from synthetic data or relative camera poses, such as from a Structure from Motion (SfM) model like COLMAP (described in the paper "Structure-from-motion revisited" by Johannes L. Schonberger and Jan-Michael Frahm in "Proceedings of the IEEE conference on computer vision and pattern recognition. 2016") can be used. In self-supervised training, the photometric loss may instead provide self-supervision.
[0177] Some embodiments may employ an autoencoder and a feature loss as described in the paper "Feature-metric Loss for Self-supervised Learning of Depth and Egomotion", which is available as an arXiv TM preprint arXiv TM :2007.10603v1, and written by Chang Shu, Kun Yu, Zhixiang Duan, and Kuiyuan Yang. Embodiments can supplement this method with differentiable fisheye back-projection and projection, for example, as described in the 2019 paper "Fisheye DistanceNet: Self-Supervised Scale-Aware Distance Estimation using Monocular Fisheye Camera for Autonomous Driving", which is available as an arXiv TM preprint arXiv TM :1910.04076v4, and written by Varun Ravi Kumar, Sandesh Athni Hiremath, Markus Bach, Stefan Milz, Christian Witt, Clément Pinard, Senthil Yogamani, and Patrick written, or in OpenCV TMImplemented in a fish-eye camera model, which can be used to calculate the back-projection of fish-eye distortion. Some embodiments also add a reflection mask during training (and inference) by thresholding the Y channel of a YUV image. During training, the loss values in these masked regions can be ignored and OpenCV TM is used for in-painting, as discussed in the paper "RNNSLAM: Reconstructing the 3D colon to visualize missing regions during a colonoscopy", which was published in "Medical image analysis 72(2021):102100" and was written by Ruibin Ma, Rui Wang, Yubo Zhang, Stephen Pizer, Sarah K. McGill, Julian Rosenman, and Jan-Michael Frahm.
[0178] Given the difficulty of obtaining real-world training data, synthetic data can be used to generate instances of some embodiments. In these example implementations, the depth loss when using synthetic data can be a "scale-invariant loss", as introduced in the 2014 paper "Depth Map Prediction from a Single Image using a Multi-Scale Deep Network", which was published as an arXiv TM preprint arXiv TM :1406.2283v1 and was written by David Eigen, Christian Puhrsch, and Rob Fergus. As mentioned above, some embodiments can adopt a general Structure from Motion (SfM) and Multi-View Stereo (MVS) pipeline COLMAP implementation to additionally learn camera intrinsics (e.g., focal length and offset) in a self-supervised manner, as described in the 2019 paper "Depth from Videos in the Wild: Unsupervised Monocular Depth Learning from Unknown Cameras", which was published as an arXiv TM preprint arXiv TM :1904.04998v1. These embodiments can also learn the distortion coefficients of the fish-eye camera.
[0179] Thus, although networks 915a and 915b are shown separately in pipeline 900a, it should be understood that a single network architecture can be used to perform variants of their two functions. Thus, for clarity, Figure 9B a variant is depicted in which a single network 915c receives all input images 910d (again, the ellipsis 960 here indicates that some embodiments may receive more than two images, but it should be understood that many embodiments will receive only two consecutive images). As previously described, such a network 915c can be configured to output depth predictions 925, pose predictions 930 at 920c, 920d, 920e, 920f, and in some embodiments, one or more uncertainty predictions 955c, 955d (e.g., determining uncertainty as in D3VO, although variants should be readily understood). Separate networks in pipeline 900a can simplify training, yet some deployments can benefit from the simplicity of a single architecture as in pipeline 900b.
[0180] Exemplary end - to - end data processing pipeline - exemplary pose and depth pipeline - exemplary training
[0181] Figure 10A is a flowchart of various operations in an illustrative example neural network training process 1000, e.g., for training each of networks 915a and 915b. At block 1005, the system can receive any synthetic images to be used in training and validation. Similarly, at block 1010, the system can receive real-world images to be used in training and validation. These datasets can be processed at blocks 1015 and 1020 in the inpainting reflection regions and fish-eye boundaries. It should be understood that once deployed, similar preprocessing can occur on images that have not been adjusted in this manner.
[0182] At block 1025, the network can be pre-trained only on synthetic images, e.g., starting from a checkpoint in the FeatDepth network of the "Feature-metric Loss for Self-supervised Learning of Depth and Egomotion" paper, or starting from a checkpoint in the Monodepth2 network of the "Digging Into self-Supervised Monocular Depth Estimation" paper cited above. It should be understood that in the case of using FeatDepth, the autoencoder and feature loss described in that paper can be used. After this pre-training, at block 1030, the network can continue to be trained with data including both synthetic and real data. In some embodiments, COLMAP sparse depth and relative camera pose supervision can be introduced into the training here.
[0183] Figure 10B is a bar plot depicting an exemplary training result set for a Figure 10A process.
[0184] Exemplary segment management
[0185] As discussed with respect to process 500, when the camera encounters sufficiently different regions, e.g., as determined at block 580, the depth frame merging process can be facilitated by organizing the frames into segments (e.g., at block 585a). In Figure 11A an example process for making such a determination at block 580 is depicted. Specifically, after receiving a new depth frame at block 1105a (e.g., as generated at block 555), the computer system can apply a set of rules or conditions to determine whether the depth frame or pose data indicates a new region (contributing to the transition to block 1105e, corresponding to a "yes" transition from block 580), or whether the frame instead indicates a continuation of an existing region (contributing to the transition to block 1105f, corresponding to a "no" transition from block 580).
[0186] In the example depicted, the determination is made by a series of conditions, the satisfaction of any one of which incurs the creation of a new segment. For example, with respect to the condition at block 1105b, if the computer system fails to estimate the pose at either block 550 or block 555 (e.g., when there are not enough values to determine, or no values with an acceptable level of uncertainty), the system can begin the creation of a new segment. Similarly, the condition at block 1105c can be satisfied when there are too few matching features (e.g., SIFT or ORB features) between consecutive frames (e.g., at block 545), e.g., less than an empirically determined threshold. In some embodiments, not only the number of matches can be evaluated, but also their distribution can be evaluated at block 1105c, e.g., by performing a singular value decomposition (SVD) of the depth values organized into a matrix, and then examining the two largest resulting eigenvalues. If one eigenvalue is not significantly larger than the other, the points may be collinear, indicating poor data acquisition. Finally, even if the pose is determined (via the pose from block 550 or from block 555), the condition at block 1105d can be used to "sanity" check whether the pose is appropriate by moving the depth values determined for that pose (at block 555) into an orientation where they can be compared with depth values from another frame. Specifically, Figure 11BThe figure illustrates an endoscope moving 1170 from a first position 1175a to a second position 1175b on a surface 1185, with corresponding fields of view 1175c and 1175d respectively. There is an overlap in depth values between the expected regions 1180, as shown by the portion 1180 of the surface 1185. The overlap in depth values can be verified by moving the values in one acquisition to their corresponding positions in another acquisition (as considered at block 1105d). The lack of similar depth values within a threshold can indicate a failure to obtain an appropriate pose or depth determination.
[0187] It should be understood that while the conditions of blocks 1105a, 1105b, and 1105c can be used to identify when the endoscope has moved into a field of view that is sufficiently different from the one it was previously in, these conditions can also indicate when smoke, biomass, body structures, etc. are obscuring the camera's field of view. For the convenience of the reader in understanding these latter situations, example situations that contribute to this result are shown in the time series of cross-sectional views in Figure 11C During an examination, the endoscope may often collide with parts inside the body. For example, initially at time 1110a, the colonoscope can be in a position 1120a with a field of view suitable for pose determination (similar to the previous discussion regarding Figures 4A - 4C ). Unfortunately, patient movement, inadvertent movement by the operator, etc. may shift the configuration 1110d to a new time state 1110b of the surface area 1115b of the main acquisition ridge, in which the camera collides with the ridge wall 1115a, resulting in a substantially blocked view. Naturally, in this orientation 1120b, the endoscope camera acquires few (if any) pixels useful for any appropriate pose determination. When the automated examination system or the operator resumes 1110e at time 1110c, the endoscope can again be in a position 1120c with a field of view suitable for pose and depth determination.
[0188] It should be understood that even if such a collision occurs only within a few seconds or less, the high frequency of the camera acquiring visual images can result in many new visual images. Thus, the system can attempt to generate many corresponding depth frames and poses, which can themselves be assembled into segments according to process 500. Such undesirable segments can be excluded through the process of global pose graph optimization at block 585b and the integration at block 570. Fortunately, this exclusion process itself can also help in detecting and identifying various adverse events during the procedure.
[0189] Specifically, Figure 11Dis a schematic collection of segments 1125a, 1125b, and 1125c. Segment 1125a may be generated when the colonoscope is at position 1110a, segment 1125b may be generated when the colonoscope is at position 1110b, and segment 1125c may be generated when the colonoscope is at position 1110c. As discussed, each of segments 1125a, 1125b, and 1125c may respectively include initial key frames 1130a, 1130e, and 1130f (here, the key frame is the first frame inserted into the segment). Thus, for clarity, the first frame of segment 1125a is key frame 1130a, frame 1130b is the next acquired frame, and so on (intermediate frames are represented by ellipsis 1130d) until the final frame 1130c is reached. During global pose estimation at block 585b, the computer system may have identified sufficient features (e.g., SIFT or ORB) or depth frame similarity between key frames 1130a and 1130f such that they can be identified as depicting a connected region of depth values (represented by link 1135c). Given the similarity of the fields of view at times 1110a and 1110c, this is not surprising. However, the radical character of the field of view at time 1110b makes key frame 1130e too different from either key frame 1130a or 1130f to form a connection (represented by non-existent links 1135a and 1135b).
[0190] Thus, as Figure 11E shown in the hypothesized graphical pose network of, viable segments 1140a, 1140b, 1140c, 1140d, 1140e, and 1125a and 1125c may form a network with reachable nodes based on their associated key frames, but segment 1125b may remain isolated. It should be understood that frame 1125b may sometimes coincidentally match other frames (e.g., in the case where there are multiple defective frames caused by the camera pressing against a flat surface, they may all be similar to each other), but these defective frames will generally form a smaller, isolated (or more isolated) network with the main network acquired corresponding to the internal body structure. Thus, at block 570, such frames can be easily identified and removed from the model generation process.
[0191] Although not shown in Figure 11D , it should be understood that in addition to depth values, each frame in the segment may also have various metadata, including for example (one or more) corresponding visual images, (one or more) estimated poses associated therewith, (one or more) timestamps at which the acquisition occurred, etc. For example, as Figure 11FAs shown, segments 1150a and 1150b are two of many segments that appear in the network (the presence of previous, subsequent, and intermediate segments represented by ellipses 1165a, 1165c, and 1165b, respectively). Segment 1150a includes frames 1150c, 1150d, and 1150f (ellipsis 1150e reflects intermediate frames), and the frame 1150c acquired at the first time is designated as the key frame. Based on the frames in segment 1150a, an intermediate model such as the TSDF representation 1155a can be generated (similarly, an intermediate model such as TSDF 1155b can be generated for the frames of segment 1150b). When such an intermediate TSDF is available, integrating the segments into a partial or complete model mesh (or remaining in TSDF form) 1160 can be performed very quickly (e.g., at box 570 or integration 780), which is useful for facilitating real-time operations during surgery.
[0192] Overview - Reference geometry for surgical instrument kinematics
[0193] Various disclosed embodiments provide precise metrics for monitoring the kinematics of a surgical instrument relative to a reference geometry, such as a geometry or manifold within the Euclidean space in which a three-dimensional model embedded within a patient lies. Although colonoscopy will generally be specifically referenced herein for ease of consistent presentation for the reader's understanding, it should be understood that, with necessary modifications, many of the disclosed systems and methods are applicable to other surgical environments, such as prostatectomy, bronchopulmonary analysis, general laparoscopic procedures, and the like. Thus, various embodiments can be applied to, for example, a surgical instrument that navigates along the lungs to detect polyps in a pre-procedure computed tomography (CT) scan. As in the colonoscopy context, the system can estimate the centerline geometry of the route to be navigated during such a lung procedure.
[0194] Figure 12A is a schematic cross-sectional view of a colonoscope 1205a and the resulting colonoscope camera view 1205e within a portion of the colon 1205d during movement of advancing 1205b (further into the colon away from the insertion point (e.g., the anus)) or withdrawing 1205c (moving the colonoscope back towards the insertion point), which can occur in some embodiments. The nature of the movement of the colonoscope 1205a within the colon can have an impact on the quality and characteristics of the surgical procedure. For example, too fast a movement of the colonoscope 1205a within the colon can result in motion blur, as shown in view 1205e.
[0195] The importance of such movement can also depend on spatial or temporal factors. For example, spatially, movement of a colonoscope too close to the lateral wall of the colon may be less desirable near an injured area of the colon than en route to that section through a healthy area. Regarding the temporal context, a higher movement speed during insertion may be appropriate, where the priority is to reach and examine the area of interest, but the same speed may be inappropriate during withdrawal past areas not fully examined during advancement. Such movement profile thresholds can be determined, for example, by key opinion leaders (KOLs). Thus, a rigorous and precise system for monitoring these and other situations will be desirable for many downstream operations, including machine learning operations as well as simply notifying the surgical operator of the current kinematic behavior of the instrument.
[0196] To provide robust kinematic data to facilitate such considerations across surgical procedures, various embodiments contemplate creating reference geometries, such as lines, circles, hemispheres, spheres, etc., within the same Euclidean space that can represent inside the patient. For example, Figure 12B is a schematic three-dimensional model 1210a of a portion of an organ having an intermediate centerline axis reference geometry 1210b (e.g., created using the positioning and mapping systems and methods described herein). Here, the centerline 1210b of the three-dimensional model of the organ can be used as a consistent reference for interpreting the kinematics of a surgical instrument. The centerline can be the central axis of all or a portion of the model along the length of the model. Movements along or with respect to the centerline 1210b and movements orthogonal or residual to the centerline 1210b can be considered. It should be understood that there are multiple methods for creating the centerline 1210b. For example, a colon model 1210a can be averaged or collapsed. However, as discussed in more detail herein, (e.g., reference Figure 14 ), some embodiments use an iterative segment-based method to determine the centerline 1210b, which can produce a centerline 1210b that is generally invariant to the complexity of the colon lateral wall surface. This invariance can be particularly useful given the wide physical differences between the lungs, colon, esophagus, etc. of different patients.
[0197] It should be understood that this comparison of the movement of the instrument relative to the centerline 1210b can occur during or after a surgical procedure (e.g., when the centerline is created at the end of a surgical procedure and then the previously recorded positions of the surgical instruments are used to determine relative kinematics). However, it is often desirable to create the centerline in real time during a surgical procedure as this can facilitate direct kinematic feedback to the surgical team. It should also be understood that, as described in more detail herein, while the same reference geometry can be used to evaluate an entire surgical procedure in some embodiments, in other embodiments the reference geometry can vary over time, e.g., as an organ deforms, as the context and requirements change, etc., and more than one geometry can be used simultaneously.
[0198] For ease of reader understanding, Figure 12C is a schematic side perspective view of a colonoscope camera 1215a in a forward orientation relative to the centerline reference geometry 1215b. Specifically, when the colonoscope moves directly along the centerline 1215b in the direction of the vector 1215c, the projection 1215d of that vector 1215c onto the centerline will be the same vector 1215e. Thus, if the colonoscope moves precisely on or parallel to the centerline forward, its velocity vector in Euclidean space can be the same vector in direction and magnitude on the centerline. In this example, the closest point on the centerline to the previous position of the camera is point 1215f, and the closest point to its new position is point 1215g.
[0199] Similarly, as Figure 12D shown, a withdrawal motion vector 1220c of a colonoscope camera 1220a above the centerline 1220b (at a distance indicated by the reference line 1220f) can result in a projection 1220d vector 1220e on the centerline 1220b that is the same as the withdrawal motion vector 1220c. Thus, in the case where the movement of the colonoscope camera is parallel to the centerline, then the movement vector of the camera, whether advancing on the centerline axis (as Figure 12C shown) or advancing away from the centerline axis, or withdrawing on the centerline axis or withdrawing away from the centerline axis (as Figure 12D shown), the speed of the camera's movement will be the same as the speed on the centerline. In this example, the closest point on the centerline to the previous position of the camera is point 1220g, and the closest point to its new position is point 1220f.
[0200] In contrast, for further clarity, it should be understood that, as Figure 12EAs shown, the movement of the colonoscope camera that is not parallel to the centerline can result in relative movement projected onto the centerline that is different from the movement of the actual camera in three-dimensional space. Specifically, in this example, camera 1225a is in a non-parallel orientation above centerline 1225b, and thus its forward movement in this orientation produces motion vector 1225c. As indicated by reference line 1225d, the projection 1225e of vector 1225c onto centerline 1225b results in a smaller movement vector 1225f. Here, point 1225g is the closest point on the centerline to the new position of the camera, and point 1225h is the closest point at the previous orientation of the camera. As will be discussed in more detail herein, in some embodiments, a portion of the projection of vector 1225c that does not appear on the centerline (e.g., movement along reference line 1225d), referred to as residual kinematics, can also be determined because this movement can be significant in various scenarios (e.g., leaving the centerline and approaching the sidewall at an inappropriate time or orientation).
[0201] Again, for further clarity, it should be understood that, as Figure 12F shown, there can be forms of residual movement that are not orthogonal to the centerline. In this example, colonoscope 1230a is neither advanced, retracted, nor moved laterally with respect to centerline 1230b, but only rotated 1230c. Thus, since there is no translational component, the projection 1230d does not produce a vector on centerline 1230b and no relative kinematic data is produced. However, some embodiments can monitor the orientation of the center of the camera's field of view relative to centerline 1230b because the relationship between these two vectors can provide information about the area of the colon sidewall being examined. Thus, even if the translational position of the camera does not change between two consecutive image acquisitions in Figure 12F and thus the closest point on the centerline for the two frames remains point 1230e, the system can note a change in the relative angle between the centerline and the center of the field of view in the two orientations (e.g., using the vector product or dot product of the two vectors) as part of the residual kinematics.
[0202] For further clarity, some embodiments determine the movement of a surgical instrument, such as a camera, based on its translation or rotation above a certain threshold relative to a previous valid frame (e.g., as described in Figure 12F ). For example, Equation 1 indicates how the translational movement of the camera can be evaluated:
[0203] Translational movement = ||T t - T 上一个有效帧 || 2 (1)
[0204] where T t is the translational vector relative to the global origin of the camera at the current time t, and T上一个有效帧 is the translation vector of the camera relative to the global origin at the time of acquisition of the previous valid frame.
[0205] Then the rotation of the camera can be determined according to Equation 2
[0206]
[0207] where
[0208]
[0209] If the translation exceeds a threshold or the rotation exceeds a threshold, then motion can be detected. Once it is detected that the rotational motion has exceeded the threshold, the system can determine the relationship between the rotated field of view and the axis vector of the center line.
[0210] It should be understood that the center line may not always be in the form of a "straight line", for example, in the case where the colon presents a curved structure. Thus, Figure 12G a curved center line 1235f is depicted. It should be understood that, as in the previously discussed figures, the combination of translation and rotation of the colonoscope relative to the center line can still produce a cumulative relative projection result on the center line. Here, for example, the translation 1235c and rotation 1235d movement from the first orientation 1235a to the second orientation 1235b can result in the cumulative projection of the vector 1235e on the center line 1235f. Thus, the point 1235g on the center line is closest to the camera in the orientation 1235a, and the point 1235h on the center line is closest to the camera in the orientation 1235b. Here, the system can record the relative translation of the projection along the center line manifold and the residual change in the rotation of the camera. It should be understood that the manifold herein refers to a three-dimensional object of a line or surface embedded in a Euclidean space having a projection on which a surgical instrument can move.
[0211] Naturally, the rate of comparison of the camera orientations can affect the granularity of the projection movement on the center line. In some embodiments, the comparison rate can be the same as the frame rate at which the camera acquires images. Generally, the acquisition rate can be fast enough such that the relative and residual kinematic data have sufficient quality. However, as Figure 12HAs shown, in some embodiments and situations, it may be desirable to interpolate the projection positions on the centerline in order to infer kinematics at a higher resolution. For example, if the projection position of the camera on the centerline 1240a at a first time corresponds to point 1240b (e.g., corresponding to point 1235g or point 1220g determined at the first time), and at the next acquisition moment, the projection on the centerline 1240a is at point 1240d (e.g., corresponding to point 1235h or point 1220f determined at a second time after the first time), rather than inferring a straight-line motion in Euclidean space from point 1240b to point 1240d, the system can interpolate the movement along the centerline manifold. Thus, the projected motion between points 1240b and 1240d can pass through point 1240c. In cases where an encoder and other mechanical sensor configurations are available, the system can compare the projected and interpolated centerline motion with the centerline motion derived from the encoder. However, it would generally be beneficial to infer motion only from the camera images, and thus a frame rate commensurate with the expected maximum speed of the surgical instrument can be selected to ensure that all desired motions are captured. Generally, motion that is too fast for the accurate determination of the reference projection may also be too fast for the proper depth frame determination.
[0212] It should be understood that this interpolation can equally occur for rotation. That is, in the case where the rotation of the camera relative to the centerline at a first acquisition time associated with point 1240b is different from the relative rotation at a later acquisition time associated with point 1240d, the system can record any intermediate value as a linear interpolation between the two (e.g., obtaining a dot product from the corresponding portion of the interpolated centerline).
[0213] In some embodiments, the proper determination of a reference geometry (such as a centerline) and the continuous orientation of the surgical instrument relative thereto can enable many useful downstream actions and evaluations. For example, Figure 12I is a schematic perspective view of an orientation on a centerline reference geometry 1245a having a radial spatial context that can occur in some embodiments. Here, the centerline 1245a passes through the region of interest, particularly the disease artifact 1245e, such as a polyp, tumor, etc. The regions around the centerline 1245a can be associated with different context functions. For example, a first region 1245g around the centerline 1245a can indicate an upper limit of movement when advancing or withdrawing the colonoscope. During these stages of the surgical procedure, moving the colonoscope outside of this region can trigger a warning or alarm. Similarly, region 1245f can be used for the same purpose in a wider region of the colon. Thus, during this "traveling" phase of the procedure, the colonoscope at position 1245h can be urged to advance along the appropriate vector 1245b.
[0214] This situational space and orientation monitoring need not be limited to regions extending radially from a centerline. Motion orthogonal or away from the centerline can equally be considered. The system can consider not only changes in orientation relative to the nearest part of the centerline, but also changes in orientation relative to parts of the centerline encountered previously in a surgical procedure or to be encountered in future operations. For example, as Figure 12I depicted in, when encountering artifact 1245e, deviating from each of centerline paths 1245c and 1245d may be better than maintaining orientation 1245b on the path, as they will provide a more direct and closer view of artifact 1245e. However, turning back in path 1245c to sense the artifact may not be as desirable as approaching the artifact with a smaller deviation from the centerline in path 1245d. Thus, high-fidelity reference geometry facilitates not only precise kinematic metrics on the geometry itself, but also situational metrics external to the geometry, such as the orthogonal-radial and situational awareness metrics described herein.
[0215] Again, while many of the embodiments disclosed herein are described consistently with reference to a colonoscopy context for clarity of understanding, it should be understood that other embodiments may be applicable to other contexts and in conjunction with other surgical instruments. For example, Figure 12J is a schematic cross-sectional view of a patient's pelvic region 1250e during a robotic surgical procedure that may occur in some embodiments. Here, a first surgical instrument 1250b and a second surgical instrument 1250c may be inserted via respective ports into a laparoscopic insufflation cavity 1250a inside the patient. A reference geometry embedding manifold may be determined within the Euclidean space of cavity 1250a, including a centerline, surfaces around regions of interest, etc. Here, a central sphere 1250d at the center of cavity 1250a provides a manifold on which the motion of one or both of instruments 1250b and 1250c is projected.
[0216] The reference geometry embedding manifold may be selected based on the structure of the modeled internal region of the patient, the nature of the surgical procedure, or both. For example, in Figure 12K a cylindrical reference geometry 1255b is located at the center of a three-dimensional model of cavity 1255a, and the elongate axis of the cylinder is oriented relative to the region of interest such that the projection of the movement of the surgical instrument can provide information relevant to the surgical procedure under consideration. Similarly, as Figure 12KAs shown, the reference geometry can be oriented with an appreciation of the structure of the internal region of the patient. For example, here, when the reference geometry is the sphere 1260b, the reference geometry can maintain a consistent orientation across the surgery relative to landmarks within the three-dimensional model of the cavity 1260a. That is, the upward-pointing axis 1260c shown here will point upward in other models as well. In this way, the motion of the surgical instrument projected onto the surface of the sphere 1260b can be easily compared across surgeries.
[0217] Exemplary mid - axis centerline estimation - system process
[0218] Naturally, a more precisely and consistently generated reference geometry (such as a centerline) can better enable more precise operations, including, for example, perimeter selection and assessment of surgical instrument kinematics. This consistency can be useful when analyzing and comparing surgical procedure performance. Thus, with specific reference to the example of creating a centerline reference geometry in the context of a colonoscopy, various embodiments envision improved methods for determining a centerline based on a localization and mapping process, as described previously herein, for example.
[0219] For the convenience of the reader, Figure 13A is a schematic three-dimensional model of the colon 1305a. As described above, during a surgical procedure, the colonoscope can start at a position and orientation 1305c within the colon 1305a and be advanced forward 1305d, collecting depth frames and iteratively generating a model (e.g., as discussed with respect to Figure 7 ), until a terminal position 1305b is reached (however, in some embodiments, the localization and mapping can occur only during withdrawal). During withdrawal 1305e, the trajectory can be largely opposite to the trajectory of the advancement 1305d, where the colonoscope starts at a position and orientation 1305b at or near the cecum and then ends at a position and orientation 1305c. During withdrawal 1305e, additional depth frame data acquisition can help improve the fidelity of the three-dimensional model of the colon (and thus any reference geometry derived from the model, such as when the centerline is estimated as the central moment of the model perimeter).
[0220] While some embodiments attempt to determine the centerline and corresponding kinematics throughout both the advancement 1305d and the withdrawal 1305e, in some embodiments, the reference geometry can be determined only during the withdrawal 1305e when at least a preliminary model is available to assist in the creation of the geometry. In other embodiments, the system can wait until after the surgery, when the model is complete, and then determine the centerline and corresponding kinematic data from a recording of the motion of the surgical instrument.
[0221] Created by approaching the centerline via an iterative method, where the centerline of the depth frame considered locally is first created and then combined with the existing global centerline estimate of the model, it may be possible to determine a reference geometry for kinematic feedback during advancement 1305d, during retraction 1305e, or during a post-surgical review. For example, during advancement 1305d or retraction 1305e, projections on the reference geometry can be used to inform the user that their movement is too fast. Such a warning can be provided and be sufficient even if the available reference geometry and model are not as accurate as they will be when the mapping is fully complete. In contrast, higher-fidelity operations, such as comparing the performance of the surgeon to that of other practitioners, can only be performed after a higher-fidelity representation of the reference geometry and model is available. Access to the lower-fidelity representation can still be sufficient for real-time feedback.
[0222] Specifically, Figure 13B is a flowchart illustrating various operations in an example intermediate centerline estimation process 1310 that can be implemented in some embodiments, which facilitates the iterative merger of local centerline determination and global centerline determination. Specifically, at block 1310a, the system can initialize a global centerline data structure. For example, at the position and orientation 1305b before retraction 1305e, if the centerline has not been created, the system can prepare the first endpoint of the centerline as the current position of the colonoscope or as the position at 1305c, which extends to the average of the model sidewalls. Conversely, if the centerline has been created during advancement 1305d, that previous centerline can be considered the current initialized global centerline. Finally, if data acquisition has just begun (e.g., before advancement 1305d) and the colonoscope is at position and rotation 1305c, the global centerline endpoint can be the current position of the colonoscope, with a small extension along the axis of the current field of view. As will be discussed in more detail with respect to Figure 14 a machine learning system for determining the local centerline from the model TSDF can be employed during initialization at block 1305a.
[0223] At block 1310b, the system can iterate over the acquired pose of the surgical camera (e.g., as they are received during advancement 1305d or retraction 1305e) until all poses have been considered, and then publish a "final" global centerline at block 1310h (however, of course, an intermediate version of the global centerline, such as that determined at block 1310i, can be used to determine kinematics). Each camera pose considered at block 1310c can be, for example, the most recent pose acquired during advancement 1305d, or the next pose to be considered in a queue of poses sorted in chronological order by their acquisition time.
[0224] At block 1310d, the system can determine the closest point on the current global centerline to the position with respect to the pose considered at block 1310c. At block 1310e, the system can consider model values (e.g., voxels in TSDF format) within a threshold distance of the closest point determined at block 1310d, herein referred to as the "segment" associated with the closest point on the centerline determined at block 1310d. In some embodiments, dividing the expected colon length by the depth resolution and multiplying by the expected examination interval (e.g., 6 minutes) can indicate an appropriate distance around the point for determining segment boundaries, as this distance corresponds to an appropriate examination "effort" for the operator to examine the area.
[0225] For clarity of illustration, referring to Figure 13C , a global centerline 1325c may have been generated for a portion 1325a of the model of the colon. The model itself may still be in TSDF format and may be represented accordingly in a "heatmap" or other voxel format. A portion 1325b of the model may not yet have a centerline, e.g., because that portion of the model does not yet exist, as during advancement 1305d, or may exist, but may not yet have been considered for centerline determination (e.g., during post-processing after the procedure).
[0226] Thus, the next pose 1325i can be considered (here, represented as an arrow in three-dimensional space, which corresponds to the position and orientation of the camera looking at the upper wall of the colon), e.g., when poses are acquired and selected chronologically at block 1310c. The point on the centerline 1325c closest to this pose 1325i, as determined at block 1310d, is point 1325d. The segment is then the portion of the TSDF model within the threshold distance of point 1325d, shown here as the TSDF values that appear in region 1325e (also shown separately for ease of understanding by the reader). Thus, the segment can include all, a portion, or none of the depth data acquired via pose 1325i. At block 1310f, the system can determine a "local" centerline 1325h of the segment in region 1325e, including its endpoints 1325f and 1325g. The global centerline (centerline 1325c) can be extended at block 1310i with this local centerline 1325h (which can cause point 1325f to now become the farthest endpoint of the global centerline relative to the starting point 1325j of the global centerline). As will be described with respect to Figure 14As discussed in more detail, in some embodiments, at block 1310g, the system may consider whether the pose-based local centerline estimate at block 1310f has failed, and if so, apply an alternative method for local centerline determination (e.g., apply a neural network and centerline determination logic) at block 1310h. Such alternative methods, while more robust and accurate than pose-based estimation, may be too computationally intensive for continuous use during real-time applications (such as during a surgical procedure).
[0227] It should be understood that there are various methods for performing the operations of block 1310f. For example, Figure 13D is a flowchart illustrating various operations in an example process 1315 for estimating such a local centerline segment. As will be described in more detail herein with reference to Figure 14 a pose-based local centerline estimate for a given segment generally may include three operations, outlined here at blocks 1315a, 1315b, and 1315c. At block 1315a, the system may construct a connectivity graph for the poses that appear in the segment (e.g., the most recent pose in front of the field of view during retraction 1305e, or the most recent pose behind the field of view during advancement 1305d). The connectivity graph may be used to determine the spatial ordering of the poses before fitting the local centerline. For each pose, a "breadth-first search" may be used to calculate the shortest distance along the graph to the "oldest" (as per the time of acquisition) pose, and then the order may be determined based on these distances. The closest pose in the graph may be selected as the first pose in the ordering, the second-closest pose in the graph may be selected as the second pose in the ordering, and so on.
[0228] Using this graph between the poses, at block 1315b, the system may then determine the extreme poses (e.g., those extreme voxels most likely corresponding to points 1325f and 1325g), the ordering of the poses along the path between these extreme points, and the corresponding weights associated with the path (e.g., weights based on the TSDF density of each voxel). The order and other factors (such as pose proximity) may also be used to determine the weights for interpolation (e.g., as constraints for fitting a spline). A least-squares fit, using B-splines, etc., may also be used to estimate the local centerline.
[0229] Finally, at block 1315c, the system may determine the local centerline 1325h based on, for example, a least-squares fit (or other suitable interpolation, such as a spline) between the extreme endpoint poses determined at block 1315b. Determining the local centerline based on such a fit can facilitate a better centerline estimate than if the process continued to be tied to the discrete orientations of the poses. The resulting local centerline may later be merged with the global centerline, as described herein (e.g., at blocks 1310i and process 1320).
[0230] Similarly, many methods can be used to implement the operations of block 1310i. For example, Figure 13E is a flowchart of various operations in an example process 1320 that can be implemented in some embodiments for extending (or, with necessary modifications, updating a pre - existing portion) a global centerline (e.g., global centerline 1325c) using a local centerline of a segment (e.g., local centerline 1325h). Here, at block 1320a, the system can determine a first "array" of points on the local centerline (a sequence of consecutive points along the longitudinal axis) and a second array of points on the global centerline, e.g., points within 0.5 mm (or other suitable threshold, e.g., adjusted according to the empirically observed colonoscope speed) of each other. Although such arrays can be determined for the entire lengths of the local and global centerlines, some embodiments determine the arrays only for portions (e.g., 1325e) that appear in or near the region under consideration. As will be described in Figure 14 an additional 1 cm - worth of points can be intentionally extended relative to the global centerline as a buffer for the array of the local centerline.
[0231] At block 1320b, the system can then identify which pair of points (one from each of the two arrays) has the spatially closest pair of points relative to the other pairs, and each of the pairs of points so identified is referred to herein as an "anchor". Thus, the anchors can be selected as those points where the local and global arrays are closest corresponding. At block 1320c, the system can then determine the weighted average between the pairs of points from the anchor point to the end of the array of the local centerline (e.g., including the 1 cm buffer). In some embodiments, the weighted average between these pairs of points can include the anchor point itself, although the anchor point can only indicate the end point for the weighted average determination. Finally, at block 1320d, the system can then determine the weighted average of the local centerline and the global centerline around that anchor point.
[0232] Exemplary mid - axis centerline estimation process - schematic pipeline
[0233] To better facilitate the reader's understanding of the Figures 13A - 13E example scenarios and processes Figure 14Many of the same operations are presented in the illustrative operation pipeline, this time in the context of an embodiment where localization, mapping, and reference geometry estimation are applied only during withdrawal. Specifically, in this example, the operator has advanced the colonoscope to the starting position without initiating centerline estimation (e.g., the examination of the colon can occur only during withdrawal, where kinematics are most relevant, and thus the operator is at least initially only concerned with maneuvering the colonoscope to the appropriate starting position), and then centerline estimation is performed throughout the withdrawal. Again, in some embodiments, model creation may have occurred during advancement, and the centerline may be created from all or only a portion of the model. However, in the depicted example, the centerline is calculated only during withdrawal, and where possible, the pose is used rather than relying on the fidelity of the model.
[0234] As shown, after the pipeline begins, the operator has advanced the colonoscope from the initial starting position 1405d within the colon 1405a to the cecum and to the final position 1405c facing the cecum. From this final position 1405c, the operator can begin to withdraw the colonoscope along path 1405e. Having reached the cecum and prior to withdrawal, the operator or other team member may manually indicate to the system (e.g., via a button press) that the current pose is at the terminal position 1405c facing the cecum. However, in some embodiments, automated system recognition (e.g., using a neural network) may be used to automatically identify the position and orientation of the colonoscope within the cecum, thus facilitating automatic initialization of the reference geometry creation process.
[0235] According to box 1310a, the system can initialize the centerline here by obtaining depth values of the cecum 1405b. These depth values (e.g., in the TSDF format and appropriately organized for input into a neural network) can be provided 1405g to the "Voxel-Complete Local Centerline Estimation" component 1470a, where a neural network 1420 is included for ensuring that the TSDF representation is in the appropriate form for centerline estimation and post-completion logic in box 1410d. Specifically, while holes can be filled by direct interpolation, planar surfaces, etc., in some embodiments, a flood-fill type neural network 1420 can be used (e.g., similar to the network described in Dai, A., Qi, C.R., Nieβner, M.'s "Shape completion using 3d-encoder-predictor cnns and shape synthesis" (In: Proc. Computer Vision and Pattern Recognition (CVPR), IEEE (2017)); it should be understood that here "conv" refers to the convolutional layer, "bn" refers to batch normalization, "relu" refers to the rectified linear unit, and the arrows indicate the concatenation of the layer output and layer input).
[0236] For example, in the TSDF voxel space 1415a (e.g., a 64x64x64 voxel grid), a segment 1415c with holes on its side is shown (e.g., a part of the colon that has not been properly observed in the field of view for mapping). Those familiar with the voxel format should understand that a larger area 1415a can be subdivided into cubes 1415b, which are referred to as voxels herein. While in some embodiments the voxel values can be binary (indicating empty space or the presence of the model), in some embodiments, the voxels can adopt a series of values similar to a heatmap, e.g., where the values can correspond to the probability that a part of the colon appears in a given voxel (e.g., between 0 (for free space) and 1 (for high confidence in the presence of the colon sidewall)).
[0237] For example, the voxels input 1470b into the voxel point cloud completion network can take values according to Equation 4:
[0238] H 输入 [v] = tanh(0.2 * d(v, S0)) (4)
[0239] And the output 1470c can take values according to Equation 5:
[0240]
[0241] In each case, where H[v] refers to the heatmap value of voxel v, d(v, S0) is the Euclidean distance between voxel v and the voxelized partial segment S0, d(v, S1) is the Euclidean distance between voxel v and the voxelized complete segment S1, and d(v, C) is the Euclidean distance between voxel v and the voxelized estimated global centerline C. In this example, the input heatmap is zero at the location of the (partial) segment surface and increases towards 1 away from it, while the output heatmap is zero at the location of the (complete) segment surface and increases towards 1 at the location of the global centerline (converging to 0.5 elsewhere).
[0242] For clarity, if the isolation plane 1415d is observed in region 1415a, it will be seen that the model 1415e is associated with many voxel values, but the region with the hole contains voxel values similar or identical to empty space. By inputting region 1415a into the neural network 1420, the system can produce 1470c an output 1415f with a filled TSDF section 1425a, including filling of the missing region. Thus, the planar cross-section 1415d of the voxel region 1415f is shown here as having filled voxels 1425b. Naturally, such a network can be trained from a dataset created by: collecting true positive model segments, excising parts according to situations frequently encountered in practice, then providing the latter as input to the network, and the former for validating the output.
[0243] Then a portion of the filled voxel representation of segment 1415f can be approximated at box 1410d corresponding to the local centerline orientation within the segment. For example, the voxel representation can be filtered to identify the centerline portion, for example, as in equation 6, by discriminating voxels having values above a threshold:
[0244] Voxel value > 1 – δ (6)
[0245] where δ is an empirically determined threshold (e.g., taking a value of approximately 0.15 cm in some embodiments).
[0246] For clarity, the operation result of the "voxel-integrity-based local centerline estimation" component 1470a (including the post-processing box 1410d) will be the local centerline 1410a of the filled segment 1425a (where, for clarity, the terminal endpoints 1410b and 1410c are explicitly shown here). During the initialization of the box 1310a, since there is no pre-existing global centerline, there is no need to integrate the local centerline determined for the cecal TSDF 1405b with the "voxel-integrity-based local centerline estimation" component 1470a via the local-to-global centerline integration operation 1490 (corresponding to the operations of the box 1310i and the process 1320). Instead, the local centerline of the cecal TSDF is the initial global centerline.
[0247] Now, as the colonoscope is withdrawn along the path 1405e, the localization and mapping operations disclosed herein can identify the colonoscope camera poses along the path 1405e. Local centerlines can be determined for these poses and then integrated with the global centerline via the local centerline integration operation 1490. In theory, each of these local centerlines can be determined by applying the "voxel-integrity"-based local centerline estimation component 1470a to each of its corresponding TSDF depth grids (and in fact, this method can be applied to some cases, such as post-surgical viewing, where computational resources are readily available). However, this method may be computationally expensive, complicating real-time applications. Similarly, certain unique grid topologies may not always be suitable for application to such components.
[0248] Thus, in some embodiments, pose-based local centerline estimation 1460 is typically performed. When complications arise or metrics indicate that the pose-based method is insufficient (e.g., the determined centerline is too close to the sidewall), as determined at block 1455b, the faulty pose-based result can be replaced with the result from component 1470a. At block 1455b, the system can, for example, determine whether the error between the interpolated centerline and the pose used to estimate the centerline exceeds a threshold. Alternatively or additionally, the system can periodically perform an alternative local centerline determination method (such as component 1470a) and check for consistency with the pose-based local centerline estimation 1460. A lack of consistency at block 1455b (e.g., the sum of the differences between centerline estimates is higher than a threshold) can then contribute to a failure determination. Although component 1470a may be more accurate than the pose-based local centerline estimation 1460, component 1470a may be computationally expensive and thus its consistency verification can be run infrequently and in parallel with the pose-based local centerline estimation 1460 (e.g., for the first lack of consistency in an estimation sequence, component 1470a can then be applied to every other frame or some other suitable interval in the sequence, and the results interpolated until the performance of the pose-based local centerline estimation 1460 improves).
[0249] Thus, for clarity, after initially applying component 1470a to the TSDF 1405b of the cecum, a withdrawal can be made along path 1405e, applying the pose-based method 1460 until region 1405f is encountered. If the pose-based local centerline estimation fails in this region 1405f, the TSDF of region 1405f and any successive faulty regions can be provided to component 1470a until the global centerline is sufficiently improved or corrected, and the pose-based estimated local centerline estimation method 1460 can resume for the remainder of the withdrawal path 1405e.
[0250] At box 1455a, which is consistent with box 1310b, as the operator withdraws and extends the global centerline along path 1405e, the system can continue to receive poses, where each local centerline is associated with each new pose. More specifically, and discussed with reference to box 1310f and process 1315, the pose-based local centerline estimation 1460 can be performed as follows. As the colonoscope is withdrawn in direction 1460a through colon 1460b, as mentioned, it will generate multiple corresponding poses during positioning, here represented as white spheres. For example, pose 1465a and pose 1465b correspond to the previous positions of the colonoscope camera when withdrawn in direction 1460a. Various of these previous poses may have been used to create the global centerline 1480a in its current form (the ellipsis at the leftmost part of centerline 1480a indicates that it can extend to the starting position of the pose corresponding to position 1405c in the cecum).
[0251] Having received a new pose (here shown as black sphere 1465h), the system can seek to determine the local centerline, here shown in exaggerated form via dashed line 1480b. Initially, the system can identify the previous poses within a threshold distance of the new pose 1465h, here represented as poses 1465c - 1465g that appear within bounding box 1470c. Although only six poses appear in the box in this schematic example, it should be understood that more poses will be considered in practice. According to process 1315, the system can construct a connectivity graph between poses 1465c - 1465g and the new pose 1465h (box 1315a), determine the extreme poses in the graph (box 1315b, here poses 1465c and the new pose 1465h), and then determine the new local centerline 1480b as a least squares fit, spline, or other suitable interpolation between the extreme poses, weighted by intermediate poses (box 1315c, i.e., as shown, the new local centerline 1480b is an interpolation line between the extreme poses 1465c and 1465h weighted based on the intermediate poses 1465d - 1465g in the order identified at box 1315b, such as a spline with poses as constraints).
[0252] Assuming that the pose-based centerline estimation of method 1460 successfully produces a viable local centerline and thus there is no fault determination at box 1455b (corresponding to decision box 1310g), the system can transition to the local and global centerline integration method 1490 (e.g., corresponding to box 1310i and process 1320). Here, in the initial state 1440a, the system can seek to integrate the local centerline 1435 (e.g., corresponding to the local centerline 1480b as determined via method 1460 or the centerline 1410a as determined by component 1470a) with the global centerline 1430 (e.g., global centerline 1480a). It should be understood that the local centerline 1435 and the global centerline 1430 are shown here with a vertical offset for ease of reader understanding and in practice can more easily overlap without such an exaggerated vertical offset.
[0253] As discussed with respect to box 1320a, the system can select points on each centerline (shown here as squares and triangles) and organize them into an array. Here, the system has produced a first array of eight points for the local centerline 1435, including points 1435a - 1435e. Similarly, the system has produced a second array of points for the global centerline 1430 (again, it should be understood that the array may not be determined for the entire global centerline 1430 but only for that terminal region near the local centerline to be integrated). Comparing the arrays, the system has identified corresponding pairs of points in their array positions, specifically, each of points 1435a - 1435d corresponds to each of points 1430a - 1430d, respectively. In this example, the correspondence is offset such that the point 1435e corresponding to the most recent point on the local centerline (e.g., corresponding to the new pose 1465h) is not included in the corresponding pair. It should be understood that since the relationship can be inherent in the array ordering, the corresponding relationship may not be explicitly identified. As mentioned, the spacing of the points in the array can be selected to ensure the desired correspondence, e.g., the spacing is such that the point 1435d before the most recent point 1435e on the local centerline 1435 will appear near the endpoint 1430d of the global centerline. Thus, after a rapid or disruptive movement of the camera, the spacing intervals on the local and global centerlines may not be the same.
[0254] As mentioned at box 1320b, the system can then identify the closest pair of points between the two centerlines as the anchor points. Here, points 1435a and 1430a are identified as the closest pair of points (e.g., nearest neighbors) and are thus identified as the anchor points, as reflected here by their representation by a triangle rather than a square.
[0255] Thus, as shown at state 1440b and according to block 1320c, the system can then use the midpoint as a weight (the new interpolation points 1445a-c fall on the weighted average 1445, shown here for clarity) to determine the weighted average 1445 from the anchor point on the centerline to the end point (the end point 1435e of the local centerline 1435 dominates at the end of the interpolation). Finally, according to block 1320d and as shown at state 1440c, the weighted average 1445 can then be appended from the anchor point 1430a in order to extend the old global centerline 1430 and create a new global centerline 1450. For clarity, points before the anchor point 1430a (such as point 1430e) will remain in the same position in the new global centerline 1450 as before the operation of the integration 1490.
[0256] Thus, in this example, via progressive local centerline estimation and integration with the growing global centerline, the global centerline can be incrementally generated during withdrawal. Once all poses have been considered at block 1455a, the final global centerline can be published for downstream operations (e.g., retrospective analysis of colonoscope kinematics). However, as described herein, since the integration affects the portion of the global centerline after the anchor point 1430a, real-time kinematic analysis can be performed on the "stable" portion of the global centerline created prior to this region. Since the stable portion of the global centerline can be only a short distance in front of or behind the current position of the colonoscope, an appropriate offset can be used such that the kinematics generally correspond to the movement of the colonoscope. Similarly, although this example has focused only on withdrawal for ease of understanding, the application during advancement (and updating a portion of the global centerline rather than extending the global centerline) can be applied with the necessary modifications as well.
[0257] Despite the complex and irregular nature of the patient's inner surface and despite the various variations between patient anatomies, a more consistent global centerline (and associated kinematic data derived from the reference geometry) can be created by using the various operations described herein. Thus, the relative kinematic data and residual kinematic data for the projection of the instrument movement can be more consistent between operations, thereby facilitating better feedback and analysis.
[0258] Situational Awareness Kinematic Assessment
[0259] Although specific examples have been provided above, once the reference geometry has been determined in any suitable manner, the system can evaluate the kinematics of the surgical instrument relative to the geometry (e.g., both relative kinematics and residual kinematics). Specifically, at a high level, Figure 15Ais a flowchart depicting various operations in an example process 1505 that can be implemented in some embodiments for updating instrument kinematics relative to a reference geometry during a surgical procedure. Generally, at block 1505a, the system can infer the reference geometry, for example, using the centerline estimation methods disclosed herein. In the case where the reference geometry is available, the system can then consider previously acquired pose information, encoder information, etc., to determine at block 1505b the relative kinematics of one or more surgical instruments projected onto the reference geometry. As mentioned, portions of the kinematic data that are not part of the relative kinematics (referred to as "residual kinematics") can also be inferred at block 1505c.
[0260] For further clarity, Figure 15B is a flowchart depicting various operations in an example process 1510 that can be implemented in some embodiments for evaluating kinematic information. During a surgical procedure, at block 1510a, the system can consider whether new kinematic information is available at block 1510b. For example, in a colonoscopy, the system can wait at block 1510b for a rotation or translation threshold of Equation 1, Equation 2, or Equation 3 to be exceeded, or for a new depth frame to have been acquired at a new pose. In the case where it is determined at block 1510b that new kinematics are available, then at block 1510c, the new kinematic information can be integrated into the kinematic data record. In some embodiments, this can involve determining relative kinematics and residual kinematics as at blocks 1505b and 1505c, but in other embodiments, this processing can be deferred.
[0261] At block 1510d, the system can consider whether situational factors and the kinematic data record indicate a need to provide feedback to the surgical team. For example, movement too close to the colon sidewall, movement too fast along the centerline near an anatomical artifact of interest, movement of an anatomical artifact in a region not suitable for viewing, etc., can each trigger the presentation of feedback at block 1510e, such as an audible warning or a graphical warning, for example, on displays 125, 150, 160a, etc.
[0262] At block 1510f, the system can consider whether refinement of the model is possible. For example, during retraction 1305e, the field of view of the camera can obtain a better view of previously encountered regions, thereby facilitating filling in holes in the model and potentially higher-resolution regional models. Improvements to these sections of the model can help improve the estimation of the centerline portions corresponding to those regions. Then, at block 1510g, the improved centerline itself can facilitate improved relative and residual kinematic data calculations. As indicated, this refinement can be possible even when new kinematic data is not available. For example, when the system selects to iterate and merge previously acquired data frames, model refinement can be possible even when there is no new kinematic data at block 1510b in order to improve the model within the patient.
[0263] Once all the data for the surgical procedure has been acquired at block 1510a, at block 1510i, the same or a different computer system can initiate an overall assessment of all the kinematic data and present feedback at block 1510j. It should be understood that in addition to or instead of presenting feedback at block 1510j, the system can store data, initiate comparisons with other instances of surgical procedures performed by the same or different surgical operators, etc.
[0264] Again, combining knowledge of the temporal and spatial orientation of the surgical instrument with relative and residual kinematic data can facilitate multiple metrics and assessments, as well as applications during and after the surgery. For example, Figure 15C is a schematic representation of a colon model having a spatial context region 1515 and a temporal context region 1525 that can be used in some embodiments. Specifically, the model can be divided into regions 1515a - 1515g associated with different context factors such as anatomical artifacts, surgical maneuvers, procedural requirements, etc. Similarly, temporal regions 1525a - 1525e can be specified between the start and end of the surgical procedure, such as time limits, surgical tasks, etc. Surgical tasks can include discrete operations within the surgical procedure that can be recognized by a machine learning system or a user (e.g., regions to be cauterized, removing a tumor, initiating a retraction, etc.).
[0265] Since localization can be performed throughout the surgical procedure, the system can consider the spatial context region 1515 and the temporal context region 1525 when considering whether to present feedback at blocks 1510d and 1510e. For example, whether early 1525a or late 1525e in the surgery, preparatory insertion and withdrawal operations in region 1515a typically can involve approaching the lateral wall of the colon, sudden changes in speed, etc. Thus, the thresholds for generating warnings can be smaller in these regions and times, then for example in region 1515e in the middle of the surgery, where encountering the lateral wall can result in greater damage or discomfort. Thus, the radial contexts 1245f, 1245g, etc. can present importance that varies with spatial and temporal context.
[0266] For further clarity, Figure 15D a set of GUI elements that can be implemented in some embodiments is provided. Such elements can be presented during or after the surgical procedure, as described in more detail herein. A representation 1520b of a three-dimensional model of the colon can be presented during the surgery in its partially created state or after the surgery in its final state, and in a TSDF format, in an exported triangular mesh representation, or other suitable representation. An indication 1520c (represented here by an arrow) of the current position and orientation of the surgical instrument relative to the representation 1520b can be used to indicate the orientation and position of the instrument at the current time in the procedure or at the current time in a playback of the procedure. Pop-up boxes (such as the excessive withdrawal speed pop-up box 1520a) can indicate the location on the representation 1520b where an undesired kinematic behavior (either relative or residual) is found to occur. Here, the timeline 1520f is likewise provided with an indication 1520l of the current playback time. Portions of the timeline 1520f can be highlighted to provide information about the kinematic data, such as with a change in luminosity or hue (e.g., green for regions completely within the kinematic metric tolerance, orange and yellow for regions approaching the tolerance boundary, and red for regions that have exceeded the tolerance boundary). Thus, the portion of the surgery that contributed to the pop-up box 1520a can also be identified by the highlighted region 1520k in the timeline (such as indicated by a red hue).
[0267] In example pop-up box 1520a, information about the time during surgery of a kinematic data event (a 40-minute interval shortly after entering the surgery), the average speed of the operator during the event ("5 cm / s"), and reference data from similar practitioners (here, the median speed of an expert during the corresponding part of their procedure was "3 cm / s"). Although this example is for withdrawal speed, it should be understood that many events can be triggered by evaluating relative and residual kinematic data from a reference geometry. Thus, an undesired approach towards a sidewall, an undesired approach towards an artifact, an undesired movement of one instrument relative to another only constitute some example events that can be identified from kinematic data and draw the attention of a surgical operator or viewer.
[0268] The current image playback corresponding to the position and orientation indicated by indicator 1520c and the time indicated by indicator 1520l can be shown in video playback area 1520d. Area 1520e can also provide information about the current kinematic assessment of the depicted frame (such as the current speed on the centerline).
[0269] In some embodiments, the GUI can include a kinematic plot 1520i depicting one of the metrics derived from kinematic data (e.g., speed along the centerline, acceleration orthogonal to the centerline, etc.). Here, the GUI includes a plot of the speed 1520h along the centerline (positive values reflect advancement and negative values reflect withdrawal) over a portion of the entire procedure (although the x-axis here indicates the time position, in some embodiments, the speed can be mapped to the length of the centerline itself and the x-axis is instead used to indicate points on the centerline), the current playback position shown by indicator 1520m (corresponding to indicator 1520l), orientation 1520c, and current playback 1520d. Here, the upper kinematic metric boundary 1520g and the lower kinematic metric boundary 1520j can vary with the orientation and the task being performed (although shown as straight lines here, it should be understood that the thresholds can vary with time and spatial context). In this example, exceeding the lower limit 1520j in area 1520n prompts an excessive withdrawal speed associated with pop-up box 1520a and area 1520k. The system can warn the operator that they are "going too fast" when the mapping results in multiple holes, the centerline movement is too fast for a reasonable assessment inside the patient, camera blur hinders proper analysis or positioning, etc. Since the speed orthogonal to the centerline can indicate additional operations performed during the procedure (e.g., adjustment of the endoscope to examine some area behind a fold), such residual kinematics above a threshold can only be allowed (or associated with a wider threshold) in the time and area where they are expected.
[0270] Specifically regarding colonoscopy, as another example of kinematic assessment, some embodiments may assume that proper examination of a portion of the colon may take approximately six minutes. Thus, each section to be examined (e.g., one or more of regions 1515a - 1515g, which may be the same as the sections used for centerline estimation) may require a dwell time of no less than six minutes, where the kinematic threshold is set based on the integrity of the mapping model in that region (e.g., higher speeds along and off the centerline may be allowed once proper mapping of the colon region is in place). In cases where such conditions are not met or there is a risk of non - fulfillment, the corresponding region in representation 1520b (e.g., one of regions 1515a - 1515g) may be highlighted.
[0271] Since the length of the colon can vary between patients, the discrimination of regions 1515a - 1515g may be based on landmarks or operator indication (e.g., the operator may have the ability to define the regions themselves when creating the model). In some cases, patients may be classified based on their body characteristics to prepare an initial estimate of colon size and corresponding regional boundaries 1515a - 1515g, as well as an adjustment to the time expectation 1525 (e.g., because the same procedure may take longer in patients with a longer colon). Any estimated uncertainty in the colon structure can then be reduced as more information becomes available during the surgical procedure, positioning, and mapping. In some embodiments, the system may require the proper creation of a colon model within a given region such that centimeter - per - second accuracy along the centerline is possible, and only then invite the operator to continue the procedure such that the operator's subsequent relative kinematic instrument motion and residual kinematic instrument motion have the desired resolution.
[0272] Consultation with KOLs to refine the kinematic threshold can be facilitated by viewing the procedure using GUI elements, such as Figure 15D those described therein. It should be understood that in cases where elements are presented during surgery, the timeline indicator 1520f and the drawing indicator 1520i may be omitted, but in some cases they may be included to describe past portions of the operation.
[0273] Intra - operative and Post - operative Kinematic Feedback - Projection Relationship
[0274] It should be understood that there are many other methods by which the kinematic data derived from the techniques disclosed herein can be presented and used by operators and viewers. Regarding colonoscopy, Figure 16AFIG. shows a schematic diagram of a three-dimensional colon model 1605a with a path graphic 1605b that can be presented in some embodiments. The model 1605a can be the same as the model derived during localization and mapping, e.g., the TSDF representation or mesh from which it was derived, or can be an idealized colon model (e.g., prepared by an artist or averaged across a set of known models). The path graphic 1605b can provide an indication of the original kinematic values of a surgical instrument during a surgical procedure (e.g., the path traveled by a camera during a colonoscopy). As Figure 16B shown, the same or a different model 1610a can present corresponding reference geometry, particularly with respect to the centerline 1610b of the model.
[0275] As in the kinematic plot 1520i, a plot 1610d of metrics derived from relative kinematic data, residual kinematic data, or a combination of both can be presented to an operator or viewer. Here, for example, an area 1605c of the original path selected, e.g., with a mouse cursor 1605d, can contribute to a corresponding indication of an associated portion 1610c of the centerline (the nearest portion of the centerline to the area 1605c) and a highlighted area 1610e (metric values derived from the motion in the area 1605c). It should be understood that the selection can occur in the reverse or other order, e.g., selection of the plot area 1610e can contribute to highlighting 1610c and 1605c. A graphic of each of the residual or relative kinematics can also be presented in the colon model 1605a.
[0276] For clarity, it should be understood that in other contexts, with necessary modifications, the embodiments discussed herein with respect to colonoscopy and similar tubular regions (such as the lungs, esophagus, etc.) can be readily implemented. For example, Figure 16C is a schematic diagram of a path graphic 1615b in a lumen model 1615a that can be presented in some embodiments, similar to Figure 16A the original kinematic representation in. Specifically, an original motion path graphic 1615b of an instrument inserted via an inlet 1615c can be shown relative to the model 1615a. In the same or a different model 1620a, a reference geometry (a sphere 1620b in this example), as well as a plot 1620d of metrics derived from relative or residual kinematic data based on the original path 1605e and the geometry 1620b, can be shown. For example, Figure 16D the plot 1620d of can depict the projected velocity of the instrument kinematics on an axis 1620f, on a meridian or latitude line of the sphere, etc.
[0277] Here, for example, selecting a portion of the drawing 1620d, geometry 1620b, or path 1615b using the cursor 1605d can cause corresponding highlighting of other GUI elements. For example, selection of the region 1605e can highlight the associated portion 1620c of the reference geometry 1620b on which the selection falls, as well as highlight the corresponding portion of the drawing 1620d, 1620e. Although not shown, it should be understood that other GUI elements among the GUI elements discussed herein (e.g., pop-up box 1520a, timeline 1520f, playback 1520d, threshold 1520g, 1520j, etc.) can equally be placed in the same GUI as the elements depicted herein.
[0278] During playback or during a surgical procedure, in some embodiments, the graphical element can provide a real-time representation of relative or residual kinematic metrics relative to a reference geometry. For example, reference Figure 16F , at a given moment during a surgical procedure, or during playback of a surgical procedure, a point 1625b corresponding to the current projected position of the orientation of a surgical instrument on a spherical reference geometry is shown (it should be understood that each spherical example herein can be applied, with necessary modifications, to any manifold surface embedded within a Euclidean geometry representing an organ model, such as a hemisphere, arbitrarily undulating surface, etc.). Similarly, as Figure 16E shown, in the context of the centerline 1625d, the current projected position of, for example, a colonoscope on the centerline at the current moment during the surgical procedure or at the current moment in the playback can be shown via a spherical indicator 1625e (or other suitable indicator, such as highlighting a rendered portion of the centerline). The length and direction of the arrows in each of the spherical case and the centerline case can indicate the current value of the relative kinematics.
[0279] For example, the arrow 1625c can indicate the projected direction and magnitude of the current projected velocity of the instrument on the reference geometry 1625a (by its length, color, brightness, etc.). Similarly, the arrow 1625f can indicate the current velocity of the projected velocity on the centerline 1625d from the position of the indicia 1625e at the current time (again, the magnitude, e.g., represented by length, color, brightness, etc.). Similarly, residual kinematics can equally be presented in the graphical element. Here, an arrow 1625g orthogonal to the surface of the sphere 1625a at the position of the indicator 1625b indicates the velocity component of the instrument orthogonal to the sphere (such a component can be useful, for example, to warn the user when a cautery or other instrument is approaching an anatomical artifact too quickly). Similarly, residual kinematics in the colonoscope context (e.g., movement away from the centerline) can be represented by an arrow 1625h also orthogonal to the centerline 1625d at the point of the indicia 1625e.
[0280] Intra - operative and Post - operative Kinematic Feedback - Graphic Elements
[0281] Figures 17A - 17F Additional graphical elements that can be used in the GUI during or after a surgical procedure during viewing are shown. Regarding Figure 17A , the GUI can include a playback or current view area 1705b for a surgical camera. In addition to presenting the current view of the camera, an indication of relative or residual kinematic metrics can also be represented. For example, the velocity on the centerline can be shown in the overlay 1705a. The "speedometer" graphic 1705d can assist the operator in recognizing how their motion relates to thresholds, such as the maximum or minimum allowable velocity for the spatial and temporal context of the surgery. In some embodiments, during or after surgery, the elements presenting the camera's field of view can be supplemented with augmented reality graphical elements. For example, in this example, an augmented reality representation of the centerline 1705c is provided. Here, the user can easily perceive that the camera is above the centerline representation 1705c. During surgery, such an overlay can be provided upon the operator's request, for example, to provide quick adjustments. In some embodiments, the augmented reality overlay can be translucent such that the operator can still perceive the original camera field of view.
[0282] As previously described, in the case of viewing a surgery after it is completed, a timeline 1705e with an indicator 1705f of the current time in the playback (e.g., the time associated with the currently depicted camera image in element 1705b) can be provided. Regions with significant kinematic events can be indicated, for example, by a change in hue or luminosity on the timeline 1705e as described herein. As previously described, pop-up boxes can also be used to annotate events. In this example, the pop-up box element 1705i indicates that the velocity along the centerline in the advancing direction in the time region 1705h exceeds the desired threshold, and the pop-up box 1705j indicates that the velocity threshold is exceeded in the retracting direction in the time region 1705g. The color of the reference geometry as reflected in the augmented reality element 1705c can change during these regions or when approaching the threshold to warn or notify the operator or viewer of a potentially undesirable condition.
[0283] As Figure 17BAs shown, the representation of the current position of the surgical instrument need not be limited to an arrow or other abstract representation, as a computer graphics model of the colon can be used. Here, for example, the model 1710b of the colonoscope is shown in the orientation of the currently depicted frame (e.g., in element 1705b) and relative to the model 1710a of the colon (e.g., a model acquired during the procedure, an idealized representation, or a combination of both). Again, a plot 1710c of the derived kinematic metric 1710f can be shown during the course of the surgery, e.g., having a current playback time indicated by indicator 1710i, and having thresholds 1710d and 1710e (here, baseline 1710j indicates zero velocity, e.g., along the centerline). As indicated, depending on the context (e.g., the various spatial and temporal regions discussed herein), the thresholds can vary over time and with orientation within the organ. The regions where kinematic events occur (such as those represented by pop-up boxes 1705i and 1705j) can be shown by corresponding highlights 1710g and 1710h.
[0284] For clarity, Figure 17C is an enlarged view of the reference geometric kinematic "speedometer" graphic 1705d depicted in the Figure 17A GUI. In this example, the range of values that the kinematic metric can assume (e.g., velocity along the centerline) can be divided into four regions 1740a - 1740d. As described herein, the "minimum" and "maximum" acceptable values can be determined by spatial and temporal context, possibly informed by the KOL. The current value of the metric can be presented, for example, by arrow indicator 1740e, by highlighted region 1740f, or by other suitable indication.
[0285] For clarity, it should be understood that there are multiple ways to represent the current kinematic metric value within the context-determined range. As another example, a bar plot 1745a can convey the same information as the format of the semi-circular speedometer 1705d. Again, the plot can be divided into regions 1745c - 1745f, where the shaded region 1745b again indicates the current value of the metric.
[0286] For a clearer understanding of further variations of the disclosed embodiments, Figure 17D a robotic surgical procedure interface 1715a (e.g., such as in display 160a) is depicted. Here, the reference geometry is a hemisphere represented by an augmented reality element 1715e. A corresponding shorthand reference 1715b showing the relative projective kinematic motion of the surgical instrument 1715f with respect to the reference geometry is also provided. The speed of the motion of the instrument 1715f (23.2 cm / s) as well as a speedometer 1715c (e.g., providing a similar range and thresholds as described above for indicator 1705d) are also superimposed for reference.
[0287] It should be understood that in some surgical procedures, more than one reference geometry may be involved. For example, in colonoscopy-based polyp removal, the reference geometry around the polyp surface and the centerline geometry of the colon can each be used, where the corresponding relative and residual kinematic data sets are collected. Such multiple references can be represented in the operator's GUI. For example, in the example of Figure 17D where augmented reality guidance 1715d is provided to indicate the path along which the instrument is expected to travel in combination with augmented reality elements 1715e around the region of interest.
[0288] In fact, a more extensive set of such different references is shown in the interface 1730a of Figure 17E . Here, one or more spherical reference geometries are associated with each of the instruments 1730g, 1730h, and 1730i. Accordingly, the corresponding shorthand references, as well as the corresponding kinematic metrics and tachometer indicators 1730b, 1730f, and 1730d (corresponding to the instruments 1730g, 1730h, and 1730i respectively), can be presented as overlays, augmented reality elements, etc. In the presence of multiple available reference geometries, the operator can cycle through their selections and presentations. For example, when the operator begins a new surgical task involving a geometry different from the previously associated geometry. For example, the identifier 1730e can be an abstract geometry created by the operator and inserted as part of the field of view as an augmented reality element. Similarly, a guidance reference 1730c can be specifically created or provided to guide the instrument along a preferred approach path. For example, as shown in more detail in Figure 17F , some reference geometries themselves can include a composite of reference geometries. The depicted guidance geometry includes a first geometry 1735b similar to the centerline geometry, which can be used to guide the instrument to an orientation associated with a second geometry 1735a. Although the second geometry 1735a shown here is a box, the second geometry 1735a can take a form suitable for a given orientation, operator, etc. For example, the geometry can take the surface profile of an anatomical artifact, auxiliary geometry for placing cautery and other tools, etc.
[0289] Overview - Image Feasibility Assessment
[0290] As discussed above with respect to block 625b and confirmation 895, various embodiments will check the field of view of the sensor before providing the sensor's data to downstream processing operations such as the localization and mapping operations discussed herein. In the case where the sensor is a camera, the field of view can be readily apparent from the camera's image data itself. Although mainly discussed here in the context of colonoscopy for ease of understanding by the reader, it should be understood that the various disclosed embodiments can be applied, with necessary modifications, to various contexts such as esophageal examinations, lung and bronchial examinations, etc.
[0291] Regarding colonoscopy image localization and mapping, many situations can render the field of view unsuitable for downstream processing. Specifically, Figure 18A A schematic collection of various surgical camera states and corresponding fields of view that may be encountered in a colonoscopy scenario is provided. In situation 1805a, the colonoscope 1805d can be in a position and orientation within the colon 1805c such that the resulting camera field of view 1805b has no significant artifacts or obstructions. In this case, downstream processing can easily be able to, for example, infer the orientation of the colonoscope relative to previously acquired data and infer accurate depth values as the colonoscope is advanced or withdrawn through the colon 1805c.
[0292] However, in situation 1810a, the colonoscope 1810d has been advanced too quickly along the central axis of the colon 1810c to be properly positioned. This can cause the camera to produce a motion-blurred image 1810b. While acceptable motion blur may vary among surgical operators, generally, their tolerance for motion blur can be different from the tolerance of downstream processing. For example, the tolerance of the localization system can be lower than that of the operator. Thus, there can be a situation where the resulting motion blur is not noticeable or not too disruptive and does not cause the operator to adjust their workflow, but where this perceived minor blur is actually destructive to downstream processing. Without notification, the operator can continue with the procedure, hardly aware that the processing "cannot" keep up with the operator's progress. Blurry images (such as image 1810b) can challenge many localization algorithms because it may be difficult, for example, to separate the smooth organ sidewall viewed statically from the smooth blur of image 1810b. For example, the application of the SIFT algorithm can produce features that are quite different from those clearly perceived in image 1805b. Thus, from the perspective of the system, a rapid withdrawal or advancement that contributes to blur can be very similar to the rotation of the camera towards a nearby sidewall. Therefore, if downstream processing is allowed, errors or other incorrect localization results can occur, resulting in errors in further downstream processes such as mapping and modeling. Similar to the advancing and withdrawing motion blur in situation 1810a, in situation 1815a, too rapid lateral movement or too rapid rotation of the camera 1815d within the colon 1815c can produce a blurry image 1815b.
[0293] As another example, when cases 1810a and 1815a produce image-wide blurring, local blurring 1820e may also occur, as in image 1820b of case 1820a, where camera 1820d does not move within colon 1820c, but fluid has accumulated on the camera lens to produce local blurring 1820e. In addition to being localized, blurring 1820e is not necessarily associated with a smooth transition vector gradient as in the blurring of cases 1810a and 1815a, since the optical properties of the fluid can vary with its density. Thus, the appearance of localized blurring 1820e can be inconsistent in terms of orientation, shape, or density, and thus it is more difficult to identify frame-wide blurring of images 1810b and 1815b (which can actually be discerned via optical flow or frequency analysis).
[0294] Even when acquiring an image frame without blurring, in some cases it may still not be suitable for downstream processing. For example, in case 1825a, camera 1825d has gotten too close to the sidewall of colon 1825c. As a result, the occluding pouch fold 1825e obscures most of the field of view, causing a significant portion of the field of view to be occluded in image 1825b, such that downstream processing will be adversely affected. For example, SIFT features determined on a small portion of the sidewall are not likely to be meaningfully different from features determined on another small portion of the sidewall, thus reducing the utility of the features for localization.
[0295] Similarly, biomass 1830e may obscure the field of view so much (as in image 1830b of case 1830a) that downstream localization of camera 1830d and mapping of colon 1830c become infeasible or at least vulnerable to erroneous results. This may be especially true in cases where the biomass is not static, but appears in various orientations at different times during a surgical procedure. If SIFT or similar features are derived from the biomass, their application in localization may have the risk of attempting to map a dynamic object to an environment that is typically static. Thus, failure to achieve proper localization usually itself indicates some other contextual condition, which can be significant for other machine learning processes (such as biomass recognition and characterization systems).
[0296] To further complicate the issue, it should be understood that the various cases discussed here can occur simultaneously, such as when there is both motion blurring and local blurring due to fluid. Thus, it should be understood that cases 1810a, 1815a, and 1820a are not necessarily mutually exclusive. Thus, the various embodiments disclosed herein can seek to identify the presence of these cases not only individually but also in combination (e.g., appropriately labeling an image as a combination of adverse cases).
[0297] For clarity, it should be understood that not all events that affect downstream processing need to be identified by the various embodiments disclosed herein. For example, an air injection system in a colonoscope 1835d can easily facilitate inflation of the colon 1835c (as in scenario 1835a) to produce a view 1835b with mostly smooth sidewalls and reduced pouch fold elongation. However, because such inflation is initiated at the operator's command, some embodiments can identify all images as not suitable for localization for some period of time after such modification. Thus, the automatic applicability determination methods described herein can sometimes be used in combination with other mechanisms to identify undesirable frames (e.g., the operator can manually disable downstream processing via an interface; an encoder can be monitored to identify motion that causes blurring; software, firmware, or hardware used to perform a field-of-view change operation (such as inflation in scenario 1835a) can cause frames to be identified as inappropriate during and for some period of time after the application of its operation, etc.).
[0298] Again, while most of the embodiments herein are disclosed in the context of colonoscopy for ease of reader understanding and for consistency of reference, it should be understood that the various embodiments can be applied in other contexts with necessary modifications. For example, in Figure 18B a laparoscopic procedure (such as a prostatectomy), instruments 1850a and 1850b have been inserted into the lumen 1850c via ports. In some procedures, it may be desirable to localize the camera (and possibly indirectly the localization of other instruments and objects within the field of view) and possibly map the lumen 1850c or structures therein. Since motion blur, occlusion, fluid blur, etc. can also occur in these cases, many of the disclosed embodiments can be applied to these contexts with appropriate variations. For example, in Figure 18C , a portion of the field of view 1840b of a laparoscopic camera 1840a is occluded by an instrument 1840c, resulting in a visible portion 1840d and an occluded portion 1840e of the tissue. As in the context of colonoscopy, identifying such a situation can help prevent downstream processing that attempts to reconcile pixel values associated with the instrument 1840c.
[0299] Similarly, as through Figure 18DAs shown in the example of, it should be understood that many of the embodiments disclosed herein can also be applied to all or less than all of the camera's field of view. For example, although positioning and mapping during colonoscopy can utilize the entire field of view of the camera, in some applications, positioning and mapping may be interested only in a specific portion of the field of view. For example, in Image 1855a, the region 1855b of the surgical interface has been designated for data acquisition (e.g., where the user wishes to generate a model of the tissue region, excise organ artifacts in the region, identify tumors in the region, model the structure of tumors in the region, etc.). Instruments in this region 1855b (as shown in Image 1855a) can complicate or impede downstream processing operations associated with data acquisition. Thus, the system can invite the user to withdraw the instrument from region 1855b (as shown in Image 1855c), or the user can do so proactively. Similarly, as will be desired in more detail herein, in some embodiments, the identification of infeasible images can prompt various response actions of the system. For example, in Image 1855d, the detection that the image within region 1855b is infeasible for downstream processing has prompted the application of a You-Only-Look-Once (YOLO) network (e.g., as described in Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi's You Only Look Once: Unified, Real-Time Object Detection (arXiv TM :1506.02640); or for example, using the method described in Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, Dhruv Batra's Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization (arXiv TM :1610.02391), or for example, as described in Chien-Yao Wang, Alexey Bochkovskiy, Hong-Yuan Mark Liao's YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors (arXiv TM: Those described in (2207.02696) can also be used) to produce a highlighting (such as highlighting 1855e) to indicate that the instrument may cause an infeasible determination. When the user is invited to withdraw the instrument from region 1855b (as shown in image 1855c), such highlighting can better inform the user. Thus, some embodiments can incorporate or anticipate downstream processing of additional "filtering" options, such as tool identification and classification, additional environmental analysis and processing (e.g., facilitating changes to the surgical procedure, tasks to be performed, priorities of operations, etc.).
[0300] Naturally, for Figure 18A such response actions can also be taken in the case of, such as when the determination of an infeasible image depicts the application of a pixel classifier to highlight biomass 1830e, occluded sidewall 1825e, or blur 1820e. Similarly, the application of Fourier transform, optical flow algorithm, etc. can be triggered for images 1810b, 1815b, and 1820b to determine the nature and orientation of the blur. Similarly, although feasibility recognition can occur on a single static image, in some embodiments, being classified as infeasible can contribute to the processing of considering an image sequence. For example, detecting a lack of feasibility in one image can trigger a reconsideration of the window of surrounding images, such as calculating the optical flow between the images in the window, inferring the motion in the images within the window (and the possible associated causal factors of the lack of feasibility), etc.
[0301] Example Data Pre - processing and Post - processing
[0302] To distinguish between feasible and infeasible images in the case of Figure 18A some embodiments employ various implementations of the general process 1905 shown in Figure 19A Specifically, at block 1905a, the system can receive a visual image acquired by an intraoperative camera. For example, in the case of applying process 1905 during the procedure itself, image 1905a can be the most recently acquired image, or the next image in the queue expected to be processed by downstream operations. However, process 1905 can also be applied offline, such as in the case of the surgical post - operative situation described herein, where it is desired to evaluate surgical data after the surgery.
[0303] At block 1905b, the system can pre - process the original visual image. For example, crop the image to an appropriate size for input into a neural network, adjust the channels to those expected by the neural network, perform contrast - limited adaptive histogram equalization (CLAHE), apply the Laplacian operator, etc., as described herein. Regarding Figure 19BAn example preprocessing process is shown. Normalizing the image via preprocessing can include transforming the image values such that the mean and standard deviation of the image become 0.0 and 1.0, respectively. For example, the channel mean can be subtracted from each input channel, and then the result can be divided by the channel standard deviation, as shown in Equation 7:
[0304]
[0305] At block 1905c, the processed image can be input into one or more neural networks, e.g., a network as described herein. In some embodiments, preliminary steps between blocks 1905b and 1905c, as described in further detail herein, can be applied to determine which network in the network corpus should be applied to the preprocessed image to evaluate feasibility. In some embodiments, after the network is determined, post-classification processing can be applied at block 1905d to produce a final feasibility determination. For example, various input edge cases can be addressed in the post-classification processing. For example, Figure 19C The process 1915 described in provides an example post-classification operation. Once the system determines the final classification value of the image, the result can be output at block 1905e (e.g., for making a decision at block 625b or for confirmation 895).
[0306] In Figure 19B An example of the various operations in the preprocessing process 1910 that can occur at block 1905b is shown. Here, at block 1910a, the system can receive a visual image, e.g., acquired by a colonoscopy camera, and at block 1910b, the channels of the image can be appropriately transformed for use by one or more neural networks, as at block 1905c. At block 1910c, the image can be similarly cropped and resized for application to one or more networks. A reflection mask can be applied at block 1910d. The image can be output at block 1910e, e.g., for consideration by the neural network at block 1905c.
[0307] Conversely, in some embodiments, processing may be applied after the neural network is applied at block 1905c, for example, to identify one or more common edge cases where the neural network is vulnerable to misclassification. For example, some neural networks may incorrectly classify a blurry image as valid when those images contain a large amount or large reflective highlights (e.g., when the colonoscope shines light on an irregular corrugated surface). Similarly, some occluded images may appear similar to regions with many homogeneous pixel groupings (e.g., large cavities, darkened holes, or sidewalls). While neural networks are generally able to identify the low-frequency characteristics of most blurry and occluded two-dimensional images, some cases (such as the presence of many highlights in the blur) may result in sufficient transitions to cause relatively consistent misclassifications (e.g., in these cases, a series of ridges with smooth contours may resemble a blurry image, at least when the highlights are similarly placed). In particular, saturated portions of the image produced by projector light reflected from the surface (which appear blurry or smeared) may consistently contribute to misclassification (such images are generally non-informative because a large number of saturated pixels similarly affect localization, as do a large number of occluding pixels).
[0308] Accordingly, in the post-processing of block 1905d, the system may apply various edge case remedies via a logical or complementary classifier. For example, remedies for addressing occlusion and blurry saturation may be accomplished by a process such as process 1915, which first determines at block 1915a whether the image is classified as valid, and if so, retains the classification at block 1915e because Figure 19C process 1915 only focuses on false positives rather than false negatives (but it should be understood that similar processes are used for various false negatives). For an image identified as viable by one or more networks, at block 1915f, the system may first assume that the image depicts an occlusion misclassification edge case. For example, hue thresholding, Euler counting, flood filling, frequency analysis, SIFT feature analysis, etc. can all be used to determine whether the image depicts an occlusion edge case. If an occlusion is found, the post-classification logic may adjust the classification at block 1915d.
[0309] At block 1915b, the system can consider whether the image contains blur. For example, while applying a neural network at block 1905c can be suitable for identifying various types of blur, such as motion blur, local blur, etc., direct analysis of the image using traditional processing techniques can reveal the presence of blur in those edge cases where a large number of reflections have caused false positives. Thus, blocks 1915b and 1915c can operate together to determine whether the image depicts the envisioned edge case. For example, the frequency content of the image in the presence of bright spots can provide a sufficiently consistent and unique profile for identification using traditional binary classifiers such as support vector machines (SVMs), logistic regression classifiers, etc. While in some embodiments a hard threshold can be based on inspection, it should be understood that a classifier such as an SVM can be easily trained to perform the operations of blocks 1915b and 1915c to distinguish truly blurred images from blurred images with a necessary number of highlights. For clarity, such an SVM can apply its own preprocessing steps to the image received at block 1905a, and such preprocessing steps can be the same as or different from the steps at block 1905b. Block 1915c can alternatively evaluate the portion of the image occupied by highly saturated pixels (such as by one or more reflections) rather than the number of reflections. In the case where the image meets the edge case condition, the classification can be adjusted accordingly at block 1915d. It should be understood that such exclusion operations can also be applied to the eliminated frames prior to the network considering the eliminated frames (e.g., an image consisting of only black pixels should clearly not even have a chance of being classified as valid by the network). However, edge case consideration after the feasibility classification process may be more suitable for edge cases that affect less than all or inconsistently affect portions of the image (such as reflection dispersion and highly saturated regions).
[0310] Although example process 1915 depicts occlusion assessment and then blur and saturation assessment, it should be understood that variations based on the present disclosure exist where each edge case is considered separately, and additional edge cases are considered. Thus, the operation of block 1905d can include only blocks 1915f, 1915b, 1915c rather than the depicted sequence (e.g., saturation can be evaluated separately without considering blur). The selection of such logic can be identified in parallel with training one or more neural networks, since the logic and corresponding thresholds can be selected to improve the overall classification result during validation. Again, for clarity, some embodiments forego process 1915, and in fact, forego all post-processing at block 1905d, instead relying solely on the classification determined by one or more networks at block 1905c. Such an approach can be suitable in cases where the network has been exposed to a sufficient variety of training samples to consider the desired edge cases.
[0311] Example Feasibility Assessment Neural Network
[0312] Figure 20A is a block diagram of an example neural network architecture that can be used to distinguish viable images from non-viable images (e.g., at block 1905c) in some embodiments. In this example architecture pipeline, the system provides the processed image 2005a (e.g., after the operations of block 1905b and process 1910) to a first stage of one or more convolutional layers 2005b. The first stage of one or more convolutional layers 2005b is itself coupled to one or more pooling layers 2005c, and one or more pooling layers 2005c can itself be coupled to a second stage of one or more convolutional layers 2005d. A linear layer 2005e and a merge layer 2005f can then follow to produce a final output classification as viable or non-viable. Although in some embodiments the final output can be a binary classification, such as "valid for downstream processing" or "not viable for downstream processing", as indicated by output 2005g, some embodiments can alternatively break out different failure states, such as Figure 18A those described in, providing an output with multiple classes, where each class other than the viable class indicates the reason 2005h why the frame is classified as non-viable. Additional knowledge of the nature and cause of the failure can better inform surrounding processes and viewing, including downstream processing. Naturally, when using a multi-class output, the labeling of the training and validation data will change.
[0313] Although it should be understood that various methods for implementing Figure 20A the network structure, as an example for ease of understanding, Figure 20B provides a partial code listing for an example implementation for creating the network topology depicted in Figure 20A In this example, it is written in the Python TM language and uses the Torch TM machine learning library, and the class Classifier_Network extends the Torch TM class module (lines 1 and 3), and creates an example implementation of the structure that appears in Figure 20A in the initialization function on lines 4 - 13.
[0314] Specifically, in this example, line 4 corresponds to the first stage of one or more convolutional layers 2005b, line 5 corresponds to one or more pooling layers 2005c, and lines 6 - 10 correspond to the second stage of one or more convolutional layers 2005d. Then, lines 11 - 12 depict an example implementation of the linear layer 2005e before being connected to the softmax layer at line 13 corresponding to the merge layer 2005f to output the result. Here, line 13 indicates a single dimension for the binary result, but a multi-class output can be produced by increasing the dimension.
[0315] For clarity, Figure 20C is a partial code listing for performing forward propagation on the exemplary network implementation of Figure 20B . Specifically, continuing the definition of the Classifier_Network class started in the listing of Figure 20B , where line 1 specifies the re-implementation of the forward propagation function of nn.Module, and lines 2 - 6 specify the connections between convolutional layers as rectified linear units (ReLU).
[0316] Similarly, line 8 indicates that the linear layer of line 11 in Figure 20B is connected via ReLU torch.nn.functional “F”, and line 8 indicates the direct pass-through output of the linear layer of line 12 in Figure 20B . Finally, at line 10, the result of softmax based on the linear result can be output. Since the softmax output presents values between 0 and 1, it is naturally suitable for a binary classifier, such as in the exemplary output 2005g between separate feasibility and infeasibility. In the case of using a multi-class network, it should be understood that the connections are adjusted accordingly. Thus, the network can learn and extract semantic information from visual images, which is suitable for determining the feasibility of the image for depth estimation and subsequent camera localization for constructing the entire three-dimensional representation of the surface. When identifying that there are no suitable cues for those operations (e.g., appropriate patterns or numbers of SIFT features), the network can thus identify that the frame is invalid for those operations.
[0317] Figure 21A is a flowchart illustrating various operations in the network training and validation process 2115 that can be implemented in some embodiments. Specifically, at block 2115a, the training system can receive a training set of labeled images, as well as images with labels that are identified for validation at block 2115b. Here, the creation of the training data set can involve applying images acquired with a colonoscope, bronchoscope, etc. to downstream processing, verifying whether the downstream processing results are within or outside an acceptable tolerance, and then labeling the images accordingly. Thus, the training and validation data sets can be created by providing a corpus of real-world images to a downstream pipeline (such as would occur in normal in-situ processing), at least some of the real-world images being considered to exhibit Figure 18A various phenomena. Images that produce results within the tolerance of the downstream processing can be labeled as “valid”, while images that produce results outside the tolerance range can be labeled as “invalid” (or e.g., according to Figure 18Aadverse situations, appropriate classes for multiple types of infeasible labels). In cases where the downstream processing is, for example, localization and mapping, the results of the image can be compared with the ground truth results, and only the localization poses within the maximum distance from the true pose are marked as "valid".
[0318] At boxes 2115c and 2115d, training epochs can be performed using this data on the neural network until the performance of the network is found to be acceptable at box 2115d. For binary classification of feasible and infeasible (e.g., as in output 2005g), binary cross-entropy on the ground truth labeled dataset can be used to evaluate the loss. For multi-class output (e.g., as in output 2005h), multi-class cross-entropy on the ground truth labeled dataset can be used.
[0319] In some embodiments, the network that has been satisfactorily executed from box 2115d can be directly provided to box 2115i for release. However, in some embodiments, a portion of the training dataset provided at box 2115f shown here can be reserved for further validation and adjustment in the second round of training 2115e. When iterating through boxes 2115g and 2115h, edge cases can be detected and appropriate post-classification processing operations can be prepared in advance (e.g., determining parameters for detecting Figure 19C blurry and reflective edge cases). Once acceptably executed and the desired edge cases and their parameters have been appropriately identified, the network can be published at box 21151 for in-situ use.
[0320] Based on the present disclosure, it will be recognized that various network architectures and corresponding training methods can be suitable for distinguishing feasible images from Figure 18A the adverse situations 1810a, 1815a, 1820a, 1825a, and 1830a depicted therein. In fact, multiple neural networks or other ensembles of classifiers can be trained to provide more robust classification or redundancy-based verification, e.g., majority voting or weighted voting based on the verification performance of the constituent networks. Thus, some embodiments can employ Figure 20A topologies as well as implementations of vision transformer (ViT) networks, mobile ViT networks, mobile nets, etc. Again, it should be understood that although particularly suitable architectures and training methods are presented herein, the architecture structures and training methods can be significantly altered while still maintaining the same functional effects.
[0321] In some embodiments, rather than applying classifiers in parallel, at least some classifiers can be applied serially. For example, a first set of one or more classifiers can be trained to generally identify feasibility or infeasibility, while a second set of one or more classifiers can be trained to identify the classes of infeasible images (e.g.,Figure 18A One of the disadvantages). When deployed, the second set can be used to determine the nature of an infeasible classification. In fact, due to their different functions, the two sets of classifiers can adopt different architectures, such as when the first classifier is not a neural network but a simple binary classifier (such as an SVM or a logistic regression classifier operating on the principal component representation of a visual image) while the second classifier is a neural network with convolutional layers.
[0322] For clarity, Figure 21B is a flowchart illustrating various operations in a serial multi-classifier classification process 2110 that can be implemented in some embodiments. At block 2110a, one or more classifiers in the first set of classifiers (e.g., the SVM or logistic regression classifier described above) can provide an initial indication of feasibility or infeasibility. In the case of applying an SVM, in the relatively controlled context of a colonoscopy, grayscale scaling the original image, applying the Laplacian operator, and then applying principal component analysis may be sufficient to produce well-separated features for classification using an SVM with a radial basis function kernel. In the case where the first set of one or more classifiers determines a high-probability feasibility classification, the system can simply output a valid classification and proceed directly to block 2110e, but in the depicted embodiment, edge case detection and classification adjustment as described herein can first be performed at block 2110d. Thus, in the case of applying an ensemble of classifiers, the agreement between the classifiers can determine the feasibility classification. However, as discussed, while some embodiments apply a combination of classifiers (e.g., in an ensemble) and substantial preprocessing operations (e.g., the application of the Laplacian operator), in many embodiments, applying only one neural network (e.g., as discussed with respect to Figures 20A - 20C is sufficient to achieve sufficiently accurate filtering results on images with minimal preprocessing (e.g., the minimal preprocessing presented in Equation 7).
[0323] At block 2110b, in the case where the first set of one or more classifiers instead classifies the image as infeasible, the system can provide the image to a second set of failure mode classifiers at block 2110c. In some embodiments, the second set of one or more failure mode classifiers can also consider specific results from the first classifier to better facilitate classification (e.g., not only the binary infeasible or feasible results of a logistic regression classifier, but also the actual numerical values returned by the classifier).
[0324] After being processed by the fault mode classifier 2110c, a final classification result can be provided at block 2110e, but again, in this embodiment, edge case handling is first performed at block 2110d. For example, edge cases between can cause misclassification of one adverse situation as another adverse situation (e.g., without considering encoder motion, frequency analysis of the original image, etc., a large fluid blur covering most of the field of view may be confused with motion blur).
[0325] It should also be understood that the training process 2115 can be used for the first and second sets of classifiers in the process 2110. For clarity, given an image corpus, tolerance verification of downstream processing can first be used to label images as viable and non-viable. Then, according to the process 2115, this data set can be used to train and validate the first set of classifiers used at block 2110a. Then a second data set can be created by manually inspecting the images in the first data set marked as non-viable and labeling the image with its corresponding class (e.g., Figure 18A the adverse situation). Then, again according to the training process 2115, this second data set can be used to train and validate the second set of classifiers used at block 2110c. When training each of the first and second sets of classifiers, edge case verification from the operations in block 2115e can be used in the post-classification processing at block 2110d (for each set of classifiers).
[0326] Example Surgical Assessment including Visibility Records
[0327] Figure 22 is a flowchart showing various operations in an example process 2205 for inferring surgical performance data from viable and non-viable image classification results that can be implemented in some embodiments. Specifically, in addition to facilitating effective downstream processing, the application of the image viability classification system and method disclosed herein can also enable novel types of surgical operator assessment and feedback. Here, until the program is found to be complete at block 2205a, the system can consider newly arrived images at block 2205b (e.g., those images arriving directly from a surgical camera during the procedure, the oldest image in the queue of images from the camera for processing, etc.). In the case where a new image is available at block 2205b, the system can prepare a preprocessed version of the image (e.g., applying the operations of process 1910) for one or more neural networks (e.g., in Figure 21A and Figure 21B the networks and first and second sets discussed in, Figure 20A the network, etc.) to consider at block 2205f.
[0328] In some embodiments of process 2205, at block 2205g, the system can record not only the final classification of the neural network, but also various intermediate results. For example, the numerical outputs of SVM or logistic regression classifiers can be recorded, rather than just the final classification. Similarly, the individual weighted votes of the networks in an ensemble configuration, the numerical values of infeasible classifications in a serial configuration, etc. can be recorded. The results of classification post-processing (such as edge case handling) can also be recorded at block 2205g.
[0329] In the case where the classification indicates an image frame suitable for use in downstream processing (here, localization and associated mapping), the system can transition from block 2205k to block 2205l to predict the placement and integration of the derived depth data. The results of this integration can also be recorded at block 2205m (e.g., the pose determined during localization, such as a large spatial distance between successive successful pose determinations).
[0330] Although the evaluation of the recorded data can be performed after the program at block 2205a is completed, in some embodiments, intermediate evaluations during the course of the program can also be performed at block 2205c. Such intermediate evaluations can be particularly suitable for cases where a surgical procedure can be considered a series of discrete tasks. Thus, the system can evaluate the performance of the surgeon during and after a given task and can provide a real-time comparison with other surgeons performing the same or similar tasks. As indicated by block 2205d, it should be understood that there can be regular periods without new images, and thus the processing can be paused.
[0331] When the program is completed, at block 2205n, the system can view the records obtained at blocks 2205g, 2205i, and 2205m, and the results can be presented at block 2205p. The evaluation at block 22050 can consider the number of feasible and infeasible images throughout the program and under specific tasks. An increased number of infeasible frames during a task where the appropriate field of view is critical (e.g., polyp inspection) can contribute to a lower evaluation compared to a case where the same number or percentage of infeasible frames occur in a less sensitive task (e.g., metastasis to a tumor).
[0332] In addition to simple occurrence counting, the pattern of infeasible results can also provide information about the behavior of the surgeon and the context of the surgical procedure. For example, an increasing number of infeasible images during a particular task (which manifests across a large number of surgical procedures by different practitioners) can indicate that some aspect of the procedure, the organ, etc., consistently produces an unfavorable view from that orientation (as opposed to the actions of any given surgeon being the cause of the infeasible frames). Conversely, a task that typically produces few infeasible images for most surgeons but consistently presents infeasible images for a given surgeon can suggest that the given surgeon's execution of the task deviates from the methods of their surgeon peers in an undesirable way. The positioning and mapping can be adjusted accordingly, or feedback can be provided to the deviating surgeon.
[0333] Example Visibility Network Real - time Intra - operative Feedback
[0334] The application of the classifier described herein can be performed so quickly that its results can be used in real time, not only to avoid incorrect application of downstream processing but also to alert the surgical team that the current view is not suitable for downstream operations. The value of the various feedback methods disclosed herein will be recognized regardless of the particular way in which an image is determined to be feasible or infeasible. For example, Figure 23AThe GUI element 2305a is a schematic visual image that can be presented to a surgical team member in a GUI (e.g., on one or more of the displays 125, 150, or 160a during a surgical procedure or on a display depicting a playback of the recorded surgical procedure). Specifically, the first indicator 2305c can inform the surgical team or viewer of how many past video frames have been classified as invalid (e.g., if the frame rate of the surgical camera is 25 frames per second, the first indicator 2305c can show the percentage of the last 125 acquired images that have been classified as infeasible). Although shown here as a bar with a solid region 2305b (e.g., indicating the percentage of frames classified as infeasible), it should be understood that a numerical value, a dial, etc. can alternatively be used. Thus, instead of the raw number of infeasible classified images, the indicator 2305c can alternatively reflect a scaled or mapped value of the number of infeasible images. For example, the indicator 2305c can indicate a value in the range from 0 to 1, where 0 indicates that the number of infeasible images in the window is fully acceptable for the current operating state, and 1 indicates that the number of infeasible images is unacceptable. Such a mapping can facilitate adjusting the feedback to the user according to the current surgical context (e.g., after the mapping reaches an equilibrium state, during a non-sensitive part of the surgical procedure, etc., more infeasible frames may be acceptable). For example, occasional infeasible images in a well-traveled and well-mapped area of the colon may not require the operator's attention. In contrast, when the colonoscope has entered a new unmapped area of the colon or an area expected to contain sensitive information (e.g., a tumor or a polyp), the same number of infeasible images may have more dire consequences for downstream processing, and thus the need to inform the surgical team may be greater. Thus, in some embodiments, the surgical context can scale the value presented in the indicator.
[0335] While some embodiments simply inform the surgical team that there are infeasible frames, in some embodiments, the element 2305a (or other parts of the GUI) can include an indicator 2305d that provides guidance on why the system believes that one or more images (e.g., the most recently acquired images) are infeasible. For example, some operators who are not accustomed to operating with the assistance of a digital system may move through the colon too quickly for the system to maintain adequate positioning or modeling. Such undesirable blurring may present in the indicator 2305d a warning that the user's actions are producing infeasible images, especially since the user's blurring prompts advancement or withdrawal.
[0336] In some embodiments, as Figure 23BAs shown, the GUI presented to the surgical team (e.g., on one or more of displays 125, 150, or 160a) or presented to a playback viewer on a display can depict the mapping results at the current moment to inform the surgical team or viewer of the orientations in the model that may have been affected by infeasible images. The GUI elements can also include information on how to confirm that there are no adverse consequences following from infeasible images (e.g., by acknowledging and removing warnings) or how to fix the 3D model (e.g., by revisiting the area inside the patient corresponding to the part of the model affected by the infeasible frames). Here, the existing model 2310a (e.g., a triangulated textured mesh exported from a TSDF representation) and a representation 2310g of the current position of the colonoscope (such as an artist's 3D model of the colonoscope) are rendered. Regions of the models 2310d and 2310c are identified (e.g., with highlighted edges, vertices, texture changes, etc.) to inform the team or viewer that infeasible images were encountered during data acquisition at those orientations. While some embodiments may simply indicate the presence of infeasible images and only on that basis invite the team to revisit the area, as shown here, the billboards 2310f and 2310e indicate the infeasible classification of the images (or the most classifications that found the image sequence to be infeasible), providing the team with context for returning and correcting the problem. Selecting the billboards 2310f and 2310e or regions 2310d and 2310c of the model with a cursor 2310b, for example, can present additional relevant information such as the time, colonoscope orientation, and other context of the event.
[0337] As Figure 23C shown, some embodiments can combine feedback on the recent number of infeasible images with correction guidance. In this example, at a first moment, the GUI image 2315a indicates via an indicator 2315d that a large number of recent frames have been classified as infeasible (e.g., a linear representation similar to indicator 2305c). In the current context of the surgery, the number of infeasible images is high enough in quantity to trigger the system to apply the YOLO network to the field of view. Here, the network is trained to identify surgical instruments, and thus, the highlight 2315e indicates that one of the surgical instruments has prematurely occluded the field of view, contributing to images that are infeasible for downstream processing (e.g., preventing proper model creation for the region of interest or before proper camera positioning can be performed). After 2315c, the user has adjusted the position of the instrument 2315e in response, and feasible images begin to accumulate. Then, as shown in GUI state 2315b, the infeasible frame indicator 2315d can begin to decline, and the system can stop applying the YOLO network and the highlighting of the previously offending instrument 2315f.
[0338] Figure 23DThe exemplary process representation 2320 illustrates an example of such feedback behavior. Specifically, in this example, if the number of infeasible images has become critical at block 2320a, at block 2320c the system can seek to determine the nature of the error and present a correction graphic at block 2320d (e.g., a highlighting on the instrument 2315e after application of the YOLO network). In some embodiments, the acceptable percentage of invalid frames can vary with different procedures (e.g., an inspection procedure requires fewer invalid frames than a simple resection procedure), at different times or orientations within the same procedure (e.g., during sensitive parts of an operation near a tumor, or during mapping, but not when leaving the anus, fewer invalid frames), or during different tasks within a procedure (e.g., "initial mapping and orientation" itself can be a task within the procedure that requires fewer invalid frames compared to a pure mechanical resection task).
[0339] Even when the infeasible count has not yet risen above the threshold, a preventive warning at block 2320b can be appropriate (e.g., if the number of infeasible frames has been slowly increasing after an action such as application of a flushing device, the system can draw attention to the temporal correlation with a warning graphic, especially if the infeasible images are classified as depicting fluid blur).
[0340] It should be understood that at block 2320a, one or more ranges can be applied rather than a single threshold, depending on the surgical context. For example, when there are no infeasible images or only a small number of occasional infeasible images, the system can take no action. However, if there is a number of infeasible images below the threshold at block 2320a but associated with a growth trend (e.g., in each successive 100 - millisecond window, the number of infeasible images is increasing), then the graphic at block 2320b can be presented. In some embodiments, the nature of the increasing number of invalids can be investigated. If, for example, it is found that the frames are caused by motion blur, the warning graphics at blocks 2320b and ultimately 2320d can each invite the surgical operator to reduce their speed in order to reduce the resulting blur.
[0341] Figure 23E is a schematic illustration of a surgical tool obscuring a portion of the field of view of a surgical camera, similar to Figure 18C and Figure 23C the case. In some embodiments, the GUI depicted on one of the displays 125, 150, or 160a during playback viewing or on a desktop display can be semi - transparently overlaid or replace the depiction of the currently acquired three - dimensional model within the current field of view as an augmented reality element. In these cases, the portion of the model rendered in the GUI can be adjusted or supplemented to indicate areas of inadequate coverage, areas with no coverage, etc. Here, in Figure 23E as inFigure 18C In this case, an obstacle (here, the instrument 2340c, but fluid obscurations can be similarly identified) can occlude a portion of the field of view 2340b of the camera 2340a of the anatomical artifact 2340h, thereby creating a visible portion 2340d and a less visible portion 2340e of the artifact. Here, the currently created model is superimposed on the GUI field of view as an augmented reality element, such that a first portion 2340f of the data-exported model is presented to the user. However, a second portion 2340g of the augmented reality element can indicate that the field of view is insufficient (e.g., based on Figure 23C the YOLO results discussed in [reference], for example, by being textured or rendered differently from the portion 2340f. Detecting an infeasible image during acquisition can trigger such a superimposition and adjustment of the model rendering. For example, under normal circumstances, the unseen portion 2340e would be processed as a natural occlusion within the patient through the localization and mapping process, such as occlusion by a pouch fold away from the camera, curvature in the side wall of an organ, etc. However, here, because the video image frame has been identified as infeasible by the methods disclosed herein, the occlusion can instead be processed as an undesirable feature to be further investigated (e.g., via the application of a YOLO network). It should be appreciated that there are various methods for rendering the augmented reality portions 2340f and 2340g as, for example, a textured mesh or a two-dimensional billboard aligned within the plane of the camera field of view in the GUI rendering pipeline.
[0342] Example Post - operative Assessment Interface
[0343] Figure 24A is a schematic illustration of an element in a GUI that can be implemented in some embodiments for evaluating surgical performance based on the recording of image feasibility data. For example, Figure 24A the GUI element shown in [reference] can be presented at box 2205p (or at box 2205p if, for example, a real-time review is desired for a previously completed task). Here, a linear timeline 2405m with a current playback position indicator 2405n provides a vehicle for the user to quickly view the acquired data. As the playback progresses, the indicator 2405n can advance along the timeline 2405m. During playback, a camera playback area 2405k can depict the field of view of the surgical camera at the current time of the playback, and the images of the surgical camera are evaluated for feasibility. Similarly, a classification element 2405l can indicate the feasibility classification of the currently depicted frame. The YOLO results and other superimpositions presented during the surgery can also be presented in the playback area 2405k, for example, such that a viewer can understand what feedback was previously provided to the surgical team.
[0344] The GUI can also present a representation of the acquired depth data in the model 2405e, which can include a representation of the position 2405f of the camera at the current point in the playback (the model 2405e can be an artist's rendition, a model created during surgical mapping, or a combination of both). Here, a field of view 2405g corresponding to the visual field in the playback region 2405k can be similarly shown.
[0345] Throughout the playback or at the current position in the playback, the system can indicate regions in the model that are affected by or have encountered infeasible images. For example, a pop-up box 2405h here indicates that an image classified as depicting blurriness was encountered in region 2405i. The pop-up box 2405h also includes a time range that indicates that blurriness was encountered approximately 20 minutes into the surgery and persisted for approximately 2 minutes and 3 seconds (as determined, for example, by the range of blurriness-classified images that have less than a threshold number of consecutive feasible classified images within that range). By, for example, clicking on the pop-up box with the cursor 2405j, the user can indicate to the system to start playback, for example, at the first invalid classification frame associated with the pop-up box 2405h and region 2405i (some embodiments can start playback at a time before the selected interval to provide context leading to the infeasible image classification).
[0346] Just as an identification such as the pop-up box 2405h can be used to indicate the location where an infeasible image was encountered in the spatial context of the model 2405e, an identification can also be provided to indicate a time orientation of interest. Here, for example, identifications 2405o, 2405p, 2405q, 2405r indicate the time along the timeline 2405m when various sequences of one or more infeasible classification images occur. For example, the identification 2405o can indicate that an infeasible image associated with occlusion may have been encountered at the corresponding time. The identification 2405p indicates that the acquired depth values failed to integrate correctly, which is an event that can occur even when a frame is classified (possibly incorrectly) as valid. Here, the presence of a failure to integrate an image classified as valid can indicate that new, previously unencountered situations are contributing to infeasible images (e.g., situations unique to the procedure and not represented in Figure 18A the case, such as when the surgeon has a particular way of moving or placing a surgical instrument). Thus, such failures can be used to identify images and new labels for future neural network training rounds (e.g., the frames associated with the identification 2405p will be labeled with a new class of the surgeon's particular behavior).
[0347] Similarly, the identifiers 2405q and 2405r indicate the time associated with the acquisition of an image classified as containing ambiguity. In some embodiments (such as those having a neural network with a binary output 2005g), rather than including any indication of the reason for the infeasible image in the identifier, the identifiers 2405o, 2405p, 2405q, 2405r, and 2405o may simply indicate that the image was found to be infeasible. Similarly, just as the indication 2405i may correspond to multiple infeasible image classes, the timeline identifier may include regions indicating a consecutive number of frames vulnerable to invalid classification, such as regions 2410a, 2410b, and 2410c. Although not all images in regions 2410a, 2410b, and 2410c may share the same infeasibility classification, a single consecutive classification may be assigned to the region, as here, where each instance within a consecutive number of other classified images is below a threshold (e.g., for an acquisition rate of 45 frames per second, fewer than 15 consecutive other classified images may be ignored).
[0348] In fact, in some embodiments, the timeline 2405m itself varies in terms of hue, brightness, intensity, etc. corresponding to the moving window average of the feasible and infeasible classifications (in some embodiments, where there are multiple infeasible classes, each class may receive a unique hue or texture). For example, when all images within the moving window are classified as feasible, the brightest value in the range may be used, and when all images within the moving window are classified as invalid, the darkest value is used.
[0349] As mentioned, some surgical procedures can be easily divided into recognizable "tasks" or relatively discrete groups of actions within the procedure (e.g., "advance to the resection site", "resect", "cauterize", "examine after cauterization", "withdraw", etc.). Selecting an appropriate task from the list 2405a can result in playback starting from the first associated image frame of that task (e.g., updating the image in region 2405k, changing the orientation of the position indicator 2405n, etc.). In some embodiments, the acquired model or reference model 2405b may be divided into regions. Here, for example, the model is divided into seven regions, including regions 2405c and 2405d. Just as selecting a task from the list 2405s starts playback in that region, so too, for example, selecting a region using the cursor 2405j may start playback at the first frame associated with that region (e.g., at the user's indication, only the withdrawal or advance encounter may be considered, both distinguished by the time between the encounter with the middle of the region and the start orientation of the encounter). It should be understood that the networks herein (such as Figure 20BThe network) can be easily configured to include orientation and other contextual information in its training and inference inputs. That is, the system can consider not only the pixel values of the image when evaluating the validity or invalidity of an image frame, but also the current orientation of the colonoscope in one of the regions of model 2405b, the task from list 2405a being performed during image acquisition, etc.
[0350] In some GUI implementations, the current surgical performance can be compared to other performances (e.g., of the same or different surgical operators). In some embodiments (e.g., those implementing binary classifier 2005g), the comparison can simply be between the number of infeasible and feasible images overall for each surgery within a particular task, and the frequency of infeasible intervals in the tasks within the surgery, or the number of consecutive infeasible images. Plotting the incidence over time can help the user identify patterns in their behavior that contribute to infeasible images, and the parts of the surgery that typically produce infeasible images. Such information can help the user adjust their behavior in the future to minimize or compensate for such events (and help technicians identify new labels and edge cases).
[0351] In the depicted example, the various infeasible image classes identified by one or more classifiers are presented to list 2405s. Here, the user has selected the occluded infeasible image class 2405t and the fluid blur infeasible image class 2405u indicated by the highlighted borders. Corresponding plots 2405v and 2405w can then be generated in plot area 2405z, indicating the occurrence of each selected infeasible image classification type over the course of the procedure for the population of surgical operators. Thus, timeline 2405x can correspond to timeline 2405m, and a similar adjustment of indicator 2405y can correspondingly adjust the playback. A bar chart or other suitable representation can also be used. Similarly, it should be understood that plots in region 2405z can be generated using cumulative counts, average counts within a sliding window, and other representations of the data (rather than the raw number of infeasible image counts for the class). In some embodiments, the average results from a corpus of similar surgeries performed by the same or other operators can be similarly overlaid to provide relative context.
[0352] The playback can be integrated across GUI elements. For example, Figure 24BA flowchart is provided that illustrates various operations in process 2410 for responding to a user playback position selection (e.g., clicking and dragging indicator 2405n). At block 2410a, the system can receive a newly selected playback position. At block 2405b, the corresponding part of the model(s) (e.g., model 2405b or 2405e) can be highlighted. For example, the representation of the camera position 2405m can be adjusted to indicate the orientation and direction of the camera at the selected playback position. Similarly, the appropriate area can be indicated in representation 2405b. In the case where the task corresponds to the current playback position, the corresponding icon can be highlighted in list 2405a. At block 2410c, the video playback in area 2405k can likewise be adjusted to the newly selected position. At blocks 2410d and 2410e, the system can retrieve the records within a threshold window of the newly selected playback time and can present them to the user. For example, the classification of the currently depicted frame can be updated in indicator 2405l, the corresponding pop-up box on the model (such as pop-up box 2405h) can be presented or highlighted, peer data can be presented in the drawing of area 2405z, and so on.
[0353] Similarly, just as a user can isolate data of interest by selecting a temporal orientation, Figure 24C A flowchart is provided that illustrates various operations in process 2415 for selecting information via a spatial selection in some embodiments. Specifically, at block 2415a, the system can receive a spatial orientation, such as by clicking on model 2405e or an area of model 2405b (e.g., using cursor 2405j), etc. Since multiple time points can correspond to the same orientation (e.g., a colonoscope can often pass through an area at least twice, once for insertion and once for withdrawal), at block 2415b, the system can highlight the portion of the timeline (e.g., timeline 2405m) that the camera passes through within a threshold distance of the selected orientation. In some embodiments, the playback can be adjusted to the first of such temporal orientations at block 2415c (or the playback can be adjusted in response to a user selection of one of the highlighted timeline regions), but again, it can be possible to select the advancement and withdrawal encounters of the entry region based only on the direction of movement and the entry point into the region. Similar to the temporal selection of records in process 2410, at blocks 2415d and 2415e, the system can retrieve the records associated with the selected orientation. For example, the classification of the currently depicted frame can be updated again in indicator 2405l, the corresponding pop-up box on the model (such as pop-up box 2405h) can be presented to highlight, peer data can be presented in the drawing of area 2405z, and so on.
[0354] Iterative Internal Body Structure Representation - Overview
[0355] Various embodiments disclosed herein contemplate a surgical navigation service that facilitates real-time navigation during a surgical procedure, e.g., as in operating room 100a or operating room 100b (it should be understood, however, that some embodiments can be readily applied with necessary modifications during a post-surgical review). Such a system can monitor the progress throughout a surgical procedure and provide guidance to a control system or a human operator in response to the status of that progress. For example, in a colonoscopy, the navigation system can guide the operator to an unexamined area within the patient, such as the colon, and can determine an estimate of the coverage of the procedure (such as the remaining percentage of the colon that is considered to remain unexamined). As described in more detail herein, the coverage can be estimated, for example, by comparing two extreme points that the colonoscopy camera has traveled to the estimated total length of the colon being examined. Also disclosed herein are various graphical feedback methods that the system can utilize to suggest the status of the progress of the procedure to the operator or viewer. While many of the examples disclosed herein are in the context of a colonoscopy, the application with necessary modifications will be readily understood in other surgical contexts (e.g., in a pulmonary context, such as an examination of a bronchial passage, an esophageal examination, an arterial context during stent delivery, etc.).
[0356] Figure 25 is a schematic sequence of states of the model, view, and projection mapping regions of the GUI during a coverage assessment process that can be implemented in some embodiments. Specifically, for example, a graphical interface presented to one or more of surgeons 105a or 105c or ancillary members 105b, 105d on a display 125, display 150, display 160a, etc. or presented to a viewer who examines surgical data after a surgery on, for example, a desktop computer can include one or more of the model, view, and projection mapping regions in a window, frame, translucent overlay, or other display area (such as a helmet). In some embodiments, the views can be displayed simultaneously, or can be displayed individually and alternately with a selector (thus, for example, different displays among displays 125, 150, 160a, etc. can display different views simultaneously during a surgery).
[0357] At the first time 2500a, the system may present to the user one or more of the following: a model area 2550a depicting a partially constructed three-dimensional model 2505a of an internal body region (here, a part of the intestine); a view area 2510a depicting a camera view of a surgical instrument, and a projection mapping area 2535a depicting a two-dimensional "flattened" image of the surface of the internal body region (here, the internal texture of the intestine, where unexamined areas that have not been assigned a surface texture may be indicated by, for example, black pixels). The projection mapping area 2535a can be used to infer a coverage state (e.g., the percentage of the entire area covered in a leak). For clarity, each of the GUI areas 2550a, 2510a, 2535a may appear on one or more of the display 125, the display 160a, the display 150, a separate computer monitoring display, etc.
[0358] While the area 2510a may depict the output from a surgical camera (such as a colonoscope), each of the areas 2550a and 2535a may depict a corresponding representation having portions reflecting inadequately examined areas inside the patient. In some embodiments, an inadequate examination may include areas that have not been directly viewed with a surgical camera (e.g., when they are obscured by intestinal folds or surgical instruments). However, in some embodiments, an inadequate area may be an area that has not been adequately viewed for a given surgical scenario (e.g., polyp search may require a minimum amount of time to view a given area, tissue recognition using a neural network may require a minimum amount of blurriness, etc.), viewed without proper filtering, for an inappropriate duration, with inappropriate laparoscopic insufflation, inappropriate staining, etc.
[0359] In the depicted example, a portion 2520a of the incomplete model 2505a has not been adequately viewed by the operator using the surgical camera. The portion 2520a in the region 2550a can be identified via the absence of model faces, faces with a specific texture or color highlighting the omission, edge contours corresponding to the omitted faces of the model, etc. Similarly, for an incomplete model, at time 2500a, a majority 2525a of the model 2505a can be outside the range of the camera or depth determination system. Each of the omissions 2525a and 2520a can have corresponding representations in the flattened image of the region 2535a, specifically regions 2540a and 2545a of the flattened image. That is, as the model 2505a is gradually generated as the surgical camera passes through the patient's interior, the corresponding texture map of the interior can be "unrolled" onto the two-dimensional surface of the region 2535a (similar to the UV mapping of texture coordinates between the faces of a three-dimensional model and a two-dimensional plane). In some embodiments, navigation arrows or other icons 2515a can be used to inform a viewer from the perspective of the model 2505a of the current relative orientation of the camera providing the view in the region 2510a (as shown in this example, the occluded faces of the model may not be rendered around the icon 2515a, but variations will be readily understood, e.g., where the icon 2515a is rendered on a billboard between the model and the viewer and the intermediate model faces are rendered semi-transparently, etc.). As indicated, the portion of the region 2535a outside the omissions 2545a and 2540a can be rendered using the bowel texture acquired using the camera. For clarity, although localization and mapping occurring during forward advancement of a colonoscope through the colon are shown here, resulting in the creation of additional model segments, it should be understood that in some colonoscopy procedures, mapping and localization can be performed only during withdrawal, or mapping during withdrawal can supplement the results from advancement.
[0360] Since the region 2550a depicts the model 2505a from a three-dimensional perspective, it may be difficult for the operator or assistant to identify the relative positions of the omissions solely from the region 2550a. While translucent faces, billboards, and other graphical methods on the model 2505a (e.g., such as the method described as rendering fiducial points 2515a) can be readily used to highlight the omissions to the operator on the opposite side of the model or in an orientation occluded by the current perspective of the model, this method can become confusing in the presence of multiple omissions. Similarly, inviting the operator or assistant to rotate or translate their perspective relative to the model 2505a to confirm the relative orientation of the omissions under the time constraints and other priorities of the surgical procedure is generally not ideal. Thus, the two-dimensional representation of 2535a facilitates quick and more intuitive guidance by which the operator or viewer can readily assess the current situation.
[0361] Thus, as time progresses from 2530a to subsequent time 2500b, each region can change its state according to the progress of the surgical procedure. Here, at time 2500b, region 2550a now depicts the supplemented partial model 2505b, region 2510a depicts the field of view of the camera in a more advanced position in the intestine, and region 2535a depicts a more textured surface. Since the operator has advanced the camera without taking time to remedy the omissions when encountered, new omissions appear in the model, including new omissions 2520b, 2520c, and 2520d. Correlatively, omission 2525b corresponding to the region not yet examined has replaced omission 2525a, and the arrow icon has advanced to the new orientation 2515b corresponding to the advanced position of the camera.
[0362] The updated representation in region 2535a will reflect the presence of the newly introduced omissions. For example, omission 2520b corresponds to the flattened region 2545b, omission 2520c corresponds to the flattened region 2545c, and omission 2520d corresponds to the flattened regions 2545d and 2545e. It should be understood that in the depicted example, the coordinates of the approximately cylindrical structure from the model are mapped to region 2535a (it should be understood that the depiction here is schematic). Thus, while the vertical dimension of region 2535a corresponds to the longitudinal progress of the colonoscope, the horizontal dimension of region 2535a is mapped to 360 degrees of the approximately cylindrical intestine.
[0363] Thus, for the convenience of the reader and to further facilitate understanding (although, as will be discussed, in some embodiments, similar superimposed notations may be provided to the operator), at time 2500b, reference 2555a is shown in the figure, which correlates the 360 degrees of the field of view of the camera at the horizontal entry of region 2535a with the 360 degrees of the field of view of the camera in region 2510a. As shown, the top of the camera is considered to be at the 180-degree azimuth in this example, corresponding to the center of the horizontal dimension of region 2535a. Conversely, the 0 and 360 positions (equivalent) in the camera view of region 2510a correspond to the leftmost edge of region 2535a. Since the mapping causes the "bottommost" position in the colon view to appear as a wrap-around at the edge of region 2535a, omissions (such as omission 2520d) that appear at the "bottom" azimuth of the camera can correspond to two regions, particularly regions 2545d and 2545e along the same one or more horizontal rows of region 2535a, but at opposite edges of the region (i.e., they refer to the same omission 2520d).
[0364] For further clarity, Figure 26A and Figure 26B respectively depict at Figure 25An enlarged view of the model and projection mapping state at time 2500b in []. Here, two perimeters 2605a and 2605b of the model 2505b corresponding to rows 2620a and 2620b in region 2535a are shown (again, it should be understood that the view is schematic and, for example, the mapping from 2520a and 2520b to regions 2545a and 2545a is not exact). Each perimeter 2605a and 2605b can be determined, for example, as the nearest points on the model 2505b in a circle including points on the center line 2650 of the model (e.g., the central axis center line, as manually determined, as programmatically determined based on model moments, as inferred from colonoscope kinematics, etc.). Here, for example, perimeter 2605b is determined by point 2650a on the center line 2650. Thus, each row in the image of region 2535a can be determined by the corresponding perimeter on the model 2505b. Thus, along the 90-degree continuous line 2555d (corresponding to reference line 2630c), portion 2610a of perimeter 2605a corresponds to point 2615a on row 2620a. Similarly, portion 2610b of perimeter 2605b corresponds to point 2615b on row 2620b. It should be understood that portions of the perimeter that encounter gaps will similarly contribute to the gap regions in the corresponding portions of the rows in the projection mapping region 2535a. For example, having determined a point on the center line and looking for the nearest vertex of a ray extending from that point in a given radial direction (e.g., 90 degrees), if no model vertex is within a suitable threshold distance of the ray, the radial direction can be associated with a non-textured or missing value (e.g., pixels in the column of the row corresponding to the perimeter associated with the radial direction can indicate a missing rather than an acquired texture value).
[0365] It should be understood that since the mapping process can be substantially temporally continuous, the system can be able to infer the orientation of the camera relative to previously mapped segments and relative to its current field of view. Similarly, since the perimeters are determined from the vertices and center line of the model rather than from the current position of the camera, the camera can assume various orientations without disrupting the generation of the three-dimensional model or the corresponding projection mapping. That is, even if the camera does not enter the region at any particular angle and even if only a portion of the perimeter is visible, the system can easily assign degrees to model the vertices in the perimeter. In this way. The presence of the perimeter can contribute to a "common" coordinate set for the operator.
[0366] For clarity, Figure 26C is a schematic representation of a pair of relatively rotating surgical camera orientations 2675a, 2675b and their corresponding fields of view 2670a, 2670b. Initially, the surgical camera can be in a first orientation 2675a slightly below and to the left of the center line 2650 of the three-dimensional model. Naturally, this can produce a field of view 2670a, similar to what has been described with respect toFigure 25 Field of view. When the camera rotates 2690 counterclockwise about its longitudinal axis 2680, naturally, the field of view 2670b will rotate correspondingly. However, as indicated by the reverse arrow 2685, the system will continue to interpret the images acquired from the camera relative to the original orientation. Thus, whether the camera advances in orientation 2675a or in orientation 2675b, the system will generate the same image in region 2535a because the same corresponding rows and columns in the image will be used to generate the same perimeter in the model.
[0367] Return to Figure 25 , as time advances again 2530b to time 2500c and the surgical procedure continues, region 2510a may depict a more advanced field of view of the camera, region 2550a may show a correspondingly more complete model 2505c, where the icons correspond to the current camera orientation at the advanced position 2515c, and region 2535a may show a mostly fully textured two-dimensional plane (e.g., where the full length of the model corresponds to the full length of the expected portion of the colon to be examined). In this example, when the end of the intestine remains open, the leak 2525c remains corresponding to the residual leak region 2540c. In this embodiment, rather than correcting the leak along the way, the operator selects to return to an earlier part of the examination after the initial progression, and then removes the leak at that orientation by examination.
[0368] Specifically, when time advances 2530c to time 2500d, the user has selected to address the leak 2520c (corresponding to region 2545c), and thus returns to the orientation of the leak 2520c and brings the missing portion of the intestine into the field of view of the camera, as shown in region 2510a. By bringing this region into view, the corresponding leak is removed from the model in view 2550a (it will be noted that the reference orientation 2515d points to the previous position of the leak 2520c at time 2500c). This results in a new model version 2505d in which the leak has been filled. Similarly, at time 2500d, the region 2545c corresponding to the leak 2520c is also omitted from view 2535a.
[0369] Iterative Internal Body Structure Representation - Example Processing Overview
[0370] Figure 27 is a flowchart illustrating the various operations in an example process 2700 for performing Figure 25 the coverage assessment process. Although the various operations may be discussed in more detail herein with respect to Figure 32 and Figures 33A - 33C nevertheless Figure 27Provides an overall overview. Specifically, during a surgical procedure, the system can be initialized at block 2705, preparing an initial model mesh (e.g., an empty space or a space with a guiding structure as discussed herein with respect to Figures 34C - 34F ). And a surface projection (e.g., where the entire image is a single monolithic void). Although a two-dimensional rectangle is depicted and discussed as the projected image region with respect to Figure 25 , it should be understood, as discussed elsewhere herein, that the mapping can be to a surface other than a two-dimensional rectangle.
[0371] At block 2710, the system can determine whether the monitoring of the surgical procedure is complete. For example, the operator may not desire void identification at all times during the entire procedure, or the procedure can end. If the monitoring has not ended, then at block 2715, the system can determine whether new depth frames and image data are available, and if such data are not yet available, wait as indicated by block 2755.
[0372] Once new data are available, the system can acquire new image and depth data at block 2720. At block 2725, the system can then update the model with the depth data, e.g., extending model 2505b to a new partial model 2505c. With the updated model, new model vertices can be used at block 2730 to extend the centerline (again, alternative methods for extending the centerline should be understood, e.g., based on camera motion, encoders, etc.). The extended centerline can in turn be used at block 2735 to determine a new perimeter (e.g., one of perimeters 2605a or 2605b). The vertices of the perimeter are themselves associated with faces, which can themselves be associated with texture coordinates from the visual image. Thus, the system has a ready set of references for inferring row pixel values (e.g., pixel values in rows 2620a and 2620b). For example, each of the 360 degrees in the perimeter can be used to identify the pixel values in the corresponding columns in the rows of the perimeter (or, as mentioned, infer the void values for a given radial direction).
[0373] In some embodiments, although the omission may be self-evident to the user from the rendering, the system may not be able to unambiguously identify the omission. However, in some embodiments, at block 2740, the system may identify the omission, for internal processes and / or for highlighting to the operator. It should be understood that omission identification may be performed in a variety of ways on the model or the projected image. For example, flood fill algorithms and blob analysis provide ready methods for determining groups of pixels associated with an omission in the projected image. Just as the perimeters in the model are used to infer texture values for filling rows in the projected image, once an omission is identified in the image, the system can review the corresponding perimeter for that row to identify the three-dimensional orientation of the omission. As another example, regions of the model that lack a threshold number of nearby vertices can likewise be interpreted as omissions.
[0374] Graphic Supplements for Navigation and Orientation
[0375] Figure 28 is something that can be implemented in some embodiments Figure 25 A schematic state sequence of the model, view, and projection mapping regions that can be implemented, but with additional graphical guidance. For example, the view region 2510a may include a direction compass 2820 that notifies the operator of nearby omissions (and in some embodiments also indicates the current rotation of the camera relative to the nearest perimeter, e.g., via a corresponding relationship from hue to radial direction). Similarly, the projection mapping region 2535a may include a local indicator 2850a that shows the current relative position and orientation on the projected image of the camera depicting the current view in the view region 2510a. The orientation of the camera can be visualized via the local indicator 2850a as, for example, a circle in a different color from the internal body texture on the mapping 2535a, or other notations (such as the arrow shown here in the first position). In some embodiments, the portion of the image that is currently in the camera view can also be indicated with a notation 2855a (e.g., with an outline, a change in brightness of the image pixels, a colored border, etc.). The lower right portion 2860 of this region may include a text overlay indicating various monitoring statistics. For example, the text overlay may indicate the percentage of the intestine that is unmapped or has been mapped relative to a standard reference, the ratio of omissions to the completed portion of the model, the length of the intestine in centimeters that the camera has traversed so far, the number of existing omissions, etc. (e.g., the insertion depth "13 cm" of the current length of the global centerline and the coverage score "67.1%" according to the ratio described herein). The length traveled as well as the lengths corresponding to the unmapped or mapped percentage of the intestine can also be displayed. The length of the global centerline can be determined at each time point, such as the geodesic distance between the two ends of the global centerline.
[0376] The depicted coverage score can be determined as the ratio of the mapped colon to the total predicted mapped area (e.g., averaged from a corpus of colon models and scaled by the patient's dimensions). Local coverage such as the current segment of the colon where only the colonoscope is present or only the magnified area (such as magnified area 2970b discussed below) can also be depicted. As yet another example, the coverage score can be calculated only for a previously surveyed area to inform the surgical team of gaps in the surface area. That is, referring to time 2500a, rectangle 2870 (also referred to as the "current surveyed area") on image region 2535a indicates the portion of image region 2535a used for coverage calculation, where the width of rectangle 2870 is the same as the width of image region 2535a, and the height of rectangle 2870 corresponds to the row furthest from the starting row containing mapped pixel values. The coverage score for this rectangle can be determined based on the ratio of missing and non-missing pixels in rectangle 2870, e.g., the number of pixels associated with the missing (including the unmapped portion in rectangle 2870 where, as here, the terminating surface of the mapped area is not flush with rectangle 2870) in the numerator and the total number of pixels in the rectangle in the denominator (or, conversely, the total number of pixels in the rectangle minus the pixels associated with the missing in the numerator and the total number of pixels in rectangle 2870 in the denominator). Thus, the fraction in the depicted example is the ratio of the sum of the pixels in missing area 2545a within rectangle 2870 and the pixels in unmapped area 2870b divided by the total number of pixels in rectangle 2870.
[0377] Local indicator 2850a or compass 2820 can guide the user to look in the direction of the gap in order to remedy the defect. Here, for example, at time 2500a, a portion 2810a of compass 2820 is highlighted to inform the user that gap 2520a is above and slightly to the left of the current field of view (in some embodiments, the missing area 2545a can likewise be highlighted). As will be discussed in more detail herein, the system can consider one or more of the following when determining which direction to recommend in compass 2820: the input camera image; the predicted depth map; the estimated pose of the camera when determining the camera's recommendation; and the centerline from the start of the sequence (e.g., in the cecum if the operation is being performed during withdrawal) to the current camera position. For example, the system can consider points along the centerline and the corresponding perimeter within a threshold distance of the current position of the camera. Gaps falling on those perimeters can then result in corresponding highlighting (such as highlighting 2810a). In some embodiments, all such gaps contribute to the same colored highlighting 2810a in compass 2820. However, in some embodiments, gaps in front of the camera can be highlighted in a first color (e.g., green), and gaps behind the camera can be highlighted in a second color (e.g., red) to provide further directional context. As will be discussed with respect toFigure 30C discussed, the color can alternatively indicate the radial position, and the forward or backward relative position can alternatively be indicated by the color or pattern around the boundary of the highlighted portion (e.g., the lack at the back and at 180 degrees can contribute to a light blue highlight and a red boundary). When the camera rotates, for example, as discussed with respect to Figure 26C the compass 2820 can rotate within the camera's field of view to inform their users of the camera's orientation relative to the model coordinates. In a further embodiment, for example, as discussed in Figure 29A the compass 2820 only indicates the lack in the perimeter (or the perimeter within a threshold distance) from the current position of the camera.
[0378] At time 2500b, the user advances the camera further into the intestine again, such that the compass 2820 and the local indicator 2850a are similarly updated (and also updated to the field of view notation 2855b in this embodiment). Here, the highlight 2810b corresponds to the lack 2520e (and region 2545f), and the highlight 2810c corresponds to the lack 2520d (and regions 2545d and 2545e; for clarity, the part of the lack 2520d wrapped under and around the model is not shown from the reader's perspective), because each of these lacks falls within the threshold distance from the current position of the camera. Similarly, in some embodiments, these lacks can be highlighted in the corresponding regions of the image area 2535a.
[0379] At time 2500c, when the user advances further into the colon, the local indicator 2850a and the notation 2855c can be updated again. In the depicted embodiment, the user has moved the cursor 2810f over the region 2545c corresponding to the lack 2520b of interest. After clicking and selecting the lack, the system can ignore the local threshold criteria for updating the compass 2820 and instead provide a highlight to guide the user to the selected lack (here, the highlight 2810d). In some embodiments, overlays, augmented reality projections, etc. can also be integrated into the region 2510a. For example, here, a three-dimensional arrow 2810e has been projected into the space of the field of view to guide the user to the selected lack 2520b.
[0380] Once the selected lack has been remedied at time 2500d, the compass 2820 can be cleared of highlights, as in the described embodiment. In some embodiments, the system can alternatively revert to depicting other lacks near within the compass 2820. For clarity, since the user is looking at the "ceiling" of the intestine, the notation 2855d only contains the corresponding central part of the two-dimensional image.
[0381] Graphic Supplements for Navigation and Orientation - Example Local References
[0382] As discussed herein, in some embodiments, the system can identify gaps that occur before and after the colonoscope position and draw the attention of the surgical team thereto. However, in some embodiments, the gap identification can be localized to specific perimeters in the model (e.g., perimeters 2605a and 2605b). That is, compared to the methods discussed hereinbelow Figure 30C (where highlighted and other regions are identified before or after the current position of the colonoscope), some embodiments can alternatively limit the notification to a specific perimeter, such as the perimeter where the colonoscope is currently located.
[0383] For example, Figure 29A depicts a pair of schematic and projected mapping regions for a local compass range that can be implemented in some embodiments. Here, the system superimposes the compass guidance element 2905b on the GUI element 2905a depicting the current field of view of the colonoscope. Here, the compass 2905b guides the operator's attention to the upper left quadrant of the colonoscope's field of view (e.g., corresponding to the presence of gap 2545f) via the highlight 2905c. This is also reflected in the two-dimensional projected mapping diagram 2910a, where only gaps that at least partially intersect a perimeter (e.g., perimeter 2950a) within a threshold distance of the current position of the colonoscope are considered for representation by corresponding highlights in the compass 2905b.
[0384] In this embodiment, the gap 2545f fully falls within one or more of the considered perimeters (here, perimeter 2950a). In cases where only a portion of the gap falls within the perimeter, only that portion of the compass corresponding to the intersecting portion of the gap can be highlighted. For clarity, the perimeter 2950a is shown relative to the incomplete three-dimensional model 2505b as in the perspective view 2960. Thus, the compass 2905b may not display a highlight, such as highlight 2905c, until the colonoscope is placed at a location (e.g., a location associated with point 2920). Such an embodiment can be useful, for example, in surgical procedures where the examination occurs during withdrawal. That is, in some surgical procedures, the initial advancement to the end near the cecum is only mainly performed to prepare for a subsequent procedure, such as the examination of the colon. Then, localization, mapping, and gap remediation occur along with the slow withdrawal and examination. Compared to relying solely on the operator's judgment, perimeter-by-perimeter verification can facilitate a more organized view. Thus, in addition to the gaps caused by the incomplete model, regions in the model and projected mapping diagram corresponding to areas that the colonoscope has not viewed for a sufficient length of time can similarly draw the operator's attention in the same manner as model gaps.
[0385] As shown in the two-dimensional projection map 2910a, arrow notations 2925 or single points 2920 can be used to indicate the current position of the camera and its relationship to the perimeter 2950a. As shown in the two-dimensional projection map 2915a, only single points or other small identifiers can be used to provide a less prominent representation of the position. Similarly, in some embodiments, neither points nor arrows are presented (i.e., there is no point-based notation), but only the rows of the two-dimensional projection map corresponding to the perimeter within a threshold distance of the current colonoscope position are recorded (e.g., via a rectangular bounding box, via the brightness change of a row relative to other rows, etc.). Naturally, the highlighting of the perimeter can also be combined with the position or orientation.
[0386] For the sake of completeness of the reader's understanding, the perimeter 2520b is also shown in this example. If the colonoscope is in position 2980a, the element 2905a can appear as shown in state 2980b, and the compass 2905b may not include any highlighting. Similarly, the two-dimensional projection map 2915c will indicate that for the current camera position indicated by the point 2915d, the nearest perimeter 2520b or the perimeter within the threshold distance of the current position does not intersect the defect.
[0387] Using the methods disclosed herein, it should be understood that the operator can sometimes benefit from GUI elements that present portions of the model and projection maps at different levels of detail. For example, Figure 29B is a projection mapping GUI element with detail level magnification that can be implemented in some embodiments. Specifically, an example projection map element 2915a of Figure 29A is shown here, where the bounding box 2970a indicates that a portion of the element 2915a appears in the magnified area 2970b. As in the magnified area 2970b, one or more increased levels of detail can be presented to the surgical team, for example, superimposed on the map element 2915a, cascaded to the mapped map element 2915a, superimposed to show different GUI elements, etc.
[0388] The magnified area 2970b can assist the operator in aligning the camera position relative to the perimeter (e.g., perimeter 2950a). This can facilitate local remediation of defects at a higher resolution than the more global representation used for the map element 2915a. As regarding Figure 29AAs described in the embodiments, the navigation compass can draw attention to the missing areas of the "surround" camera. In fact, in some embodiments, portions of the rectangle 2950a corresponding to one or more perimeters can have areas that are highlighted corresponding to the highlighting of the compass (e.g., in the case of highlighting the top 180-degree position of the compass, the center of the rectangle 2950a can be similarly highlighted in the magnified area, but the missing rendering may have made the correspondence clear). Thus, the operator can compare the feedback from the compass with the representation of the perimeter (e.g., the perimeter 2950a in the magnified area 2970b).
[0389] In some embodiments, the magnified area can automatically follow the area around the current position of the colonoscope. This behavior can be useful in cases where extending and projecting the image in only one resolution can cause various details to degrade (making it difficult for the surgical team to identify small missing areas, for example).
[0390] Example Variations in Graphic Supplements
[0391] For further clarification, Figure 30A is a schematic representation of a continuous navigation compass 3005 (shown here in circular form) that can be implemented in some embodiments. Here, the exemplary missing area 3010a is in front of the camera's field of view and in the upper left quadrant of the camera's field of view, and has a corresponding highlighted area 3015a on the compass 3005. Similarly, the missing area 3010b is in front of the camera's field of view and in the upper right quadrant of the camera's field of view, and has a corresponding highlighted area 3015b on the compass 3005. In the depicted embodiment, the sizes of the highlighted areas 3015a and 3015b correspond to the projected sizes of the corresponding missing areas.
[0392] Although there are many ways to correspond the highlighting and the missing areas, Figure 30CProvides a schematic representation of a series of states in determining the relative position for display highlighting that can occur in some embodiments. Specifically, as shown in state 3000a, by envisioning a cylinder 3040 around the centerline 3045 of the field of view of the colonoscope 3050, the system can infer the corresponding part of the compass to highlight a given omission (although only a part of the cylinder in front of the camera is shown, but a variant where the cylinder extends behind the camera to accommodate the projection of the omission behind the camera can be easily understood). For example, the cylinder 3040 can have the same radius as the navigation compass (e.g., compass 3005). The system can project the omission 3010b 3020a and can project the omission 3010a 3020b onto the surface of the cylinder 3040. This can result in a projected shape 3030a on the cylinder surface, as shown in state 3000b. By compressing the shape along the perimeter of the cylinder, or by considering the radial boundaries in the cylinder, a restricted shape 3030b can be inferred from the shape 3030a, as shown in state 3000c. Then the restricted shape 3030b or the corresponding boundary can be mapped to the dimensions of the compass 3035a to determine the highlighted part 3035b, as shown in state 3000d. It should be understood that this is only an example, and highlighting can be easily determined by other projections (e.g., direct projection on a circle in the plane of the camera's field of view).
[0393] As mentioned, in some embodiments such as those discussed in conjunction with Figures 29A - 29B where only the omissions intersecting the radial direction of the perimeter where the colonoscope camera currently resides can be depicted on the compass, it may not be necessary to distinguish the omissions along the longitudinal axis. In contrast, in embodiments corresponding to Figure 30C the example where the omissions in front of or behind the current position of the colonoscope can be represented on the compass, overlapping omissions can be distinguished by different colors, boundary profiles, indication of the number of omissions in the region, etc. In some embodiments, each omission or at least each considered omission can be assigned a different color or other unique discriminator, and then these unique discriminators are used to distinguish the omission representations within the compass.
[0394] Returning to Figure 30A , the highlights 3015a and 3015b can thus be determined in the manner described with respect to Figure 30C As mentioned, the highlights may not be the same color and can actually depict various colors to enhance the relative position of the camera and the omissions. For example, as Figure 30DAs shown, colors can be assigned to the 360 radial direction range in the model (it should be understood that the granularity can vary, and thus the mapping from hue to angle can be continuous or discretized at different levels). In the case of representing colors using hue, saturation, and lightness (HSL), it should be understood that when represented using eight bits, the hue can thus have a value between 0 and 255 (thus, red can correspond to the value 0, green can correspond to the value 90, blue can correspond to the value 170, etc.). Figure 30D Illustrated is the correspondence between the radial degrees from 0 to 360 via reference 3090a and the range of 0 to 256 hue values shown by reference 3090b. Thus, the top part of the compass can adopt light blue and green values, while the bottom part of the compass can adopt slightly red values.
[0395] Therefore, it should be understood that even when the camera rotates, as discussed with respect to Figure 26C the compass 3005 will also rotate in the reverse direction within the field of view to maintain an appropriate orientation within the three-dimensional model and the space of region 2535a. Similarly, as previously discussed, the compass can include additional indicators to assist the user in identifying the missing relative positions. Here, in Figure 30A additional relative references 3025a and 3025b are shown outside the corresponding highlights 3015a and 3015b. References 3025a and 3025b can be colored or textured to indicate whether the missing highlights are in front of or behind the current field of view of the camera (here, since both missing highlights are in front of the camera, the references can share the same forward indication). Variations will be readily understood, for example, instead of references 3025a and 3025b, the boundaries of highlights 3015a and 3015b can indicate the position of the missing highlights via color, transparency, brightness, animation, etc.
[0396] As mentioned, in some embodiments, the compass 3005 and its highlights can be translucent to facilitate the operator in obtaining an appropriate field of view. However, the limited lighting inside the body can make it difficult to distinguish the state of the compass during operation. Therefore, some embodiments contemplate Figure 30B a compass in the form shown in the example of Figure 30ALike the compass, the indicators can be color-coded according to their radial position (e.g., the topmost indicator takes a value of 128 in the 0 - 255 hue range, and the bottommost indicator takes a hue value of 0 in the 0 - 255 range). Here, indicators 3055a and 3055b are shown as highlighted (corresponding to highlights 3015a and 3015b). Although references 3060a and 3060b can also be provided in the same manner as references 3025a and 3025b, it will be understood again that, in addition to the separate longitudinal references, the indicators can alternatively be bolded with a color or pattern indicating the missing relative longitudinal position.
[0397] It should be understood that multiple cameras can be used in some surgical procedures. In some embodiments, when the field of view of some cameras includes another camera with a compass, augmented reality graphics can be introduced into the displays of these cameras. For example, Figure 30E A perspective view of the first colonoscopy camera 3075 as seen from the perspective of the second surgical camera is depicted. To easily facilitate cross-referencing between views, an augmented reality overlay 3070 of the compass seen in the display of camera 3075 can be presented in this image of the second camera, including corresponding highlights 3065. In this way, the user can easily cross-reference the compass as depicted in the field of view to achieve an overall understanding of the surgical field, including the cameras and any selected missing relative orientations.
[0398] Example GUI Update Process
[0399] Figure 31 is a flowchart illustrating various operations in an example process 3100 for rendering Figure 28 various graphical guides (process 3100 can run, for example, as part of the visualization thread 3215 discussed below). At block 3105, the system can determine whether the monitoring is complete, and if so, process 3100 can end. In the case where the monitoring is ongoing, at block 3110, the system can determine whether a new image frame is available (e.g., from cache 3205o), and if so, process their data (e.g., the operations of block 3215a) at block 3115. Since rendering and data processing can occur at different rates and / or can occur in different threads, the waits at blocks 3120 and 3130 indicate that the process may be delayed due to different rates (e.g., anticipating updating cache 3205o and preparing the projection mapping image 3210j and the updated mesh 3210e).
[0400] Thus, before rendering, at block 3125, the system can consider whether the user has selected any specific omission (e.g., based on the selection of region 2545c by cursor 2810f). In the case of an omission being selected, at block 3130, the system can determine the relative position of the omission with respect to the current instrument position (e.g., using techniques such as those described with respect to Figure 30C . At block 3135, the system can then update the GUI and the projection surface representation to reflect the relative position, such as by indicating a highlighted area of a compass, emphasizing the boundaries of regions in the map 2535a, etc.).
[0401] In the case where no specific omission is determined to be selected at block 3125, at block 3140, the system can determine whether one or more omissions are close to the current position of the camera (e.g., by considering a perimeter generated from centerline positions within a threshold distance from the current camera position). If not, the system can clear the GUI and the HUD at block 3145, e.g., to avoid distracting the user. However, in the case where one or more omissions are close, at block 3150, the system can determine the relative position of the one or more omissions with respect to the current instrument position (again, e.g., using Figure 30C 's method, by directly projecting the grid positions onto the azimuth of the camera, etc.). In the case where more than one omission is nearby, the system can prioritize the omissions at block 3155 (e.g., larger omissions or omissions in sensitive areas of the surgery can be presented to the user before smaller or less - attended areas, or highlighted more strongly than smaller or less - attended areas). At block 3160, the system can then update the GUI with an appropriate overlay based on the one or more relative positions.
[0402] Iterative Internal Body Structure Representation - Example Processing Pipeline
[0403] Figure 32 FIG. is a schematic block diagram showing various components and their relationships in an example processing pipeline for iterative internal structure representation and navigation that can be implemented in some embodiments. In the depicted embodiment, the processing can generally be divided into three parts: visualization, mapping, and tracking. Here, different computational threads are assigned to each part, namely: visualization thread 3215, mapping thread 3210, and tracker thread 3205. It should be understood that the threads can be programmed to run in parallel on one or more processors and communicate with each other, e.g., using appropriate semaphore flags, queues, etc. Similarly, it should be understood that each thread can contain sub - threads, as in the case where multiple trackers associated with tracker thread 3205 operate in their own threads.
[0404] Starting from the tracker thread 3205, during a surgical procedure, new camera images 3205a (e.g., RGB images, grayscale images, indexed images, etc.) can arrive for processing. At block 3205b, the tracker thread can apply filters to determine whether the visual image is suitable for downstream processing (e.g., the localization and mapping operations disclosed herein). For example, blurred images, images occluded by biological matter or organ walls, etc. may not be usable for localization. In the case where a frame is unavailable, it can be discarded (however, in some embodiments, it should be understood that interpolation or prediction methods between frames can be used to correct some defective frames).
[0405] In contrast, in the case where an image is found to be available at block 3205b, a first copy of the available image can be provided to the pose and depth estimation block 3205e, and a second copy of the available frame can be provided to the feature extraction block 3205d. For example, the features extracted at block 3205d can be scale-invariant feature transform (SIFT) features of the visual image. Again, it should be understood that, for example, each of blocks 3205d and 3205e can operate in separate threads or sequentially in the same thread according to the methods described in, for example, Posner, Erez et al.'s "C3Fusion: Consistent Contrastive Colon Fusion, Towards Deep SLAM in Colonoscopy" (arXiv TM Preprint arXiv TM :2206.01961(2022)). The extracted features 3205c can be stored in the record 3205h. As discussed elsewhere herein, the image can be used for pose and depth determination at block 3205e. These determinations from block 3205e, the previously stored image itself 3205f, and the features extracted at block 3205d can be used to determine the sequential pose estimate of the camera at block 3205g, for example, as described elsewhere herein. Specifically, the system can use the previous frame features 3205f and the latest frame features 3205d to calculate the sequential pose estimate. Thus, block 3205g can verify and refine (if necessary) the pose estimated by, for example, one or more convolutional neural networks in block 3205e.
[0406] A positioning filter (such as Kalman filter 3205j) can be used to further refine the positioning and attitude estimation and store the results in data cache 3205o. As indicated, Kalman filter 3205j can consider the results of the previous Kalman filter 3205i in its analysis. Before performing global attitude optimization at block 3205n, further refinement can be achieved by providing the matches to the corresponding matcher 3205k and sending the resulting matching frames to the local attitude optimization block 3205m (which itself can update record 3205h with the modified Kalman filter result 3205l) (it should be understood that the modified Kalman filter result 3205l itself can be used as the latest result 3205i in subsequent iterations).
[0407] In the case where the data cache 3205o is filled with new attitude estimation results, the mapping thread 3210 can now start using the determined attitude information to integrate the newly acquired data. Specifically, the frames and attitude 3210a can be extracted from the cache 3205o and used to update the centerline determination at block 3210c (again, it should be understood that various alternative methods are used to determine the centerline). Similarly, at block 3210b, the previous depth frame can be acquired for integration with the TSDF structure. Thereby, the system can then extract the surface of the mesh at block 3210d (e.g., using marching cubes, convex hulls, or other suitable methods) to create an updated mesh 3210e. The mesh surface can also be used to update the depth map rendering at block 3210f to produce a refined depth map 3210h. Then, each of the updated mesh centerline 3210g, the refined depth map 3210h, and the latest camera attitude and image 3210k can be used to perform surface parameterization at block 3210i to produce a projected surface flattening image 3210j. For example, the system can determine the perimeter corresponding to the active area (e.g., the surrounding area where the colonoscopy camera is currently active) and the corresponding row pixel values. Thus, instead of recreating the entire projected image surface with each iterative adjustment to the mesh model, the system can alternatively only update the active portion of the projected surface flattening image 3210j corresponding to the most recently acquired, integrated, and updated portion of the entire 3D model.
[0408] Since the estimated depth map may sometimes be noisy, some embodiments can create multiple depth maps from the same position, re-render the mesh from the same position of the camera, so as to create a new, refined depth map. This local iterative refinement can be applied, and the system encourages the operator (e.g., via GUI feedback) to stay in the areas where gaps occur or the flattening image or model is poorly constructed.
[0409] Then, the visualization thread 3215 can obtain the updated projection image 3210j and the mesh 3210e, provide the latter for display at box 3215e and potentially for storage at box 3215c. Similarly, the image 3210j along with the latest pose and image information 3215d can be used to determine the position and orientation of the navigation compass (e.g., compass 2820) at box 3215a, as described herein. In some embodiments, at box 3215a, the system will ensure that the compass representation remains properly oriented, e.g., to maintain the proper global orientation as described herein Figure 26C regardless of the camera's roll angle or other orientation changes. Then the updated compass can be rendered on the two-dimensional display 3215b together with the new image 3210j. In some cases, at box 3215c, the rendered two-dimensional image can also be stored to a file, e.g., for subsequent viewing.
[0410] Iterative Internal Body Structure Representation - Example Sub - processing Pipeline
[0411] As described above, in some embodiments, the navigation itself can be established in two stages based on prior pose determination: e.g., surface parameterization corresponding to box 3210i, and determination of the navigation compass corresponding to box 3215a, for example.
[0412] Figure 33A is a schematic block diagram of the various operational relationships between the components of the surface parameterization process 3305a that can be implemented, for example, at box 3210i in some embodiments. As in box 3210i, surface parameterization maps the three-dimensional reconstructed surface of the mesh to a two-dimensional image (such as the projection map of region 2535a). For example, in the case where the mesh is the colon, surface parameterization can "unroll" the colon along the centerline of the model, as Figure 25 shown, where each horizontal row of the image is derived from the mesh vertices along the corresponding perimeter derived from the centerline. Similarly, the columns can correspond to angles uniformly sampled by flattening the colon wall on the planar image using a conformal structure.
[0413] For example, as regarding Figures 26A - 26CAs described, the surface parameterization algorithm can determine the perimeter by taking cross-sections of the current estimated point cloud at points along the centerline of the mesh model. For each vertex that appears in the cross-section or perimeter, the system can assign an angle. Thus, given the centerline 3305g (e.g., centerline 2650) at time t, a K-dimensional tree (KD-tree) can be generated on the centerline point cloud sample at box 3305h. The input depth map 3305c (e.g., refined according to box 3210f) and the camera image 3305b can be downsampled at box 3305d (in the prototype implementation, it takes about 1 ms to complete). The results can then be back-projected to the 3D coordinates of the mesh at box 3305e (in the prototype implementation, it takes about 1.5 ms to complete) to produce a point cloud representing the current estimate of the scene. In box 3305f, the system can, for example, perform cross-sectioning of the estimated point cloud along the centerline, query the centerline from all estimated point cloud vertices from the KD-tree, and consider the previously generated flattened image 3305i (in the prototype implementation, it takes about 75 ms to complete). Then, box 3305f can produce a new flattened image 3305i for rendering.
[0414] Figure 33B is a schematic block diagram of various operational relationships between components of a surface flattening image update process that can be implemented, for example, at box 3305f in some embodiments. The centerline KD-tree 3310d can be the output from box 3305h, the point cloud 3320c can correspond to the back-projected point cloud from box 3305e, and the centerline 3320o can correspond to the centerline 3305g (corresponding to each input to Figure 33A box 3305f in
[0415] Indices can be assigned to each vertex x of the currently constructed mesh or the back-projected point cloud from the estimated refined depth map (e.g., where the system only performs this operation on the "active" area around the camera), and each index represents the nearest point c on the centerline with K samples via the KD-tree k , as shown in equations 8 and 9:
[0416]
[0417] {S k} = {argmin k∈{1,...,K} ||x i - c k || 2} (9)
[0418] where {S k} is the set of all vertices assigned to the k-th point on the center line. After this index assignment, an angle around the center line can be assigned to each vertex (e.g., representing the radial relationship of the vertex to the corresponding perimeter). In some embodiments, the vertices can be grouped into segments according to their angles (e.g., based on their presence in one of a set of radial ranges), e.g., grouped into groups of angular "bins" of approximately equal angular width (e.g., 0.5 degrees). In some embodiments, the system can employ an axis-angle representation, where the axis-angle is along the forward direction of the center line (e.g., calculated between two adjacent samples along the center line). As discussed, for each row in the projected image, the associated columns (starting from 0 - 359 degrees) can be colored according to the assigned angles of the corresponding estimated point cloud vertices, e.g., as indicated by the "paint row" box 3310l (in the prototype implementation, it takes approximately 65 ms to complete). Thus, the cross-sectional perimeter of the points along the center line at box 3310i can be transformed into a row of the image at box 3310j, and the updated pixels (e.g., those that have changed or are newly encountered) are written at box 3310k to produce the updated surface image 3310m. In some embodiments, rather than rewriting the entire image, the old pixels in the previous image 3310n can remain unchanged. Thus, the system can query the KD tree at box 3310e (in the prototype implementation, it takes approximately 7.5 ms to complete) to determine the center line 3310f and the active row index 3310g. The center line 3305g can be used to extend the surface flattened image, as discussed at box 3310h (in the prototype implementation, it takes approximately 0 ms (i.e., a negligible amount) to complete).
[0419] Regarding the navigation compass, for example, as discussed with respect to box 3215a, Figure 33C is a schematic block diagram depicting the various operational relationships between the components of a navigation compass update process 3315a that can be implemented in some embodiments. As mentioned, a compass (such as compass 2820) can be used to guide the user to areas that have not been inspected or have not been adequately inspected. The navigation compass can be linked to the surface parameterization box 3305a via the updated surface flattened image 3310m, and the updated surface flattened image 3310m can be used to create the compass, e.g., by facilitating the identification of omissions.
[0420] Process 3315a may receive the surface flattening image 3305i and the camera position 3315d as inputs. Before aligning the camera pose and the 3D model coordinate system (defined by the centerline axis) at block 3315j (which takes approximately 0.7 ms in some embodiments), the overlapping navigation assist color box 3315h may use the current camera pose 3315i and the corresponding portion of the flattening image 3315f to determine the appropriate radial coloring (e.g., where the hue corresponds to a consistent global radial degree, as for example with respect to Figure 26C and Figure 30B discussed). After the alignment between the camera pose and the centerline forward vector (e.g., as discussed herein with respect to Figure 36B view 3650c), the offset angle between the x-axes can be calculated for the offset compass visualization such that the navigation will be invariant to any camera roll (again, as for example with respect to Figure 26C discussed). In other words, for clarity, given a first "x-axis" vector v x1 of a first coordinate system defined by the camera pose and a second "x-axis" vector v x2 of a second coordinate system defined by the centerline, the system can use the axis-angle representation to calculate the angle, as shown in Equation 10:
[0421]
[0422] Then, a navigation compass can be created at block 3315k (which takes approximately 7 ms in the prototype implementation). Then, before creating the navigation image 3315m depicting the field of view image overlaid with the compass, upsampling for display can occur at block 3315l (which takes approximately 2 ms to complete in the prototype implementation). Then, the system can combine this image 3315m with the surface flattening image 3315e and the original image 3315g at block 3315n to output to the display at block 3315b.
[0423] Example Central Axis Centerline Estimation - System Process
[0424] Naturally, a more precisely and consistently generated centerline can better enable a more precise perimeter selection for mapping. While it should be understood that there are multiple methods for dividing the model into perimeters (with or without a centerline), this section provides example centerline estimation methods and procedures for the reader's understanding. Consistent centerline estimation as can be achieved with these methods can be particularly useful when analyzing and comparing surgical procedure performance. Thus, although specifically presented with reference to an example of creating a centerline in a colonoscopy context, various embodiments envision improved methods for determining a centerline based on localization and mapping processes, e.g., as previously referenced herein Figure 14 and as applied to various anatomies.
[0425] Iterative Internal Body Structure Representation - Variations
[0426] Variations of the above-described embodiments should be understood. For example, in some embodiments, the vertical dimension of region 2535a is not fixed but will grow as the examination continues. This can be appropriate when the size of the organ is unknown. In the case of limited space in the given GUI, the user can be invited to scroll along the vertical dimension of region 2535a, or the vertical dimension can be scaled so that the available data always fits the vertical dimension of region 2535a (e.g., there is no region 2540b corresponding to the open end of grid 2525a, but rather available texture extending to the top of region 2535a; in such an embodiment, magnifying regions (such as magnifying region 2970b) can facilitate local viewing). However, in many surgical procedures, at least the approximate size of the internal region of the patient's body to be examined can be known. Accordingly, some embodiments can adjust region 2535a based on these expectations. By doing so, the operator and other members of the surgical team can anticipate the future state of the surgery and understand the current state and viewing scope.
[0427] For example, similar to Figure 25 and Figure 28 , Figure 34A and Figure 34B depict successive schematic representations of various GUI panels that can be presented to a user during a surgical procedure. Here, the partial complete model 3405e shown in views 3405b and 3405c can again include various omissions having corresponding regions in the mapped representation 3405d (e.g., corresponding to partial models 2505a - 2505c). As the camera advances, the field of view 3405a (corresponding to region 2510a) can be updated accordingly. It should be understood that the two views 3405b and 3405c can be two simultaneously presented views of the same model from different orientations, or a single view rotated using, for example, the mouse cursor 3405g. However, different from the Figure 25 embodiment, the three-dimensional view includes a non-data-exported reference grid 3405f of the current model 3405e. The reference grid 3405f can be determined by averaging the organ models of other patients, the convex hull of their cumulative models (then scaled by the size of the current patient), the artistic reproduction of the organ, etc.
[0428] As the surgical procedure progresses and more current patient data is acquired, the reference grid can be replaced with corresponding portions of the data-created grids 3405e and 3405h. The reference grid 3405f can be rendered without texture or otherwise clearly distinguished from the data-acquisition grids 3405e and 3405h. By anticipating the full length of the organ, the row placement from the centerline perimeter can be managed accordingly within the limited vertical dimension of the mapped representation 3405d. That is, if 30% of the reference grid is retained, the acquired data can be scaled such that approximately 30% of the vertical dimension in region 3405d remains available and is identified as "missing". For example, in Figure 34A it, approximately 30% of the reference grid 3405f is retained and thus the acquired data can be placed in region 3405d such that the unexplored region 3490 includes approximately 30% of the vertical dimension of region 3405d. Thus, corresponding to the anticipated length of the organ, the non-data-derived grid 3405f can anticipate the presence of regions not yet explored in the surgical procedure. As Figure 34B shown, at a later time, views 3405b, 3405c can replace the non-data-derived grid 3405f with corresponding portions of the updated partial grid 3405h.
[0429] In some embodiments, the non-data-derived grid 3405f can be an idealized geometry corresponding to the relevant anatomy. For example, as Figure 34C shown in it, it can be assumed that the cylindrical reference geometry grid 3415a corresponds to the actual data-derived intestinal grid geometry 3415b of 3415c. While the reference geometry grid 3415a can be created by hand by an artist, it should be understood that the dimensions can be determined by various methods. For example, the extraction of vertices on a grid generated from the average data-derived grid can result in a near-cylindrical structure that can itself be used, or an idealized cylinder with substantially similar radius and length. Thus, the idealized reference grid can be generated based on the accumulation of real-world data.
[0430] Similarly, while intestinal examinations are presented herein primarily for ease of reader understanding, it should be understood that many embodiments need not be limited to this context. For example, Figure 34D and Figure 34EPerspective views of a reference spherical geometry mesh 3420a, a cumulative convex hull reference geometry mesh 3420b, and an exemplary cavity mesh geometry 3425 collected from a current surgical procedure (e.g., prostatectomy) are depicted respectively. As indicated, each of the reference meshes 3420a, 3420b can be used to map 3420c, 3420d corresponding to the mesh 3425 generated during the current surgical procedure. Just as the iterative "consumption" of the data meshes 3405e and 3405h by the reference mesh 3405f facilitated the reference for determining the overall progress, the iterative consumption of the meshes 3420a and 3420d during the exploration of the cavity generating the mesh 3425 can similarly facilitate the generation of a mapping region similar to region 3405d depicting the relative overall progress during the surgery.
[0431] To determine which portions of the reference mesh should be replaced by portions of the data-derived mesh, the system can employ a process as Figure 34F shown. Here, the edges and vertices of the reference mesh are schematically presented by a series of edges and vertices 3430a on the two-dimensional plane of the page. For example, the edges and vertices 3430a can be the inner or outer surface of the cylinder 3415a, or the inner or outer surface of the mesh 3420a or 3420b. As the data-derived mesh 3430b grows, the system can compare the 3440 meshes. Vertices within a threshold distance of the nearest neighbor vertices in the corresponding mesh can be interpreted as "associated" with the other mesh (represented here by arrows such as arrow 3445), while vertices outside the threshold remain unassociated. In some cases, a finer granularity can alternatively be achieved by ta...
Claims
1. A computer-implemented method for evaluating the progress of a surgical instrument within a patient, the method comprising: Determining pose data associated with the surgical instrument; Determining depth data associated with the pose data; Constructing at least a portion of a three-dimensional model of at least a portion within the patient based on the pose data and the depth data; And Determining a centerline position associated with the at least a portion of the three-dimensional model.
2. The computer-implemented method according to claim 1, wherein constructing at least a portion of the three-dimensional model of at least a portion within the patient based on the pose data and the depth data comprises: Determining features between two images; Generating segments based on differences between the two features; And Merging the segments to form at least a portion of the three-dimensional model.
3. The computer-implemented method according to claim 1, wherein determining the centerline position comprises: Providing the depth data to a neural network configured to fill in portions of at least a portion of the three-dimensional model.
4. The computer-implemented method according to claim 1, claim 2, or claim 3, the method further comprising: Determining a kinematic threshold, wherein The kinematic threshold is one of the following: A velocity threshold of movement of at least a portion of the surgical instrument projected on the centerline; A velocity threshold of movement of at least a portion of the surgical instrument radially protruding from the centerline; and The distance of at least a portion of the surgical instrument from the centerline.
5. The computer-implemented method according to claim 4, wherein The kinematic threshold is a velocity threshold of movement of at least a portion of the surgical instrument projected on the centerline, and wherein The method further comprises: Determining that the surgical instrument is being withdrawn; and Determining that the velocity of the surgical instrument projected on the centerline exceeds the velocity threshold.
6. The computer-implemented method according to claim 5, wherein Determining the kinematic threshold comprises: Consulting a database that includes surgical instrument kinematic data projected on a reference geometry for a plurality of surgical procedures; and Determining the threshold based on a kinematic data value in the database corresponding to the time when the surgical instrument is being withdrawn.
7. The computer-implemented method according to claim 6, wherein determining a centerline position associated with at least a portion of the three-dimensional model comprises: Filtering at least a portion of the three-dimensional model to produce a filtered portion; Determining centerline endpoints based on the filtered portion; Generating a new local centerline from the pose of the surgical instrument; And Extending the centerline using the new local centerline.
8. The computer-implemented method according to claim 7, wherein extending the centerline using the new local centerline comprises: Determining a first array of points on the centerline; Determining a second array of points on the new local centerline; And Determine a weighted average between paired points in the first array and the second array.
9. The computer-implemented method according to claim 1, claim 2, or claim 3, wherein the method further comprises: Providing an image to at least one neural network; And Receiving, from the at least one neural network, a classification of the image, the classification indicating the feasibility of the image for determining the pose, and wherein, Determining the pose is based on the image.
10. The computer-implemented method according to claim 9, wherein the at least one neural network is configured to determine the image as infeasible if the image depicts at least one of the following: Motion blur; Fluid blur; A number of reflections exceeding a threshold; and Occlusion.
11. The computer-implemented method according to claim 10, wherein the method further comprises: Preprocessing the image before providing the image to the at least one neural network, wherein the preprocessing comprises at least one of the following: Transforming at least one channel of the image; Cropping the image; Resizing the image; and Applying a reflection mask to the image.
12. The computer-implemented method according to claim 11, wherein the method further comprises: Applying classification post-processing to the output of the at least one neural network to determine whether the image depicts an edge case.
13. The computer-implemented method according to claim 12, wherein the at least one neural network comprises a neural network comprising: A first set of one or more convolutional layers; A set of one or more pooling layers; A second set of one or more convolutional layers; A set of one or more linear layers; and A set of one or more merging layers.
14. The computer-implemented method according to claim 13, wherein the method further comprises: Using the at least one neural network to determine several images of the field of view of the surgical instrument that are infeasible for downstream processing; And Causing a spatial indication of the orientation within the patient's interior to be displayed, at the orientation where the images determined to be infeasible for downstream processing were acquired.
15. The computer-implemented method according to claim 1, claim 2, or claim 3, wherein the method further comprises: Determining a two-dimensional mapping of the patient's interior corresponding to the three-dimensional model, wherein, The two-dimensional mapping includes gaps, and wherein, The gaps correspond to one or more unexamined regions within the patient's interior.
16. The computer-implemented method according to claim 15, the method further comprises: Causing a field of view of the surgical instrument to be displayed, the field of view depicting a portion of the patient's interior; And Causing a compass to be displayed within the field of view, the compass comprising a circular representation of radial directions, wherein, The compass includes a portion indicating the position of the gaps relative to the pose of the surgical instrument.
17. The computer-implemented method according to claim 16, wherein determining the two-dimensional mapping comprises: Determining the perimeter of the three-dimensional model based on the centerline; And Update a row of pixels in the two-dimensional mapping to correspond to texture pixels associated with the perimeter.
18. The computer-implemented method according to claim 17, wherein the indication that there is at least one omission includes a portion of the two-dimensional representation that is not associated with the textured surface of the three-dimensional model.
19. The computer-implemented method according to claim 18, the method further comprising: Scaling the size of the two-dimensional mapping based on a proportion completion value.
20. The computer-implemented method according to claim 19, wherein the method further comprises: Determining the proportion completion value at least in part by: Determining a first portion of a reference grid within a threshold distance of the three-dimensional model; Determining a second portion of the reference grid not within the threshold distance of the three-dimensional model; and And Determining the proportion completion value by comparing the first portion and the second portion.
21. A non-transitory computer-readable medium, the non-transitory computer-readable medium comprising instructions configured to cause a computer system to execute a method for evaluating the progress of a surgical instrument within a patient's interior, the method comprising: Determining pose data associated with the surgical instrument; Determining depth data associated with the pose data; Constructing at least a portion of a three-dimensional model of at least a portion of the patient's interior based on the pose data and the depth data; and And Determining a centerline position associated with the at least a portion of the three-dimensional model.
22. The non-transitory computer-readable medium according to claim 21, wherein constructing at least a portion of the three-dimensional model of at least a portion of the patient's interior based on the pose data and the depth data includes: Determining features between two images; Generating segments based on the differences between the two features; and And Merging the segments to form at least a portion of the three-dimensional model.
23. The non-transitory computer-readable medium according to claim 21, wherein determining the centerline position includes: Providing the depth data to a neural network configured to fill a portion of at least a portion of the three-dimensional model.
24. The non-transitory computer-readable medium according to claim 21, claim 22, or claim 23, the method further comprising: Determining a kinematic threshold, wherein the kinematic threshold is one of the following: A speed threshold for the movement of at least a portion of the surgical instrument projected on the centerline; A speed threshold for the movement of at least a portion of the surgical instrument radially protruding from the centerline; and The distance of at least a portion of the surgical instrument from the centerline.
25. The non-transitory computer-readable medium according to claim 24, wherein the kinematic threshold is a speed threshold for the movement of at least a portion of the surgical instrument projected on the centerline, and wherein the method further comprises: Determining that the surgical instrument is being withdrawn; and Determining that the speed of the surgical instrument projected on the centerline exceeds the speed threshold.
26. The non-transitory computer-readable medium according to claim 25, wherein determining the kinematic threshold includes: consulting a database that includes kinematic data of surgical instruments projected onto a reference geometry for a plurality of surgical procedures; and determining the threshold based on kinematic data values in the database corresponding to the time when the surgical instrument is being withdrawn.
27. The non-transitory computer-readable medium according to claim 26, wherein determining the centerline position associated with at least a portion of the three-dimensional model includes: filtering at least a portion of the three-dimensional model to produce a filtered portion; determining centerline endpoints based on the filtered portion; generating a new local centerline from the pose of the surgical instrument; and extending the centerline using the new local centerline.
28. The non-transitory computer-readable medium according to claim 27, wherein extending the centerline using the new local centerline includes: determining a first array of points on the centerline; determining a second array of points on the new local centerline; and determining a weighted average between paired points in the first array and the second array.
29. The non-transitory computer-readable medium according to claim 21, claim 22, or claim 23, wherein the method further includes: providing an image to at least one neural network; and receiving a classification of the image from the at least one neural network, the classification indicating the feasibility of the image for determining the pose, and wherein determining the pose is based on the image.
30. The non-transitory computer-readable medium according to claim 29, wherein the at least one neural network is configured to determine the image as infeasible if the image depicts at least one of the following: motion blur; fluid blur; a number of reflections exceeding a threshold; and occlusion.
31. The non-transitory computer-readable medium according to claim 30, wherein the method further includes: preprocessing the image before providing the image to the at least one neural network, where the preprocessing includes at least one of the following: transforming at least one channel of the image; cropping the image; resizing the image; and applying a reflection mask to the image.
32. The non-transitory computer-readable medium according to claim 31, wherein the method further includes: applying classification post-processing to the output of the at least one neural network to determine whether the image depicts an edge case.
33. The non-transitory computer-readable medium according to claim 32, wherein the at least one neural network includes a neural network comprising: a first set of one or more convolutional layers; a set of one or more pooling layers; a second set of one or more convolutional layers; a set of one or more linear layers; and a set of one or more merging layers.
34. The non-transitory computer-readable medium according to claim 33, wherein the method further includes: Determine, using the at least one neural network, several images of the field of view of the surgical instrument that are not feasible for downstream processing; and Cause a spatial indication of the orientation within the patient's interior to be displayed, at which orientation the images determined to be not feasible for downstream processing are acquired.
35. The non-transitory computer-readable medium according to claim 21, claim 22, or claim 23, wherein the method further comprises: Determine a two-dimensional mapping of the patient's interior corresponding to the three-dimensional model, wherein, the two-dimensional mapping includes gaps, and wherein, the gaps correspond to one or more unexamined regions of the patient's interior.
36. The non-transitory computer-readable medium according to claim 35, the method further comprises: Cause a field of view of the surgical instrument to be displayed, the field of view depicting a portion of the patient's interior; and Cause a compass to be displayed within the field of view, the compass including a circular representation of radial directions, wherein, the compass includes a portion indicating the position of the gaps relative to the orientation of the surgical instrument.
37. The non-transitory computer-readable medium according to claim 36, wherein determining the two-dimensional mapping comprises: Determine the perimeter of the three-dimensional model based on the centerline; and Update a row of pixels in the two-dimensional mapping to correspond to texture pixels associated with the perimeter.
38. The non-transitory computer-readable medium according to claim 37, wherein, the indication that there is at least one gap includes a portion of the two-dimensional representation that is not associated with the textured surface of the three-dimensional model.
39. The non-transitory computer-readable medium according to claim 38, the method further comprises: Scale the size of the two-dimensional mapping based on a scale completion value.
40. The non-transitory computer-readable medium according to claim 39, wherein the method further comprises: Determine the scale completion value at least in part by: Determine a first portion of a reference grid within a threshold distance of the three-dimensional model; Determine a second portion of the reference grid that is not within the threshold distance of the three-dimensional model; and Determine the scale completion value by comparing the first portion and the second portion.
41. A computer system, comprising: At least one processor; and At least one memory, the at least one memory including instructions configured to cause the computer system to execute a method for evaluating the progress of a surgical instrument within a patient's interior, the method comprising: Determine pose data associated with the surgical instrument; Determine depth data associated with the pose data; Construct at least a portion of a three-dimensional model of at least a portion of the patient's interior based on the pose data and the depth data; and Determine a centerline position associated with the at least a portion of the three-dimensional model.
42. The computer system according to claim 41, wherein constructing at least a portion of the three-dimensional model of at least a portion of the patient's interior based on the pose data and the depth data comprises: Determine features between two images; Generate segments based on the differences between the two features; and Merge the segments to form at least a part of the three-dimensional model.
43. The computer system according to claim 41, wherein determining the centerline position includes: Providing the depth data to a neural network, the neural network being configured to fill a part of at least a part of the three-dimensional model.
44. The computer system according to claim 41, claim 42 or claim 43, the method further includes: Determine a kinematic threshold, wherein, The kinematic threshold is one of the following: A speed threshold for the movement of at least a part of the surgical instrument projected on the centerline; A speed threshold for the movement of at least a part of the surgical instrument radially protruding from the centerline; and The distance of at least a part of the surgical instrument from the centerline.
45. The computer system according to claim 44, wherein, The kinematic threshold is a speed threshold for the movement of at least a part of the surgical instrument projected on the centerline, and wherein, The method further includes: Determine that the surgical instrument is being withdrawn; and Determine that the speed of the surgical instrument projected on the centerline exceeds the speed threshold.
46. The computer system according to claim 45, wherein, Determining the kinematic threshold includes: Consulting a database, the database including surgical instrument kinematic data projected on a reference geometry for a plurality of surgical procedures; and Determine the threshold based on the kinematic data value corresponding to the time when the surgical instrument is being withdrawn in the database.
47. The computer system according to claim 46, wherein determining the centerline position associated with at least a part of the three-dimensional model includes: Filter at least a part of the three-dimensional model to produce a filtered part; Determine centerline endpoints based on the filtered part; Generate a new local centerline from the pose of the surgical instrument; and Extend the centerline using the new local centerline.
48. The computer system according to claim 47, wherein extending the centerline using the new local centerline includes: Determine a first array of points on the centerline; Determine a second array of points on the new local centerline; and Determine the weighted average between the paired points in the first array and the second array.
49. The computer system according to claim 41, claim 42 or claim 43, wherein the method further includes: Provide an image to at least one neural network; and Receive a classification of the image from the at least one neural network, the classification indicating the feasibility of using the image to determine the pose, and wherein, Determine the pose is based on the image.
50. The computer system according to claim 49, wherein, The at least one neural network is configured to determine the image as infeasible if the image depicts at least one of the following: Motion blur; Fluid blur; Number of reflections exceeding a threshold; and Occlusion.
51. The computer system according to claim 50, wherein the method further comprises: Preprocessing the image before providing the image to the at least one neural network, wherein the preprocessing comprises at least one of the following: Transforming at least one channel of the image; Cropping the image; Resizing the image; and Applying a reflection mask to the image.
52. The computer system according to claim 51, wherein the method further comprises: Applying a classification post - processing to the output of the at least one neural network to determine whether the image depicts an edge case.
53. The computer system according to claim 52, wherein the at least one neural network comprises a neural network comprising: A first set of one or more convolutional layers; A set of one or more pooling layers; A second set of one or more convolutional layers; A set of one or more linear layers; and A set of one or more merging layers.
54. The computer system according to claim 53, wherein the method further comprises: Using the at least one neural network to determine a number of images of the field of view of the surgical instrument that are not feasible for downstream processing; And Causing a spatial indication of the orientation within the patient's interior to be displayed, at the orientation where the images determined to be not feasible for downstream processing were acquired.
55. The computer system according to claim 41, claim 42 or claim 43, wherein the method further comprises: Determining a two - dimensional mapping of the patient's interior corresponding to the three - dimensional model, wherein, The two - dimensional mapping includes gaps, and wherein, The gaps correspond to one or more unexamined regions within the patient's interior.
56. The computer system according to claim 55, the method further comprises: Causing a field of view of the surgical instrument to be displayed, the field of view depicting a part of the patient's interior; And Causing a compass to be displayed within the field of view, the compass comprising a circular representation of radial directions, wherein, The compass includes a portion indicating the position of the gap relative to the pose of the surgical instrument.
57. The computer system according to claim 56, wherein determining the two - dimensional mapping comprises: Determining the perimeter of the three - dimensional model based on the centerline; And Updating a row of pixels in the two - dimensional mapping to correspond to texture pixels associated with the perimeter.
58. The computer system according to claim 57, wherein, The indication that there is at least one gap includes the portion of the two - dimensional representation that is not associated with the textured surface of the three - dimensional model.
59. The computer system according to claim 58, the method further comprises: Scaling the size of the two - dimensional mapping based on a scale completion value.
60. The computer system according to claim 59, wherein the method further comprises: Determining the scale completion value at least in part by: Determining a first portion of a reference grid within a threshold distance of the three - dimensional model; Determining a second portion of the reference grid not within the threshold distance of the three - dimensional model; And The ratio completion value is determined by comparing the first part and the second part.