Systems and methods for universal phase recognition for intraoperative and postoperative applications

EP4736178A1Pending Publication Date: 2026-05-06INTUITIVE SURGICAL OPERATIONS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
INTUITIVE SURGICAL OPERATIONS INC
Filing Date
2024-06-27
Publication Date
2026-05-06

AI Technical Summary

Technical Problem

Current surgical data processing systems face challenges in efficiently and accurately recognizing surgical phases in robot-assisted surgeries due to high intra-/inter-class variances in patient anatomy, surgeon skills, and workflows, leading to excessive latency and resource consumption.

Method used

The system employs machine learning-based phase recognition using surgical videos, system events, and kinematics data, generating multiple sets of consecutive frames with varying temporal resolutions and utilizing attention-based deep learning models to predict phases, incorporating uncertainty analysis and phase transition maps for real-time recognition.

Benefits of technology

This approach enables accurate and efficient recognition of full-length surgical phases, providing real-time performance metrics and alerts, improving patient care by reducing computational latency and resource usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024035866_02012025_PF_FP_ABST
    Figure US2024035866_02012025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for universal phase recognition for intraoperative and postoperative applications are provided. The system receives a video stream that captures a procedure over a time interval with a robotic medical system. The system generates, from the video stream, sets of consecutive frames. A first set of consecutive frames can include a first temporal resolution, and a second set of consecutive frames can include a second temporal resolution. The system determines, via the sets of consecutive frames input into a first model trained with machine learning, phases of the procedure on a moment-to-moment basis over the time interval. The system inputs the phases of the procedure into a second model to generate phase segment over the time interval. The second model can be trained with machine learning based on historical workflows. The system provides an action based on a metric of the phase segment.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR UNIVERSAL PHASE RECOGNITION FOR INTRAOPERATIVE AND POSTOPERATIVE APPLICATIONSCROSS-REFERENCES TO RELATED APPLICATIONS

[0001] This application claims the benefit of, and priority to, under 35 U.S.C. § 119, U.S. Provisional Patent Application No. 63 / 511,592, filed June 30, 2023, which is hereby incorporated by reference herein in its entirety.BACKGROUND

[0002] Surgical procedures can involve capturing imagery such as video feeds from a variety of viewpoints. For example, in some instances, at least part of the surgical procedure can be performed with a computer-assisted robotic medical system. A medical tool, such as an imaging device, can be used in the robotic medical system to provide imagery. Data sources such as cameras, sensors, etc. can be located at various viewpoints in the surgical facility to capture and provide imagery of various aspects of the surgical procedure. The captured imagery from the surgical procedure can be processed in various ways.SUMMARY

[0003] This technical solution is generally related to systems and methods for universal phase recognition for intraoperative and postoperative applications. The technology can automatically recognize full-length surgical phases that take place at any moment in a procedure of robot-assisted surgery. The phases can refer to high-level, universal activities that can occur in different types of procedures, and can include: exposure, dissection, transection, extraction, and reconstruction.

[0004] This technical solution can recognize surgical phases using one or more approaches. For example, the technology can perform phase recognition using machine learning based on surgical videos, system events, or kinematics data. In another example, this technical solution can perform phase recognition using low-level surgical task annotations.

[0005] To perform phase recognition using machine learning based on surgical videos, the technology can generate multiple sets of consecutive frames (e.g., 3) with different temporal resolutions (e.g., short, long, and varying). The technology can use a machine learning model to generate moment-to-moment phase predictions using the multiple sets of consecutive frames. For example, the technology can use a multi-pathway spatial -temporal decoding unitconfigured with an attention-based deep learning model. The model can be based on or utilize a vision transformer architecture with self-attention over space and time to allow for jointed spatial and temporal feature learning from the video stream. The technology can use multiple parallel streams to process the multiple sets of consecutive frames in a simultaneous or overlapping manner in order to extract features. The technology can perform feature-level fusion by aggregating numerical feature outputs of the multiple parallel streams to make moment-wise predictions of phases generated based on an aggregated feature vector.

[0006] The technology can then recognize full-length phases from the moment-to-moment phase prediction using a second machine learning model that is trained with surgical workflows. The technology can utilize a phase transition map with priors and probability to transition to each phase label in order to recognize the full-length phases from the moment-to- moment phase predictions. For example, the technology can identify or find boundaries of each surgical phase and generate full-length phase recognition for a procedure from moment-to- moment phase predictions. In some cases, the technology can quantify the uncertainty in the boundaries based on the variance in the moment-to-moment predictions. For example, the technology can use a long-range temporal module based on analysis of surgical workflows to model distribution over phases at each moment in time. The model can combine information about both the average duration of phases and the temporal ordering of phases in order to define a prior belief of likelihood of staying within the same phase, or transitioning to each of the other phases. The technology can model the likelihood of each phase label for a given timestamp to generate uncertainty-aware phase boundaries throughout an entire case in order to provide real-time phase predictions. Thus, the technology can transition from moment-to- moment to full length phase predictions that can incorporate a variety of different information sources in different ways, including, for example: average predictions; transition probabilities refined by additional info about procedure type, hospital site; unique events to identify boundaries of certain phases, such as installment and unmount of needle driver; decision-tree methods; or multi-modal.

[0007] The technology an adapt this approach to an event stream (e.g., information about tool installation or uninstallation) or kinematics stream (e.g., for task detection), as well as a combination of the video stream, event stream, and kinematics stream.

[0008] To perform phase recognition based on low-level surgical task annotations, this technology can input task annotations and use a many-to-one task-to-phase converter toperform temporal sequential decoding in order to recognizes the phases for the entire procedure.

[0009] The technology can use the recognized full-length phases to generate actions, notifications, or alerts in various intraoperative and postoperative applications. For example, the technology can provide real-time dynamic updates on performance or efficiency of each phase, average duration of each phase, influence scheduling, or provide alerts if tools are being used incorrectly.

[0010] At least one aspect is directed to a system. The system can include one or more processors, coupled with memory. The system can receive a video stream that captures a procedure over a time interval with a robotic medical system. The system can generate, from the video stream, multiple sets of consecutive frames. The sets of consecutive frames can include a first set of consecutive frames with a first temporal resolution and a second set of consecutive frames with a second temporal resolution. The system can determine, via the multiple sets of consecutive frames input into a first model trained with machine learning, phases of the procedure on a moment-to-moment basis over the time interval. The system can input the phases of the procedure determined on the moment-to-moment basis into a second model to generate at least one phase segment over the time interval. The second model can be trained with machine learning based on historical workflows. The system can provide an action based on a metric of the at least one phase segment.

[0011] In some implementations, the system can execute, prior to generation of the plurality of sets of consecutive frames, one or more pre-processing functions on the video stream. The one or more pre-processing functions can include at least one of a central crop transform, frame resizing, filtering of non-surgical frames, or filtering of noisy frames.

[0012] The system can generate the first set of consecutive frames with the first temporal resolution that is greater than or equal to a first threshold. The system can generate the second set of consecutive frames with the second temporal resolution that is less than or equal to a second threshold that is less than the first threshold. The system can generate a third set of consecutive frames with a varying temporal threshold that varies based at least in part on a function.

[0013] The phases of the procedure can include at least one of exposure, dissection, transection, extraction, or reconstruction. The system can determine the phases of theprocedure on the moment-to-moment basis via a the first model configured with a multipathway spatial-temporal decoding unit comprising an attention-based deep learning model. The system can input the multiple sets of consecutive frames into a corresponding plurality of parallel streams to generate corresponding numerical feature outputs. The system can fuse the numerical feature outputs to generate the phases of the procedure on the moment-to-moment basis.

[0014] The system can determine a variance in the phases of the procedure determined on the moment-to-moment basis. The system can generate, based on the variance and a phase transition map configured with a plurality of prior probabilities, a plurality of uncertainty- aware phase boundaries throughout the time interval. The system can generate the at least one phase segment that corresponds to a phase boundary of the plurality of uncertainty-aware phase boundaries. The system can generate the at least one phase segment based at least in part on one or more of an average duration of phases, an ordering of phases, a type of the procedure, or a site location of the procedure.

[0015] The system can determine the phases of the procedure on the moment-to-moment basis based on a combination of the video stream and at least one of a stream of system events data or a stream of kinematics data. The system can provide the action indicating a level of performance of the procedure during a phase segment of the at least one phase segment determined over the time interval.

[0016] The system can determine a phase segment of the at least one phase segment at a current time. The system can identify a tool used in the phase segment based on a stream of system events. The system can determine that the tool does not match any of a predetermined set of tools configured for the phase segment. The system can provide an alert during the phase segment responsive to the determination that the tool does not match any of the predetermined set of tools configured for the phase segment.

[0017] An aspect can be directed to a non-transitory computer-readable medium storing processor executable instructions that, when executed by one or more processors, cause the one or more processors to receive a video stream that captures a procedure over a time interval with a robotic medical system. The instructions can include instructions to generate, from the video stream, a plurality of sets of consecutive frames comprising a first set of consecutive frames with a first temporal resolution and a second set of consecutive frames with a second temporal resolution. The instructions can include instructions to determine, via the plurality of sets ofconsecutive frames input into a first model trained with machine learning, phases of the procedure on a moment-to-moment basis over the time interval. The instructions can include instructions to input the phases of the procedure determined on the moment-to-moment basis into a second model to generate at least one phase segment over the time interval. The second model can be trained with machine learning based on historical workflows. The instructions can include instructions to provide an action based on a metric of the at least one phase segment.

[0018] An aspect can be directed to a method. The method can be performed by one or more processors coupled with memory. The method can include the one or more processors receiving a video stream that captures a procedure over a time interval with a robotic medical system. The method can include the one or more processors generating, from the video stream, a plurality of sets of consecutive frames comprising a first set of consecutive frames with a first temporal resolution and a second set of consecutive frames with a second temporal resolution. The method can include the one or more processors determining, via the plurality of sets of consecutive frames input into a first model trained with machine learning, phases of the procedure on a moment-to-moment basis over the time interval. The method can include the one or more processors inputting the phases of the procedure determined on the moment-to- moment basis into a second model to generate at least one phase segment over the time interval. The second model can be trained with machine learning based on historical workflows. The method can include the one or more processors providing an action based on a metric of the at least one phase segment.

[0019] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of various aspects and implementations and provide an overview or framework for understanding the nature and character of the claimed aspects and implementations. The drawings provide illustration and a further understanding of the various aspects and implementations and are incorporated in and constitute a part of this specification. The foregoing information and the following detailed description and drawings include illustrative examples and should not be considered as limiting.BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are not intended to be drawn to scale. Like reference numbers and designations in the various drawings indicate like elements. For purposes of clarity, not every component can be labeled in every drawing. In the drawings:

[0021] FIG. 1 depicts an example system for universal phase recognition for intraoperative and postoperative applications.

[0022] FIG. 2 depicts an example process for universal phase recognition for intraoperative and postoperative applications.

[0023] FIG. 3 depicts an example process for universal phase recognition for intraoperative and postoperative applications.

[0024] FIG. 4 depicts an example user interface for universal phase recognition for intraoperative and postoperative applications.

[0025] FIG. 5 depicts an example process for universal phase recognition for intraoperative and postoperative applications.

[0026] FIG. 6 depicts an example process for universal phase recognition for intraoperative and postoperative applications.

[0027] FIG. 7 depicts an example method for universal phase recognition for intraoperative and postoperative applications.

[0028] FIG. 8 depicts an example medical environment.

[0029] FIG. 9 is a block diagram depicting an architecture for a computer system that can be employed to implement elements of the systems and methods described and illustrated herein, including aspects of the systems depicted in FIG. 1 and FIG. 8, the user interface depicted in FIG. 4, and the methods or processes depicted in FIGS. 2, 3, and 5-7.DETAILED DESCRIPTION

[0030] Following below are more detailed descriptions of various concepts related to, and implementations of, methods, apparatuses, and systems of universal phase recognition for intraoperative and postoperative applications. The various concepts introduced above and discussed in greater detail below can be implemented in any of numerous ways.

[0031] Although the present disclosure is discussed in the context of a surgical procedure, in some embodiments, the present disclosure can be applicable to other medical sessions or environments or activities, as well as non-medical activities where removal of irrelevant information is desired.

[0032] This technical solution is directed to automatically recognizing universal surgical activity that can occur at any moment in robot-assisted surgery for intraoperative and postoperative applications. For example, a large amount of data can be captured from minimally invasive and robot assisted surgery. With the significant amount of information, it can be challenging to efficiently, reliably, and accurately process the surgical data without introducing excessive latency to recognize surgical activities and provide context-awareness and decision support for improving patient care. Processing surgical data can utilize significant computing resources, which can be time consuming, cost-intensive, and not scalable. Further, it can be a challenge to recognize activities that have high intra- / inter-class variances due to the variability of patient anatomy, surgeon technical skills, as well as workflows across different procedures.

[0033] This technical solution can automatically recognize full-length surgical phases that take place at any moment in a procedure of robot-assisted surgery. The surgical phases can correspond to high-level, universal activities that constitute an entire surgical procedure and occur in different types of procedures. For example, this technical solution can recognize, classify, or otherwise identify the following five phases: exposure, dissection, transection, extraction, and reconstruction.

[0034] A data processing system of this technical solution can recognize phases using machine learning based on one or more of surgical videos, system events, or kinematics data. For example, the imagery that is captured by this technical solution from a surgical procedure can be in the form of a data stream or video stream. A data stream can refer to or include any sequence of digital encoded data or analog data captured from a data source. Each unit of transmission of a data stream can be considered a frame. The data stream can be considered to include a sequence of elements, and each element can be considered a frame.

[0035] The data processing system can collect the data and then reformat the data using one or more pre-processing techniques. The data processing system can construct multiple sets of consecutive frames with different resolutions (e.g., time intervals). Using different temporal intervals, can result in the data processing system generating different views in videos at variedtemporal resolutions, which can facilitate or improve the ability for a machine learning model to capture temporal relationships at different levels of granularity that can result in more accurate, generalized predictions. The data processing system can generate moment-to-moment predictions using the different sets of consecutive frames. Further, by processing the different streams using parallel processing, this technical solution can accurately recognize phases in an efficient manner without introducing unnecessary delays or latencies in the computing workflow.

[0036] The data processing system can input the moment-to-moment phase predictions generated using the multiple sets of consecutive frames into a second machine learning model that is configured to recognize full-length phases. The data processing system can identify or find boundaries of each surgical phase and generate full-length phase recognition for a procedure from the moment-to-moment phase predictions. To do so, the data processing system can utilize the second model that is trained or configured with contextual information about surgical workflows, which can inform the moment-to-moment predictions. For example, the data processing system can use a long-range temporal module constructed based on an analysis of surgical workflows to model a distribution over phases at each moment in time that serves as a prior belief or probability or likelihood. The long-range temporal module can incorporate information about logical workflows to augment the information gained from the vision-based model, which is used to update the belief about the phase at each timestep into a posterior distribution. The model can combine information about both the average duration of phases and the temporal ordering of phases, so that at each timestep the data processing system can define a prior belief of the likelihood of staying within the same phase, or transitioning to each of the other phases.

[0037] The data processing system can use this combination model to leverage uncertainty by modeling the likelihood of each phase label for a given timestamp rather than just making a single prediction. The data processing system can analyze the variance in the moment-to- moment predictions to facilitate quantifying uncertainty in the boundaries. For example, when the data processing system determines, based on the surgical workflow analysis, that there is a strong prior belief that a surgery will enter a given phase next, an uncertain prediction from the moment-to-moment model will be less likely to override this prior belief. However, when there are multiple potential next phases in surgery, the prior beliefs can be more evenly distributed between phase classes, and the data processing system can rely more heavily on the moment-to-moment model. This structure allows the data processing system to generate uncertainty- aware phase boundaries throughout an entire case, while also utilizing previous information, which makes real-time phase predictions possible.

[0038] FIG. 1 depicts an example system 100 for universal phase recognition for intraoperative and postoperative applications, in accordance with implementations. The system 100 can be associated with a medical environment 105. The medical environment 105 can be a medical surgical environment. A medical surgical environment can include a surgical facility such as an operating room in which a surgical procedure, whether, invasive, non-invasive, inpatient, or out-patient, can be performed on a patient. The system 100 can be associated with other types of medical sessions or activities, or non-medical environments that can require removal of non-surgical information from a data stream captured from that environment. The system 100 can include one or more data capture devices 110. Data capture devices 110 can collect images, kinematics data, or system events, for example. For example, the data capture device can include an image capture devices designed, constructed and operational to capture images from a particular viewpoint within the medical environment 105. The data capture devices 110 can be positioned, mounted, or otherwise located to capture content from any viewpoint that facilitates the data processing system recognizing phases of a procedure.

[0039] For example, in some embodiments, a first data capture devices can be positioned to capture one or more images of an area where a patient is located within the medical environment 105. A second data capture device 110 can be positioned to capture one or more images of an area where one or more medical professionals are located within the medical environment 105. A third data capture device 110 can be configured to capture one or more images of other designated areas within the medical environment 105. The data capture devices 110 can include any of a variety of sensors, cameras, video imaging devices, infrared imaging devices, visible light imaging devices, intensity imaging devices (e.g., black, color, grayscale imaging devices, etc.), depth imaging devices (e.g., stereoscopic imaging devices, time-of- flight imaging devices, etc.), medical imaging devices such as endoscopic imaging devices, ultrasound imaging devices, etc., non-visible light imaging devices, any combination or subcombination of the above mentioned imaging devices, or any other type of imaging devices that can be suitable for the purposes described herein.

[0040] The images that are captures by the data capture device 110 can include still images, video images, vector images, bitmap images, other types of images, or combinations thereof. Insome embodiments, one or more of the data capture devices 110 can be configured to capture other parameters (e.g., sound, motion, pressure, or temperature) within the medical environment 105. The data capture devices 110 can capture the images at any suitable predetermined capture rate or frequency. Other settings, such as zoom settings or resolution, of each of the data capture devices 110 can vary as desired to capture suitable images from a particular viewpoint. The data capture devices 110 can have fixed locations, positions, or orientations. The data capture devices 110 can be portable, or otherwise configured to change orientation or telescope in various directions. The data capture devices 110 can be part of a multi-sensor architecture including multiple sensors, with each sensor being configured to detect, measure, or otherwise capture a particular parameter (e.g., sound, images, or pressure).

[0041] The images captured by the data capture devices 110 can be sent as a data stream component to a visualization tool 170. A data stream component can be considered any sequence of digital encoded data or analog data from a data source such as the data capture devices 110. The visualization tool 170 can be configured to receive a plurality of data stream components and combine the plurality of data stream components into a single data stream.

[0042] The visualization tool 170 can receive a data stream component from a medical tool 120. The medical tool 120 can be any type and form of tool used for surgery, medical procedures or a tool in an operating room or environment associated with or having an image capture device. The medical tool 120 can be an endoscope for visualizing organs or tissues, for example, within a body of the patient. The medical tool 120 can include other or additional types of therapeutic or diagnostic medical imaging implements. The medical tool 120 can be configured to be installed in a robotic medical system 125.

[0043] The robotic medical system 125 can be a computer-assisted system configured to perform a surgical or medical procedure or activity on a patient via or using or with the assistance of one or more robotic components or medical tools. The robotic medical system 125 can include one or more manipulator arms that perform one or more computer-assisted medical tasks. The medical tool 120 can be installed on a manipulator arm of the robotic medical system 125 to perform a surgical task. The images (e.g., video images) captured by the medical tool 125 can be sent to the visualization tool 170. The robotic medical system 125 can include one or more input ports to receive direct or indirect connection of one or more auxiliary devices. For example, the visualization tool 170 can be connected to the robotic medical system 125 to receive the images from the medical tool when the medical tool is installed in therobotic medical system (e.g., on a manipulator arm of the robotic medical system). The visualization tool 170 can combine the data stream components from the data capture devices 110 and the medical tool 120 into a single combined data stream for presenting on a display 172 (e.g., display 930 depicted in FIG. 9). The display 172 can be associated with a client device 174, user control system or other type of display system, whether within the medical environment 105 or remote, to view the single combined data stream. The client device 174 can refer to or include a laptop computer, desktop computer, tablet, smartphone, portable computing device, or wearable device, for example.

[0044] The system 100 can include a data processing system 130 associated with the medical environment 105. The data processing system 130 can include an interface 132 designed, constructed and operational to communicate with one or more component of system 100 via network 101, including, for example, the robotic medical system 125 or client device 174. The data processing system 130 can include a data collector 134 to capture or otherwise receive or obtain data from one or more component or system associated with the medical environment 105 or via the network 101. The data processing system 130 can include a frame reformatter 136 to perform one or more pre-processing techniques on the captured data. The data processing system 130 can include a moment predictor 138 to predict a phase associated with the captured data on a moment-by-moment basis. The data processing system 130 can include a phase logit unit 140 to determine or predict full-length phases based on the moment- by-moment basis phase predictions. The data processing system 130 can include an action generator 176 to perform an action or provide an alert based on the full-length phase predictions. The data processing system 130 can include a model manager 142 that can maintain, manage, operate, or otherwise provide one or more models for utilization by the data processing system 130.

[0045] The interface 132, data collector 134, frame reformatter 136, moment predictor 138, phase logit unit 140, action generator 176, or model manager 142 can each communicate with the data repository 150 or database. The data processing system 130 can include or otherwise access the data repository 150. The data repository 150 can include one or more data files, data structures, arrays, values, or other information that facilitates operation of the data processing system 130. The data repository 150 can include one or more local or distributed databases, and can include a database management system. The data repository 150 can include, maintain, or manage video stream data 152. The data stream received by the data collector 134 can bereferred to as a data stream 152 or a video stream 152. Video stream 152 can include a captured medical procedure (e.g., surgery) as viewed from one or more sources (e.g., data capture devices 110 or medical tools 120). Video stream 152 can include a series of image frames from a variety of angles or vintage points with respect to the procedure activity (e.g., point or area of surgery), as well as any sound data, temperature data, pressure data, patient’s vital signs data or any other data corresponding to the procedure).

[0046] The data repository 150 can include, maintain, or manage event stream 154 data. The event stream 154 can include a stream of event data or information, such as packets, that identify or convey a state of the robotic medical system 125 or an event that occurred in association with the robotic medical system 125 or surgical or medical surgery being performed with the robotic medical system. Data of the event stream 154 can be captured by the robotic medical system 125 or a data capture device 110. An example state of the robotic medical system 125 can indicate whether the medical tool 120 is installed on a manipulator arm of the robotic medical system or not, whether it was calibrated, or whether it was fully functional (e.g., without errors) during the procedure. For example, when the medical tool 120 is installed on a manipulator arm of the robotic medical system 125, a signal or data packet(s) can be generated indicating that the medical tool has been installed on the manipulator arm of the robotic medical system 125. The signal or data packet(s) can be sent to the data collector 134 as the event stream 154. Another example state of the robotic medical system 125 can indicate whether the visualization tool 170 is connected, whether directly to the robotic medical system or indirectly through another auxiliary system that is connected to the robotic medical system.

[0047] The data repository 150 can include, maintain, or manage kinematics stream 156 data. For example, one or more of the manipulator arms or medical tools 120 attached to manipulator arms can include one or more displacement transducers, orientational sensors, positional sensors, or other types of sensors and devices to measure parameters or generate kinematics information. The kinematics data 156 can include sensor data along with time stamps and an indication of the medical tool 120 or type of medical tool 120 associated with the sensor data.

[0048] The data repository 150 can include resolution 158 data. The resolution 158 can refer to or include the time intervals or resolutions used by the frame reformatter 136 togenerate one or more sets of consecutive frames from the video stream 152. For example, the resolution 158 data structure can include a first time interval that is higher than a second time interval. The resolution 158 data structure can include a function that can output a random or varying time interval. The resolution 158 data structure can include a predetermined list of time intervals that can corresponding to the varying time interval. In an illustrative example, a high- interval sampling that is evenly-spaced can include sampling at: t, t+1, and t+2. In an illustrative example, a low-interval sampling that is evenly-spaced can include sampling at: t, t+5, and t+10. In an illustrative example, a random or varying interval sampling with varied spacing can include sampling at: t, t+1, and t+50.

[0049] The data repository 150 can store phases 160 data. Phases 160 can include, for example, exposure, dissection, transection, reconstruction, and extraction. Exposure can refer to or include the process of visualizing and accessing a surgical site by creating a clear and adequate field of view, dissection can refer to or include cutting, separating and removing tissues or anatomical structures to gain access to specific areas, identify structures, or perform surgical procedures. Transection can refer to or include severing or cutting a structure, such as a blood vessel, nerve, or organ using a surgical instrument. Extraction can refer to or include the removal of a tissue, organ, foreign object, or other anatomical structure from the body. Reconstruction can refer to or include the process of restoring or rebuilding a damaged or missing tissue, organ, or body part, and can include techniques or tasks such as grafting, suturing, or using prosthetic materials to recreate the structure and restore form and function.

[0050] The data repository 150 can include, manage or maintain historical data 162. Historical data 162 can include prior video stream, event stream, or kinematic stream data. Historical data 162 can include data associated with a location of the medical environment 105, workflows, facilities, or other information that can facilitate automatic full-length phase recognition. Historical data 162 can include training data used to train or update the models using machine learning.

[0051] The data processing system 130 can interface with, communicate with, or otherwise receive or provide information with one or more component of system 100 via network 101, including, for example, the robotic medical system 125 or client device 174. The data processing system 130, robotic medical system 125 or client device 174 can each include at least one logic device such as a computing device having a processor to communicate via thenetwork 101. The data processing system 130, robotic medical system 125 or client device 174 can include at least one computation resource, server, processor or memory. For example, the data processing system 130 can include a plurality of computation resources or processors coupled with memory.

[0052] The data processing system 130 can be part of or include a cloud computing environment. The data processing system 130 can include multiple, logically-grouped servers and facilitate distributed computing techniques. The logical group of servers may be referred to as a data center, server farm or a machine farm. The servers can also be geographically dispersed. A data center or machine farm may be administered as a single entity, or the machine farm can include a plurality of machine farms. The servers within each machine farm can be heterogeneous - one or more of the servers or machines can operate according to one or more type of operating system platform.

[0053] The data processing system 130, or components thereof can include a physical or virtual computer system operatively coupled, or associated with, the medical environment 105. In some embodiments, the data processing system 130, or components thereof can be coupled, or associated with, the medical environment 105 via a network 101, either directly or directly through an intermediate computing device or system. The network 101 can be any type or form of network. The geographical scope of the network can vary widely and can include a body area network (BAN), a personal area network (PAN), a local-area network (LAN) (e.g., Intranet), a metropolitan area network (MAN), a wide area network (WAN), or the Internet. The topology of the network 101 can assume any form such as point-to-point, bus, star, ring, mesh, tree, etc. The network 101 can utilize different techniques and layers or stacks of protocols, including, for example, the Ethernet protocol, the internet protocol suite (TCP / IP), the ATM (Asynchronous Transfer Mode) technique, the SONET (Synchronous Optical Networking) protocol, the SDH (Synchronous Digital Hierarchy) protocol, etc. The TCP / IP internet protocol suite can include application layer, transport layer, internet layer (including, e.g., IPv6), or the link layer. The network 101 can be a type of a broadcast network, a telecommunications network, a data communication network, a computer network, a Bluetooth network, or other types of wired and wireless networks.

[0054] The data processing system 130, or components thereof, can be located at least partially at the location of the surgical facility associated with the medical environment 105 orremotely therefrom. Elements of the data processing system 130, or components thereof can be accessible via portable devices such as laptops, mobile devices, wearable smart devices, etc. The data processing system 130, the data collector 134, or components thereof, can include other or additional elements that can be considered desirable to have in performing the functions described herein. The data processing system 130, or components thereof, can include, or be associated with, one or more components or functionality of computing system 900 depicted in FIG. 9, including, for example, one or more processors coupled with memory.

[0055] The data processing system 130 can include an interface 132 designed, constructed and operational to communicate with one or more component of system 100 via network 101, including, for example, the robotic medical system 125 or client device 174. The interface 132 can include a network interface. The interface 132 can include or provide a user interface, such as a graphical user interface.

[0056] The data processing system 130 can include a data collector 134 designed, constructed and operational to receive a video stream 152 that captures a procedure over a time interval with a robotic medical system 125. The data collector 134 can access the data stream from the visualization tool 170 or the display 172. The data collector 134 can receive an event stream 154 from the robotic medical system 125 or kinematic stream 156. The event stream 154 can include a stream of event data or information, such as packets, that identify or convey a state of the robotic medical system 125 or an event that occurred in association with the robotic medical system or surgical or medical surgery being performed with the robotic medical system. An example state of the robotic medical system 125 can indicate whether the medical tool 120 is installed on a manipulator arm of the robotic medical system or not. For example, when the medical tool 120 is installed on a manipulator arm of the robotic medical system 125, a signal or data packet(s) can be generated indicating that the medical tool has been installed on the manipulator arm of the robotic medical system 125. The signal or data packet(s) can be sent to the data collector 134 as the event stream 154. Another example state of the robotic medical system 125 can indicate whether the visualization tool 170 is connected, whether directly to the robotic medical system or indirectly through another auxiliary system that is connected to the robotic medical system.

[0057] The robotic medical system 125 can have other states, which can be detected by the data collector 134. The data collector 134 can determine (e.g., record) or otherwise receive the event stream 154 through an Application Programming Interface (API) of the robotic medicalsystem 125. In other embodiments, the data collector 134 can determine or otherwise receive the event stream 154 via other suitable mechanisms. The data collector 134 can poll the robotic medical system 125 to determine the state of the robotic medical system 125.

[0058] The data processing system 130 can include a frame reformatter 136 designed, constructed and operational to generate, from the video stream, sets of consecutive frames. The frame reformatter 136 can execute, prior to generation of the sets of consecutive frames, one or more pre-processing functions on the video stream. Pre-processing functions can include at least one of a central crop transform, frame resizing, filtering of non-surgical frames, or filtering of noisy frames. The frame reformatter 136 can include or perform the functionality of the frame sampler 604 depicted in FIG. 6.

[0059] The frame reformatter 136 can apply one or more pre-processing techniques to the video stream. For example, given a set of video frames from an endoscopic video stream 152, the frame reformatter 136 can perform a set of image-based operations to pre-process individual raw frames, such as central crop transform or frame resizing. Central crop transform can refer to or include a technique used in video editing to extract a smaller portion of a frame. The central crop transform can focus or detect portions of the frame that includes content corresponding to surgical activity that is indicative of a phase of the surgery. For example, the frame reformatter 136 can be configured to preserve, emphasize, more heavily weight, or otherwise focus on portions of a frame that include the color red or shades thereof, while deweighting or cropping out portions of the frame that are not red or have wavelengths that are further from the color red.

[0060] The frame reformatter 136 can extract the smaller portion of the frame from the center of the frame, and discard the outer areas. For example, the frame reformatter 136 can determine the crop dimensions, extract the desired region, and create a new frame with the cropped content. In some cases, the frame reformatter 136 can remove or crop portions of the frame that are blurred or contain images or content that is not discernable or has a low signal- to-noise ratio. By performing central crop transforms, the frame reformatter 136 can allow the downstream components or processing of the data processing system 130 to focus on regions of interest in the video stream, which can improve the efficiency by minimizing or reducing unnecessary computing resource utilizations on portions of the video frame that may contain data that is not of interest.

[0061] The frame reformatter 136 can pre-process the video stream by performing frame resizing, which can include the process of changing the dimensions of the video frame by altering the width or height of the frame to either increase or decrease the size of the frame. The data processing system 130 can perform this transformation to adapt the video to different display resolutions, aspect ratios, or to satisfy other constraints or improve efficiency, reliability or accuracy or downstream components or functions of the data processing system 130. For example, to the extent raw video frames have different sizes before or after executing a central crop transform, the frame reformatter 136 can resize the frames to a predetermined or default frame size so that the consecutive sets of frames generated from the input video stream have a consistent frame size.

[0062] The frame reformatter 136 can remove or filter out non-surgical frames, noisy frames (or portions of frames with excessive noise), or corrupted frames. To remove the noisy frames, or portions of the frame with excessive noise, the frame reformatter 136 can be configured with one or more functions to automatically detect frames with high levels of noise (e.g., above a threshold or frames with a low signal-to-noise ratio) based on analyzing pixel variations, statistical properties, or utilizing machine learning-based noise detection techniques. The frame reformatter 136 can attempt to reduce the noise in the frames by applying noise reductions techniques, or can determine to filter out or remove the frame. The frame reformatter 136 can detect and remove corrupted frames based on comparing a frame with a previous frame to determine an amount of inconsistency or variation between corresponding pixels or regions between frames.

[0063] The frame reformatter 136 can remove non-surgical frames using various techniques. For example, the frame reformatter 136 can remove frames that correspond to timestamps before a start time of the surgery, or after an end time of the surgery. The frame reformatter 136 can remove frames that were captured by cameras, medical tools, or data capture devices that were not used to perform the surgical procedure.

[0064] Upon pre-processing the raw video stream 152, the frame reformatter 136 can construct or generate one or more sets of consecutive frames by sub-sampling from the set of video frames based on a temporal resolution. The frame reformatter 136 can sample the pre- processed video stream using different sampling rates or intervals in order to generate the different sets of consecutive frames. The frame reformatter 136 can generate more than one set of consecutive frames by using different temporal resolutions. The time resolution can refer toa measure of how finely the video frames or samples are spaced or timestamped in relation to the duration of the video. A higher time resolution means that changes in the video signal can be captured with more precision, allowing for finer temporal details to be observed. Conversely, a lower time resolution may result in a loss of temporal fidelity, where rapid changes or subtle variations in the video may not be accurately represented. The vision-based data reformatting by the frame reformatter 136 can sample the input videos and extract sufficient frames to capture surgical activities that may have varied motion characteristics. The number of sampled frame sets and frame interval settings can be based on configuration settings.

[0065] The time resolutions or time intervals can be evenly spaced or have varying spacing. The frame reformatter 136 can retrieve the temporal resolutions from the resolutions data structure 158 stored in data repository 150. The frame reformatter 136 can determine the temporal resolution using a resolution function stored in the resolution data structure 158. For example, the frame reformatter 136 can generate a first set of consecutive frames from the video stream 152 based on a short temporal interval that is evenly spaced. The frame reformatter 136 can use a high temporal resolution for the first set to capture the potentially fast changing surgical motions that occur swiftly or evolve within a short time span. The frame reformatter 136 can generate the first set of consecutive frames with the first temporal resolution that is less than or equal to a first threshold. For example, short, evenly spaced temporal intervals the frame reformatter 136 can use can include, for example, 0.25 seconds, 0.5 seconds, 1 second, 1.5 seconds, 2 seconds or other time interval. The time resolution can be a reciprocal of the time interval, and be in units of Hertz (Hz). For example, the corresponding high time resolutions can include 4 Hz, 2 Hz, 1 Hz, 0.67 Hz, or 0.5 Hz.

[0066] The frame reformatter 136 can generate a second set of consecutive frames from the video stream 152 based on a long temporal interval that is evenly spaced. The frame reformatter 136 can generate the second set of consecutive frames with the second temporal resolution that is greater than or equal to a second threshold that is greater than the first threshold. The frame reformatter 136 can use a low temporal resolution to capture relatively slow changing surgical operations that occur or evolve slowly. Long, evenly spaced temporal intervals the frame reformatter 136 can use can include, for example, 3 seconds, 4 seconds, 5 seconds, 6 seconds, or 10 seconds. The time resolution can be a reciprocal of the time interval,and be in units of Hertz (Hz). For example, the corresponding low time resolutions can include 0.3 Hz, 0.25 Hz, 0.2 Hz, 0.16 Hz, or 0.1 Hz.

[0067] The frame reformatter 136 can generate a third set of consecutive frames from the video stream by sub-sampling the video stream based on or using varying sampling rates, time intervals, or temporal resolutions. The third set of consecutive frames may not be evenly spaced, or lack even spacing, from one sample to the next sample. The varying time intervals or temporal resolutions can be predetermined or retrieved from the resolutions data structure 158. In some cases, the frame reformatter 136 can compute or determine the varying temporal resolutions or time intervals using a functions, program, or random number generator. For example, the varying time intervals can be predetermined or stored as a list. The frame reformatter 136 can repeat the list of varying time intervals, or the list of varying time intervals can include sufficient time intervals for the duration of the video stream. For example, the varying time intervals can include time intervals: 1 second, 10 seconds, 3 seconds, 50 seconds, 25 seconds, etc. In some cases, the frame reformatter 136 can include or utilize a random number generator to determine the varying time intervals. The random number generator can generate a sequence of numbers with an unpredictable pattern. The frame reformatter 136 can use an initial seed value, which serves as the starting point, and subsequent iterations based on predefined rules to generate a sequence of seemingly random numbers to use for time intervals or temporal resolutions to generate the third set of consecutive frames. Thus, the frame reformatter 136 can generate the third set of consecutive frames to increase temporal complexity while preserving the sequential relations in the set of video frames.

[0068] By multiple sets of consecutive frames with different or varying temporal resolutions, the frame reformatter 136 can provide different views in videos at varied temporal resolutions, which can improve the reliably, accuracy or efficiently with which the data processing system 130 (e.g., using one or more machine learning models) can capture temporal relationships at different levels of granularity that can lead to more accurate predictions.

[0069] The data processing system 130 can include a moment predictor 138 designed, constructed and operational to determine phases of the procedure on a moment-to-moment basis over the time interval. The moment predictor 138 can determine the phases on the moment-to-moment basis using the multiple sets of consecutive frames generated by the frame reformatter 136. For example, the moment predictor 138 can input the multiple sets of consecutive frames into a first model 144 trained with machine learning. The first model 144can be maintained or managed by a model manager 142. The moment predictor 138 can interface with model manager 142 to utilize or otherwise access the first model 144. The moment predictor 138 can include or perform the functionality of a temporal + spatial feature extractor 606 as depicted in FIG. 6. The moment predictor 138 can include or perform the functionality of a frame-wise predictor 608 depicted in FIG. 6.

[0070] The moment predictor 138 can include, utilize or be configured with a multipathway spatial -temporal decoding unit based on attention-based deep learning model. The first model 144 can include or correspond to a model trained using attention-based deep learning. The moment predictor 138 can use the multi -pathway spatial -temporal decoding unit based on the attention-based deep learning model (e.g., first model 144) to generate moment-to-moment predictions of phase for the current moment t. The first model 144 can include a multi-pathway spatial-temporal decoding unit comprising an attention-based deep learning model.

[0071] For example, the frame reformatter 136 can generate three sets of consecutive frames at a current moment, which can be referred to as current moment “f ’. For the current moment t, the moment predictor 138 can utilize the multi-pathway spatial -temporal decoding unit based on the attention-based deep learning model to generate moment-to-moment predictions of phase for the current moment t. The model 144 can be based on or leverage a vision transformer architecture with self-attention over space and time to provide jointed spatial and temporal feature learning from the videos.

[0072] The moment predictor 138 can include parallel processing streams. The moment predictor 138 can include three processing streams that are in parallel or otherwise configured to process the sets of consecutive frames output by the frame reformatter 136 in parallel. Processing the different sets of consecutive frames in parallel can refer to processing a frame from each set of consecutive frames at the same time, substantially the same time, or in an overlapping manner. The moment predictor 138 can include three streams in parallel, and each stream can take as input a set of frames and extract or output features from the input videos.

[0073] To do so, the moment predictor 138 can use a first model 144 trained or configured with a vision transformer (“ViT”). A vision transformer can refer to a transformer-type model configured to handle or process vision processing tasks. The moment predictor 138 using the first model 144 based on a vision transformer can capture global dependencies and interactions between image patches or pixels by leveraging self-attention mechanisms. The moment predictor 138 can break down an image into smaller patches and process them using multi-headself-attention layers, allowing the moment predictor 138 to learn spatial relationships and extract meaningful representations from the image data.

[0074] In an illustrative example, the moment predictor 138 can obtain a frame from a set of consecutive frames. The moment predictor 138 can divide the input frame into smaller fixed-size regions or portions, where each portion can represent a region of the frame. The moment predictor 138 can transform the portion of the frame into an embedding vector through a linear projection (e.g., linear projection 312 depicted in FIG.3). The moment predictor 138 can input the embedding vector (e.g., vector 316 depicted in FIG. 3) into a first model 144. The first model 144 can include or be based on the vision transformer model, such as the space-time timesformer model or encoder 320 depicted in FIG. 3. The first model 144 can include a stack of transformer encoder layers, where each layer can include spatial or temporal self-attention mechanisms and feed-forward neural networks. The moment predictor 138 (e.g., via first model 144) can use the self-attention mechanism to attend to different portions of the frame and capture relationships between the different portions by determining attention weights for each portion by considering relationships between the portions in comparison to other portions. The moment predictor 138 can determine global dependencies within the image based on this comparison.

[0075] In some cases, the moment predictor 138 can additional apply a temporal attention mechanism, which can include a transformer encoder layer with a self-attention mechanism configured to capture long-range temporal dependencies within portions of the frame. With this technique, the moment predictor 138 can determine or account for temporal dynamics and relationships between frames in a set of consecutive frames. For example, the moment predictor 138 can use a timesformer architecture to extract spatial and temporal features, with attention in both space and time. The timesformer architecture can include applying selfattention over space and time by adapting the transformer architecture to generate spatiotemporal features from a sequence of portions of frames. The timesformer architecture can adapt the ViT to video by extending the self-attention mechanism from the image space to the space-time 3D volume. Each portion can be linearly mapped into an embedding and augmented with positional information to allow the moment predictor 138 to interpret the resulting sequence of vectors. The moment predictor 138 can combine the outputs of the selfattention mechanism with the embedding vector to create context-aware representations of the frame, which can facilitate capturing both local and global information within the image frame.

[0076] The moment predictor 138 can input the context-aware spatial -temporal representations into a classifier or classification layer. The classifier or classification layer can be part of the first model 144 or moment predictor 138. The moment predictor 138, configured with the classifier or otherwise utilizing a classification layer, can generate output probabilities corresponding to image classification or pixel-wise predictions for segmentation. The moment predictor 138 can generate a numerical feature output from each of the parallel processing streams. In some cases, the moment predictor 138 can utilize additional transformations, pooling operations, or fully connected layers to generate the output probabilities.

[0077] The moment predictor 138 can generate an output based on the outputs of each of the parallel streams. For example, the moment predictor 138 can aggregate or otherwise combine the numerical feature outputs from each of the parallel streams to generate an aggregated numerical feature output. For example, the moment predictor 138 can perform feature-level fusion by aggregating the numerical feature outputs of the multiple parallel streams (e.g., three parallel streams corresponding to the different sets of consecutive frames). The moment predictor 138 can aggregate the numerical feature outputs to generate an aggregated feature vector, for example. The aggregated feature vector can include numerical feature values for each of the parallel streams. The aggregated feature vector can be based on a combination of the numerical feature values for each of the parallel streams. The moment predictor 138 can use one or more techniques for feature-level fusion to generation the aggregated feature vector, including, for example, concatenation, summation, averaging, weighted combination, principal component analysis, or deep learning-based fusion networks. For example, the moment predictor 138 can concatenate the numerical feature outputs from the parallel streams along a dimension. In some cases, the moment predictor 138 can normal the fused feature based on feature scaling or feature-wise normalization to generate the aggregated feature vector.

[0078] The moment predictor 138 can use the aggregated numerical feature vector to generate the moment-wise prediction of phases for the current moment t. For example, the moment predictor 138 can generate the phase on a moment-by-moment basis (e.g., the phase for the current moment t) based on the aggregated feature vector. The data processing system 130 can generate a phase prediction for each time stamp or current moment t (e.g., phase predictions on a moment-by-moment basis or a moment-wise phase prediction. Thus, the data processing system 130 can input the multiple sets of consecutive frames into a correspondingplurality of parallel streams to generate corresponding numerical feature outputs, and fuse the numerical feature outputs to generate the phases of the procedure on the moment-to-moment basis.

[0079] The data processing system 130 can include a phase logit unit 140 designed, constructed and operational to generate a full-length phase prediction. The full-length phase prediction can refer to a phase segment. A phase segment can by bounded by either different phases at before or after the phase segment, or a start time or end time of the procedure. A phase at a current moment t or a moment-wise phase prediction can refer to the phase at a specific timestamp or moment, whereas a phase segment (e.g., full-length phase) can refer to the phase over multiple timestamps that spans or covers the duration of the phase from the beginning of the phase to the end of the phase. For example, a particular phase such as exposure may occur at the beginning of the procedure and have a duration of 30 minutes. Thus, the phase logit unit 140 can determine, based on the mom ent- wise phase predictions output by the moment predictor 138, that the first phase segment in the procedure corresponds to the exposure phase and has a duration of 30 minutes, and is followed by a second phase in the procedure, such as dissection in a cholecystectomy procedure, for example.

[0080] The phase logit unit 140 can input the phases of the procedure determined on the moment-to-moment basis (e.g., the moment-wise phase predictions generated by the moment predictor 138) into a second model 146. The second model 146 can be trained with machine learning based on historical workflows (e.g., historical data 162). The phase logit unit 140 can use the second model 146 to generate at least one phase segment over the time interval of the procedure based on the moment-to-moment phase predictions.

[0081] To do so, the phase logit unit 140 can find the boundaries of each surgical phase and generate full-length phase recognition (e.g., identify, recognize, or otherwise determine each phase segment) for a procedure from the moment-to-moment phase predictions. The phase logit unit 140 can determine the phase boundaries using context about the surgical workflow to inform moment-to-moment predictions. The phase logit unit 140 can identify the phase boundaries by analyzing variance in the moment-to-moment predictions to facilitate quantifying uncertainty in the boundaries. The phase logit unit 140 can identify the phase boundaries using both the surgical workflow context and the variance in the moment-to- moment predictions.

[0082] The phase logit unit 140 can include or leverage a second model 146. The secondmodel 146 can include a long-range temporal module that is trained, configured or otherwise generated based on an analysis of surgical workflows to model a distribution over phases at each moment in time that serves as a prior belief. This second model 146 can configured with or incorporate information about logical workflows to augment the information gained from the first model 144 (e.g., the vision-based model). The data processing system 130 (e.g., phase logit unit 140) can use the second model 146 to update the belief or probability about the phase at each timestep (e.g., the moment- wise phase predictions output by the moment predictor 138) into a posterior distribution. The second model 146 can combine information about both the average duration of phases and the temporal ordering of phases, so that at each timestep the data processing system 130 can define, identify or otherwise establish a prior belief of the likelihood of staying within the same phase, or transitioning to each of the other phases.

[0083] The data processing system 130 can use this combination model to leverage uncertainty by modeling the likelihood of each phase label for a given timestamp. The phase logit unit 140 can use the uncertainty in the phase label at each timestamp rather than, or in addition to, making a single phase prediction.

[0084] For example, the data processing system 130 can determine based on surgical workflow analysis that there is a strong prior belief that a surgery may enter a given phase next. For example, in the example of a cholecystectomy procedure, the data processing system 130 can determine, based on the surgical workflow context incorporated into the second model 146, that there is a strong likelihood or probability that the dissection phase will follow the first exposure phase. Due the strong likelihood or probability that the dissection phase will follow the first exposure phase based on the workflow context of the second model 146, any uncertainty prediction from the moment-to-moment model may be less likely to override this prior belief.

[0085] However, when there are multiple potential next phases in the surgery procedure (e.g., after dissection, the potential next phases could be transection or extraction), the prior beliefs may be more evenly distributed between phase classes. If the prior beliefs are more evenly distributed between phase classes, then moment-to-moment model 144 may be more heavily relied upon. The data processing system 130 can utilize this structure to generate uncertainty-aware phase boundaries throughout the procedure time interval, while using previous information, thereby facilitating or allowing the data processing system 130 to make real-time phase predictions.

[0086] The phase logit unit 140 can use a phase transition map to generate the full-length phase recognition or phase segments of the procedure. The phase transition map can be stored in the phases 160 data structure in the data repository 150. The phase logit unit 140 can retrieve or otherwise access the phase transition map from the phases data structure 160 to determine the full-length phases. The phase transition map can be configured, tuned, calibrated, customized, or otherwise generated for each type of surgical procedure to include prior probabilities of phase transitions from one moment t to next moment t. The phase transition map can be customized, generated or modified for particular types of procedures, a site location of the surgery, or a site location and type of procedure. The phase transition map can be generated based on the first model 144 or second model 146.

[0087] For example, the phase transition map can indicate, for a particular type of procedure, a first likelihood of a first phase at a current moment t being followed by the same first phase at a next moment t+1; a second likelihood of the first phase at the current moment t being followed by a second phase at a next moment t+1; a third likelihood of the first phase at the current moment t being followed by a third phase at a next moment t+1; a fourth likelihood of the first phase at the current moment t being followed by a fourth phase at a next moment t+1; and a fifth likelihood of the first phase at the current moment t being followed by a fifth phase at a next moment t+1. An illustrative phase transition map 638 is depicted in FIG. 6.

[0088] The phase logit unit 140 can determine a variance in the moment-wise phase predictions of the procedure. The phase logit unit 140 can generate, based on the variance and a phase transition map configured with a plurality of prior probabilities, uncertainty-aware phase boundaries throughout the time interval. The phase logit unit 140 can generate the at least one phase segment that corresponds to a phase boundary of the plurality of uncertainty-aware phase boundaries. The phase segment can be bounded by two phase boundaries. The phase boundaries can correspond to different types of phases, the same type of phase, a procedure start time, or a procedure end time.

[0089] The phase logit unit 140 can generate the at least one phase segment based at least in part on one or more of an average duration of phases, an ordering of phases (e.g., based on a phase transition map), a type of the procedure (e.g., prior probabilities of ordering of phases for a particular type of procedure), or a site location of the procedure (e.g., a phase transition map can be customized for a site location and a type of procedure).

[0090] The phase logit unit 140 can incorporate a variety of different information sourcesto transition from moment-to-moment predictions to full-length phase predictions (e.g., phase segments). For example, the phase logit unit 140 can average predictions over time (e.g., a sliding window with the majority vote of predicted phase labels) to create smooth phase boundaries. The data processing system 130 can incorporate the phase transition map transition probabilities, which can include updated or modified probabilities based on the procedure type or hospital site. This modified phase transition map can provide further context, since phase workflows may differ by procedure, and the workflows and techniques may vary from surgeon to surgeon. The second model 146, by including this additional information, can be both generalizable and specialized or tuned or calibrated to particular input data.

[0091] In some cases, the data processing system 130 can leverage unique events to identify the boundaries of certain phases. For example, the data processing system 130 can use event stream 154 data, such as the installment and unmount of a needle driver, to facilitate identification of the phase segment boundaries of a reconstruction phase.

[0092] For full-length phase recognition, the phase logit unit 140 can implement a decision tree by combining the video-based model prediction, kinematic patterns (e.g., from the kinematics stream 156), and event knowledge (e.g., from the event stream 154). Based on multi-modal information, the system can improvement the performance of full-length phase recognition.

[0093] In some cases, the data processing system 130 can determine the phases of the procedure on the moment-to-moment basis based on a combination of a video stream 152 and at least one of a stream of system events data (e.g., event stream 154) or a stream of kinematics data (e.g., kinematics stream 156). The data processing system 130 can receive the event stream 154 and kinematics stream 156 from the robotic medical system 125. These multimodal information can also be used to recognize surgical phases in combination with video, or as a replacement for video.

[0094] The data processing system 130 can use the kinematic stream 156 to recreate the position of the medical tools 120 tools at one or more points in time to generate a time series of tool positions. With the time series of these tool positions, the data processing system 130 can determine the motions of each tool, and recognize different gestures that are involved with different surgical techniques or tasks (e.g., suturing or cutting). For example, the data processing system 130 can detect a suturing motion through kinematic patterns. Since these actions, such as suturing, cutting, and others can be associated with surgical intent, the dataprocessing system 130 can use the detected tasks to recognize the current phase of surgery (e.g., the needle driving can be uniquely identified in the phase of reconstruction).

[0095] The data processing system 130 can use the event stream 154 data, such as information about tool installations and the use of special system features. The data processing system 130 can determine additional context to the surgical intent at any given moment. For example, the data processing system 130 can determine that a predetermined set of tools are configured for the phase segment (e.g., such as a stapler, clip applier, or needle driver, are useful or configured for specific operations within a surgery, and these can be mapped to the current phase of surgery). Thus, the data processing system 130 can combine information or aspects of the video stream, kinematics stream, and event stream to inform each other to make a moment-wise phase prediction or recognize a phase segment.

[0096] The data processing system 130 can include an action generator 176 designed, constructed and operational to provide an action based on a metric of the at least one phase segment. The action generator 176 can generate a metric for the recognized at least one phase segment. The metric can be a performance metric associated with the phase segment. The metric can be a measure of performance of the phase segment. The metric can indicate a level of performance of the procedure during one or more phase segments determined by the phase logit unit 140 over the time interval of the procedure. The action generator 176 can output the metrics in real-time (e.g., before the procedure is complete, during the phase segment itself, after the beginning phase segment boundary is detected or after the end phase segment boundary is detected). The metric can include one or more numerical values, scores, grades, percentages, text, colors, or symbols. The action generator 176 can provide an action that includes generating and presenting the metric using one or more of visual output, audio output, or haptic output or feedback.

[0097] For example, the metric can be the duration of the recognize phase segment. The metric can be based on a comparison of the duration of the phase segment for the procedure with an average duration of the phase segment for this type of procedure. The average duration of the phase segment for the particular type of procedure can be determined for a particular surgeon, time interval (e.g., last 24 hours, 48 hours, week, month, or quarter), site location, or geographic region (e.g., city, state, country). Thus, the data processing system 130 can first identify the phase segment, and then generate a performance metric for the phase segment. The data processing system 130 can provide the metric upon detecting a phase boundary at the endof the phase segment, or upon transitioning to the next phase segment. The data processing system 130 can provide an action that includes outputting the metric to a display 172, or client device 174.

[0098] The data processing system 130 can output the metric in real-time. For example, the data processing system 130 can determine that the duration of the phase segment exceeds a threshold duration for the type of phase segment. Responsive to determining that the duration of the phase segment has exceeded the threshold, the data processing system 130 can generate an alert, alarm, warning or other notification indicating that the duration of the phase segment has exceeded the threshold. The threshold can be based on the average duration for this type of segment for this type of procedure, for the site location, or for the surgeon, for example.

[0099] The action generator 176 can generate performance metrics corresponding to efficiency of the recognized phase segment. Efficiency can be based on the duration, kinematics data for the phase segment, or types of number of medical tools used during the phase segment. For example, the data processing system can use the kinematics stream 156 to determine the efficiency of movement associated with tasks during the phase segment to determine the metric. The data processing system 130 can use the event stream 154 to determine the number or types of tools used during the phase segment or the number of tool installation and uninstallations during the phase segment. The data processing system 130 can compare these values with average or expected or historical values to determine a performance metric for the phase segment.

[0100] The action generator 176 can facilitate scheduling an operating room based on full- length phase recognition. For example, the action generator 176 can generate an action to modify a start time of a next procedure scheduled for the operating room based on recognizing a current phase segment. The action generator 176 can determine based on the recognized phase segment that the current procedure is behind schedule or ahead of schedule. Based on the determination, the action generator 176 can generate a notification, instruction, command, or alert to push back or push forward subsequent procedures planned for the operating room. In some cases, the action generator 176 can generate an action or notification to reserve a different operating room in the same site location, or in a different site location.

[0101] The data processing system 130 can include a model manager 142 designed, constructed and operational to maintain, manage, update, train, or otherwise provide components of the data processing system 130 with access to one or more models, such as thefirst model 144 and the second model 146. The model manager 142 can include one or more functions, such as machine learning functions. The model manager 142 can receive input requests from one or more components of the data processing system 130, and provide output from the models 144 or 146 to the components of the data processing system 130. The model manager 142 can train the models, or can receive the models from a remote source, such as a cloud computing environment or one or more servers via the network. The first model 144 can be or include a multi-pathway spatial-temporal decoding unit comprising an attention-based deep learning model. The second model 146 can be trained with machine learning based on historical workflows (e.g., historical data 162). Historical surgical workflows can refer to or include historical information about the order or phase segments in different types of surgeries or procedures. The order of phase segments can be specific to one or more of a type of procedure, a site location (e.g., type of hospital, size of hospital, unique identifier of the hospital), a surgeon, or a geographic region (e.g., town, city, state, or country). The phase transition map can indicate the order of phases segments.

[0102] Thus, the data processing system 130 can determine a phase segment of the at least one phase segment that extends over a current time. The data processing system can identify a tool used in the phase segment based on a stream of system events. The data processing system 130 can determine that the tool does not match any of a predetermined set of tools configured for the phase segment. The data processing system 130 can provide an alert during the phase segment responsive to the determination that the tool does not match any of the predetermined set of tools configured for the phase segment.

[0103] FIG. 2 depicts an example process for universal phase recognition for intraoperative and postoperative applications. The process 200 can be performed by one or more system or component depicted in FIG. 2 or FIG. 9, including, for example, a data processing system. The procedure can begin at 202 and end at 204. The time interval or duration of the procedure can extend from 202 to 204. The procedure can refer to a surgical procedure or medical procedure. The procedure can be a specific type of procedure.

[0104] At each current moment t during the procedure, the process 200 can include performing one or more of ACTs 206, 208, 210, 212, 214, or 216 to recognize phase segments 218. At ACT 206, the data processing system can perform data collection and reformatting. For example, the data processing system can capture or receive video stream data 152 from an image capture device. The data processing system can pre-process and reformat (e.g.,subsample at different temporal resolutions) the data to generate sets of consecutive frames with different temporal resolutions. For example, the video stream 152 can reflect the raw video stream received from the robotic medical system, or the pre-processed or reformatted image frames of the video stream. The data processing system can execute, prior to generation of the plurality of sets of consecutive frames, one or more pre-processing functions on the video stream. The one or more pre-processing functions can include at least one of a central crop transform, frame resizing, filtering of non-surgical frames, or filtering of noisy frames.

[0105] At ACT 208, the data processing system can input the sets of consecutive frames into a first model, such as a multi-pathway spatial-temporal decoding unit comprising an attention-based deep learning model. The first model can provide output 210, which can correspond to the pre-processed and reformatted image frames 152 input into the first model.

[0106] At ACT 212, the data processing system can use the output 210 to predict phases at a current moment t (e.g., a moment-wise phase prediction or a phase prediction determined on a moment-to-moment basis). At ACT 214, the data processing system can input the moment- to-moment phase predictions into a second model which can include a long-range temporal module 214 to facilitate generating full-length phase predictions. For example, the data processing system can use a long-range temporal module 214 constructed based on an analysis of surgical workflows to model a distribution over phases at each moment in time that serves as a prior belief or probability or likelihood. The long-range temporal module 214 can incorporate information about logical workflows to augment the information gained from the vision-based model, which is used to update the belief about the phase at each timestep into a posterior distribution. The model 214 can combine information about both the average duration of phases and the temporal ordering of phases, so that at each timestep the data processing system can define a prior belief of the likelihood of staying within the same phase, or transitioning to each of the other phases.

[0107] At ACT 216, the data processing system can perform phase recognition for the entire procedure. The data processing system can utilize a full-length phase logit unit to identify the phase segments 218 for the entire procedure, from 202 to 204. The full-length phase logit unit can identify phase boundaries 220 between each phase segment.

[0108] FIG. 3 depicts an example process for universal phase recognition for intraoperative and postoperative applications. The process 300 can be performed by one or more system or component depicted in FIG. 2 or FIG. 9, including, for example, a data processing system. Theprocess 300 can include temporal deep learning with two or more streams in parallel. The process 300 can include obtaining a video stream 152. The video stream can have a dimension of T x N, which can correspond to a resolution of the video stream 152. The data processing system can generate sets of consecutive frames with different frame intervals. For example, the data processing system can generate a set of consecutive frames with a high frame interval 304 with a set of consecutive frames 308 that includes frame #1, frame #2, up to frame #m. The data processing system can generate a set of consecutive frames with a low frame interval 306 with a set of consecutive frames 310.

[0109] The data processing system can use a linear projection 312 to generate an embedding vector 316 from the high frame interval set of frames. The data processing system can use a linear projection 314 to generate an embedding vector 318 from the low frame interval set of frames. The data processing system can input the embedding vectors 316 and 318 into respective space-time timesformer encoders 320 and 322. At 324, the data processing system can perform feature-level fusing of the numerical feature outputs from the space-time timesformer encoders 320 and 322 to generate an aggregated feature vector. The data processing system can make a prediction 326 as to the moment-wise phase or full-length phase segment based on the aggregated feature vector output at 324.

[0110] FIG. 4 depicts an example user interface for universal phase recognition for intraoperative and postoperative applications. The graphical user interface (GUI) 400 can be generated by one or more system or component depicted in FIG. 2 or FIG. 9, including, for example, a data processing system. FIG. 4 depicts example moment-to-moment phase predictions from machine learning models based on endoscopic video streams. The GUI 400 can depict an image frame of a video stream at one or more moments, such as moment A 402 or moment B 418. For moment A 402, the data processing system can output a confidence 406 in a moment-wise phase prediction or phase segment prediction. The data processing system can display the probability that the image frame corresponds to each of the different types of phases (408, 410, 412, 414 and 416) as follows: exposure 0.788; dissection 0.066; transection 0.110; reconstruction 0.036, and extraction 0.000. Based on these probabilities, the data processing system can determine that the phase name 404 is exposure with a confidence value of 0.7888, and an uncertainty of 0.454. At moment B 418, the data processing system can display the probability that the image frame corresponds to each of the different types of phases as follows: exposure 0.101; dissection 0.002; transection 0.010; reconstruction 0.887, andextraction 0.000. Based on these probabilities, the data processing system can determine for moment B that the phase name 420 is reconstruction with a confidence value of 0.3887, and an uncertainty of 0.246. The data processing system can output or display the GUI for moment A 402 or moment B 418 in real-time.

[0111] FIG. 5 depicts an example process for universal phase recognition for intraoperative and postoperative applications. The process 500 can be performed by one or more system or component depicted in FIG. 2 or FIG. 9, including, for example, a data processing system. FIG. 5 illustrates an overall workflow or process of automatic generated phase recognition based on task annotations. The procedure can begin at 502 and end at 504. At 506, the data processing system can obtain or parse task-level human annotations. The annotations can correspond to actions, steps, or procedures performed during the surgical procedure. The task-level annotations can include identifying and labeling the surgical tasks or activities, such as incision, dissection, suturing, ligation, hemostasis, retraction, or closure. At ACT 508, the data processing system can identify the various task-level annotations from moment-to-moment throughout the procedure. At ACT 510, the data processing system can use a many-to-one task- to-phase convertor to generate a phase corresponding to the task-level annotations. At ACT 512 can input the moment-wise phase predictions 512 into a temporal sequencing decoder 514 to generate phase segments 516.

[0112] FIG. 6 depicts an example process for universal phase recognition for intraoperative and postoperative applications. The process 600 can be performed by one or more system or component depicted in FIG. 2 or FIG. 9, including, for example, a data processing system. The method 600 can include a data stream 602, frame sampler 604, temporal + spatial feature extractor, frame-wise predictor 608 and a phase transition map 610. The process 600 can include providing a video 152 data stream to a frame sampler 604. The frame sampler 604 can include one or more component or functionality of the frame reformatter 136 depicted in FIG.1. The frame sampler 604 can perform evenly-spaced high-interval sampling 614 (e.g., at t, t+1, t+2, etc.) to generate a first frame sequence 622. The frame sampler 604 can perform a low- interval evenly spaced frame sampling 616 (e.g., at t, t+5, t+10, etc.) to generate a second frame sequence 624. The frame sampler 604 can perform a varying interval sampling 618 (e.g., at t, t+1, t+50) to generate a third frame sequence 626.

[0113] The frame sampler 604 can perform structured sampling 620 that is focused on preferred surgical content or aspects. The structured sampling 620 can include cropping orremoving portions of the image that are not relevant or may detract or negatively impact downstream processing. The structured sampling 620 can include removing non-surgical content.

[0114] The temporal + spatial feature extractor 606 can include one or more component or functionality of the moment predictor 138 depicted in FIG. 1. The feature extractor 606 can input the frame sequences 622, 624 and 626 into a space-time vision transformer encoder 628 to generate output 630, 632, and 634 for the corresponding input frame sequences. The framewise predictor 608 can include one or more component or functionality of the moment predictor 138 in order to generate moment-to-moment predictions 636 from the output 630, 632 and 634.

[0115] The data processing system can perform phase transition mapping 610 on the moment-to-moment predictions. The phase transition map 638 can indicate the probability of current phase and next phase, and can be generated using phase data. The data processing system can provide the output of the phase transition map to a fusion 654 function. The fusion 654 can include a feature-level fusion. The fusion 654 can combine or aggregate information generated from system events 640 and clinical knowledge 648.

[0116] For example, system events 640 can be based on an event stream 154, and include information 642 regarding the installation or uninstallation of one or more tools A or B. The data processing system can identify the installation of tool A 644 during the procedure, and the uninstallation of tool A 646 during the procedure using the event stream data. The data processing system can provide the tool installation status information to a tool install converter 648 to generate phase labels based on the types of tools used. In some cases, the data processing system can use clinical knowledge 650 to facilitate tool-phase mapping at 650, and feed the tool phase mapping into the tool install converter 648. The data processing system can utilize a heuristic phase converter 656 to identify phases based on the clinical knowledge 650.

[0117] The data processing system can fuse 654 the input from the transition map 638, tool install converter 648, and heuristic phase converter 656 to recognize full-length phase start / stops (e.g., phase segments) 658.

[0118] FIG. 7 depicts an example method for universal phase recognition for intraoperative and postoperative applications. The process 300 can be performed by one or more system or component depicted in FIG. 2 or FIG. 9, including, for example, a data processing system. Thedata processing system can automatically identify the universal surgical phase activities of procedures across a variety of procedures in robot-assisted surgery and can flexibly generate full-length phase recognition for either intraoperative or postoperative applications and use cases. The data processing system can use a multi-pathway spatial and temporal encoding to efficiently extract and reveal the temporal relationship from the input video sequence. The extracted video representations can provide more accurate performance when generating moment-to-moment phase predictions. The real-time phase recognition can be deployed by the data processing system in intraoperative clinical applications to improve the context-awareness of an on-going procedure.

[0119] At ACT 702, the data processing system can receive a video stream. The data processing system can receive the video stream from one or more cameras or a robotic medical system. The video stream can include raw video frames or images. The video stream can capture at least a portion of a procedure over a time interval with a robotic medical system.

[0120] At ACT 704, the data processing system can generate sets of consecutive frames from the video stream. The data processing system can generate, from the video stream, a plurality of sets of consecutive frames. The sets of consecutive frames can include one or more of: a first set of consecutive frames with a first temporal resolution; a second set of consecutive frames with a second temporal resolution; or a third set of consecutive frames with a varying temporal threshold that varies based at least in part on a function. In some cases, the data processing system can execute, prior to generation of the sets of consecutive frames, one or more pre-processing functions on the video stream (e.g., a central crop transform, frame resizing, filtering of non-surgical frames, or filtering of noisy frames.)

[0121] At ACT 706, the data processing system can determine phases on a moment-to- moment basis. The data processing system can input the sets of frames into a model trained with machine learning, such as a space-time vision transformer encoder, to extract temporal and spatial features from the sets of frames. The data processing system can determine moment-wise phases based on the extracted features.

[0122] At ACT 708, the data processing system can input the phases into a second model to generate or recognize a phase boundary or phase start / stop to determine a phase segment. The second model can include a model trained on a procedure workflow data or historical workflows. The second model can include a long-range temporal module. At ACT 710, the data processing system can provide an action based on a metric of the at least one phasesegment. The action can include displaying an indication or label of the phase segment, or adjusting scheduling of the operating room or medical environment based on the recognized phase segment.

[0123] FIG. 8 depicts an example medical environment 105, in accordance with embodiments. The medical environment 105 can refer to or include a surgical environment or surgical system. The medical environment 105 can include a robotic medical system 125, a user control system 810, and an auxiliary system 815 communicatively coupled one to another. A visualization tool 820 (e.g., the visualization tool 170) can be connected to the auxiliary system 815, which in turn can be connected to the robotic medical system 125. Thus, when the visualization tool 820 is connected to the auxiliary system 815 and this auxiliary system is connected to the robotic medical system 125, the visualization tool can be considered connected to the robotic medical system. In some embodiments, the visualization tool 820 can additionally or alternatively be directly connected to the robotic medical system 125.

[0124] The medical environment 105 can be used to perform a computer-assisted medical procedure on a patient 825. In some embodiments, surgical team can include a surgeon 830 A and additional medical personnel 830B-830D such as a medical assistant, nurse, and anesthesiologist, and other suitable team members who can assist with the surgical procedure or medical session. The medical session can include the surgical procedure being performed on the patient 825, as well as any pre-operative (e.g., which can include setup of the medical environment 105, including preparation of the patient 825 for the procedure), and postoperative (e.g., which can include clean up or post care of the patient), or other processes during the medical session. Although described in the context of a surgical procedure, the medical environment 105 can be implemented in a non-surgical procedure, or other types of medical procedures or diagnostics that can benefit from the accuracy and convenience of the surgical system.

[0125] The robotic medical system 125 can include a plurality of manipulator arms 835 A- 835D to which a plurality of medical tools (e.g., the medical tool 120) can be coupled or installed. Each medical tool can be any suitable surgical tool (e.g., a tool having tissueinteraction functions), imaging device (e.g., an endoscope, an ultrasound tool, etc.), sensing instrument (e.g., a force-sensing surgical instrument), diagnostic instrument, or other suitable instrument that can be used for a computer-assisted surgical procedure on the patient 825 (e.g., by being at least partially inserted into the patient and manipulated to perform a computer-assisted surgical procedure on the patient). Although the robotic medical system 125 is shown as including four manipulator arms (e.g., the manipulator arms 835A-835D), in other embodiments, the robotic medical system can include greater than or fewer than four manipulator arms. Further, not all manipulator arms can have a medical tool installed thereto at all times of the medical session. Moreover, in some embodiments, a medical tool installed on a manipulator arm can be replaced with another medical tool as suitable.

[0126] One or more of the manipulator arms 835A-835D or the medical tools attached to manipulator arms can include one or more displacement transducers, orientational sensors, positional sensors, or other types of sensors and devices to measure parameters or generate kinematics information. One or more components of the medical environment 105 can be configured to use the measured parameters or the kinematics information to track (e.g., determine poses of) or control the medical tools, as well as anything connected to the medical tools or the manipulator arms 835A-835D.

[0127] The user control system 810 can be used by the surgeon 830A to control (e.g., move) one or more of the manipulator arms 835A-835D or the medical tools connected to the manipulator arms. To facilitate control of the manipulator arms 835A-835D and track progression of the medical session, the user control system 810 can include a display (e.g., the display 172) that can provide the surgeon 830A with imagery (e.g., high-definition 3D imagery) of a surgical site associated with the patient 825 as captured by a medical tool (e.g., the medical tool 120, which can be an endoscope) installed to one of the manipulator arms 835A-835D. The user control system 810 can include a stereo viewer having two or more displays where stereoscopic images of a surgical site associated with the patient 825 and generated by a stereoscopic imaging system can be viewed by the surgeon 830 A. In some embodiments, the user control system 810 can also receive images from the auxiliary system 815 and the visualization tool 820.

[0128] The surgeon 830A can use the imagery displayed by the user control system 810 to perform one or more procedures with one or more medical tools attached to the manipulator arms 835A-835D. To facilitate control of the manipulator arms 835A-835D or the medical tools installed thereto, the user control system 810 can include a set of controls. These controls can be manipulated by the surgeon 830 A to control movement of the manipulator arms 83 SA- 835D or the medical tools installed thereto. The controls can be configured to detect a wide variety of hand, wrist, and finger movements by the surgeon 830 A to allow the surgeon tointuitively perform a procedure on the patient 825 using one or more medical tools installed to the manipulator arms 835A-835D.

[0129] The auxiliary system 815 can include one or more computing devices configured to perform processing operations within the medical environment 105. For example, the one or more computing devices can control or coordinate operations performed by various other components (e.g., the robotic medical system 125, the user control system 810) of the medical environment 105. A computing device included in the user control system 810 can transmit instructions to the robotic medical system 125 by way of the one or more computing devices of the auxiliary system 815. The auxiliary system 815 can receive and process image data representative of imagery captured by one or more imaging devices (e.g., medical tools) attached to the robotic medical system 125, as well as other data stream sources received from the visualization tool. For example, one or more image capture devices (e.g., the data capture devices 110) can be located within the medical environment 105. These image capture devices can capture images from various viewpoints within the medical environment 105. These images (e.g., video streams) can be transmitted to the visualization tool 820, which can then passthrough those images to the auxiliary system 815 as a single combined data stream. The auxiliary system 815 can then transmit the single video stream (including any data stream received from the medical tool(s) of the robotic medical system 125) to present on a display (e.g., the display 172) of the user control system 810.

[0130] In some embodiments, the auxiliary system 815 can be configured to present visual content (e.g., the single combined data stream) to other team members (e.g., the medical personnel 830B-830D) who may not have access to the user control system 810. Thus, the auxiliary system 815 can include a display 840 configured to display one or more user interfaces, such as images of the surgical site, information associated with the patient 825 or the surgical procedure, or any other visual content (e.g., the single combined data stream). In some embodiments, display 840 can be a touchscreen display or include other features to allow the medical personnel 830A-830D to interact with the auxiliary system 815.

[0131] The robotic medical system 125, the user control system 810, and the auxiliary system 815 can be communicatively coupled one to another in any suitable manner. For example, in some embodiments, the robotic medical system 125, the user control system 810, and the auxiliary system 815 can be communicatively coupled by way of control lines 845, which can represent any wired or wireless communication link as can serve a particularimplementation. Thus, the robotic medical system 125, the user control system 810, and the auxiliary system 815 can each include one or more wired or wireless communication interfaces, such as one or more local area network interfaces, Wi-Fi network interfaces, cellular interfaces, etc.

[0132] It is to be understood that the medical environment 105 can include other or additional components or elements that can be needed or considered desirable to have for the medical session for which the surgical system is being used.

[0133] FIG. 9 is a block diagram depicting an architecture for a computer system 900 that can be employed to implement elements of the systems and methods described and illustrated herein, including aspects of the systems depicted in FIG. 1 and FIG. 8, the user interface depicted in FIG. 4, and the methods or processes depicted in FIGS. 2, 3, and 5-7. For example, the data processing system 130, robotic medical system 125, or client device 174 can include one or more component or functionality of computing system 900. The computer system 900 can be any computing device used herein and can include or be used to implement a data processing system or its components. The computer system 900 includes at least one bus 905 or other communication component or interface for communicating information between various elements of the computer system. The computer system further includes at least one processor 910 or processing circuit coupled to the bus 905 for processing information. The computer system 900 also includes at least one main memory 915, such as a random-access memory (RAM) or other dynamic storage device, coupled to the bus 905 for storing information, and instructions to be executed by the processor 910. The main memory 915 can be used for storing information during execution of instructions by the processor 910. The computer system 900 can further include at least one read only memory (ROM) 920 or other static storage device coupled to the bus 905 for storing static information and instructions for the processor 910. A storage device 925, such as a solid-state device, magnetic disk or optical disk, can be coupled to the bus 905 to persistently store information and instructions.

[0134] The computer system 900 can be coupled via the bus 905 to a display 930, such as a liquid crystal display, or active-matrix display, for displaying information. An input device 935, such as a keyboard or voice interface can be coupled to the bus 905 for communicating information and commands to the processor 910. The input device 935 can include a touch screen display (e.g., the display 930). The input device 935 can also include a cursor control, such as a mouse, a trackball, or cursor direction keys, for communicating direction informationand command selections to the processor 910 and for controlling cursor movement on the display 930.

[0135] The processes, systems and methods described herein can be implemented by the computer system 900 in response to the processor 910 executing an arrangement of instructions contained in the main memory 915. Such instructions can be read into the main memory 915 from another computer-readable medium, such as the storage device 925. Execution of the arrangement of instructions contained in the main memory 915 causes the computer system 900 to perform the illustrative processes described herein. One or more processors in a multiprocessing arrangement can also be employed to execute the instructions contained in the main memory 915. Hard-wired circuitry can be used in place of or in combination with software instructions together with the systems and methods described herein. Systems and methods described herein are not limited to any specific combination of hardware circuitry and software.

[0136] The processor 910 can execute one or more instructions associated with the system 100. The processor 910 can include an electronic processor, an integrated circuit, or the like including one or more of digital logic, analog logic, digital sensors, analog sensors, communication buses, volatile memory, nonvolatile memory, and the like. The processor 910 can include, but is not limited to, at least one microcontroller unit (MCU), microprocessor unit (MPU), central processing unit (CPU), graphics processing unit (GPU), physics processing unit (PPU), embedded controller (EC), or the like. The processor 910 can include, or be associated with, a memory 915 operable to store or storing one or more non-transitory computer-readable instructions for operating components of the system 100 and operating components operably coupled to the processor 910. The one or more instructions can include at least one of firmware, software, hardware, operating systems, or embedded operating systems, for example. The processor 910 or the system 100 generally can include at least one communication bus controller to effect communication between the system processor and the other elements of the system 100.

[0137] The memory 915 can include one or more hardware memory devices to store binary data, digital data, or the like. The memory 915 can include one or more electrical components, electronic components, programmable electronic components, reprogrammable electronic components, integrated circuits, semiconductor devices, flip flops, arithmetic units, or the like. The memory 915 can include at least one of a non-volatile memory device, a solid-state memory device, a flash memory device, a NAND memory device, a volatile memory device,etc. The memory 915 can include one or more addressable memory regions disposed on one or more physical memory arrays.

[0138] Although an example computing system has been described in FIG. 9, the subject matter including the operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.

[0139] The herein described subject matter sometimes illustrates different components contained within, or connected with, different other components. It is to be understood that such depicted architectures are illustrative, and that in fact many other architectures can be implemented which achieve the same functionality. In a conceptual sense, any arrangement of components to achieve the same functionality is effectively “associated” such that the desired functionality is achieved. Hence, any two components herein combined to achieve a particular functionality can be seen as “associated with” each other such that the desired functionality is achieved, irrespective of architectures or intermedial components. Likewise, any two components so associated can also be viewed as being “operably connected,” or “operably coupled,” to each other to achieve the desired functionality, and any two components capable of being so associated can also be viewed as being “operably couplable,” to each other to achieve the desired functionality. Specific examples of operably couplable include but are not limited to physically mateable or physically interacting components or wirelessly interactable or wirelessly interacting components or logically interacting or logically interactable components.

[0140] With respect to the use of plural or singular terms herein, those having skill in the art can translate from the plural to the singular or from the singular to the plural as is appropriate to the context or application. The various singular / plural permutations can be expressly set forth herein for sake of clarity.

[0141] It will be understood by those within the art that, in general, terms used herein, and especially in the appended claims (e.g., bodies of the appended claims) are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” etc.).

[0142] Although the figures and description can illustrate a specific order of method steps, the order of such steps can differ from what is depicted and described, unless specified differently above. Also, two or more steps can be performed concurrently or with partial concurrence, unless specified differently above. Such variation can depend, for example, on the software and hardware systems chosen and on designer choice. All such variations are within the scope of the disclosure. Likewise, software implementations of the described methods can be accomplished with standard programming techniques with rule-based logic and other logic to accomplish the various connection steps, processing steps, comparison steps, and decision steps.

[0143] It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation, no such intent is present. For example, as an aid to understanding, the following appended claims can contain usage of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles “a” or “an” limits any particular claim containing such introduced claim recitation to inventions containing only one such recitation, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” or “an” should typically be interpreted to mean “at least one” or “one or more”); the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should typically be interpreted to mean at least the recited number (e.g., the bare recitation of “two recitations,” without other modifiers, typically means at least two recitations, or two or more recitations).

[0144] Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together, etc.). In those instances where a convention analogous to “at least one of A, B, or C, etc.” is used, in general, such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, orA, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B.”

[0145] Further, unless otherwise noted, the use of the words “approximate,” “about,” “around,” “substantially,” etc., mean plus or minus ten percent.

[0146] The foregoing description of illustrative implementations has been presented for purposes of illustration and of description. It is not intended to be exhaustive or limiting with respect to the precise form disclosed, and modifications and variations are possible in light of the above teachings or can be acquired from practice of the disclosed implementations. It is intended that the scope of the invention be defined by the claims appended hereto and their equivalents.

Claims

CLAIMSWhat is claimed is:

1. A system, comprising: one or more processors, coupled with memory, to: receive a video stream that captures a procedure over a time interval with a robotic medical system; generate, from the video stream, a plurality of sets of consecutive frames comprising a first set of consecutive frames with a first temporal resolution and a second set of consecutive frames with a second temporal resolution; determine, via the plurality of sets of consecutive frames input into a first model trained with machine learning, phases of the procedure on a moment-to-moment basis over the time interval; input the phases of the procedure determined on the moment-to-moment basis into a second model, trained with machine learning based on historical workflows, to generate at least one phase segment over the time interval; and provide an action based on a metric of the at least one phase segment.

2. The system of claim 1, comprising the one or more processors to: execute, prior to generation of the plurality of sets of consecutive frames, one or more pre-processing functions on the video stream, the one or more pre-processing functions comprising at least one of a central crop transform, frame resizing, filtering of non-surgical frames, or filtering of noisy frames.

3. The system of claims 1 or 2, comprising the one or more processors to: generate the first set of consecutive frames with the first temporal resolution that is greater than or equal to a first threshold; and generate the second set of consecutive frames with the second temporal resolution that is less than or equal to a second threshold that is less than the first threshold.

4. The system of any one of claims 1-3, comprising the one or more processors to: generate a third set of consecutive frames with a varying temporal threshold that varies based at least in part on a function.

5. The system of any one of claims 1-4, wherein the phases of the procedure comprise at least one of exposure, dissection, transection, extraction, or reconstruction.

6. The system of any one of claims 1-5, comprising the one or more processors to: determine the phases of the procedure on the moment-to-moment basis via a the first model configured with a multi-pathway spatial-temporal decoding unit comprising an attention-based deep learning model.

7. The system of any one of claims 1-6, comprising the one or more processors to: input the plurality of sets of consecutive frames into a corresponding plurality of parallel streams to generate corresponding numerical feature outputs; and fuse the numerical feature outputs to generate the phases of the procedure on the moment-to-moment basis.

8. The system of any one of claims 1-7, comprising the one or more processors to: determine a variance in the phases of the procedure determined on the moment-to- moment basis; generate, based on the variance and a phase transition map configured with a plurality of prior probabilities, a plurality of uncertainty-aware phase boundaries throughout the time interval; and generate the at least one phase segment that corresponds to a phase boundary of the plurality of uncertainty-aware phase boundaries.

9. The system of any one of claims 1-8, comprising the one or more processors to: generate the at least one phase segment based at least in part on one or more of an average duration of phases, an ordering of phases, a type of the procedure, or a site location of the procedure.

10. The system of any one of claims 1-9, comprising: the one or more processors to determine the phases of the procedure on the moment-to- moment basis based on a combination of the video stream and at least one of a stream of system events data or a stream of kinematics data.

11. The system of any one of claims 1-10, comprising: the one or more processors to provide the action indicating a level of performance of the procedure during a phase segment of the at least one phase segment determined over the time interval.

12. The system of any one of claims 1-11, comprising the one or more processors to: determine a phase segment of the at least one phase segment at a current time; identify a tool used in the phase segment based on a stream of system events; determine that the tool does not match any of a predetermined set of tools configured for the phase segment; and provide an alert during the phase segment responsive to the determination that the tool does not match any of the predetermined set of tools configured for the phase segment.

13. A non-transitory computer-readable medium storing processor executable instructions that, when executed by one or more processors, cause the one or more processors to: receive a video stream that captures a procedure over a time interval with a robotic medical system; generate, from the video stream, a plurality of sets of consecutive frames comprising a first set of consecutive frames with a first temporal resolution and a second set of consecutive frames with a second temporal resolution; determine, via the plurality of sets of consecutive frames input into a first model trained with machine learning, phases of the procedure on a moment-to-moment basis over the time interval; input the phases of the procedure determined on the moment-to-moment basis into a second model, trained with machine learning based on historical workflows, to generate at least one phase segment over the time interval; and provide an action based on a metric of the at least one phase segment.

14. The non-transitory computer-readable medium of claim 13, wherein the instructions further include instructions to: execute, prior to generation of the plurality of sets of consecutive frames, one or more pre-processing functions on the video stream, the one or more pre-processing functionscomprising at least one of a central crop transform, frame resizing, filtering of non-surgical frames, or filtering of noisy frames.

15. The non-transitory computer-readable medium of claims 13 or 14, wherein the instructions further include instructions to: generate the first set of consecutive frames with the first temporal resolution that is greater than or equal to a first threshold; and generate the second set of consecutive frames with the second temporal resolution that is less than or equal to a second threshold that is less than the first threshold.

16. The non-transitory computer-readable medium of any one of claims 13-15, wherein the instructions further include instructions to: generate a third set of consecutive frames with a varying temporal threshold that varies based at least in part on a function.

17. The non-transitory computer-readable medium of any one of claims 13-16, wherein the phases of the procedure comprise at least one of exposure, dissection, transection, extraction, or reconstruction.

18. The non-transitory computer-readable medium of any one of claims 13-17, wherein the instructions further include instructions to: determine the phases of the procedure on the moment-to-moment basis via a the first model configured with a multi-pathway spatial-temporal decoding unit comprising an attention-based deep learning model.

19. A method, comprising: receiving, by one or more processors coupled with memory, a video stream that captures a procedure over a time interval with a robotic medical system; generating, by the one or more processors from the video stream, a plurality of sets of consecutive frames comprising a first set of consecutive frames with a first temporal resolution and a second set of consecutive frames with a second temporal resolution; determining, by the one or more processors, via the plurality of sets of consecutive frames input into a first model trained with machine learning, phases of the procedure on amoment-to-moment basis over the time interval; inputting, by the one or more processors, the phases of the procedure determined on the moment-to-moment basis into a second model, trained with machine learning based on historical workflows, to generate at least one phase segment over the time interval; and providing, by the one or more processors, an action based on a metric of the at least one phase segment.

20. The method of claim 19, comprising: executing, by the one or more processors, prior to generating the plurality of sets of consecutive frames, one or more pre-processing functions on the video stream, the one or more pre-processing functions comprising at least one of a central crop transform, frame resizing, filtering of non-surgical frames, or filtering of noisy frames.