End to end automated data quality checking
An automated end-to-end quality control system addresses data errors in robotic medical systems by using timestamping and machine learning to reorder data packets and improve synchronization, enhancing the reliability and confidence in surgical performance metrics.
Patent Information
- Application Number
- PCT/US2025/017699
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2025-02-27
- Publication Date
- 2025-09-04
AI Technical Summary
The challenge of maintaining efficient and reliable operation of robotic medical systems in complex medical environments is exacerbated by the difficulty in handling large and diverse data streams, leading to errors such as delays, jitter, and data loss, which affect the quality of surgeon performance analysis.
An automated end-to-end quality control system using timestamping and machine learning to reorder out-of-order data packets and identify system configurations, improving data synchronization and reducing errors in robotic system pipelined data streams.
Enhances the reliability and confidence in data-based surgical performance metrics by correcting data errors and minimizing data loss, thereby improving the overall performance of robotic medical systems.
Smart Images

Figure US2025017699_04092025_PF_FP_ABST
Abstract
Description
END TO END AUTOMATED DATA QUALITY CHECKINGCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 559,130, filed February 28, 2024, which is incorporated herein by reference in its entirety and for all purposes.BACKGROUND
[0002] Medical procedures can be performed in a medical environment, such as an operating room. As the amount and variety of equipment in the operating room increases, or medical procedures become increasingly complex, it can be challenging to maintain the equipment operating efficiently, reliably, or without incident.SUMMARY
[0003] The technical solutions of this disclosure provide an automated end-to-end quality control for various types of data streams generated by robotic medical systems. The data generated by robotic medical systems used in robot-assisted medical procedures can be used to provide metrics on surgeon performance. However, processing various data streams for different system tasks can lead to errors, such as delays, jitter or data loss, adversely affecting the quality of data used for the surgeon performance analysis. Moreover, handling of large data streams from diverse sources can be difficult, time consuming as well as energy and compute intensive. The disclosed solutions overcome such challenges by providing automated quality checking for robotic system pipelined data streams, such as data streams of kinematics data, events data and sensor (e.g., video) data. Using timestamping and machine learning, these solutions can reduce data errors, improve synchronization and confidence in data streams used to compute performance metrics and identify system configurations to improve system reliability and resource and energy consumption.
[0004] At least one aspect of the technical solutions relates to a system. The system can include one or more processors, coupled with memory. The one or more processors can receive, for a medical session with a robotic medical system, a first one or more data packets of a first stream of data at a first time. The one or more processors can receive, for the medical session, a second one or more data packets of a second stream of data at a second time subsequent to the first time. The one or more processors can identify a first feature in the firstone or more data packets and a second feature in the second one or more data packets. The one or more processors can detect, based at least on the first feature and the second feature, that the first one or more data packets and the second one or more data packets are out of order. The one or more processors can reorder, responsive to the detection, the first one or more data packets and the second one or more data packets to cause the first one or more data packets to be subsequent to the second one or more data packets. The one or more processors can determine, using the reordered first one or more data packets and second one or more data packets, a metric indicative of performance of the medical session. The one or more processors can generate, using the reordered first one or more data packets and second one or more data packets, a composite video of at least a portion of the medical session with an indication of the metric.
[0005] The one or more processors can identify that the first stream corresponds to a first of a data generated by a sensor of the robotic medical system, a data on kinematics of one or more instruments of the robotic medical system used during the medical session, and a data of an event at the robotic medical system. The one or more processors can identify that the second stream corresponds to a second one of the data generated by the sensor, the data on kinematics of the one or more instruments and the data on the event at the robotic medical system.
[0006] The one or more processors can detect, based at least on one of the first feature and the second feature, at least one of a delay or a jitter corresponding to at least one of the first stream and the second stream. The one or more processors can determine, responsive to the detection of the at least one of the delay or the jitter, a portion the robotic medical system causing the at least one of the delay or the jitter.
[0007] The one or more processors can identify the first feature comprising a first timestamp for the first one or more data packets and the second feature comprising a second timestamp for the second one or more data packets. The one or more processors can identify, based at least one the first timestamp and the second timestamp, a time at which at least one of the first one or more data packets and the second one or more data packets are at least one of generated, transmitted, received or stored by at least a portion of the robotic medical system. The one or more processors can detect, based at least on the time, that the first one or more data packets and the second one or more data packets are out of order.
[0008] The one or more processors can identify, based at least on the first feature, the first time and the second time, that a likelihood of the second feature corresponding to an event ofthe robotic medical system exceeds a threshold. The one or more processors can detect, based at least on the likelihood exceeding the threshold, that the first one or more data packets and the second one or more data packets are out of order.
[0009] At least one of the first feature and the second feature can indicate at least one of an event, a task of a medical procedure, a phase of a medical procedure, an object used in the medical procedure or a workflow of the medical procedure. The one or more processors can identify, based at least on the reordered first one or more data packets and second one or more data packets, a data missing from at least one of the first stream or the second stream. The one or more processors can determine, based at least on the identified data, a confidence score for the metric. The one or more processors can provide the confidence score for display with the composite video.
[0010] The one or more processors can identify a machine learning (ML) model trained on a plurality of features of a plurality of data packets of the robotic medical system. The one or more processors can detect that the first one or more data packets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the ML model.
[0011] The one or more processors can identify one or more machine learning (ML) models utilizing one or more neural networks trained on sensor data from a plurality of streams of sensors of a robotic medical system, events data corresponding to a plurality of events of a plurality of medical procedures and kinematics data corresponding to data on kinematics relating the plurality of medical procedures implemented on the robotic medical system. The one or more processors can detect that the first one or more data packets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the one or more ML models.
[0012] The one or more processors can identify a first ML model trained on the sensor data and corresponding to a first encoder and a first one or more classification blocks to process representations on the sensor data. The one or more processors can identify a second ML model trained on the events data and corresponding to a second encoder and a second one or more classification blocks to process representations on the events data. The one or more processors can identify a third ML model trained on the kinematics data and corresponding to a third encoder and a third one or more classification blocks to process representations on the kinematics sensor data. The one or more processors can detect that the first one or more datapackets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the first ML model, the second ML model and the third ML model.
[0013] The one or more processors can determine, using the reordered first one or more data packets and second one or more data packets, that a portion of at least one of the first stream and the second stream is missing. The one or more processors can generate an alert to indicate that the portion of at least the one of the first stream and the second stream is missing. The one or more processors can overlay the alert on the composite video. The one or more processors can determine, responsive to the detection, that a portion of at least one of the first stream over a time interval and the second stream over the time interval is missing. The one or more processors can modify, prior to display of the composite video, the composite video to exclude a portion of the composite video corresponding to the time interval.
[0014] The one or more processors can display, via a graphical user interface, the composite video generated using the reordered first one or more data packets and the second one or more data packets. The one or more processors can store, in a data repository, a data file comprising the composite video generated using the reordered first one or more data packets and the second one or more data packets.
[0015] At least one aspect of the technical solutions relates to a method. The method can include one or more processors coupled with memory receiving a first one or more data packets of a first stream of data at a first time. The method can include the one or more processors receiving a second one or more data packets of a second stream of data at a second time subsequent to the first time. The first stream and the second stream can correspond to a medical session implemented using a robotic medical system. The method can include the one or more processors detecting, based at least on a first feature of the first one or more data packets and a second feature of the second one or more data packets, that the first one or more data packets and the second one or more data packets are out of order. The method can include the one or more processors reordering, responsive to the detection, the first one or more data packets and the second one or more data packets to cause the first one or more data packets to be subsequent to the second one or more data packets. The method can include the one or more processors determining, using the reordered first one or more data packets and second one or more data packets, a metric indicative of performance of the medical session. The method can include generating, by the one or more processors using the reordered first one or more datapackets and second one or more data packets, a composite video of at least a portion of the medical session with an indication of the metric.
[0016] The method can include the one or more processors identifying that the first stream corresponds to a first of a data generated by a sensor of the robotic medical system, a data on kinematics of one or more instruments of the robotic medical system used during the medical session, and a data of an event at the robotic medical system. The method can include identifying, by the one or more processors, that the second stream corresponds to a second one of the data generated by the sensor, the data on kinematics of the one or more instruments and the data on the event at the robotic medical system.
[0017] The method can include detecting, by the one or more processors based at least on one of the first feature and the second feature, at least one of a delay or a jitter corresponding to at least one of the first stream and the second stream. The method can include determining, by the one or more processors responsive to the detection of the at least one of the delay or the jitter, a portion the robotic medical system causing the at least one of the delay or the jitter.
[0018] The method can include identifying, by the one or more processors, the first feature comprising a first timestamp for the first one or more data packets and the second feature comprising a second timestamp for the second one or more data packets. The method can include identifying, by the one or more processors, based at least one the first timestamp and the second timestamp, a time at which at least one of the first one or more data packets and the second one or more data packets are at least one of generated, transmitted, received or stored by at least a portion of the robotic medical system. The method can include detecting, by the one or more processors, based at least on the time, that the first one or more data packets and the second one or more data packets are out of order.
[0019] The method can include the one or more processors identifying, based at least on the first feature, the first time and the second time, that a likelihood of the second feature corresponding to an event of the robotic medical system exceeds a threshold. The method can include detecting, by the one or more processors, based at least on the likelihood exceeding the threshold, that the first one or more data packets and the second one or more data packets are out of order. At least one of the first feature and the second feature can indicate at least one of: an event, a task of a medical procedure, a phase of a medical procedure, an object used in the medical procedure or a workflow of the medical procedure.
[0020] The method can include identifying, by the one or more processors, a machine learning (ML) model trained on a plurality of features of a plurality of data packets of the robotic medical system. The method can include detecting, by the one or more processors, that the first one or more data packets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the ML model. The method can include identifying, by the one or more processors, one or more machine learning (ML) models utilizing one or more neural networks trained on sensor data from a plurality of streams of sensors of a robotic medical system, events data corresponding to a plurality of events of a plurality of medical procedures and kinematics data corresponding to data on kinematics relating the plurality of medical procedures implemented on the robotic medical system. The method can include detecting, by the one or more processors, that the first one or more data packets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the one or more ML models.
[0021] The method can include determining, by the one or more processors, responsive to the detection, that a portion of at least one of the first stream over a time interval and the second stream over the time interval is missing. The method can include modifying, by the one or more processors, prior to display of the composite video, the composite video to exclude a portion of the composite video corresponding to the time interval.
[0022] At least an aspect of the technical solutions relates to a non-transitory computer- readable medium storing processor-executable instructions. The instructions, when executed by one or more processors, can cause the one or more processors to receive, for a medical session with a robotic medical system, a first one or more data packets of a first stream of data at a first time and a second one or more data packets of a second stream of data at a second time subsequent to the first time. The instructions, when executed, can cause the one or more processors to identify one or more machine learning (ML) models including one or more neural networks trained on sensor data from a plurality of streams of sensors of a robotic medical system, events data corresponding to a plurality of events of a plurality of medical procedures and kinematics data corresponding to data on kinematics relating the plurality of medical procedures implemented on the robotic medical system. The instructions, when executed, can cause the one or more processors to detect, based at least on a first feature in the one or more data packets and a second feature in the second one or more data packets input into the one or more ML models, that the first one or more data packets and the second one or more data packets are out of order. The instructions, when executed, can cause the one or more processorsto reorder, responsive to the detection, the first one or more data packets and the second one or more data packets to cause the first one or more data packets to be subsequent to the second one or more data packets. The instructions, when executed, can cause the one or more processors to determine, using the reordered first one or more data packets and second one or more data packets, a metric indicative of performance of the medical session. The instructions, when executed, can cause the one or more processors to generate, using the reordered first one or more data packets and second one or more data packets, a composite video of at least a portion of the medical session with an indication of the metric.BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings are not intended to be drawn to scale. Like reference numbers and designations in the various drawings indicate like elements. For purposes of clarity, not every component can be labeled in every drawing. In the drawings:
[0024] FIG. 1 depicts an example system to provide an end-to-end data quality check and control in robotic systems.
[0025] FIG. 2 illustrates an example block diagram of a system in which ML models trained for different types of data streams detect features indicative of errors in the data.
[0026] FIG. 3 illustrates an example block diagram of an example computer system.
[0027] FIG. 4 illustrates an example flow diagram of a method for providing an end-to-end automated data quality checking in robotic systems is illustrated.
[0028] FIG. 5 illustrates an example of a surgical system, in accordance with some aspects of the technical solutions.DETAILED DESCRIPTION
[0029] Following below are more detailed descriptions of various concepts related to, and implementations of, systems, methods, apparatuses for automated end-to-end data quality checking in robotic systems. The various concepts introduced above and discussed in greater detail below can be implemented in any of numerous ways.
[0030] Although the present disclosure is discussed in the context of a surgical procedure, in various aspects, the technical solutions of this disclosure can be applicable to other medical treatments, sessions, environments or activities, as well as non-medical activities where dataquality monitoring and checking is desired. For instance, technical solutions can be applied in any environment, application or industry in which data quality monitoring and correction can be used to improve the reliability and quality of machine learning based robotic system performance analysis.
[0031] Large scale analysis of data streams generated by robotic systems used in robot- assisted medical surgeries can be used to generate and provide metrics relevant to surgeon performance and efficiency. However, as this analysis can be based on various types of data streams whose data packets can experience different types of errors (e.g., delays jitter or packet loss), the quality of the data used for performance analysis can be impacted, reducing the reliability and the confidence in the performance analysis and the metrics generated. Given that a single robot-assisted medical procedure can result in a large amount of generated data (e.g., several gigabytes of data streams), manually troubleshooting and cleaning this data can be challenging as well as time and compute resource and energy intensive. In addition, the rate at which the data is generated can make it difficult to take timely corrective action to reduce the data corruption, all of which can impact the ability of the system to reliably analyze the surgeon’s performance and determine the performance metrics. The inability to generate accurate, reliable performance metrics may impact the performance of a medical procedure performed using the robotic medical system.
[0032] The technical solutions overcome these challenges by providing systems and methods for automated quality checking and error correction of pipelined robotic system data streams. The technical solutions can monitor and analyze streams of kinematics data, robotic system events data and sensor (e.g., video) data using data packet timestamping and machine learning to identify relations between different various data stream features in the context of specific robotic procedures. In doing so, the technical solutions can detect and correct various data errors, reducing the instances of out of order data packets or delays and jitter between different data streams. The technical solutions can identify specific robotic system configurations and operating room setups to operate the robotic system to improve data synchronization and minimize data loss. These solutions can improve the ability of engineers to debug integrated data pipelines and identify causes of data errors across the system. In doing so, the technical solutions can improve the reliability and confidence in data-based surgical performance metrics.
[0033] FIG. 1 depicts an example system 100 to provide an end-to-end data quality check and control in robotic systems, such as robotic medical systems used in robot-assistedsurgeries. Example system 100 can include a combination of hardware and software for providing autonomous and automated quality data analysis across different types of data streams of a robotic medical system. Example system 100 can provide quality control analysis on data packets across several types of robotic system data streams, including sensors data streams (e.g., video data), kinematics data streams (e.g., robotic arm movement data) and events data streams (e.g., installation, selection and configuration data) as well as provide issue detection, correction, mitigation and indication functionalities to address identified issues.
[0034] Example system 100 can include a medical environment 102 including one or more data capture devices 110, medical instruments 112, visualization tools 114, displays 116 and robotic medical systems (RMSs) 120. RMS 120 can include or generate various types of data (e.g., kinematics data 172, sensor data 174, events data 176, state data 178 or mapping data 179) any of which can be processed or transmitted individually or combined into data streams 162. The RMS 120 can utilize the generated data along with system configurations 122 to perform RMS operations (e.g., surgical tasks of a medical procedure). One or more RMSs 120 can be communicatively coupled (e.g., via one or more networks 101) with one or more data processing systems (DPSs) 130.
[0035] DPS 130 can include one or more machine learning (ML) trainers 132 having one or more training datasets 134 to train and generate one or more ML models 140. Training database 134 can include any number of data streams 162 corresponding to any number of medical sessions or procedures performed by any number of RMSs 120. ML models 140 can include various types of machine learning models for processing various types of data, including one or more kinematics data ML models (KDML) models 142, sensors data ML models (SDML) models 144 and events data ML (EDML) models 146, mapping data ML models 148 and any other types of ML models 140 that can be trained or configured for analyzing data from medical procedures or any surgical performance based on data from data streams 162. ML models 140 can include one or more encoders 150, classifiers 152, loss functions 154 and weights 156 that can be utilized for the operation of ML models 140. DPS 130 can include one or more depth mappers 158 for generating one or more depth maps providing depth representation (e.g., mapping data 179) of a portion of a patient’s body being imaged. DPS 130 can include one or more data repositories 160, processing functions 180 and performance analyzers 190. Data repository 160 can include data streams 162 that can include or correspond to data packets 164, features 166, timestamps 168, kinematics data 172, sensor data 174 and events data 176. Processing functions 180 can include one or more detectionfunctions 182, corrective functions 184, mitigating functions 186 and indication functions 188. Performance analyzer 190 can include or generate composite videos 196 for the analyzed medical sessions or procedure, as well as one or more metrics 192 and scores 194, such as confidence scores 194 for the metrics 192 as determined using data streams 162 quality controlled by the DPS 130.
[0036] The data processing system 130 can be configured to utilize machine learning across various different data modalities to provide end-to-end automated data quality checking. For example, data processing system 130 can monitor data from any combination of kinematics data 172, sensor or video data 174, events data 176, state data 178 or mapping data 179, which can be either retrieved from a data repository 160 or received in real-time from a robotic medical system 120. Responsive to monitoring, the data processing system 130 can detect, using detection functions 182, that some of the data packets from various data streams 162 are out of their chronological order. The data processing system 130 can utilize corrective functions 184 or mitigating functions 186 to take actions and rearrange the data packets that were out of order into their respective correct order, or take mitigating actions to, such as providing indications or warnings of the data errors detected. In doing so, the data processing system 130 can improve or maintain the data integrity and quality end-to-end, across the process.
[0037] For instance, a data processing system 130 can identify or receive a first set of data packets of a first data stream 162 (e.g., any data of any type such as kinematics, sensor, events, state or mapping data). The first set of data can be received, marked or associated with a first time (e.g., a first timestamp 168) and be associated with a medical session with a robotic medical system (e.g., a particular surgery performed via an RMS 120). The data processing system 130 can also identify or receive a second set of data packets of a second data stream 162 (e.g., a different data of the same or a different type, gathered from any one or more of: kinematics, sensor, events, state or mapping data). The second set of data can be received, marked or associated with a second time (e.g., a second timestamp 168) and be associated with the same medical sessions with the same robotic medical system (e.g., the same particular surgery via the RMS 120).
[0038] The data processing system 130 can process the first and the second data sets and identify a first feature 166 in the first one or more data packets (e.g., of the first data set) and a second feature 166 in the second one or more data packets (e.g., of the second data set). For instance, the data processing system 130 can utilize on or more ML models 140 trained toidentify the features 166. The data processing system 130 can utilize a detection function 182, along with one or more ML models 140, to detect, based at least on the first feature 166 and the second feature 166, that the first one or more data packets and the second one or more data packets are out of order. The data processing system 130 can utilize a corrective function 184 to reorder the data packets that are out of order into their correct order. For instance, the corrective function 184 can utilize one or more ML models 140 to reorder, responsive to the detection, the first one or more data packets and the second one or more data packets to cause the first one or more data packets to be subsequent to the second one or more data packets. The data processing system 130 can utilize the reordered data for any range of processes. For instance, a performance analyzer 190 can be used (e.g., via one or more ML models 140) to determine, using the reordered first one or more data packets and second one or more data packets, a metric 192 indicative of performance of the medical session. The data processing system 130 can utilize the processing functions 180 (e.g., corrective functions 184, mitigating functions 186 or indication functions 188) to generate, using the reordered first one or more data packets and second one or more data packets, a composite video (e.g., updated sensor data 174 of a subtitled video) of at least a portion of the medical session with an indication of the metric 192.
[0039] Robotic medical system 120, also referred to as an RMS 120, can be deployed in any medical environment 102. Medical environment 102 can include any space or facility for performing medical procedures, such as a surgical facility, or an operating room. Medical environment 102 can include medical instruments 112 (e.g., surgical tools used for specialized tasks) that the RMS 120 can use for performing operational procedures, such as surgical patient procedures, whether invasive, non-invasive, or any in-patient or out-patient procedures. Robotic medical system 120 can be centralized or distributed across a plurality of computing devices or systems, such as computing devices 300 (e.g., used on servers, network devices or cloud computing products) to implement various functionalities of the RMS 120, including communicating or processing data streams 162 across various devices via the network 101.
[0040] The medical environment 102 can include one or more data capture devices 110 (e.g., optical devices, such as cameras or sensors or other types of sensors or detectors) for capturing data streams 162. Data streams 162 can include any sensor data 174, such as images or videos of a surgery, kinematics data 172 on any movement of medical instruments 112, or any events data 176, such as installation, configuration or selection events corresponding to medical instruments 112. Data streams 162 can include one or more video streams, such as forexample, a first stream of video frames from a first camera or imaging device external to a patient’s body and a second stream of video frames from a second camera attached to an endoscopic medical instrument 112 (e.g., a stereo camera of an endoscope) that can be inserted within the patient’s body.
[0041] The medical environment 102 can include one or more visualization tools 114 to gather the captured data streams 162 and process it for display to the user (e.g., a surgeon, a medical professional or an engineer or a technician configuring RMS) via one or more (e.g., touchscreen) displays 116. A display 116 can present data stream 162 (e.g., images or video frames) of a medical procedure (e.g., surgery) being performed using the robotic medical system 120 handling, manipulating, holding or otherwise utilizing medical instruments (e.g., also referred to as medical tools 112) to perform surgical tasks at the surgical site. RMS 120 can include system configurations 122 based on which RMS 120 can operate, and the functionality of which can impact the data flow of the data streams 162.
[0042] Data processing system 130 can include any combination of hardware and software for providing quality control of data of a robotic medical system 120. DPS 130 can include any computing device (e.g., computing device 300) and can include one or more servers, virtual machines or can be part of or include a cloud computing environment. The data processing system 130 can be provided via a centralized computing device (e.g., 300) or be provided via distributed computing components, such as including multiple, logically grouped servers and facilitating distributed computing techniques. The logical group of servers may be referred to as a data center, server farm or a machine farm. The servers, which can include virtual machines, can also be geographically dispersed. A data center or machine farm may be administered as a single entity, or the machine farm can include a plurality of machine farms. The servers within each machine farm can be heterogeneous - one or more of the servers or machines can operate according to one or more type of operating system platform.
[0043] The data processing system 130, or components thereof can include a physical or virtual computer system operatively coupled, or associated with, the medical environment 102. The data processing system 130, or components thereof can be coupled, or associated with, the medical environment 102 via a network 101, either directly or directly through an intermediate computing device or system. The network 101 can be any type or form of network. The geographical scope of the network can vary widely and can include a body area network (BAN), a personal area network (PAN), a local-area network (LAN) (e.g., Intranet), a metropolitan area network (MAN), a wide area network (WAN), or the Internet. The topologyof the network 101 can assume any form such as point-to-point, bus, star, ring, mesh, tree, etc. The network 101 can utilize different techniques and layers or stacks of protocols, including, for example, the Ethernet protocol, the internet protocol suite (TCP / IP), the ATM (Asynchronous Transfer Mode) technique, the SONET (Synchronous Optical Networking) protocol, the SDH (Synchronous Digital Hierarchy) protocol, etc. The TCP / IP internet protocol suite can include application layer, transport layer, internet layer (including, e.g., IPv6), or the link layer. The network 101 can be a type of a broadcast network, a telecommunications network, a data communication network, a computer network, a Bluetooth network, or other types of wired and wireless networks.
[0044] The data processing system 130, or components thereof, can be located at least partially at the location of the surgical facility associated with the medical environment 102 or remotely therefrom. Elements of the data processing system 130, or components thereof can be accessible via portable devices such as laptops, mobile devices, wearable smart devices, etc. The data processing system 130, or components thereof, can include other or additional elements that can be considered desirable to have in performing the functions described herein. The data processing system 130, or components thereof, can include, or be associated with, one or more components or functionality of a computing including, for example, one or more processors coupled with memory that can store instructions, data or commands for implementing the functionalities of the DPS 130 discussed herein.
[0045] Data repository 160 of the DPS 130 can include one or more different types of data that can be arranged in various data streams 162. Data repository 160 can include or be implemented in a storage device, such as a storage device 325. Data streams 162 can include any series of data packets 164 of a particular type or a form, which can be generated by a particular device (e.g., a sensor, a camera, a robotic device, a detector or an event detection function). Data stream 162 can include one or more types of data (e.g., only kinematics data 172, sensor data 174, events data 176, state data 178 or mapping data 179), including a combination of any number of types of data. For instance, data stream 162 can include a stream of data packets 164 corresponding to a sensor data 174, such as packets of a video frame, video image of a video captured by a video camera or an endoscopic device, as well as measurements or sensors (e.g., force, torque or biometric data, haptic feedback data, endoscopic images or data, ultrasound images or videos). Data stream 162 can include a stream of data packets 164 corresponding to kinematics data 172, including any data indicative of temporal positional coordinates of a device, or indicative of movement of a medical instrument 112 on an RMS120. Data stream 162 can include a stream of data packets 164 corresponding to a stream of events, such as events indicative of, or corresponding to, installation, uninstallation, engagement or disengagement, setting or unsetting of any medical instrument 112 on an RMS 120. Data stream 162 can include state data 178, such as data on a state of any medical instrument 112, including states such as: coupled or engaged with an RMS 120, uncoupled or disengaged from the RMS 120, turned on or turned off, calibrated or uncalibrated or any other state or setting of a medical instrument 112. Data stream 162 can include mapping data 179 including any data pertaining to a map of distance between a camera and points in a scene imaged by an image or a video frame, thereby providing a spatial representation of a portion of a patient’s body (e.g., a three-dimensional map of a cavity in which a medical procedure is performed).
[0046] The system 100 can include one or more data capture devices 110 (e.g., video cameras, sensors or detectors) for collecting any data stream 162, that can be used for machine learning, including detection of objects from sensor data 174 (e.g., video frames or force or feedback data), detection of particular events (e.g., user interface selection of, or a surgeon’s engaging of, a medical instrument 112) or detection of kinematics (e.g., movements of the medical instrument 112). Data capture devices 110 can include cameras or other image capture devices for capturing videos or images from a particular viewpoint within the medical environment 102. The data capture devices 110 can be positioned, mounted, or otherwise located to capture content from any viewpoint that facilitates the data processing system capturing various surgical tasks or actions.
[0047] Data capture devices 110 can include any of a variety of detectors, sensors, cameras, video imaging devices, infrared imaging devices, visible light imaging devices, intensity imaging devices (e.g., black, color, grayscale imaging devices, etc.), depth imaging devices (e.g., stereoscopic imaging devices, time-of-flight imaging devices, etc.), medical imaging devices such as endoscopic imaging devices, ultrasound imaging devices, etc., non-visible light imaging devices, any combination or sub-combination of the above mentioned imaging devices, or any other type of imaging devices that can be suitable for the purposes described herein. Data capture devices 110 can include cameras that a surgeon can use to perform a surgery and observe manipulation components within a purview of field of view suitable for the given task performance. Data capture devices can output any type of data streams 162, including data streams 162 of kinematics data 172 (e.g., kinematics data stream), data streams 162 of events data 176 (e.g., events data stream), data streams 162 of sensor data 174 (e.g.,sensors data stream), state data 178 (e.g., data or stream of data indicative of a state of an instrument 112), or mapping data 179 (e.g., depth map representing metric distance from a camera to a point in a scene or an video frame).
[0048] For instance, data capture devices 110 can capture, detect, or acquire sensor data 174, such as videos or images, including for example, still images, video images, vector images, bitmap images, other types of images, or combinations thereof. The data capture devices 110 can capture the images at any suitable predetermined capture rate or frequency. Settings, such as zoom settings or resolution, of each of the data capture devices 110 can vary as desired to capture suitable images from any viewpoint. For instance, data capture devices 110 can have fixed viewpoints, locations, positions, or orientations. The data capture devices 110 can be portable, or otherwise configured to change orientation or telescope in various directions. The data capture devices 110 can be part of a multi-sensor architecture including multiple sensors, with each sensor being configured to detect, measure, or otherwise capture a particular parameter (e.g., sound, images, or pressure).
[0049] Data capture devices 110 can generate sensor data 174 from any type and form of a sensor, such as a positioning sensor, a biometric sensor, a velocity sensor, an acceleration sensor, a vibration sensor, a motion sensor, a pressure sensor, a light sensor, a distance sensor, a current sensor, a focus sensor, a temperature or pressure sensor or any other type and form of sensor used for providing data on medical tools 112, or data capture devices (e.g., optical devices). For example, a data capture device 110 can include a location sensor, a distance sensor or a positioning sensor providing coordinate locations of a medical tool 112 (e.g., kinematics data 172). Data capture device 110 can include a sensor providing information or data on a location, position or spatial orientation of an object (e.g., medical tool 112 or a lens of data capture device 110) with respect to a reference point for kinematics data 172. The reference point can include any fixed, defined location used as the starting point for measuring distances and positions in a specific direction, serving as the origin from which all other points or locations can be determined.
[0050] Display 116 can show, illustrate or play data stream 162, such as a video stream, in which medical tools 112 at or near surgical sites are shown. For example, display 116 can display a rectangular image of a surgical site along with at least a portion of medical tools 112 (e.g., instruments) being used to perform surgical tasks. Display 116 can provide compiled or composite images generated by the visualization tool 114 from a plurality of data capture devices 110 to provide visual feedback from one or more points of view.
[0051] The visualization tool 114 that can be configured or designed to receive any number of different data streams 162 from any number of data capture devices 110 and combine them into a single data stream displayed on a display 116. The visualization tool 114 can be configured to receive a plurality of data stream components and combine the plurality of data stream components into a single data stream 162. For instance, the visualization tool 114 can receive a visual sensor data from one or more medical tools 112, sensors or cameras with respect to a surgical site or an area in which a surgery is performed. The visualization tool 114 can incorporate, combine or utilize multiple types of data (e.g., positioning data of a medical tool 112 along sensor readings of pressure, temperature, vibration or any other data) to generate an output to present on a display 116. Visualization tool 114 can present locations of medical tools 112 along with locations of any reference points or surgical sites, including locations of anatomical parts of the patient (e.g., organs, glands or bones).
[0052] Medical tools 112 can be any type and form of tool or instrument used for surgery, medical procedures or a tool in an operating room or environment. Medical tool 112 can be imaged by, associated with or include an image capture device. For instance, a medical tool 112 can be a tool for making incisions, a tool for suturing a wound, an endoscope for visualizing organs or tissues, an imaging device, a needle and a thread for stitching a wound, a surgical scalpel, forceps, scissors, retractors, graspers, or any other tool or instrument to be used during a surgery. Medical tools 112 can include hemostats, trocars, surgical drills, suction devices or any instruments for use during a surgery. The medical tool 112 can include other or additional types of therapeutic or diagnostic medical imaging implements. The medical tool 112 can be configured to be installed in, coupled with, or manipulated by an RMS 120, such as by manipulator arms or other components for holding, using and manipulating the medical instruments or tools 112.
[0053] RMS 120 can be a computer-assisted system configured to perform a surgical or medical procedure or activity on a patient via or using or with the assistance of one or more robotic components or medical tools 112. RMS 120 can include any number of manipulator arms for grasping, holding or manipulating various medical tools 112 and performing computer-assisted medical tasks using medical tools 112 controlled by the manipulator arms.
[0054] Data streams 162 can be generated by the RMS 120. For instance, sensor data 174 can include images (e.g., video images) captured by a medical tool 112 can be sent to the visualization tool 114. For instance, a touchscreen display 116 can be used by a surgeon to select, engage or configure a particular medical instrument 112, thereby triggering an event thatcan be indicated or included in data packets 164 of a data stream of events data 176. RMS 120 can include one or more input ports to receive direct or indirect connection of one or more auxiliary devices. For example, the visualization tool 114 can be connected to the RMS 120 to receive the images from the medical tool when the medical tool is installed in the RMS 120 (e.g., on a manipulator arm for handing medical instruments 112). For example, data stream 162 can include data indicative of positioning and movement of medical instruments 112 that can be captured or identified by data packets 164 of a kinematics data 172, as well as state data178 indicative of a state of the medical instruments 112 or mapping data 179 indicative of a depth map or depth representation of portion of a patient’s body at which the medical procedure is performed. The visualization tool 114 can combine the data stream components from the data capture devices 110 and the medical tool 112 into a single combined data stream 162 including multiple streams for kinematics data 172, sensor data 174, events data 176, state data 178 or mapping data 179, any of which can be indicated or presented on a display 116.
[0055] Data packet 164 can include a unit of data in a data stream 162. Data packet 164 can include the actual information being sent and metadata, such as a source and a destination address, a port identifier or any other information for transmitting data. Data packets 164 can include a data (e.g., a payload) corresponding to an event (e.g., installation, uninstallation, engagement or setup of a medical instrument 112). Data packet 164 can include a data corresponding to a sensor information (e.g., a video frame captured by a camera), or a data on movement of a medical instrument 112. Data packets 164 can be transmitted in data streams 162 that can be separated or combined. For instance, a data stream 162 for kinematics data 172 (e.g., a kinematics data stream), state data 178 or mapping data 179 can include a plurality of data packets 164. The data packets 164 can be indicative of movement of robotic system components or features, state of medical instruments 112 being utilized or depth mapping data179 indicating depth (e.g., distance from a camera to any point) with respect to different tissues or organs being imaged. Data stream 162 for sensor data 174 (e.g., a sensor or video data stream) can include a plurality of data packets 164 of video frames of a video of a medical procedure being performed and analyzed by the analyzer 190 to generate metrics 192 and scores 194. Data stream 162 for events data 176 (e.g., an events data stream) can include a plurality of data packets 164 indicative of events on an RMS 120, such as selection, calibration or configuration of a medical instrument 112, an action on a graphical user interface of an RMS 120 (e.g., a surgeon performing a selection on a menu of a touchscreen display 116).
[0056] Data packets 164 can include one or more timestamps 168, which can indicate a particular time when particular events took place. Timestamps 168 can be included as metadata or associated with any type of data, including kinematics data 172, sensor or video data 174, events data 176, state data 178 or mapping data 179, Timestamps 168 can include time indications expressed in any combination of nanoseconds, microseconds, milliseconds, seconds, hours, days, months or years. Timestamps 168 can also be used in CPU clock counters that can be synchronized across the robotic medical system. For example, timestamps 168 can have multiple different formats and sources which can be used for synchronization of various data packets 164 at various stages of the data path. Timestamps 168 can be included in the payload or metadata of data packets 164 and can indicate the time when a data packet 164 was generated, the time when the data packet 164 was transmitted from the device that generated the data packet 164, the time when the data packet 164 was received by another device (e.g., a system within the RMS 120, or another device on a network) or a time when the data packet 164 is stored into a data repository 160.
[0057] Data packets 164 can include, form or indicate one or more features 166 that can be utilized to identify issues with the data in one or more data streams 162. Feature 166 can include any information indicative or used to identify, detect or infer an issue with data in one or more data streams 162. Feature 166 can include information about a location, position, movement, action, state or status of a particular medical instrument 112, data capture device 110 or an action of a surgeon. Feature 166 can correspond to artifacts, such as video data artifacts or distortions detected during processing of different data of various data streams 162. Feature 166 can include data indicative of a state of medical procedure or surgery, including a task or a portion of a task being implemented during a procedure, such as a location or a movement of a scalpel, selection by a surgeon or a measurement by a sensor at a particular time, which can be indicated by one or more timestamps 168.
[0058] Feature 166 can include or be associated with a timestamp 168. Feature 166 can include a payload indicative of an event, a sensor measurement or a kinematics action (e.g., a motion). Feature 166 can include a combination of a timestamp 168 and data, such an event data 176 in which a surgeon used a graphical user interface (GUI) to select or click on a medical instrument 112 to use. Feature 166 can be indicated by kinematics data 172 (e.g., selection of timestamped coordinates) indicative of a movement of a medical instrument 112. This can occur at a particular time (e.g., a single data packet 164 or single timestamp 168 or over a period of time or over a plurality of timestamps 168. Feature 166 can be indicated bysensor data (e.g., one or more video frames captured by a camera) which can include one or more timestamps 168. Feature 166 can be indicated by events data 176 indicating an event (e.g., engagement or disengagement of a particular medical instrument 112 with respect to the RMS 120). Feature 166 can be indicated by state data 178 indicating a state of a particular medical instrument 112 (e.g., used or unused by the RMS 120). Feature 166 can be indicated by mapping data 179 which can show different depth of various portions of the surgical location or portion of the patient’s body at which the medical procedure is performed. Feature 166 can be defined or included within a single data packet 164 or can be indicated by a plurality (e.g., a series) of data packets 164. Features 166 can indicate events, movements, actions or occurrences that happen or occur in a particular chronological order (e.g., a scalpel has to touch a tissue before the tissue is cut), which can be used by the detection functions 182 to detect if some of the data packets 164 indicative of particular features 166 are in order or out of their respective order. Features 166 from various data packets 164 from various data streams 162 can be interrelated, coordinated and corresponding to one another, which can be used by the DPS 130 to identify time offsets (e.g., delays or jitters) between various data streams 162.
[0059] State data 178 can be any data that represents the status or condition of medical instruments and components within a robotic medical system 120. State data 178 may include data that is not presented in a stream, such as individual data measurements indicative of a periodic updated state of instrument reading. For instance, medical instruments 112 can be tracked through state (e.g., a first instrument 112 is currently in a first arm of an RMS 120) instead of an event stream which can indicate that the same instrument was installed at a first time and uninstalled at a second time. State data 178 can include information on whether instruments are engaged, calibrated, or in use during a medical session. State data 178 can be used to monitor states of various medical instruments 112 throughout a medical procedure and help identify discrepancies, such as out of order data that may occur during a processing of the data streams 162. For instance, state data 178 can provide information on an instrument 112 to allow for detection of data being out of order, such as when events data 176 or sensor data 174 is not synchronized with the state data 178 for a given instrument 112.
[0060] Mapping data 179 can be any data that provides a spatial representation of a patient's body during a medical procedure, such as depth information. Mapping data 179 can include a two-dimensional or three-dimensional representation of a patient’s body portion at which the medical procedure is taking place. Mapping data 179 can include information on depth from the location of camera for each individual point (e.g., a pixel in 2D image)providing information on depth from the camera for various tissues, organs, vessels or other features being imaged at the scene. The mapping data 179 can be generated by the depth mapper 158 and can be used to create three-dimensional maps of the surgical site, which can be used to detect if various data is out of order to trigger corrective actions.
[0061] The data repository 160 can include one or more data files, data structures, arrays, values, or other information that facilitates operation of the data processing system 130. The data repository 160 can include one or more local or distributed databases and can include a database management system. The data repository 160 can include, maintain, or manage one or more data streams 162. The data stream 162 can include or be formed from one or more of a video stream, image stream, stream of sensor measurements, event stream, kinematics stream, on or more instances of state data 178 or any mapping data 179. The data stream 162 can include data collected by one or more data capture devices 110, such as a set of 3D sensors from a variety of angles or vantage points with respect to the procedure activity (e.g., point or area of surgery).
[0062] Data stream 162 can include any stream of data. Data stream 162 can include a video stream, including a series of video frames or organized into video fragments, such as video fragments of about 1, 2, 3, 4, 5, 10 or 15 seconds of a video. Each second of the video can include, for example, 30, 45, 60, 90 or 120 video frames per second. Video fragments can be used to form a composite video 196. Data streams 162 can include an event stream which can include a stream of event data 176 or information, such as packets, which identify or convey a state of the robotic medical system 120 or an event that occurred in association with the robotic medical system 120. For example, data stream 162 can include any portion of system configuration 122, including information on operations on data streams 162, data on installation, uninstallation, calibration, set up, attachment, detachment or any other action performed by or on an RMS 120 with respect to medical instruments 112.
[0063] Data stream 162 can include data about an event, such as a state of the robotic medical system 120 indicating whether the medical tool or instrument 112 is calibrated, adjusted or includes a manipulator arm installed on a robotic medical system 120. Stream of event data 176 (e.g., event data stream) can include data on whether a robotic medical system 120 was fully functional (e.g., without errors) during the procedure. For example, when a medical instrument 112 is installed on a manipulator arm of the robotic medical system 120, a signal or data packet(s) can be generated indicating that the medical instrument 112 has been installed on the manipulator arm of the robotic medical system 120.
[0064] Data stream 162 can include a stream of kinematics data 172 which can refer to or include data associated with one or more of the manipulator arms or medical tools 112 (e.g., instruments) attached to the manipulator arms, such as arm locations or positioning. Data corresponding to medical tools 112 can be captured or detected by one or more displacement transducers, orientational sensors, positional sensors, or other types of sensors and devices to measure parameters or generate kinematics information. The kinematics data 172 can include sensor data along with time stamps and an indication of the medical tool 112 or type of medical tool 112 associated with the data stream 162.
[0065] Data repository 160 can store sensor data 174 having video frames that can include one or more static images or frames extracted from a sequence of images of a video file. Video frame can represent a specific moment in time and can be identified by a metadata including a timestamp 168. Video frame can display a visual content of the video of a medical procedure being analyzed by performance analyzer 190 to form a composite video 196 along with performance metrics 192 indicative of the performance of the surgeon performing the procedure. For example, in a video file capturing a robotic surgical procedure, a video frame can depict a snapshot of the surgical task, illustrating a movement or usage of a medical instrument 112 such as a robotic arm manipulating a surgical tool within the patient's body.
[0066] The system 100 can include a DPS 130 that can be deployed in or associated with the medical environment 102, or it can be provided by a remote server or be cloud-based. DPS 130 can include an interface designed, constructed and operational to communicate with one or more component of system 100 via network 101, including, for example, the RMS 120 or other network devices (e.g., edge servers, processing servers or cloud services). DPS 130 can be implemented using instructions, commands, computer code or data stored in memory locations and processed by one or more processors, controllers or integrated circuitry. DPS 130 can include functionalities, computer codes or programs for executing or implementing ML trainer 132 any of the ML models 140 and processing functions 180 to detect and take any actions responsive to data quality issues (e.g., data packets 164 falling out of order) and performance analyzer 190 for determining metrics 192 according to confidence scores 194 based on, or responsive to, the corrective actions from processing functions 180.
[0067] The ML trainer 132 can any combination of hardware and software for training ML models 140. The ML trainer 132 can train any type and form of ML models including kinematics data ML models 142, sensors data ML models 144 and events data ML models 146. The ML trainer 132 can train and generate ML models 140 for analyzing the performance of asurgeon based on data streams 162 and providing metrics 192 and confidence scores 194 based on the performance analysis. ML model trainer 132 can generate ML models 140 to generate and output composite videos 196 based on the processed analysis and data streams 162.
[0068] ML trainer 132 can include a framework or functionality for training any machine learning models, such as neural network models identifying or detecting errors in various data streams of an RMS 120. ML trainer 132 can train ML models 140 to utilize attention mechanisms for cross-referencing features between various data streams 162 (e.g., kinematic data stream, sensor data stream and events data stream) to detect errors in the data streams 162. For instance, ML trainer 132 can train one or more ML models 140 to utilize attention mechanism to correlate events data 176 (e.g., installation and selection events for particular medical instruments 112) with particular kinematics data 172 (e.g., expected motions of the same medical instruments 112), sensor data 174 (e.g., haptic or feedback sensor data or video images of medical instruments), state data 178 or mapping data 179 to determine that one or more sections of data are missing, out of order, delayed, lagged, or otherwise impacted during the transit from the source devices generating the data stream 162 to the DPS 130.
[0069] ML trainer 132 can include any combination of hardware and software for training or generating ML models 140. ML trainer 132 can include the functionality to train ML models 140 using supervised or unsupervised techniques, including for example self-supervision techniques applied to neural network models, such as feedforward neural networks (FNN), long short-term memory networks (LSTM), gated recurrent units (GRU) and generative adversarial networks (GANs) and transformers. ML trainer 132 can utilize training datasets 134 to train any of the ML models 140. ML trainer 132 can use the training dataset 134 for label-based or non-label based (e.g., self-supervision) training and can develop or establish encoders 150 and classifiers (e.g., classification blocks) for any of the ML models 140, utilizing weights 156 to adjust and tune loss functions 154 for different ML models 140. For instance, ML trainer 132 can improve the performance of the ML models 140 using weights 156 to more accurately detect and identify determinations or predictions of ML models 140 to one or more embedding layers or attention mechanisms, depending on the implementation.
[0070] ML models 140 can include any machine learning models for detecting or identifying errors or issues in data streams 162. ML models 140 can include neural network models, including RNN, LSTM, GAN, GRU, FNN, convolutional neural network, autoencoder, capsule networks or any other type of machine learning models. ML models 140 can include one or more kinematics data ML model 142 for identifying errors in a data streamof kinematics data 172 (e.g., a kinematics data stream). ML models 140 can include one or more sensors data ML models 144 for identifying errors in a data stream of sensor data 174 (e.g., a sensor data stream), such as a stream of video frames of a video recording of a medical procedure. ML models 140 can include one or more events data ML models 146 for identifying errors in a data stream of events data 176 (e.g., a stream of data on installation, configuration, selection, engagement or activation of medical instruments 112). ML models 140 can include a single ML model 140 including the functionality of each one of the KDML model 142, SDML model 144 and EDML model 146 to process all of the data streams 162 from all types of data (e.g., kinematics data 172, sensor data 174, events data 176, state data 178 or mapping data 179).
[0071] Encoder 150 can include any neural network component for transforming raw input data into an operational (e.g., meaningful) contextual representation of data to be used by the ML models 140. Encoder 150 can transform input into the ML model by extracting features (e.g., values or parameters) and capturing hierarchical patterns. Encoder 150 can include one or more layers for processing and transforming data, to extract relevant information from the input to be used in subsequent analysis. For instance, encoder 150 can facilitate the ML model 140 to create a consolidated representation of the relationships within a particular data stream 162 or between different data streams 162. For instance, encoder 150 can provide hierarchies of relations between kinematics data, sensor data, events data, state data or mapping data to allow the ML model 140 to discern the patterns and variations and identify errors or anomalies based on the data in the data packets 164 of each of the data streams 162.
[0072] Classifier 152, also referred to as a classification block, can include any function of a ML model 140 taking learned features from one layer of a neural network and assigning them to specific categories. Classifier 152 can use the categories assigned by the classifier 152 to make predictions or decisions based on the input data. Classifier 152 can include a plurality of connected layers and activation functions to transform encoded features into probability distributions across different classes. For instance, classifier 152 can utilize the encoded features generated by the encoder 150, and in collaboration with the loss function 154 can guide the optimization process by assessing the disparity between predicted and actual outputs, which can prompt the adjustment of weights 156 within the ML model 140 to minimize this disparity.
[0073] To train ML models 140, the ML trainer 132 can avoid rule-based synchronizations between the three data streams 162 (e.g., kinematics stream, sensor stream, event stream, statedata or mapping data). For instance, ML trainer 132 can automatically learn alignments between the three data streams, such as by learning mutual information between the three data streams 162. For instance, ML trainer 132 can train three neural networks separately (e.g., KDML model 142, SDML model 144 and EDML model 146) or can train a single ML model 140 combining the functionalities of all three of these models to process each data stream 162 (e.g., kinematic, sensor, events, state or mapping data). Each of the ML models 140 can be trained to include or utilize one or more encoders 150 (e.g., encoder blocks) and classifiers 152 (e.g., classification blocks).
[0074] ML models 140 can be trained on independent self-supervision tasks to learn to identify the features 166 indicative of errors in the data streams 162. ML models 140 can be trained to identify the features 166 in temporal windows (e.g., 1 to 5 seconds) that can be derived from each data stream 162 independently. ML trainer 132 can utilize a learning signal from this portion of training as a self-supervision loss that can be applied to each of the ML models 140 (e.g., 142, 144 and 146) in combination or individually. For instance, in the context of ML models 140 engaged in independent self-supervision tasks, the objectives can include tasks of recognizing generic features inherent to one or more temporal segments. For example, by aligning different features in a common latent space through a shared loss function, ML models 140 can be used to detect or identify errors by identifying time synchronization issues between features generated by separate data streams over the same time interval. ML models 140 can be trained to identify patterns within a current state of a data stream and implement similar pattern interpolation with respect to analyses from multiple data streams to converge these features into a unified feature space. In doing so, the technical solutions can infer or detect issues or errors from various data streams by detecting temporal synchronization discrepancies in the data patterns. For example, ML models 140 can be trained to recognize a feature (e.g. summarize current state of the data stream), and then using generated information from one or more ML models 140 on one or more data streams or types of data analyses, the ML models 140 can provide or plot these features or determinations onto the same feature space. These determinations on the same feature space can be used to identify patterns indicative of various issues, such as data flow issues or data flow synchronization issues.
[0075] ML models 140 can include mapping data ML models 148. Mapping data ML models 148 can be any machine learning models trained to process and analyze mapping data generated by robotic medical systems 120. Mapping data ML models 148 can utilize neural networks to interpret spatial representations of a patient's body, captured during medicalprocedures, such as metric distances from a camera lens to various points in a field of view. A mapping ML model 148 can be configured or trained to generate, from a video frame or an image, a two-dimensional image where each pixel corresponds to a depth value from a camera from which the image or video frame was captured to a point in the image or video frame corresponding to a tissue, organ or medical instrument 112 captured in the image (e.g., a depth map). A depth mapper 158 can utilize a mapping data ML model 148 can generate depth maps for various video frames or images (e.g., sensor data 174). The mapping data ML model 148 can be trained on extensive datasets that include sensor data (e.g., video frames from a stereo camera of an endoscope), events data, and kinematics data from various medical sessions. The depth maps generated using a mapping data ML model 148 can be used to detect and correct errors in the data streams 162, including reordering the data packets 164 that fell out of their intended original (e.g., chronological) order.
[0076] Loss function 154 can include any determination or calculation to quantify discrepancies between predicted outcomes and measured outcomes of ML models 140. Loss function 154 can correspond to the alignment of information from three distinct data streams: kinematics, sensors or video, and event data. Loss function 154 applied during training can correspond to a knowledge distillation loss, which can be applied to an embedding layer of each network or can be applied across all pairs of networks. Loss function 154 can be designed to encourage the networks to learn similar features to represent the data, thus creating embeddings that represent information in multiple data streams.
[0077] Weights 156 can include any parameters determined and assigned to relations or connections of an ML model 140. For instance, weight 156 can include a parameter to indicate an importance level of a connection of a particular neuron in a neural network ML models 140. Weights 156 can be provided or assigned by ML trainer 132 to represent the strength or importance of particular connections or relations. During the training process, weights 156 can be iteratively adjusted to minimize the difference between predicted and actual outputs, allowing the ML model 140 to learn and adapt to the underlying patterns in the data. Through the training process of the ML models 140, the loss can be scheduled, using for example, semisupervised learning (SSL) in which a model can be trained on a dataset including both labeled and unlabeled examples and knowledge distillation (KD) techniques in which a smaller model is trained to mimic the predictions of a larger, more complex model. The loss can be scheduled by starting with higher values or steps sizes for weights 156 on a first loss function 154 (e.g., the SSL loss function 154) to encourage the learning of meaningful representations andreducing the updated step size for weight 156 to balance between both loss functions 154. Weight 156 can be placed on the second loss function (e.g., KD loss function 154) to encourage the networks to learn mutual information, thereby training ML models 140 to detect discrepancies using data from each of the data streams 162. ML models 140 can be trained to produce features (e.g., 166) at every timestamp 168 and measure the distance between the features 166 (e.g., e.g., between features’ respective embeddings in feature space and time) in order to determine time synchronization issues or existence of errors within a data stream 162 or between the data streams 162.
[0078] Depth mapper 158 can be any combination of hardware and software for generating mapping data 179, such as data used to generate depth maps representing depth in images of a portion of patient's body in which a medical procedure is performed. Depth mapper 158 can utilize mapping data ML models 148 to determine depth representation for each pixel of an image or a video frame. Depth mapper 158 can generate two-dimensional array of depth values, where each pixel is represented by a distance between a camera lens and a point in an image corresponding to a particular organ, tissue, bone or other anatomy of the patient. A depth mapper 158 can capture spatial information using imaging devices, such as stereo cameras or time-of-fhght sensors. The depth mapper 158 can process the video stream or image data (e.g., sensor data 174) to create depth maps. A depth map can include a two-dimensional (e.g., array) representation of pixels of an image or video frame in terms of their respective metric distance from a camera lens to each of the pixel locations (e.g., corresponding to various tissues, organs or objects in a surgical scene). A depth map can include a two-dimensional representation of depths from the camera lens to each of the various tissues, capillaries, organs, blood streams, medical instruments 112 or other objects captured in a corresponding image from which the depth map is generated. The depth mapper 158 can generate mapping data 179 corresponding to one or more depth maps which can be associated with timestamps 168 and used to synchronize data by minimizing or reducing the difference in time intervals the data (e.g., synchronizing kinematics data 172 and mapping data 179 of a depth map according to the matching timestamps 168).
[0079] Performance analyzer 190 can include any combination of hardware and software for analyzing the performance of a surgeon with respect to a medical procedure captured by the data streams 162. Performance analyzer 190 can generate metrics 192 for various surgical tasks, phases or procedures, based on the performance of the surgeon with respect to such tasks. Performance analyzer 190 can generate metrics 192 corresponding to quality of data,such as the quality of video sensor data 174. For instance, metrics 192 can be indicative of presence of artifacts, such as distortions in the video quality. For example, performance analyzer 190 can detect blurry or unclear images beyond a predetermined threshold, or images in which the visibility due to glare and other obstacles is reduced below a predetermined threshold. In response to such detection, the performance analyzer 190 can determine that such images or video frames include distortions or are distorted. Performance analyzer 190 can utilize data streams 162 to determine the performance metrics 192. Performance analyzer can generate scores 194 indicative of the confidence or the level of certainty in the metrics 192. Confidence score 194 can be higher when no errors in data streams 162 are detected by the processing functions or ML models 140. Confidence score 194 can be lower when errors are detected and not addressed. Performance analyzer 190 can generate the metrics 192 and confidence scores 194 based on the operation and feedback from processing functions.
[0080] Composite video 196 can be generated by the performance analyzer 190 utilizing the data streams 162. Composite video 196 can include a video file generated from, based on, or using, any combination of sensor data 174 (e.g., video frames), kinematics data 172, events data 176, state data 178 or mapping data 179. Composite video 196 can include metrics 192 and confidence scores 194 overlaid over the composite video 196 during display of the video. Composite video 196 can include indications or alerts, including indications or alerts about any errors in the data streams 162 detected by the processing functions 180 or ML models 140.
[0081] Processing functions 180 can include any combination of hardware and software for processing data streams 162, identify errors in the data and address the errors with either corrective actions, mitigation of the impact of the errors or indications or alerts to be provided to the end user. Processing functions 180 can include or utilize ML models 140 to detect issues in the data streams 162. Processing functions 180 can include or utilize detection functions 182 for detecting errors or issues in the data, including data packets 164 out of order (e.g., based on timestamps 168), delays or jitters in the data stream 162 or missing or lost data packets 164. Processing functions 180 can detect presence of lower quality or corrupted data, such as video frame data that includes distortions. Processing functions 180 can include or utilize corrective functions 184 for taking corrective actions on the errors in the data, including rearrangement or reordering of data packets 164 identified to be out of order or eliminating corrupted data and preventing performance analyzer 190 from using corrupted data. Processing functions 180 can include or utilize mitigating functions 186 for mitigating effects of the corrupted data, such as informing the performance analyzer 190 to reduce the confidence score 194 in response to acorrupted or missing portion of data being used. Processing functions 180 can include or utilize indication functions 188 for providing indications or alerts that can be displayed or overlaid over the composite video 196.
[0082] Processing functions 180 can include multiple stages for detecting, correcting, mitigating and communicating various errors or issues with data in the system. For instance, processing functions 180 can include a detection function 182 or a stage at which errors or erroneous data elements (e.g., data packets 164) are detected and identified. For instance, processing functions 180 can include a correction function 184 or a stage at which corrections are made to identified data errors or issues. For instance, processing functions 180 can include a mitigation function 186 or a stage at which impacts of any data errors that cannot fixed by the corrective functions 184 can be mitigated to reduce their effects on the downstream process (e.g., performance analysis and metrics 192). For instance, processing functions 180 can include an indication function 188 or a communication stage at which the processing functions 180 can communicate to the system (e.g., performance analyzer 190) or to end user (e.g., via indications) any errors and their confidence impacts on various data elements.
[0083] Processing functions 180, can be implemented at any stage of the data flow. For example, any of the processing function 180 as well as data timestamping can be implemented when data is generated at the RMS 120, when data reaches the API server of the RMS 120, when data is recorded on an edge device on a network 101, when data is uploaded to cloud for pre-processing, when data is transferred to cloud for main data pipeline operations (e.g., DPS 130 operations) or when data is ready for metrics 192 computations by the performance analyzer 190. Processing functions 180 can be implemented at any of these devices or any stages of the process. In doing so, the processing functions 180 can more accurately identify or narrow down the portions of the system at which data errors are caused.
[0084] Detection functions 182 can include any combination of hardware and software for detecting and identifying data quality issues in data streams 162. Detection function 182 can include the functionality to compare timestamps 168 of various data packets 164 in the same or different data streams. Detection function 182 can include the functionality to identify data packets 164 that fell out of order during the transmission through the data pipeline. Detection function 182 can detect missing or lost data. Detection function 182 can utilize ML models 140 to identify or detect errors, including delays, jitters, misplacement, data packets getting out of order or any other issues with the data packets 164. For instance, detection function 182 can identify errors across various types of data, such as within a single stream (e.g., a kinematicsstream or a sensor stream alone) or across various data streams 162 including any combination of kinematics data 172, data stream 162 of sensor data 174, data stream 162 of events data 176 or any state data 178 or mapping data 179.
[0085] Detection function 182 can determine data presence and test for data consistency. For instance, kinematics data 172 can be sampled at a set rate (e.g., a sample rate of 100ms) from each moving component of the robot. The presence and consistency of the data stream 162 can be measured and checked deterministically. Metadata can be analyzed in a deterministic pattern and expected sequence order can be verified. For example, as values in certain kinematics actions are usually grouped, data can be checked against expected kinematics data patterns.
[0086] Detection function 182 can utilize timestamps 168 applied to data packets 164 at various stages in the data pipeline, including when the data packet is created, before being transmitted from the device, at it the time of the packet’s arrival or at the time of the recording of the packet into a data storage. In the instances in which system buffers (e.g., at the RMS 120) cause delays, jitters or data loss, timestamped data can be used to identify the faulty buffers at the RMS 120 allowing for changes in RSM configurations 122 in order to remedy the issue. Since a centralized clock can be used to timestamp the data packets 164 from various data streams 162 that are pipelined in a known pattern, the corrective functions 184 can recreate the order of the data packets 164 upon receipt.
[0087] Detection function 182 can track data stream 162 of events data 176 by order and source location and to check that events occur in an order of the events that are expected in a particular medical or surgical task or stage. Timestamps 168 of the data packets 164 of the events stream can mark the timing at several locations along the data pipeline that can be used to validate data streams 162. Events data 176 can be used to check for missed events and detect abnormalities. For example, when installing a tool (e.g., medical instrument 112) on a particular robotic arm, installation should occur on the same arm at which tool was docked. Events updating the user interface for the new tool can be observed and so several events can be expected to occur in same order or within a particular time frame. For instance, the solutions can identify that a medical instrument 112 is installed and ready to be used, but that no user interface actions (e.g., updates) occurred, which can indicate that there is an error.
[0088] Detection function 182 can monitor the sensor data 174 (e.g., video stream) for presence and consistency. For example, video files can include timestamps 168 for each videoframe, or a group of frames (e.g., video segment). These timestamps can be used to track consistent video recording across the duration of a procedure. Detection function 182 can also use video data timestamps 168 to check that all sections of the video are present in the recording (e.g., a composite video 196). Detection function 182 can detect distortions in video frames to identify frames of reduced quality (e.g., quality below a predetermined threshold metric 192 indicative of the video quality). For instance, through a combination of computer vision and log metadata tracking, the detection function 182 can determine when and how frames are recorded. For example, as video data can be timestamped at its generation, by observing the timestamps 168 at their expected rate at the destination, buffering issues (e.g., delays or jitters) and data loss can be detected when timestamped video data packets arrive inconsistent with their timestamped order.
[0089] Data streams 162 corresponding to sensor data 174 (e.g., videos), events 176, kinematics 172, state data 178 or mapping data 179 can include related, corresponding or duplicate information that can be used for cross-data comparisons and verification that all three data sources are in agreement. For instance, the detection function 182 can implement a check for consistency between diverse data types and data sources by mapping and comparing timestamps between different data types to facilitate if they consistently progress over time, such as in accordance with expected flow and correlation of events, video stream details and kinematics values.
[0090] For example, an installation of a medical instrument 112 can be recorded as a system event and provided in a data stream 162 of events data 176. At the same or similar expected time frame, the installed medical instrument 112 can shows up in a sensor data 174 (e.g., in a video) which can be detected with a SDML model 144, which can include a computer vision model. Kinematics data 172 can confirm movements of the medical instrument 112 according to the movements detected by the KDML model 142. Using these cross-data stream correlation techniques, detection functions 182 can verify time synchronization across the three data sources (e.g., three data streams 162).
[0091] For instance, detection function 182 can include a rule that camera movements (e.g., recorded as kinematic data 172 in a kinematic data stream) can only happen following a camera clutch system event, and therefore these two data streams can be used to verify their coordination or detect data issues. For example, detection function 182 can compare sensor data 174 of the video frames to kinematics data 172, such as when an ML model 140 is used to identify a camera motion from video, which then can be compared to extracted camera motionfrom kinematics to verify the movements, identify delays (e.g., time difference), jitters or any data loss.
[0092] Detection function 182 can pair events from the events data 176 according to expected trends. For instance, some events can be paired together based on their expected data, such as a tool install and uninstall data that can be done according to particular procedure. These patterns can be programmed manually, or mined through several pattern detection techniques, such as association rules, which can be used to detect outlier sessions. For example, a medical procedure can be broken up in five second chunks. A determination can be made to identify events, such as tool install event or tool ready event. In such instances, there can be a set threshold probability (e.g., 95% chance) that user interface data for the same event arrives. For instance, when a procedure or action is detected and this procedure or action is expected within a period of time (e.g., 5 seconds), then an end of that procedure can be expected to occur within the same time period.
[0093] Detection function 182 can operate on kinematics data 172. Kinematics data 172 can have expected patterns of motion and speed for particular detected actions or movements. Outlier kinematic data 172 samples can be detected through movements and data patterns that are improbable (e.g., beyond a probability range threshold). Metadata checking can be used for various procedures, including metadata with a procedure type information and surgeon information. These data can be reconciled with the logs from the system events data stream and can be used to validate data and check if the procedure recording is complete and includes all the expected information.
[0094] Corrective functions 184 can include any combination of hardware and software for making corrections or taking corrective actions on the errors in the data streams 162. Corrective functions 184 can include the functionality to reorder the data packets 164 identified to have gotten out of order during transmission. Corrective functions 184 can rearrange or reorder data packets 164 based on their timestamps 168 (e.g., timestamps 168 at any point or location in the transmission, processing or storage). Corrective function 184 can identify and remove portions of data streams 162 that are corrupted in response to the data streams not exceeding a threshold for minimal amount of data to be used for performance analysis by the performance analyzer 190. Corrective function 184 can remove outliers in the data and include the functionality to smooth data, as well as make adjustments or corrections to data streams 162, such as locate missing data packets 164 and incorporate them into the data stream 162.
[0095] Corrective functions 184 can provide time synchronization when two data streams diverge by looking for common data or known or expected data patterns. For instance, when a tool installation is detected in a video stream, system events can be synchronized to provide the alignment, such as by identifying and minimizing the time difference between the specific events. Similarly, user interface events (e.g., selection of particular actions by a user, such as tool installation actions) can be identified in correlation with installation events or kinematics (e.g., motion of the robotic arm on which the tool is installed), or in correlation with state data 178 identifying states of medical instruments 178 or any mapping data 179 indicating depth map of a surgical scene at which a medical procedure is taking place. For example, a selfsupervised sequence ML model 140 can be utilized to identify a sequence of events based on the data. A supervised model can be used to detect the sequence of events directly. For instance, self-supervised learning can be used for label-free generic time synchronization issues, while supervised leaming / models can be used for detecting time synchronization issues when there is a specific event (e.g. tool install) to synchronize around using multiple sources of data. When data streams diverge in time, corrective function 184 can use these determinations to synchronize (e.g., offset) the data streams 162 temporally to reduce their respective temporal mismatch.
[0096] Correction functions 184 can implement any data corrections. For instance, as kinematics and event-based data streams 162 can be timestamped at creation, this information can be used to reorder the events when they are recorded. Because there can be multiple timestamps 168 used to match kinematic, mapping data, state data or event data streams 162 to video, these can be smoothed to find the closest timestamps 168 that match the video and system data smoothly. Corrective functions 184 can generate configuration settings or correct system configurations 122 and adjust the operation of the RMS 120 minimize errors in data streams 162. Corrective actions 184 can provide, facilitate or encourage software updates or system configurations 122 and suggest certain adjustments to operation. Correction functions 184 can include interpolating kinematic data that is missing, or smoothing paths of kinematics that are noisy.
[0097] Corrective functions 184 can provide alert for portions of data streams that are identified as missing. For instance, a locally weighted scatterplot smoothing (LOWESS) regression approach can be used as a baseline to identify a long-term trend. LOWESS can include a non-parametric regression technique in which weighted averages are used to create a smooth curve representing trends in a scatterplot. For instance, when a clock drifts or changes,there can be some noise as a result. Corrective functions 184 can use the LOWESS regression to find baseline drift and make appropriate corrections, accounting for the known clock issues, at each stage.
[0098] Detection functions 182 and corrective functions 184 can operate on data packets 164 at one or more points in a data pipeline. This can allow the detection functions 182 to identify locations or portions of the system in which errors (e.g., delays, jitter or data loss) are occurring, facilitating identification of configuration issues and solutions for different stages. For instance, the corrective functions 184 can debug issues internally, such as by location the exact source of error. The corrective functions 184 can debug issues externally, including identifying and providing configuration for customer products, such as the robotic systems in which a recording device and various other sensors, events data and kinematics data can have settings that can affect data quality. Detection functions 182 can include specific rules of kinematics / events patterns that are expected to be present for high data quality (e.g., data above a particular quality threshold).
[0099] Mitigating functions 186 can include any combination of hardware and software for mitigating effects of errors in the data streams 162. Mitigating functions 186 can include functions to mitigate the effects of erroneous data. For instance, when a section of data stream 162 is known to be corrupted or include missing data packets 164, mitigating functions 186 can avoid downstream computations on the identified erroneous sections of data to mitigate any incorrect metrics 192 computed by the performance analyzer 190. For instance, when a sufficiently large section of data (e.g., greater than a threshold) is identified as missing, the entire data pipeline can be removed or omitted (e.g., short circuited) other than procedure-level metrics as there may not be enough information to implement metrics. For example, when more than a threshold amount of kinematic data 172, or sensor data 162 or events data 176 or state data 178 or mapping data 179 is missing, the mitigating function 186 can provide or issue instructions to the performance analyzer 190 to not compute metrics 194 for the manipulators on the system for which data is missing. For example, when event data 176 is missing, the mitigating functions 186 can instruct the system to not compute event-based determinations for this data. For example, when an amount of missing data exceeds a threshold for the minimum amount of data relied on to determine the metrics 192, the mitigating function 186 can cause the system to not include the procedure recording as part of video recommendations or other tools for comparison that require generally clean data.
[0100] Indication functions 188 can include any functionality for generating indications or alerts, such as indications or alerts indicating errors in data, corrections made to data, or alerts that corrections were made to data stream 162. Indication functions 188 can generate messages to be displayed along with the composite video 196 indicating that data for a portion of the composite video 196 is missing. For instance, indication function 188 can alert end users that a portion of data stream 162 is corrupted or has a data quality that is below a particular threshold. Indication functions 188 can overlay alerts and indications to the end user (e.g., surgeon or user analyzing medical procedure and its metrics). Indications or alerts can be provided to systems engineers to facilitate improving of system configurations 122 and addressing cleaner data sources or to correct data that may be missing.
[0101] FIG. 2 illustrates an example block diagram 200 of a system in which ML models 140 trained for different types of data streams 162 are utilized to detect features 166 indicative of errors in the data. Block diagram 200 can include a data stream 162 including kinematics data 172 being input into a KDML model 142, another data stream 162 having events data 176 can be input into EDML model 146, while a third data stream 162 having sensor data 174 (e.g., video feed) can be input into SDML model 144. Each of the ML models 140 can process their respective input data streams 162 to output features 166a, 166b and 166c. Detection function 182 can analyze the detected features 166a-c using rules 202 to determine whether an error in the data streams 162 has occurred.
[0102] Rules 202 can include any one or more rules for determining presence of an error in one or more data streams 162. Rules 202 can include a rule that a particular arrangement of features 166 can trigger a determination that there is an error in the data. Rules 202 can include commands or instructions to determine existence of an error when particular inconsistencies between features 166 from various types of data streams 162 are encountered. For instance, rule 202 can state that when features of two data streams 162 indicate a same operation, while the third data stream has a feature 166 that is inconsistent with the features of the other two streams 162, an error in the third data stream 162 exists.
[0103] Based on the respective data streams 162 input into the three ML models 140 (e.g., 142, 144 and 146), the ML models 140 can identify respective features 166. From the data stream 162 of data packets 164 on kinematics data 172 KDML model 142 can identify feature 166a. From the data stream 162 of data packets 164 on events data 176 EDML model 146 can identify feature 166b. From the data stream 162 of data packets 164 on sensor data 174, SDMLmodel 144 can identify feature 166c. Models 142-146 can identify features 166a-c, based on a relationship, correlation or issues identified using these three features.
[0104] For example, feature 166b can include an engagement, selection or configuration of a first medical instrument 112, while feature 166c includes a video frame of a second (e.g., different) medical instrument 112 being used, whereas feature 166a can include kinematics data 172 indication movement of the second medical instrument 112, and not the first. Based on the features 166a-c, detection function 182 can determine, using a rule 202, that there are missing data packets 164 in events data 176. Events 176 can be mined for rules 202 in small temporal chunks, such as 5 seconds of data, to find common groupings of events that can be compared with data from the same time frame (e.g., based on the timestamps 168) from other data streams 162. Using for example, association rule mining, patterns and probabilities of those patterns can be determined by any combination of detection functions 182 and ML models 140.
[0105] For example, feature 166b can correspond to, indicate or denotes a vector in a latent or feature space, corresponding to engagement, selection, or configuration actions associated with a first medical instrument (112). For example, feature 166c can be represented as a vector, capturing a video frame depicting the usage of a second, distinct medical instrument (112). For example, feature 166a can be expressed as a vector in the latent space, embodying kinematics data (172) indicative of a movement related to the second medical instrument, excluding the first. Such vectors in the latent space can allow for a mathematical representation of the features, facilitating their analysis and interpretation. Using a rule (202) on such vector representations, the detection function (182) can identify patterns within the latent space, such as am occurrence of missing data packets (164) in the events data (176). This approach can include, utilize or provide an example for various aspects of Semi-Supervised Learning (SSL), in which the features are transformed into a vector space, enabling generalization for diverse applications, which can be used by the present solutions for various (e.g., ML-based) determinations.
[0106] For instance, if events data 176 includes a feature 166b indicative that a medical instrument 112 is installed and ready, based on a rule 202, detection function 182 can determine that there is 95% probability that within the given time frame (e.g., 5 seconds), another feature 166b will appear indicating that the medical instrument 112 is now visible in the user interface. Patterns such as this one can be mined automatically and then applied to analyze the probability or likelihood of a data stream 162 for a video recording to be used by the performance analyzer 190 to include all of the events and data used for a reliabledetermination of a performance metrics 192. For example, DPS 130 can compute the likelihood that a recorded medical session with 100 examples of this pairing only has encountered a third event at a rate of 85% and can in response determine that it is unlikely that the recording has complete data. Responsive to this determination, an error can be identified. Such groupings mined automatically can help be used for identifying “unlikely” sessions that break normal patterns and detecting errors.
[0107] FIG. 3 depicts an example block diagram of an example computer system 300 is shown, in accordance with some embodiments. The computer system 300 can be any computing device used herein and can include or be used to implement a data processing system or its components. The computer system 300 includes at least one bus 305 or other communication component or interface for communicating information between various elements of the computer system. The computer system further includes at least one processor 310 or processing circuit coupled to the bus 305 for processing information. The computer system 300 also includes at least one main memory 315, such as a random-access memory (RAM) or other dynamic storage device, coupled to the bus 305 for storing information, and instructions to be executed by the processor 310. The main memory 315 can be used for storing information during execution of instructions by the processor 310. The computer system 300 can further include at least one read only memory (ROM) 320 or other static storage device coupled to the bus 305 for storing static information and instructions for the processor 310. A storage device 325, such as a solid-state device, magnetic disk or optical disk, can be coupled to the bus 305 to persistently store information and instructions.
[0108] The computer system 300 can be coupled via the bus 305 to a display 330, such as a liquid crystal display, or active-matrix display, for displaying information. An input device 335, such as a keyboard or voice interface can be coupled to the bus 305 for communicating information and commands to the processor 310. The input device 335 can include a touch screen display (e.g., the display 330). The input device 335 can also include a cursor control, such as a mouse, a trackball, or cursor direction keys, for communicating direction information and command selections to the processor 310 and for controlling cursor movement on the display 330.
[0109] The processes, systems and methods described herein can be implemented by the computer system 300 in response to the processor 310 executing an arrangement of instructions contained in the main memory 315. Such instructions can be read into the main memory 315 from another computer-readable medium, such as the storage device 325. Execution of thearrangement of instructions contained in the main memory 315 causes the computer system 300 to perform the illustrative processes described herein. One or more processors in a multiprocessing arrangement can also be employed to execute the instructions contained in the main memory 315. Hard-wired circuitry can be used in place of or in combination with software instructions together with the systems and methods described herein. Systems and methods described herein are not limited to any specific combination of hardware circuitry and software.
[0110] Although an example computing system has been described in FIG. 3, the subject matter including the operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.
[0111] System 100 can include and combine any features illustrated herein, including in FIGs. 1-3. For instance, system 100 can implement the functionalities of the DPS 100, including example block diagram 200, utilizing computing system 300. System 100 can include or utilize one or more processors 310 that can be coupled with memory 315 or utilize data stored in ROM 320 or storage devices 325 to implement any functionalities herein. System 100 can use one or more processors 310 implementing instructions stored in a non-transitory computer-readable medium (e.g., 315). For instance, the instructions stored in the non- transitory computer-readable medium, when executed by the one or more processors 310, can cause the one or more processors to implement various functionalities of the DPS 130.
[0112] The one or more processors 310 can receive, for a medical session with a robotic medical system 120, a first one or more data packets 164 of a first stream 162 of data at a first time. The data can include kinematics data 172, sensor data 174 or events data 176. The data can be received at a first time that can be identified or defined by a timestamp 168. The medical session of the RMS 120 can include a medical surgery performed using the RMS 120, which can be recorded via data capture devices 110 to generate a composite video 196 along with metrics 192 indicating the performance of the surgeon based on the data.
[0113] The one or more processors 310 can receive, for the medical session, a second one or more data packets 164 of a second stream 162 of data at a second time subsequent to the first time. The one or more processors 310 can receive both the first packet 164 and the second packet 164 from a single (e.g., a same) stream 162 of data or from different streams 162 of data. For instance, a first packet 164 can be from a stream of data of a first sensor generatingsensor data 174, while a second packet 164 can be from the same stream of data of the first sensor, or from a stream of sensor data 174 from a different sensor, or from a stream of kinematics data 172 or events data 176 from different devices.
[0114] For instance, the second time can correspond to a second timestamp 168 that can indicate a time of the timestamping of the second data packet 164 that is following the first timestamp 168 for the first data packet 164. The first and the second timestamps can be generated at the same phase of the data pipeline, such when data packets 164 were generated, when data packets 164 were transmitted to an API server for the RMS 120, when the data packets 164 arrived at the API server, when the data packets were recorded on an edge device of the system, when the data packets 164 were uploaded to cloud for preprocessing, when the data packets 164 were preprocessed, when the data packets 164 were transmitted to the DPS 130 or when the data packets 164 were indicated as ready for performance analysis by the performance analyzer 190.
[0115] The one or more processors 310 can identify a first feature 166 in the first one or more data packets 164 and a second feature 166 in the second one or more data packets 164. The first and the second features 166 can correspond to features 166 of any one or more of the kinematics data 172 (e.g., movement of medical instruments 112), sensor data 174 (e.g., video frames of a video of the medical session or procedure, or pressure, temperature or vibration measurements taken by sensors) or events data 176 (e.g., calibration, installation, engagement, setup or activation of medical instruments 112).
[0116] The one or more processors 310 can detect, based at least on the first feature 166 and the second feature 166, that the first one or more data packets 164 and the second one or more data packets 164 are out of order. For instance, a detection function 182 can utilize one or more ML models 140 to detect or determine that the one or more data packets 164 are out of order. For example, KDML model 142 processing kinematics data 172 along with SDML model 144 processing sensor data 174 and EDML model 146 processing events data 176 can determine that the data packets 164 are out of order.
[0117] The one or more processors 310 can, responsive to the detection that the data packets 164 are out of order, rearrange or reorder the first one or more data packets and the second one or more data packets. Reordering or rearrangement of the data packets 164 can cause the first one or more data packets 164 to be subsequent to the second one or more data packets 164. The reordering can be implemented, for example, by a corrective function 184.The corrective function 184 can rearrange or reorder the data packets 164 responsive to timestamps 168 for the first one or more data packets 164 and the second one or more data packets 164.
[0118] The one or more processors can determine, using the reordered first one or more data packets 164 and second one or more data packets 164, a metric 192 indicative of performance of the medical session. Performance analyzer 190 can receive an indication from the processing functions 180 that the first and second one or more data packets 164 have passed the quality and reliability analysis (e.g., likelihood of errors is below a threshold) and are ready to be used for assessment and metrics 192 determination. Responsive to this indication, performance analyzer 190 can utilize the first and the second one or more data packets 164 to analyze the performance of the surgeon in the composite video 196 and determine the metrics 192 and the corresponding confidence scores 194.
[0119] The one or more processors 310 can generate, using the reordered first one or more data packets 164 and second one or more data packets 164, a composite video 196 of at least a portion of the medical session with an indication of the metric 192. The performance analyzer 190 can generate the composite video 196 from a plurality of portions of data streams 162 of video clips captured by data capture devices 110 or medical instruments 112 during the medical session. Performance analyzer 190 can generate metrics 192 using the kinematics data 172, sensor data 174 and events data 176 according to expected data for the actions performed by the surgeon during the particular type of medical session.
[0120] The one or more processors 310 can identify that the first stream 162 corresponds to a first of a data generated by a sensor (e.g., 174) of the robotic medical system 120, a data on kinematics (e.g., 172) of one or more instruments of the robotic medical system 120 used during the medical session, and a data of an event (e.g., 176) at the robotic medical system 120. The one or more processors 310 can identify that the second stream 162 corresponds to a second one of the data generated by the sensor (e.g., 174), the data on kinematics (e.g., 172) of the one or more instruments and the data on the event (e.g., 176) at the robotic medical system 120.
[0121] The one or more processors 310 can detect, based at least on one of the first feature 166 and the second feature 166, at least one of a delay or a jitter corresponding to at least one of the first stream 162 and the second stream 162. For example, a delay or a jitter can be detected by comparing a plurality of timestamps 168 for each of the first one or more datapackets 164 and the second one or more data packets 164. Variations in the timestamps 168 for the data packets 164 that covered the same data path can indicate a delay or jitter between the first one or more data packets 164 and the second one or more data packets 164, or their delays or jitters with respect to the origin and the destination. The one or more processors 310 can determine, responsive to the detection of the at least one of the delay or the jitter, a portion the robotic medical system 120 causing the at least one of the delay or the jitter. For instance, detection function 182 can identify a particular buffer at the RMS 120, a server (e.g., edge server) coupled with the RMS 120 via a network 101, or at a cloud system, as the cause of the data error (e.g., delay or jitter).
[0122] The one or more processors 310 can identify the first feature 166 comprising a first timestamp 168 for the first one or more data packets 164 and the second feature 166 comprising a second timestamp 168 for the second one or more data packets 164. The one or more processors 310 can identify, based at least one the first timestamp 168 and the second timestamp 168, a time at which at least one of the first one or more data packets 164 and the second one or more data packets 164 are at least one of generated, transmitted, received or stored by at least a portion of the robotic medical system 120. The one or more processors 310 can detect, based at least on the time (e.g., 168), that the first one or more data packets 164 and the second one or more data packets 164 are out of order.
[0123] The one or more processors 310 can identify, based at least on the first feature 166, the first time (e.g., 168) and the second time (e.g., 168), that a likelihood of the second feature 166 corresponding to an event (e.g., 176) of the robotic medical system exceeds a threshold. The threshold can include a likelihood or probability threshold that event data 176 is over 95% likely to indicate a particular event (e.g., medical instrument 112 displayed on the graphical user interface of the display 116) responsive to another event (e.g., medical instrument 112 being activated) having occurred within 5 seconds. The one or more processors 310 can detect, based at least on the likelihood exceeding the threshold, that the first one or more data packets 164 and the second one or more data packets 164 are out of order.
[0124] At least one of the first feature 166 and the second feature 166 can indicate at least one of: an event (e.g., 176), a task of a medical procedure of a medical session being recorded, a phase of a medical procedure, an object used in the medical procedure or a workflow of the medical procedure. For instance, the features 166 can indicate or correspond to any particular action of a surgeon corresponding to any particular operation of a medical procedure implemented during the medical session. The one or more processors 310 can identify, based atleast on the reordered first one or more data packets 164 and second one or more data packets 164, a data missing from at least one of the first stream 162 or the second stream 162. The one or more processors 310 can determine, based at least on the identified data, a confidence score 194 for the metric 192. The one or more processors 310 can provide the confidence score 194 for display with the composite video 196.
[0125] The one or more processors 310 can identify a machine learning (ML) model 140 trained on a plurality of features of a plurality of data packets 164 of the robotic medical system 120. The ML model 140 can include a KDML model 142, a SDML model 144 or an EDML model 146. The ML model 140 can include a single ML Model 140 including any combination of models 142-146. The one or more processors 310 can detect that the first one or more data packets 164 and the second one or more data packets 164 are out of order based at least on the first feature 166 and the second feature 166 input into the ML model 140.
[0126] The one or more processors 310 can identify one or more machine learning (ML) models 140 utilizing one or more neural networks trained on sensor data 174 from a plurality of streams 162 of sensors of a robotic medical system 120, events data 176 corresponding to a plurality of events of a plurality of medical procedures and kinematics data 172 corresponding to data on kinematics relating the plurality of medical procedures implemented on the robotic medical system 120. The one or more processors 310 can detect that the first one or more data packets 164 and the second one or more data packets 164 are out of order based at least on the first feature 166 and the second feature 166 input into the one or more ML models 140.
[0127] The one or more processors 310 can identify a first ML model (e.g., 144) trained on the sensor data 174 and corresponding to a first encoder 150 and a first one or more classifiers or classification blocks 152 to process representations on the sensor data 174. The one or more processors 310 can identify a second ML model (e.g., 146) trained on the events data 176 and corresponding to a second encoder 150 and a second one or more classification blocks 152 to process representations on the events data 176. The one or more processors 310 can identify a third ML model 142 trained on the kinematics data 172 and corresponding to a third encoder 150 and a third one or more classification blocks 152 to process representations on the kinematics data 172. The one or more processors 310 can detect that the first one or more data packets 164 and the second one or more data packets 164 are out of order based at least on the first feature 166 and the second feature 166 input into the first ML model 144, the second ML model 146 and the third ML model 142.
[0128] The one or more processors 310 can determine, using the reordered first one or more data packets 164 and second one or more data packets 164, that a portion of at least one of the first stream 162 and the second stream 162 is missing. The one or more processors 310 can generate an alert or an indication to indicate that the portion of at least the one of the first stream 162 and the second stream 162 is missing. The one or more processors 310 can overlay the alert on the composite video 196. The one or more processors 310 can determine, responsive to the detection, that a portion of at least one of the first stream 162 over a time interval (e.g., up to 0.1, 0.5, 2, 4, 5, 6 or 10 seconds) and the second stream 162 over the time interval (e.g., up to 0.1, 0.5, 2, 4, 5, 6 or 10 seconds) is missing. The one or more processors 310 can modify, prior to display of the composite video 196, the composite video 196 to exclude a portion of the composite video 196 corresponding to the time interval.
[0129] Turning now to FIG. 4, an example flow diagram of a method 400 for providing an end-to-end automated data quality checking in robotic systems is illustrated. The method 400 can be performed by a system having one or more processors executing computer-readable instructions stored on a memory. The method 400 can be performed, for example, by system 100 and in accordance with any features or techniques discussed in connection with FIGS. 1-3 and 5. For instance, the method 400 can be implemented one or more processors 310 of a computing system 300 executing non-transitory computer-readable instructions stored on a memory (e.g., the memory 315, 320 or 325) and using data from a data repository 160 (e.g., storage device 325).
[0130] The method 400 can be used provide quality control of a plurality of data streams in a robotic system, such as a robotic medical system. At operation 405, the method can receive data packets from data streams. At operation 410, the method can detect features using the data packets. At operation 415, the method can determine if the features indicate error in the data. At operation 420, the method can take corrective action on the error. At 425, the method can determine if the error has been corrected. At operation 430, the method can take mitigating action if the error is not corrected. At operation 435, the method can generate the composite video.
[0131] At operation 405, the method can receive data packets from data streams. The method can include one or more processors coupled with memory receiving a first one or more data packets of a first stream of data at a first time. The method can include receiving a second one or more data packets of a second stream of data at a second time. The second time can be subsequent to the first time. The method can compare the timestamps of arrival at the dataprocessing system of the first one or more data packets with the timestamps of arrival at the data processing system of the second one or more data packets. The method can determine that the second one or more data packets arrived subsequent to the first one or more data packets responsive to the timestamps indicating the arrival of the second one or more data packets subsequent to the arrival of the first one or more data packets.
[0132] The first stream and the second stream can correspond to a medical session implemented using a robotic medical system. Each of the first and the second data streams can include or correspond to any of: a data stream including kinematics data (e.g., kinematics data stream), a data stream including sensor data (e.g., sensor data stream), a data stream including events data (e.g., events data stream), state data (e.g., indicative of a state of a medical instrument) or mapping data (e.g., indicative of a spatial representation nor a depth representation of a view from a camera).. The medical session can include a medical operation performed using a robotic medical system and recorded by data captured devices and using a plurality of streams of data, including kinematics data stream, sensor (e.g., video) data stream and events data stream.
[0133] The method can include the one or more processors identifying that the first stream corresponds to a first of a data generated by a sensor of the robotic medical system, a data on kinematics of one or more instruments of the robotic medical system used during the medical session, and a data of an event at the robotic medical system. The method can identify that the second stream corresponds to a second one of the data generated by the sensor, the data on kinematics of the one or more instruments and the data on the event at the robotic medical system. This identification can be implemented by a detection function or one or more ML models, such as the kinematics data ML model, sensors data ML model or events data ML model.
[0134] The method can include receiving a third one or more data packets of a third stream of data at a third time. The third time can be subsequent to the second time. The third one or more data packets can be data packets of a state data indicating a state of a medical instrument. The third one or more data packets can be data packets of a mapping data indicating a depth map of a surgical location within a patient’s body in which a medical procedure is being performed. The third time can be determined based on a third timestamp of the third one or more data packets, which can be compared with the timestamps of the first and the second one or more data packets. The method can identify that the third stream corresponds to a third oneof the data generated by the sensor, the data on kinematics of the one or more instruments and the data on the event at the robotic medical system.
[0135] At operation 410, the method can detect features using the data packets. The method can include identifying or detecting a first feature of the first one or more data packets and a second feature of the second one or more data packets. The method can identify or detect the features by parsing the contents of the data packets or by analyzing a plurality of data packets. For example, a feature can include or be indicated by kinematics data, events data, sensor data, state data or mapping data included into or represented by one or more data packets. For example, a feature can be determined, detected or recognized from a plurality of payloads from a plurality of data packets, including any combination of a plurality of kinematics data, sensor data, events data, state data or mapping data.
[0136] Data packets can be identified using one or more ML models. For instance, the first feature can be recognized or identified using any one of a kinematics data ML model, a sensors data ML model, a mapping data ML model and an events data ML model. The second feature can be recognized or identified using any one of a remaining one of the kinematics data ML model, the sensors data ML model, the mapping data ML model and the events data ML model. The third feature can be recognized or identified using the last remaining one of the kinematics data ML model, the sensors data ML model, the mapping data ML model and the events data ML model.
[0137] Detecting or identify at least on a first feature of the first one or more data packets and a second feature of the second one or more data packets that can be indicative of errors in data. The method can include the one or more processors identifying, based at least one the first timestamp and the second timestamp, a time at which at least one of the first one or more data packets and the second one or more data packets are at least one of generated, transmitted, received or stored by at least a portion of the robotic medical system.
[0138] The features can include or indicate one or more timestamps of the data packets. The features can include or indicate differences between the timestamps of the one or more data packets. The features can include or indicate one or more events that are expected to be correlated or done together with (e.g., within a particular time interval of) another event. The time interval can be up to 0.1s, 0.5s, Is, 2s, 3s, 4s, 5s or more than 5 seconds. The event can include, for example an installation, configuration, setup, activation or engagement of a medical instrument or GUI selection or display of an action. The features can include orindicate one or more kinematics actions (e.g., movements or motions of a medical instrument) that are expected to be correlated or done together with (e.g., within a particular time interval of) another kinematic action. The features can be indicated by a state data indicative of a state of a medical instrument or by mapping data indicating a depth map determined from one or more video or image frames. The features can include or indicate one or more sensor data (e.g., video frames displaying an action of a surgeon) that are expected to be correlated or done together with (e.g., within a particular time interval of) another sensor data (e.g., another action within a next video frame). These data, such as a correlated sensor data (e.g., video recorded actions), can be recognized as or learned by an ML model detecting tasks or actions of surgeon based on video data as parts of a task or a surgical phase.
[0139] The features can include, correspond to or indicate one or more events that are expected to be correlated or done together with (e.g., within a particular time interval of) a kinematics data (e.g., a particular movement). The features can include or indicate one or more kinematics actions (e.g., movements or motions of a medical instrument) that are expected to be correlated or done together with (e.g., within a particular time interval of) a sensor data (e.g., a video recorded action of a surgeon). The features can include or indicate one or more sensor data (e.g., video frames displaying an action of a surgeon) that are expected to be correlated or done together with (e.g., within a particular time interval of) an event data. The features can include or indicate one or more event data that are expected to be correlated or done together with (e.g., within a particular time interval of) a kinematics data.
[0140] At operation 415, the method can determine if the features indicate an error in the data. The method can include the one or more processors detecting that at least one of the first feature and the second feature are indicative of an error (e.g., unsynchronized data or data out of its chronological order) in one or more data streams. The first feature or the second feature can indicate a delay, a jitter or a temporal mismatch between two data streams such as any two of kinematics data stream, sensor data stream and events data stream. The first feature or the second feature can indicate a probable or likely mismatch between a kinematics movement or data and a sensor data. The first feature or the second feature can indicate a probable or likely mismatch between an event and a kinematic data. The first feature or the second feature can indicate a probable or likely mismatch between a sensor data and an event data, between state data and a mapping data, or between sensor data and the mapping data.
[0141] The first feature or the second feature can indicate a mismatch between a prior kinematic data and a kinematic data of the present data packet. The first feature or the secondfeature can indicate a mismatch between a prior sensor data and a sensor data of the present data packet. The first feature or the second feature can indicate a mismatch between a prior events data and an events data of the present data packet. The mismatch can include a ML model learned analysis of any combination of event data, sensor data and kinematics data whose probability of occurrence at the same time or within a time interval (e.g., 5 seconds) is less than a particular threshold of probability (e.g., less than 1%, 3% or 5%)
[0142] The method can include the one or more processors detecting, based at least on a first feature of the first one or more data packets and a second feature of the second one or more data packets, that the first one or more data packets and the second one or more data packets are out of order. The method can identify a time at which at least one of the first one or more data packets and the second one or more data packets are at least one of generated, transmitted, received or stored by at least a portion of the robotic medical system. The method can detect, based at least on the time, that the first one or more data packets and the second one or more data packets are out of order.
[0143] The method can include the one or more processors detecting, based at least on one of the first feature and the second feature, at least one of a delay or a jitter corresponding to at least one of the first stream and the second stream. For example, a detection function can utilize ML models to detect delay, jitter or a temporal mismatch or offset between data packets of different data streams The method can include identifying the first feature comprising a first timestamp for the first one or more data packets and the second feature comprising a second timestamp for the second one or more data packets. The method can include detecting artifacts or distortions in a video frame. This detection can be implemented using ML models processing log metadata tracking video frame data.
[0144] The method can detect the first feature and the second feature and determine, responsive to, or based on, the detection of the features, that a portion of at least one of the first stream over a time interval and the second stream over the time interval is missing. The method can include identifying, based at least on the first feature, the first time and the second time, that a likelihood of the second feature corresponding to an event of the robotic medical system exceeds a threshold. The method can include identifying, based at least on the first feature, the first time and the second time, that a likelihood of the second feature corresponding to an event of the robotic medical system is below a threshold.
[0145] The method can include the one or more processors detecting, based at least on the likelihood exceeding the threshold, that the first one or more data packets and the second one or more data packets are out of order. The at least one of the first feature and the second feature indicate at least one of an event, a task of a medical procedure of the medical session, a phase of the medical procedure, an object used in the medical procedure or a workflow of the medical procedure.
[0146] The method can include identifying, by the one or more processors, a machine learning (ML) model trained on a plurality of features of a plurality of data packets of the robotic medical system. The method can include the one or more processors detecting that the first one or more data packets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the ML model. For instance, an error in the data (e.g., data packets out of order, delay, jitter or data loss) can be identified based on a portion of the first one or more data packets can be input into a first ML model and a portion of the second one or more data packets can be input into a second ML model.
[0147] The one or more processors can identify one or more machine learning (ML) models utilizing one or more neural networks. The one or more ML models can be trained on sensor data from a plurality of streams of sensors of a robotic medical system, events data corresponding to a plurality of events of a plurality of medical procedures and kinematics data corresponding to data on kinematics relating the plurality of medical procedures implemented on the robotic medical system. The method can include detecting that the first one or more data packets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the one or more ML models.
[0148] If at operation 415, the method determines that the first feature and the second feature do not indicate error in the data, the operation can go back to operation 405 to receive a new set of a first one or more data packets and a second one or more data packets. If at operation 415, the method determines that the features indicate the error in the data, the method can proceed to operation 420.
[0149] At operation 420, the method can take corrective action on the error. The method can include the one or more processors implementing corrective functions. A corrective function can implement a time shift or temporal offset between the first data stream and the second data stream to reduce the delay between the first data stream and the second data stream. A corrective function can include reordering of the first one or more data packets andthe second one or more data packets. The reordering can be implemented using timestamps of the data packets. The method can include identifying a part of the first data stream or the second data stream as corrupted or lost and removing this part of the data stream from the first data stream or the second data stream.
[0150] The method can include the one or more processors reordering, responsive to the detection, the first one or more data packets and the second one or more data packets to cause the first one or more data packets to be subsequent to the second one or more data packets. The method can include the one or more processors determining, responsive to the detection of the at least one of the delay or the jitter, a portion the robotic medical system causing the at least one of the delay or the jitter.
[0151] At 425, the method can determine if the error has been corrected. The method can include the one or more processors determining if the corrective action has addressed the error. For example, the one or more processors can determine that the first one or more data packets and the second one or more data packets have been reordered or rearranged into the correct order. For example, the one or more processors can determine that the first one or more data packets and the second one or more data packets have been time shifted or temporally shifted with respect to each other to reduce the effects of delays or jitter. For example, the one or more processors can determine that a portion of the first data stream and the second data stream that were identified as corrupted or lost have been removed from data and the performance metrics by the performance analyzer can be determined with a level of confidence exceeding a threshold.
[0152] The method can include the one or more processors that the corrective action was not successful. For example, the method can determine that the amount of data missing or corrupted exceeds a threshold for disregarding the entire data stream. The method can determine that at least a portion of the data stream is not to be used by the performance analyzer to determine performance metrics or use in composite video. The method can determine that the confidence score for any metrics for a time period for which the data stream is corrupted is below a threshold for reliable performance analysis. Responsive to such determinations, the method can determine that the error in the data is not corrected. If at operation 425, the method determines that the error is not corrected, the method can proceed to operation 430 to take mitigating action. If at operation 425, the method determines that the error is corrected, the method can proceed to generate composite video at operation 435.
[0153] At operation 430, the method can take mitigating action if the error is not corrected. The method can include the one or more processors determining, responsive to the detection, that a portion of at least one of the first stream over a time interval and the second stream over the time interval is missing. The method can include the one or more processors modifying, prior to display of the composite video, the composite video to exclude a portion of the composite video corresponding to the time interval.
[0154] The method can include one or more indications or alerts generated by the indication functions to provide indications to the performance analyzer. The indications or alerts can indicate that the first or the second data streams corresponding to the time intervals for which the data has been removed (e.g., due to its being identified as corrupted or lost) is below a reliability threshold. The method can provide indications to performance analyzer that confidence score for any metrics determined for a given data stream over the time period during which errors were detected is to be reduced below a set threshold.
[0155] At operation 435, the method can generate the compositive video. The method can include one or more processors determining, using the reordered first one or more data packets and second one or more data packets, a metric. The method can be indicative of performance of the medical session. The one or more processors can generate, using the reordered first one or more data packets and second one or more data packets, a composite video of at least a portion of the medical session with an indication of the metric.
[0156] Performance analyzer can analyze the performance of a surgeon of an account associated with the medical session for the medical procedure represented by the data streams. The performance can be determined using the data streams. The performance can be affected by the quality of the data streams. Performance analyzer can generate a confidence score for the metrics based at least on the presence or absence of the errors in the data streams. For the duration of the data stream during which unaddressed errors have occurred or persisted, the performance analyzer can reduce the confidence score.
[0157] Composite video for the medical session can be generated from a plurality of video fragments. The composite video can include overlayered metrics and confidence scores. The metrics and the scores can vary over time, based on the data streams and the analysis. Confidence scores can vary based on the quality (e.g., absence of uncorrected errors) in the data streams. Confidence scores can be reduced during the time intervals for which a portion of a data stream is corrupted or missing. Composite videos can be displayed for users. Any one ormore of kinematics data, sensor data, events data, state data or mapping data can be provided, such as via overlayered indications, in the composite video.
[0158] FIG. 5 depicts a surgical system 500, in accordance with some embodiments. The surgical system 500 may be an example of the medical environment 102. The surgical system 500 may include a robotic medical system 505 (e.g., the robotic medical system 120), a user control system 510, and an auxiliary system 515 communicatively coupled one to another. A visualization tool 520 (e.g., the visualization tool 114) may be connected to the auxiliary system 515, which in turn may be connected to the robotic medical system 505. Thus, when the visualization tool 520 is connected to the auxiliary system 515 and this auxiliary system is connected to the robotic medical system 505, the visualization tool may be considered connected to the robotic medical system. In some embodiments, the visualization tool 520 may additionally or alternatively be directly connected to the robotic medical system 505.
[0159] The surgical system 500 may be used to perform a computer-assisted medical procedure on a patient 525. In some embodiments, surgical team may include a surgeon 530A and additional medical personnel 530B-530D such as a medical assistant, nurse, and anesthesiologist, and other suitable team members who may assist with the surgical procedure or medical session. The medical session may include the surgical procedure being performed on the patient 525, as well as any pre-operative (e.g., which may include setup of the surgical system 500, including preparation of the patient 525 for the procedure), and post-operative (e.g., which may include clean up or post care of the patient), or other processes during the medical session. Although described in the context of a surgical procedure, the surgical system 500 may be implemented in a non-surgical procedure, or other types of medical procedures or diagnostics that may benefit from the accuracy and convenience of the surgical system.
[0160] The robotic medical system 505 can include a plurality of manipulator arms 535 A- 535D to which a plurality of medical tools (e.g., the medical tool 112) can be coupled or installed. Each medical tool can be any suitable surgical tool (e.g., a tool having tissueinteraction functions), imaging device (e.g., an endoscope, an ultrasound tool, etc.), sensing instrument (e.g., a force-sensing surgical instrument), diagnostic instrument, or other suitable instrument that can be used for a computer-assisted surgical procedure on the patient 525 (e.g., by being at least partially inserted into the patient and manipulated to perform a computer- assisted surgical procedure on the patient). Although the robotic medical system 505 is shown as including four manipulator arms (e.g., the manipulator arms 535A-535D), in other embodiments, the robotic medical system can include greater than or fewer than fourmanipulator arms. Further, not all manipulator arms can have a medical tool installed thereto at all times of the medical session. Moreover, in some embodiments, a medical tool installed on a manipulator arm can be replaced with another medical tool as suitable.
[0161] One or more of the manipulator arms 535A-535D and / or the medical tools attached to manipulator arms can include one or more displacement transducers, orientational sensors, positional sensors, and / or other types of sensors and devices to measure parameters and / or generate kinematics information. One or more components of the surgical system 500 can be configured to use the measured parameters and / or the kinematics information to track (e.g., determine poses of) and / or control the medical tools, as well as anything connected to the medical tools and / or the manipulator arms 535A-535D.
[0162] The user control system 510 can be used by the surgeon 530A to control (e.g., move) one or more of the manipulator arms 535A-535D and / or the medical tools connected to the manipulator arms. To facilitate control of the manipulator arms 535A-535D and track progression of the medical session, the user control system 510 can include a display (e.g., the display 116 or 1130) that can provide the surgeon 530A with imagery (e.g., high-definition 3D imagery) of a surgical site associated with the patient 525 as captured by a medical tool (e.g., the medical tool 112, which can be an endoscope) installed to one of the manipulator arms 535A-535D. The user control system 510 can include a stereo viewer having two or more displays where stereoscopic images of a surgical site associated with the patient 525 and generated by a stereoscopic imaging system can be viewed by the surgeon 530A. In some embodiments, the user control system 510 can also receive images from the auxiliary system 515 and the visualization tool 520.
[0163] The surgeon 530A can use the imagery displayed by the user control system 510 to perform one or more procedures with one or more medical tools attached to the manipulator arms 535A-535D. To facilitate control of the manipulator arms 535A-535D and / or the medical tools installed thereto, the user control system 510 can include a set of controls. These controls can be manipulated by the surgeon 530A to control movement of the manipulator arms 535A- 535D and / or the medical tools installed thereto. The controls can be configured to detect a wide variety of hand, wrist, and finger movements by the surgeon 530A to allow the surgeon to intuitively perform a procedure on the patient 525 using one or more medical tools installed to the manipulator arms 535A-535D.
[0164] The auxiliary system 515 can include one or more computing devices configured to perform processing operations within the surgical system 500. For example, the one or more computing devices can control and / or coordinate operations performed by various other components (e.g., the robotic medical system 505, the user control system 510) of the surgical system 500. A computing device included in the user control system 510 can transmit instructions to the robotic medical system 505 by way of the one or more computing devices of the auxiliary system 515. The auxiliary system 515 can receive and process image data representative of imagery captured by one or more imaging devices (e.g., medical tools) attached to the robotic medical system 505, as well as other data stream sources received from the visualization tool. For example, one or more image capture devices (e.g., the image capture devices 110) can be located within the surgical system 500. These image capture devices can capture images from various viewpoints within the surgical system 500. These images (e.g., video streams) can be transmitted to the visualization tool 520, which can then passthrough those images to the auxiliary system 515 as a single combined data stream. The auxiliary system 515 can then transmit the single video stream (including any data stream received from the medical tool(s) of the robotic medical system 505) to present on a display (e.g., the display 116) of the user control system 510.
[0165] In some embodiments, the auxiliary system 515 can be configured to present visual content (e.g., the single combined data stream) to other team members (e.g., the medical personnel 530B-530D) who might not have access to the user control system 510. Thus, the auxiliary system 515 can include a display 540 configured to display one or more user interfaces, such as images of the surgical site, information associated with the patient 525 and / or the surgical procedure, and / or any other visual content (e.g., the single combined data stream). In some embodiments, display 540 can be a touchscreen display and / or include other features to allow the medical personnel 530A-530D to interact with the auxiliary system 515.
[0166] The robotic medical system 505, the user control system 510, and the auxiliary system 515 can be communicatively coupled one to another in any suitable manner. For example, in some embodiments, the robotic medical system 505, the user control system 510, and the auxiliary system 515 can be communicatively coupled by way of control lines 545, which can represent any wired or wireless communication link that can serve a particular implementation. Thus, the robotic medical system 505, the user control system 510, and the auxiliary system 515 can each include one or more wired or wireless communication interfaces, such as one or more local area network interfaces, Wi-Fi network interfaces, cellular interfaces,etc. It is to be understood that the surgical system 500 can include other or additional components or elements that can be needed or considered desirable to have for the medical session for which the surgical system is being used.
[0167] The herein described subject matter sometimes illustrates different components contained within, or connected with, different other components. It is to be understood that such depicted architectures are illustrative, and that in fact many other architectures can be implemented which achieve the same functionality. In a conceptual sense, any arrangement of components to achieve the same functionality is effectively “associated” such that the desired functionality is achieved. Hence, any two components herein combined to achieve a particular functionality can be seen as “associated with” each other such that the desired functionality is achieved, irrespective of architectures or intermedial components. Likewise, any two components so associated can also be viewed as being “operably connected,” or “operably coupled,” to each other to achieve the desired functionality, and any two components capable of being so associated can also be viewed as being “operably couplable,” to each other to achieve the desired functionality. Specific examples of operably couplable include but are not limited to physically mateable or physically interacting components or wirelessly interactable or wirelessly interacting components or logically interacting or logically interactable components.
[0168] With respect to the use of plural or singular terms herein, those having skill in the art can translate from the plural to the singular or from the singular to the plural as is appropriate to the context or application. The various singular / plural permutations can be expressly set forth herein for sake of clarity.
[0169] It will be understood by those within the art that, in general, terms used herein, and especially in the appended claims (e.g., bodies of the appended claims) are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” etc.).
[0170] Although the figures and description can illustrate a specific order of method steps, the order of such steps can differ from what is depicted and described, unless specified differently above. Also, two or more steps can be performed concurrently or with partial concurrence, unless specified differently above. Such variation can depend, for example, on the software and hardware systems chosen and on designer choice. All such variations are withinthe scope of the disclosure. Likewise, software implementations of the described methods can be accomplished with standard programming techniques with rule-based logic and other logic to accomplish the various connection steps, processing steps, comparison steps, and decision steps.
[0171] It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation, no such intent is present. For example, as an aid to understanding, the following appended claims can contain usage of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles “a” or “an” limits any particular claim containing such introduced claim recitation to inventions containing only one such recitation, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” or “an” should typically be interpreted to mean “at least one” or “one or more”); the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should typically be interpreted to mean at least the recited number (e.g., the bare recitation of “two recitations,” without other modifiers, typically means at least two recitations, or two or more recitations).
[0172] Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together, etc.). In those instances where a convention analogous to “at least one of A, B, or C, etc.” is used, in general, such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B.”
[0173] Further, unless otherwise noted, the use of the words “approximate,” “about,”“around,” “substantially,” etc., mean plus or minus ten percent.
[0174] The foregoing description of illustrative implementations has been presented for purposes of illustration and of description. It is not intended to be exhaustive or limiting with respect to the precise form disclosed, and modifications and variations are possible in light of the above teachings or can be acquired from practice of the disclosed implementations. It is intended that the scope of the invention be defined by the claims appended hereto and their equivalents.
Claims
CLAIMSWhat is claimed is:
1. A system, comprising: one or more processors, coupled with memory, to: receive, for a medical session with a robotic medical system, a first one or more data packets of a first stream of data at a first time; receive, for the medical session, a second one or more data packets of a second stream of data at a second time subsequent to the first time; identify a first feature in the first one or more data packets and a second feature in the second one or more data packets; detect, based at least on the first feature and the second feature, that the first one or more data packets and the second one or more data packets are out of order; reorder, responsive to the detection, the first one or more data packets and the second one or more data packets to cause the first one or more data packets to be subsequent to the second one or more data packets; determine, using the reordered first one or more data packets and second one or more data packets, a metric indicative of performance of the medical session; and generate, using the reordered first one or more data packets and second one or more data packets, a composite video of at least a portion of the medical session with an indication of the metric.
2. The system of claim 1, comprising the one or more processors to: identify that the first stream corresponds to a first of a data generated by a sensor of the robotic medical system, a data on kinematics of one or more instruments of the robotic medical system used during the medical session, and a data of an event at the robotic medical system; and identify that the second stream corresponds to a second one of the data generated by the sensor, the data on kinematics of the one or more instruments and the data on the event at the robotic medical system.
3. The system of claim 1, comprising the one or more processors to: detect, based at least on one of the first feature and the second feature, at least one of a delay or a jitter corresponding to at least one of the first stream and the second stream; anddetermine, responsive to the detection of the at least one of the delay or the jitter, a portion the robotic medical system causing the at least one of the delay or the jitter.
4. The system of claim 1, comprising the one or more processors to: identify the first feature comprising a first timestamp for the first one or more data packets and the second feature comprising a second timestamp for the second one or more data packets; identify, based at least one the first timestamp and the second timestamp, a time at which at least one of the first one or more data packets and the second one or more data packets are at least one of generated, transmitted, received or stored by at least a portion of the robotic medical system; and detect, based at least on the time, that the first one or more data packets and the second one or more data packets are out of order.
5. The system of claim 1, comprising the one or more processors to: identify that a likelihood of the second feature corresponding to an event of the robotic medical system exceeds a threshold based at least on the first feature, the first time and the second time; and detect, based at least on the likelihood exceeding the threshold, that the first one or more data packets and the second one or more data packets are out of order.
6. The system of any one of claims 1-5, wherein at least one of the first feature and the second feature indicate at least one of: an event , a task of a medical procedure of the medical session, a phase of the medical procedure, an object used in the medical procedure or a workflow of the medical procedure.
7. The system of claim 1, comprising the one or more processors to: identify, based at least on the reordered first one or more data packets and second one or more data packets, a data missing from at least one of the first stream or the second stream; and determine, based at least on the identified data, a confidence score for the metric; and provide the confidence score for display with the composite video.
8. The system of claim 1, comprising the one or more processors to: identify a machine learning (ML) model trained on a plurality of features of a pluralityof data packets of the robotic medical system; and detect that the first one or more data packets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the ML model.
9. The system of claim 1, comprising the one or more processors to: identify one or more machine learning (ML) models utilizing one or more neural networks trained on sensor data from a plurality of streams of sensors of a robotic medical system, events data corresponding to a plurality of events of a plurality of medical procedures and kinematics data corresponding to data on kinematics relating the plurality of medical procedures implemented on the robotic medical system; and detect that the first one or more data packets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the one or more ML models.
10. The system of claim 9, comprising the one or more processors to: identify a first ML model trained on the sensor data and corresponding to a first encoder and a first one or more classification blocks to process representations on the sensor data; identify a second ML model trained on the events data and corresponding to a second encoder and a second one or more classification blocks to process representations on the events data; identify a third ML model trained on the kinematics data and corresponding to a third encoder and a third one or more classification blocks to process representations on the kinematics data; and detect that the first one or more data packets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the first ML model, the second ML model and the third ML model.
11. The system of claim 1, comprising the one or more processors to: determine, using the reordered first one or more data packets and second one or more data packets, that a portion of at least one of the first stream and the second stream is missing; generate an alert to indicate that the portion of at least the one of the first stream and the second stream is missing; andoverlay the alert on the composite video.
12. The system of claim 1, comprising the one or more processors to: determine, responsive to the detection, that a portion of at least one of the first stream over a time interval and the second stream over the time interval is missing; and modify, prior to display of the composite video, the composite video to exclude a portion of the composite video corresponding to the time interval.
13. The system of claim 1, 11 or 12, comprising the one or more processors to: display, via a graphical user interface, the composite video generated using the reordered first one or more data packets and the second one or more data packets.
14. The system of claim 1, 11 or 12, comprising the one or more processors to: store, in a data repository, a data file comprising the composite video generated using the reordered first one or more data packets and the second one or more data packets.
15. A method, comprising: receiving, by one or more processors coupled with memory, a first one or more data packets of a first stream of data at a first time and a second one or more data packets of a second stream of data at a second time subsequent to the first time, the first stream and the second stream corresponding to a medical session implemented using a robotic medical system; detecting, by the one or more processors, based at least on a first feature of the first one or more data packets and a second feature of the second one or more data packets, that the first one or more data packets and the second one or more data packets are out of order; reordering, by the one or more processors responsive to the detection, the first one or more data packets and the second one or more data packets to cause the first one or more data packets to be subsequent to the second one or more data packets; determining, by the one or more processors, using the reordered first one or more data packets and second one or more data packets, a metric indicative of performance of the medical session; and generating, by the one or more processors using the reordered first one or more data packets and second one or more data packets, a composite video of at least aportion of the medical session with an indication of the metric.
16. The method of claim 15, comprising: identifying, by the one or more processors, that the first stream corresponds to a first of a data generated by a sensor of the robotic medical system, a data on kinematics of one or more instruments of the robotic medical system used during the medical session, and a data of an event at the robotic medical system; and identifying, by the one or more processors, that the second stream corresponds to a second one of the data generated by the sensor, the data on kinematics of the one or more instruments and the data on the event at the robotic medical system.
17. The method of claim 15, comprising: detecting, by the one or more processors based at least on one of the first feature and the second feature, at least one of a delay or a jitter corresponding to at least one of the first stream and the second stream; and determining, by the one or more processors responsive to the detection of the at least one of the delay or the jitter, a portion the robotic medical system causing the at least one of the delay or the jitter.
18. The method of claim 15, comprising: identifying, by the one or more processors, the first feature comprising a first timestamp for the first one or more data packets and the second feature comprising a second timestamp for the second one or more data packets; identifying, by the one or more processors, based at least one the first timestamp and the second timestamp, a time at which at least one of the first one or more data packets and the second one or more data packets are at least one of generated, transmitted, received or stored by at least a portion of the robotic medical system; and detecting, by the one or more processors, based at least on the time, that the first one or more data packets and the second one or more data packets are out of order.
19. The method of claim 15, comprising: identifying, by the one or more processors, based at least on the first feature, the first time and the second time, that a likelihood of the second feature corresponding to an event of the robotic medical system exceeds a threshold; anddetecting, by the one or more processors, based at least on the likelihood exceeding the threshold, that the first one or more data packets and the second one or more data packets are out of order, wherein at least one of the first feature and the second feature indicate at least one of an event , a task of a medical procedure of the medical session, a phase of the medical procedure, an object used in the medical procedure or a workflow of the medical procedure.
20. The method of claim 15, comprising: identifying, by the one or more processors, a machine learning (ML) model trained on a plurality of features of a plurality of data packets of the robotic medical system; and detecting, by the one or more processors, that the first one or more data packets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the ML model.
21. The method of claim 15, comprising: identifying, by the one or more processors, one or more machine learning (ML) models utilizing one or more neural networks trained on sensor data from a plurality of streams of sensors of a robotic medical system, events data corresponding to a plurality of events of a plurality of medical procedures and kinematics data corresponding to data on kinematics relating the plurality of medical procedures implemented on the robotic medical system; and detecting, by the one or more processors, that the first one or more data packets and the second one or more data packets are out of order based at least on the first feature and the second feature input into the one or more ML models.
22. The method of claim 15, comprising: determining, by the one or more processors, responsive to the detection, that a portion of at least one of the first stream over a time interval and the second stream over the time interval is missing; and modifying, by the one or more processors, prior to display of the composite video, the composite video to exclude a portion of the composite video corresponding to the time interval.
23. A non-transitory computer-readable medium storing processor executable instructions, that when executed by one or more processors, cause the one or more processors to: receive, for a medical session with a robotic medical system, a first one or more data packets of a first stream of data at a first time and a second one or more datapackets of a second stream of data at a second time subsequent to the first time; identify one or more machine learning (ML) models including one or more neural networks trained on sensor data from a plurality of streams of sensors of a robotic medical system, events data corresponding to a plurality of events of a plurality of medical procedures and kinematics data corresponding to data on kinematics relating the plurality of medical procedures implemented on the robotic medical system; detect, based at least on a first feature in the one or more data packets and a second feature in the second one or more data packets input into the one or more ML models, that the first one or more data packets and the second one or more data packets are out of order; reorder, responsive to the detection, the first one or more data packets and the second one or more data packets to cause the first one or more data packets to be subsequent to the second one or more data packets; determine, using the reordered first one or more data packets and second one or more data packets, a metric indicative of performance of the medical session; and generate, using the reordered first one or more data packets and second one or more data packets, a composite video of at least a portion of the medical session with an indication of the metric.
Citation Information
Patent Citations
Operating room black-box device, system, method and computer readable medium for event and error prediction
US20220270750A1