Surgical instrument presence detection with noise-labeled machine learning
Patent Information
- Application Number
- CN202480087779.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-12-17
- Publication Date
- 2026-09-08
AI Technical Summary
随着手术室中设备的数量和种类的增加,或医疗程序变得越来越复杂,高效、可靠或无事故地执行此类医疗程序可能具有挑战性
Smart Images

Figure CN122720003A_ABST
Abstract
Description
[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 611,636, filed December 18, 2023, pursuant to 35 USC §119, the entire contents of which are hereby incorporated by reference. Background Technology
[0002] Medical procedures can be performed in the operating room. However, as the number and types of equipment in the operating room increase, or as medical procedures become more complex, performing such procedures efficiently, reliably, or without accidents can become challenging. Summary of the Invention
[0003] This disclosure provides a noisy label-tolerant machine learning (ML) model for detecting objects in robotic procedure videos. For example, it provides a framework for identifying surgical instruments captured in surgical videos using an ML model trained on a noisy label dataset based on robotic system instrument installation logs, instead of using a supervised learning process. Therefore, the solution can scale up the available training dataset to provide highly accurate identification of surgical instruments with reduced training resources. For instance, it can provide a spatiotemporal transformer neural network model that is tolerant to noisy label datasets and identifies and tags medical instruments captured in robotic surgery videos.
[0004] At least one aspect of the technical solution relates to a system. The system may include one or more processors coupled to memory. The one or more processors can identify a series of frames from a video of a medical procedure captured by a robotic medical system. The one or more processors can identify a model trained at least in part on frames from a plurality of videos captured for the one or more medical procedures, the frames of which are labeled with data identifying the installation of one or more instruments of the one or more robotic medical systems. The one or more processors can use the model to determine a frame-by-frame label for each frame in the series of frames, the frame-by-frame label indicating the probability of the presence of one or more types of instruments. The one or more processors can display an indication of the presence of such a type of instrument via a graphical user interface, the indication being at least in part based on timestamps in the video on the frame-by-frame labels of the series of frames determined by the model.
[0005] One or more processors can be configured to receive a set of frames from multiple videos captured for one or more medical procedures. The one or more processors can identify a label for the set of frames based on the last frame in the set. The label may include data indicating the installation time of one or more devices and the last timestamp of the last frame. The label can be used to train a model.
[0006] One or more processors can be configured to identify multiple labels for frames of multiple videos. Each of the multiple labels can include a vector of one or more values corresponding to one or more devices. One or more processors can be configured to train a model using the multiple labels.
[0007] One or more processors can be configured to determine multiple labels for frames of multiple videos, each label having a value indicating whether one or more instruments are mounted at one or more robotic medical systems at time in each corresponding frame of the video. One or more processors can be configured to train a model using the multiple labels.
[0008] One or more processors may be configured to identify one or more logs from one or more robotic medical systems. Each log in the one or more logs indicates the installation time of one or more instruments for a corresponding video in a plurality of videos. The one or more processors may be configured to assign a label from a plurality of tags indicating the installation time to each frame in the plurality of videos, the installation time being derived from the corresponding log in the one or more logs corresponding to the corresponding video in the plurality of videos.
[0009] The model may include a transformer neural network model that applies first one or more weights to one or more spatial dimensions of an image region within a frame of multiple videos, and applies second one or more weights to a temporal dimension of a set of frames across the multiple videos. One or more processors may be configured to generate heatmaps based at least on frames of the multiple videos, the heatmaps indicating regions within a subset of frames where the probability of the presence of a device of that type exceeds a threshold of the heatmap.
[0010] One or more processors may be configured to identify a second region within a subset of frames, based on at least a plurality of video frames, where the probability of the presence of such a device exceeds a second threshold, which exceeds a first threshold. The one or more processors may be configured to display at least two of the subset of frames, a heatmap, the first region, and the second region via a graphical user interface.
[0011] One or more processors may be configured to compare the probability of presence of each corresponding frame in a series of frames with a threshold for the presence of that type of device. One or more processors may be configured to determine, at least in part, a second per-frame label for each corresponding frame in the series of frames based on this comparison, the second per-frame label indicating the presence of that type of device at the robotic medical system.
[0012] One or more processors can be configured to identify a subset of frames from a series of frames in a video that correspond to a portion of the video capturing an instrument of that type used in a medical procedure. One or more processors can be configured to determine a corresponding per-frame label for the subset of frames based on the subset of frames input into the model.
[0013] One or more processors may be configured to receive a file from a robotic medical system, the file including an indication of the time at which one or more instruments are installed at the robotic medical system. One or more processors may be configured to determine a corresponding per-frame label for at least one frame in the frames, based at least on the indication of the time input into the model.
[0014] One or more processors may be configured to generate a series of per-frame labels based at least on per-frame labels for each frame in a series of frames. One or more processors may be configured to adjust the value of a first per-frame label of a first frame in a series of frames using at least a second value of a second per-frame label of a second frame adjacent to the first frame.
[0015] One or more processors can be configured to determine the label for each frame using a model based at least on the installation time of one or more instruments at one or more robotic medical systems. One or more processors can be configured to determine the probability of presence based on a comparison of timestamps and installation times.
[0016] One or more processors can be configured to identify a device of that type based at least on a probability of presence exceeding a threshold for that type of device. One or more processors can be configured to display an indication that the device of that type has been identified. The indication can be overlaid on a subset of a series of frames displayed on a graphical user interface, the subset of which has a probability of presence exceeding the threshold for that type of device.
[0017] At least one aspect of the technical solution relates to a method. The method may include one or more processors coupled to a memory that uses data identifying the installation of one or more types of instruments in one or more robotic medical systems to label frames of multiple videos captured for one or more medical procedures. The method may include one or more processors using the labeled frames to train a model. The method may include one or more processors using the model to determine a frame-by-frame label for each frame in a series of frames. The frame-by-frame label may indicate the probability of the presence of one or more types of instruments. The method may include one or more processors displaying an indication of the presence of one type of instrument via a graphical user interface.
[0018] This method may include one or more processors determining the presence of a device of that type based at least in part on timestamps in the video on frame labels of a series of frames determined via a model. The method may include one or more processors receiving a set of frames from multiple videos captured for one or more medical procedures. The method may include one or more processors identifying a label for the frame set based on the last frame in the frame set. The label may include data indicating the installation time of one or more devices and the last timestamp of the last frame. The method may include one or more processors using the labels of the frame set to train a model.
[0019] At least one aspect of the technical solution relates to a non-transitory computer-readable medium storing processor-executable instructions, which, when executed by one or more processors, cause the one or more processors to identify a series of frames from a video of a medical procedure captured by a robotic medical system. When executed by the one or more processors, the instructions can cause the one or more processors to identify a model trained at least in part on frames from a plurality of video captured for the one or more medical procedures, the frames of which are labeled with data identifying the installation of one or more instruments of the one or more robotic medical systems. When executed by the one or more processors, the instructions can cause the one or more processors to use the model to determine a frame-by-frame label for each frame in the series of frames, the frame-by-frame label indicating the probability of the presence of one or more types of instruments. When executed by the one or more processors, the instructions can cause the one or more processors to display an indication of the presence of such an instrument via a graphical user interface, the indication being at least in part based on timestamps in the video on the frame-by-frame labels of the series of frames determined by the model. The indication may overlay a subset of the series of frames displayed on the graphical user interface. The probability of the presence of this subset of the series may exceed a threshold for such an instrument.
[0020] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of various aspects and implementations, and provide an overview or framework for understanding the nature and characteristics of the claimed aspects and implementations. The accompanying drawings provide illustrations and further understanding of the various aspects and implementations, and are incorporated in and constitute a part of this specification. The foregoing information and the following detailed description and drawings include illustrative examples and should not be considered limiting. Attached Figure Description
[0021] The accompanying drawings are not intended to be drawn to scale. The same reference numerals and names in the various drawings indicate the same elements. For clarity, not every part may be labeled in every drawing. In the drawings: Figure 1 An example system for generating and deploying noisy label-tolerant ML models to detect objects in videos of robotic programs is described.
[0022] Figure 2 The illustration shows an example of a graphical user interface that provides instructions for capturing and identifying video frames processed by an ML model of a medical device.
[0023] Figure 3 The illustration shows an example system configuration for generating and deploying noise-label-tolerant ML models to detect the presence of instruments and spatial context.
[0024] Figure 4 The diagram illustrates an example flowchart of a method for generating and using a noise-label-tolerant ML model to detect objects in a robot program video.
[0025] Figure 5 The illustration shows an example of a surgical system based on some aspects of the technical solution.
[0026] Figure 6 The illustration shows an example block diagram of an example computer system according to some aspects of the technical solution. Detailed Implementation
[0027] The following is a more detailed description of the various concepts and implementations of systems, methods, and devices for detecting the presence of surgical instruments using machine learning with noise labels. The various concepts introduced above and discussed in more detail below can be implemented in any of a number of ways.
[0028] Although this disclosure is discussed in the context of surgical procedures, the technical solutions of this disclosure can be applied in various aspects to other medical treatments, sessions, environments, or activities, as well as non-medical activities in which object-based procedural identification is desired. For example, the technical solutions can be applied to any environment, application, or industry in which activities, operations, processes, or actions are performed using tools or instruments that can be captured on video, and for such environments, applications, or industries, ML modeling can be used to identify or distinguish tools or instruments in the video.
[0029] Training ML models to identify surgical instruments from recorded procedural videos can be challenging and relies on large, manually annotated training datasets and supervised learning. However, training ML models with such datasets can be time-consuming and computationally intensive. Furthermore, such training processes may fail to represent the full diversity of scenarios the model might encounter and are prone to introducing annotator bias, all of which can negatively impact model performance. Training ML models with large or manually annotated datasets can be difficult to modify or scale up because the annotation work on large datasets can introduce additional time delays, adversely affecting the system's ability to be promptly corrected through data retraining.
[0030] However, in robotic medical systems, system logs can be used to record surgical instrument installation events, capturing instances where surgical instruments are installed in the robotic arm of the robotic surgical system. Using such installation logs in combination with surgical videos to provide training samples for machine learning models can be challenging because the system logs may not match the timing of instrument appearance in the video recordings (e.g., there may be a temporal offset). For example, there may be a delay between the time of surgical instrument installation in the installation log and the moment the installed instrument appears in the video frame. Such delays can vary depending on the circumstances, making it difficult to estimate the duration of the delay. Furthermore, video recordings of medical procedures may include instances where the visibility or appearance of surgical instruments is affected, such as visual obstructions from the camera's perspective or occlusion (e.g., by another object, the patient's body, or the surgeon's hand). Therefore, instrument installation logs can introduce noise into the dataset, potentially leading to noisy labeling of the training dataset, including, for example, missing labels or mislabeling of data.
[0031] This technical solution overcomes these challenges by providing a noise-label-tolerant ML neural network model trained to detect surgical instruments in robotic surgical videos using a noisy-labeled dataset with video and instrument installation logs. This solution can be implemented using a pipeline for training an ML model with multiple training phases. For example, the input framing phase may include discretizing the surgical video into consecutive frame images. The frame rate can be varied, and the number of images used to formalize the input batch can be selected, such that the trained video clips capture a sufficient amount of temporal information (e.g., sufficient duration) to provide adequate surgical context (e.g., the surgical task performed in the clip). Labels associated with image fragments, clips, or batches may include vectors containing several (n) binary elements representing several (n) classes of instruments. The label vector may indicate the presence of an instrument in the last frame of an image clip or batch. The binary labels may come from or be created using system tool installation logs and therefore do not include any manual annotations or markings.
[0032] Technical solutions may include a spatiotemporal transformer neural network model, which may include an ML core of a type of neural network employing sequential frames. This model can perform numerical operations on sequential inputs and transform the operations into compact feature vectors representing the inputs. More specifically, the model can process or determine any spatial and temporal relevance of the representative semantics of identifying surgical events by applying an "attention" mechanism. The attention mechanism may include selective weighting techniques, which can apply different weights to emphasize different parts of the information in the data, thereby identifying the most efficient compact representation for completing the task. The same weighting process may allow for the inclusion of any errors in the labeling (e.g., noise) during the training process. For example, the model's attention mechanism can use weight allocation to focus more on relevant spatial and temporal features while downplaying the impact of noisy or mislabeled data. For example, the model can prioritize the informational parts of the data and reduce the impact of less relevant or erroneous information. The attention mechanism may include weighting capabilities in both spatial and temporal dimensions. ML model architectures may include, for example, neural network architectures that process and understand 3D visual data using transformer-based models, or non-convolutional methods for video classification using spatial and temporal self-attention.
[0033] The technical solution may include a classification function that treats the identification of surgical instruments in a video as a classification problem. The classification function can be designed or configured to regress feature vectors to a length (n) vector. The classification function may include a learning objective that compares elements with data labels and minimizes the difference between them. The technical solution may include a model training module that can acquire training data from a database and process the data in a format compatible with the neural network module configuration.
[0034] During the inference or deployment phase, the learned configuration and parameters of the neural network model can be transferred to the processing unit. The inference program video can be discretized in the same or similar manner as the video in the training phase and can be fed into the ML model. The output from the model can include a numerical vector of length n, representing the probability of the presence of each type of device. By comparing these probabilities with a threshold used to determine the presence of each type of device, the vector of length n can be binarized at each element, thus indicating the presence of the corresponding device.
[0035] The technical solution allows for the detection or identification of surgical instruments on each inference frame. Smoothing post-processing can be used and applied to the temporal sequence labels of the image sequence. The labels present in each frame can then be converted into instrument tags, recording the instrument type and start or end time.
[0036] Benefiting from the attention mechanism embedded in the neural network structure, this technique can spatially utilize "attention" in the form of a heatmap during the inference phase. A heatmap can include hot zones on image frames of a video clip corresponding to the most likely location of the device. For example, a heatmap can be helpful in providing additional spatial information to label images, such as distinguishing the left or right position or end of a device.
[0037] The technical solution may include components in a neural network model, which may include different implementations. For example, alternatively or additionally, the instrument presence recognition model may be trained using any combination of labeled data based on noisy system logs or manually annotated tags. For example, a detailed spatiotemporal neural network may be configured differently depending on the use case, such as convolutional and recurrent neural network architectures, two-stream network architectures with separate streams for spatial and temporal processing, or graph neural networks. By doing so, the technical solution can train an accurate surgical instrument presence recognition model with the aid of a large number of noisy labels. Visual attention may correspond to an active region in the spatial-temporal dimension, which can indicate the spatial location of the identified instrument. For example, the technical solution may provide a heatmap as a highlighted presence area by visualizing the model attention, thereby providing the spatial context or indication of the image in the user interface. For example, the technical solution may include post-processing of the visualized heatmap, which may facilitate the differentiation of instruments mounted on the left or right arm.
[0038] Technical solutions may include user interface or user experience functionality, where identified devices can be displayed in a timeline showing their presentation time (e.g., start and end times) and duration. Such functionality can provide support to guide users in checking and reviewing device usage status. The user interface or user experience functionality may include a visual procedural video with a highlighted heatmap around the device and recognition results, which may include a label on the left or right arm of the identified device. Device identification may include detailed performance evaluations, individual status, and confidence scores.
[0039] Figure 1 An example system 100 is depicted for generating and deploying noise-label-tolerant ML models to detect objects in a video of a robotic procedure. Example system 100 may include a robotic system for performing tasks using tools or instruments, such as a robotic medical system 120 used by a surgeon to perform surgery on a patient. The robotic medical system 120 (also referred to as RMS 120) may be deployed in a medical environment 102. Medical environment 102 may include any space or facility for performing medical procedures, such as a surgical facility or operating room. Medical environment 102 may include medical devices 112 that the RMS 120 may use to perform surgical patient procedures, whether invasive, non-invasive, inpatient, or outpatient.
[0040] Medical environment 102 may include one or more data capture devices 110 (e.g., optical devices, such as cameras or sensors, or other types of sensors or detectors) for capturing data streams 162 (e.g., images or videos of surgery). Medical environment 102 may include one or more visualization tools 114 to acquire and process the captured data streams 162 for display to a user (e.g., a surgeon or other medical professional) at one or more displays 116. Displays 116 may present data streams 162 (e.g., images or video frames) of medical procedures (e.g., surgical procedures) being performed by a robotic medical system 120, which handles, manipulates, holds, or otherwise utilizes medical instruments 112 to perform surgical tasks at a surgical site. RMS 120 may include installation data 122, which may include system logs indicating the installation time of medical instruments 112 on various manipulator arms of the robotic medical system 120. Data processing system (DPS) 130 may be coupled to RMS 120 via network 101. DPS 130 may include one or more machine learning (ML) trainers 140, data storage 160, processing functions 170, and interfaces 180.
[0041] The machine learning (ML) trainer 140 may include or generate an instrument model 144, which can be trained using a training dataset 142, which may include video frames 164 from RMS 120 and installation data 122. The ML trainer 140 can use the training dataset 142 to label the video frames 164 with labels 148 and use weights 146 to improve the performance of the instrument model 144 to more accurately detect and identify instrument predictions 152 based on an attention mechanism 154.
[0042] The data storage 160 of the DPS 130 may include one or more data streams 162, such as video frame streams 164. Data streams 162 may include measurements or sensor data, such as force, torque, or biometric data, haptic feedback data, endoscopic images or data, ultrasound images or video, or communication and command data streams. The data storage 160 may include installation data, such as system files or logs, including timestamps and data regarding the installation, activation, calibration, or use of a particular medical device 112.
[0043] Processing functionality 170 may include functionality for processing data, including, for example, functionality for generating a heatmap 172 and performing frame smoothing 174. Heatmap 172 may include hot zones or highlighted areas in video frame 164, where medical device 112 is determined by device model 144 to be located within video frame 164 of data stream 162. Frame smoothing 174 may include correcting labels 148 of video frame 164 based on labels 148 surrounding other frames of a given video frame 164. Interface 180 may include, for example, a graphical user interface for providing indications 182, which may indicate features such as label 148, device prediction 152, or device 112 identified in heatmap 172.
[0044] System 100 may include one or more data capture devices 110 (e.g., video cameras, sensors, or detectors) for collecting any data stream 162, which can be used for machine learning and detection of objects such as medical devices or tools. Data capture devices 110 may include cameras or other image capture devices for capturing video or images from specific viewpoints within the medical environment 102. Data capture devices 110 may be positioned, mounted, or otherwise positioned to capture content from any viewpoint that facilitates the data processing system's capture of various surgical tasks or actions.
[0045] The data capture device 110 may include any of the following: a variety of sensors, cameras, video imaging devices, infrared imaging devices, visible light imaging devices, intensity imaging devices (e.g., black and white, color, grayscale imaging devices, etc.), depth imaging devices (e.g., stereo imaging devices, time-of-flight imaging devices, etc.), medical imaging devices (such as endoscopic imaging devices, ultrasound imaging devices, etc.), non-visible light imaging devices, any combination or sub-combination of the above imaging devices, or any other type of imaging device suitable for the purposes described herein. The data capture device 110 may include a camera that a surgeon can use to perform surgery and observe manipulating components within a field of view suitable for the performance of a given task.
[0046] Data capture device 110 can capture, detect, or acquire sensor data, such as video or images, including, for example, still images, video images, vector images, bitmap images, other types of images, or combinations thereof. Data capture device 110 can capture images at any suitable predetermined capture rate or frequency. Settings of each data capture device 110 (such as scaling settings or resolution) can be varied as needed to capture a suitable image from any viewpoint. For example, data capture device 110 may have a fixed viewpoint, position, orientation, or orientation. Data capture device 110 may be portable or otherwise configured to change orientation or extend and retract in various directions. Data capture device 110 may be part of a multi-sensor architecture comprising multiple sensors, each configured to detect, measure, or otherwise capture a specific parameter (e.g., sound, image, or pressure).
[0047] Data capture device 110 may include any type and form of sensor, such as a positioning sensor, biometric sensor, velocity sensor, acceleration sensor, vibration sensor, motion sensor, pressure sensor, light sensor, distance sensor, current sensor, focus sensor, temperature or pressure sensor, or any other type and form of sensor used to provide data about medical tool 112 or data capture device (e.g., optical device). For example, data capture device 110 may include a position sensor, distance sensor, or positioning sensor that provides the coordinate position of medical tool 112 or data capture device 110. Data capture device 110 may include sensors that provide information or data about the position, orientation, or spatial orientation of an object (e.g., a lens of medical tool 112 or data capture device 110) relative to a reference point. A reference point may include any fixed, defined location used as a starting point for measuring distance and orientation in a specific direction, serving as an origin from which all other points or locations can be determined.
[0048] Display 116 may show, illustrate, or play data stream 162 (e.g., video frame 164) showing medical tools 112 at or near a surgical site. For example, display 116 may display a rectangular image of the surgical site (e.g., video frame 164) and at least a portion of the medical tools 112 (e.g., instruments) used to perform surgical tasks. Display 116 may provide compiled or synthesized images generated by visualization tools 114 from multiple data capture devices 110 to provide visual feedback from one or more viewpoints.
[0049] The visualization tool 114 can be configured or designed to receive any number of different data streams 162 from any number of data capture devices 110 and combine them into a single data stream displayed on the display 116. The visualization tool 114 can be configured to receive multiple data stream components and combine them into a single data stream 162. For example, the visualization tool 114 can receive visual sensor data about the surgical site or the area where surgery is performed from one or more medical instruments 112, sensors, or cameras. The visualization tool 114 can incorporate, combine, or utilize multiple types of data (e.g., positioning data of the medical instrument 112 along sensor readings of pressure, temperature, vibration, or any other data) to generate output to be presented on the display 116. The visualization tool 114 can present the position of the medical instrument 112 as well as the position of any reference points or surgical sites, including the position of anatomical parts of the patient (e.g., organs, glands, or bones).
[0050] Medical tool 112 can be any type and form of tool or instrument used in surgery, medical procedures, or operating rooms or environments. Medical tool 112 can be imaged by, associated with, or include an image capturing device. For example, medical tool 112 can be a tool for cutting, a tool for suturing wounds, an endoscope for visualizing organs or tissues, an imaging device, needles and sutures for stitching wounds, a scalpel, forceps, scissors, a retractor, a grasper, or any other tool or instrument used during surgery. Medical tool 112 can include hemostats, cannulas, surgical drills, suction devices, or any other instruments used during surgery. Medical tool 112 can include other or additional types of therapeutic or diagnostic medical imaging devices. Medical tool 112 can be configured to be mounted in, coupled to, or manipulated by RMS 120, such as through a manipulator arm or other component for holding, using, and manipulating the medical device or tool 112.
[0051] RMS 120 can be a computer-aided system configured to perform surgical or medical procedures or activities on a patient via or using one or more robotic components or medical tools 112 or with their assistance. RMS 120 may include any number of manipulator arms for grasping, holding, or manipulating various medical tools 112 and using the medical tools 112 controlled by the manipulator arms to perform computer-aided medical tasks.
[0052] Images captured by medical tool 112 (e.g., video images) can be sent to visualization tool 114. Robotic medical system 120 may include one or more input ports to receive direct or indirect connections to one or more assistive devices. For example, visualization tool 114 may be connected to RMS 120 to receive images from the medical tool when it is mounted in RMS 120 (e.g., on a manipulator arm used to support medical device 112). Visualization tool 114 can combine data stream components from data capture device 110 and medical tool 112 into a single combined data stream for presentation on display 116.
[0053] System 100 may include a data processing system 130. The data processing system 130 may be deployed in or associated with a medical environment 102, or it may be provided by a remote server or be cloud-based. The data processing system 130 may include an interface 180 designed, constructed, and operated to communicate with one or more components of system 100, including, for example, a robotic medical system 120, via a network 101. The data processing system 130 may be implemented using instructions stored in a memory location and processed by one or more processors, controllers, or integrated circuit systems. The data processing system 130 may include functional, computer code, or programs for executing or implementing the ML trainer 140 and the instrument model 144 to identify, discriminate, detect, or indicate the location of the medical device 112 in video frames 164 of a surgical video recording.
[0054] ML trainer 140 can be any combination of hardware and software for training ML models. ML trainer 140 can include frameworks or functionalities for training noise-label-tolerant machine learning models, such as neural network spatiotemporal attention mechanisms designed to detect medical device 112 or any other tool in a video (e.g., a video of robotic surgery). ML trainer 140 can utilize or leverage video recordings of robotic surgery performed with RMS 120, paired with label data derived from device installation logs (e.g., installation data 122) from RMS 120. ML trainer 140 can use installation data 122 to train ML models, such as device model 144, which may include noise or discrepancies, such as a time mismatch between the installation time of medical device 112 and the time when medical device 112 appears in the video file.
[0055] The ML trainer 140 may include an attention mechanism 154, which can be used to address the noise challenge in the data. The attention mechanism 154 may include a scheme or functionality that utilizes weights 146 to configure the instrument model 144 to selectively focus on certain portions of the input data (e.g., video frames 164), thereby assigning different levels of importance to each portion of the input data during the learning process. For example, the attention mechanism 154 may include a spatial-temporal attention mechanism 154 within a neural network architecture, which configures the model to selectively focus on relevant spatial and temporal features in the video data. By assigning weights 146 to different segments of the input video, the attention mechanism 154 allows the model to attenuate the influence of noise labels 148, emphasizing more reliable cues for more accurate instrument detection. In doing so, the attention mechanism 154 can mitigate or reduce the influence of noise in the training dataset 142, thereby enhancing the ability of the instrument model 144 to generalize and accurately predict previously unseen surgical videos.
[0056] For example, attention mechanism 154 can facilitate assigning higher weights 146 to specific parts of the surgical procedure where the device has a defined procedure, thereby reducing the impact of potential mislabeling during less obvious phases. For instance, attention mechanism 154 can assign higher weights 146 to portions of video frames 164 of the data stream 162 that include timestamped labels 148, with the timestamps falling within a timeframe corresponding to the installation time identified in the installation data 122 of the RMS 120. The result of such weight allocation could be a noise-tolerant ML neural network capable of accurately identifying the medical device 112. This training strategy demonstrates the adaptability of ML methods to the complexity of the noisy training dataset 142.
[0057] The medical device model 144 may include any kind or combination of machine learning architectures. For example, the medical device model 144 may include a support vector machine (SVM) which may be beneficial for predictions about class boundaries (e.g., anatomy, instruments, objects, actions, or any other), a random forest for classification and regression tasks, a decision tree for predicting trees about different decision points, a K-nearest neighbor (KNN) model that can use similarity measures to make predictions based on the characteristics of neighboring data points, a Naive Bayes function for probabilistic classification, a logistic regression or linear regression or gradient boosting model. The medical device model 144 may include neural networks, such as deep neural networks configured for hierarchical representations of features, convolutional neural networks (CNNs) for image-based classification and prediction as well as spatial relationships and hierarchies, recurrent neural networks (RNNs) and long short-term memory (LSTM) networks for determining structures and processes unfolding over time, or where medical images may be integrated with multimodal data of patient data or historical combinations.
[0058] Device model 144 may include or utilize a transformer or transformer-based architecture, such as a spatiotemporal transformer or a graphical neural network with a transformer, which may be configured to perform device prediction 152. Device prediction 152 may include any identification, prediction, determination, or discrimination (e.g., device type) of medical device 112 captured by video. The spatiotemporal transformer may facilitate the determination of heatmap 172 or highlight specific regions of interest in video frame 164 corresponding to the location in which medical device 112 is being identified or detected. For example, the transformer may be used for multimodal ensemble, where data streams 162 from multiple types of sources (e.g., data from various detectors, sensors, and cameras) may be combined for prediction. A spatiotemporal transformer neural network may be applied to video frame 164 to facilitate spatial relationships of features across different images or data sources (e.g., 110). Device model 144 may include any one or more machine learning (e.g., deep neural network) models trained on different datasets to learn to discern complex details of objects, devices, or tools, such as the edges or shape of the device or the specific device type.
[0059] The device model 144 may be stored in the data store 160 along with the training dataset 142, video frames 164, or installation data 122. The device model 144 may be trained, created, configured, updated, or otherwise provided by the ML trainer 140. The device model 144 may be configured to identify, predict, classify, categorize, or otherwise score various performance aspects. For example, the device model 144 may be configured to determine a confidence score regarding device prediction 152 (e.g., device type). For example, the confidence score may indicate a score representing the level of certainty or confidence that the device model 144 has relative to a particular device prediction 152 (e.g., a confidence percentage from 0 to 100%).
[0060] Device model 144 can be configured to perform device prediction 152 (e.g., prediction of any object the model is trained to identify). Device prediction 152 can include any determination, identification, recognition, or prediction of an object (such as medical device 112 (e.g., device type)) or any other object the model can be trained to identify. Device prediction 152 can include or correspond to label 148. Label 148 can be used to indicate the presence or absence of an identified or recognized object (e.g., medical device 112). For example, label 148 can be used to indicate the location of medical device 112 in video frame 164. Label 148 can indicate the probability of identifying an object (e.g., medical device 112) within video frame 164 of incoming (e.g., live streaming) video that can be input into device model 144 to determine device prediction 152. Label 148 can include a vector of multiple values, each of which can correspond to the probability of the presence, identification, or recognition of a particular medical device 112 (e.g., device type).
[0061] Instrument model 144 may include, for example, a deep learning model configured to identify, detect, or distinguish instrument prediction 152 for a specific medical or surgical instrument used in the procedure, such as any one or more of the following: scissors, needles, thread, scalpel, clamps, rings, bone screws, grippers, retractors, saws, clamps, imaging devices, or any other medical instrument 112 or tool used in the medical procedure. Instrument model 144 may be configured to detect or distinguish any tool or object, depending on the design, such as a machine tool, electrical or mechanical tool, robotic machine or feature, or any other object or device.
[0062] Data storage 160 may include one or more data files, data structures, arrays, values, or other information that facilitates the operation of data processing system 130. Data storage 160 may include one or more local or distributed databases and may include a database management system. Data storage 160 may include, maintain, or manage data stream 162. Data stream 162 may include, or be formed from, one or more of video streams, image streams, streams of sensor measurements, event streams, or kinematic streams. Data stream 162 may include data collected by one or more data acquisition devices 110 (such as a set of 3D sensors) from various angles or advantageous locations regarding procedural activities (e.g., points or areas of surgery).
[0063] Data stream 162 may include any data stream. Data stream 162 may include a video stream comprising a series of video frames 164. The video frames 164 may be formed or organized into video segments, such as video segments of approximately 1, 2, 3, 4, 5, 10, or 15 seconds. The video per second may include, for example, 30, 45, 60, 90, or 120 video frames 164 per second. Data stream 162 may include an event stream, which may include event data or information streams, such as groupings, that identify or communicate the status of the robotic medical system 120 or events occurring in association with the robotic medical system 120. For example, data stream 162 may include any portion of installation data 122, including information or data regarding the installation, unloading, calibration, setup, attachment, disassembly, or any other action performed by or on the RMS 120 of the medical device 112.
[0064] Data stream 162 may include data about events, such as indications of whether medical instrument or device 112 has been calibrated, adjusted, or the status of the robotic medical system 120, including a manipulator arm mounted on it. Event streams may include data about whether the robotic medical system 120 functioned fully during the procedure (e.g., without errors). For example, when medical instrument 112 is mounted on the manipulator arm of the robotic medical system 120, signals or one or more data packets indicating that medical instrument 112 has been mounted on the manipulator arm of the robotic medical system 120 may be generated. Signals may be recorded in the mounting data 122 along with a timestamp of the event.
[0065] Data stream 162 may include kinematic flow data, which may refer to or include data associated with one or more manipulator arms or medical instruments 112 (e.g., devices) attached to the manipulator arm, such as arm position or positioning. Data corresponding to medical instrument 112 may be captured or detected by one or more displacement transducers, orientation sensors, positioning sensors, or other types of sensors and devices to measure parameters or generate kinematic information. Kinematic data may include sensor data, as well as timestamps and indications of the type of medical instrument 112 associated with data stream 162.
[0066] Data storage 160 can store video frames 164. Video frames 164 can include a single still image extracted from an image sequence of a video file. Video frames 164 can represent a specific moment in time and can be identified by metadata including a timestamp. Video frames 164 can display the visual content of the video at a specific moment. For example, in a video file capturing a robotic surgical procedure, video frames 164 can depict a snapshot of the surgical task, illustrating the movement or use of medical devices 112, such as a robotic arm manipulating surgical tools inside a patient.
[0067] The data repository may store installation data 112. Installation data 122 of the RMS 120 may include any data or information recording the setup, calibration, configuration, or attachment of the RMS 120 to the medical device 112. Installation data 122 may include system or installation files that may include or indicate events related to the installation, attachment, calibration, connection, or setup of the medical device 112 to the manipulator arm of the RMS 120. For example, the installation data file may include one or more lists with timestamps and details indicating the timing (e.g., seconds, minutes, hours, or date) when the medical device 112 was attached, calibrated, or configured by the RMS 120.
[0068] Processing function 170 may include any combination of hardware and software for processing data or output from device model 144. Processing function 170 may include functionality or a framework for processing data generated or determined by device model 144 and may be used to refine and enhance the information determined by the model. Processing function 170 may provide post-processing adjustments or operate simultaneously with and together with device model 144. For example, processing function 170 may generate a heatmap 172 illustrating the predicted location of medical device 112 as determined by device model 144. Heatmap 172 may provide a visual representation highlighting areas where medical device 112 is predicted by device prediction 152 to be present in video frame 164. Processing function 170 may generate heatmap 172 relative to a specific threshold level corresponding to certain threshold levels of certainty or confidence that medical device 112 will be found at a given location. Heatmap 172 may include multiple layers corresponding to multiple threshold (e.g., confidence) levels that are met.
[0069] Processing function 170 may include a framework or functionality for generating or implementing post-processing frame smoothing 174 for the model output. Frame smoothing 174 may include a function or functionality for adjusting the values of features or characteristics (e.g., label 148 or instrument prediction 152) of a particular video frame 164 based on values of the same features or characteristics on previous or subsequent video frames 164. For example, frame smoothing function 174 may perform any correction across or given multiple frames to facilitate that a given video frame 164 is processed in accordance with the label determined by the model relative to other video frames 164. For example, frame smoothing function 174 may determine that multiple previous video frames 164 and multiple subsequent frames 164 have a specific determination (e.g., label 148 or instrument prediction 152), and that video frames 164 in between differ in this respect from all other video frames 164. In response to this determination, frame smoothing 174 may correct the instrument prediction 152 or label 148 of a given video frame 164 to make it conform to adjacent video frames 164. In doing so, frame smoothing 174 can improve performance by reducing noise or inconsistencies in predictions across frames, resulting in a more coherent and accurate depiction of the presence and movement of medical devices across a sequence of 164 video frames.
[0070] DPS 130 may include an interface 180 designed, constructed, and operated to communicate with one or more components of system 100 via network 101, including, for example, robotic medical system 120 or another device such as a client's personal computer. Interface 180 may include a network interface. Interface 180 may include or provide a user interface, such as a graphical user interface. A graphical user interface may include, for example, a window for displaying video frames 164. Interface 180 may provide data for presentation via a display (such as display 116) and may depict, describe, render, present, or otherwise provide instructions 182 indicating the determination (e.g., output) of device model 144, such as device prediction 152 or a tag 148 identifying medical device 112.
[0071] Data processing system 130 may interface with, communicate with, or otherwise receive or provide information to one or more components of system 100 (including, for example, robotic medical system 120) via network 101. Devices in data processing system 130, robotic medical system 120, and medical environment 102 may each include at least one logical device, such as a computing device having a processor for communication via network 101. Data processing system 130, robotic medical system 120, or client devices coupled to network 101 may include at least one computing resource, server, processor, or memory. For example, data processing system 130 may include multiple computing resources or processors coupled to memory.
[0072] Data processing system 130 may be part of or include a cloud computing environment. Data processing system 130 may include multiple logically grouped servers and is conducive to distributed computing technologies. The logical group of servers may be referred to as a data center, server cluster, or machine cluster. Servers may also be geographically distributed. A data center or machine cluster may be managed as a single entity, or a machine cluster may include multiple machine clusters. Servers within each machine cluster may be heterogeneous—one or more servers or machines may operate according to one or more types of operating system platforms.
[0073] Data processing system 130 or its components may include physical or virtual computer systems operatively coupled to or associated with medical environment 102. In some embodiments, data processing system 130 or its components may be coupled to or associated with medical environment 102 directly or indirectly via network 101 through an intermediate computing device or system. Network 101 may be any type or form of network. The geographical extent of the network can vary widely and may include body area networks (BANs), personal area networks (PANs), local area networks (LANs) (e.g., the Internet), metropolitan area networks (MANs), wide area networks (WANs), or the Internet. The topology of network 101 can take any form, such as point-to-point, bus, star, ring, mesh, tree, etc. Network 101 may utilize different technologies and protocol layers or stacks, including, for example, Ethernet protocol, Internet Protocol Suite (TCP / IP), ATM (Asynchronous Transfer Mode) technology, SONET (Synchronous Optical Networking) protocol, SDH (Synchronous Digital Hierarchy) protocol, etc. TCP / IP may include application layer, transport layer, Internet layer (including, for example, IPv6), or link layer. Network 101 can be a broadcast network, telecommunications network, data communication network, computer network, Bluetooth network, or other types of wired and wireless networks.
[0074] Data processing system 130 or its components may be located at least partially in or away from the surgical facility associated with medical environment 102. Elements of data processing system 130 or its components may be accessible via portable devices such as laptops, mobile devices, or wearable smart devices. Data processing system 130 or its components may include other or additional elements that may be considered desirable in performing the functions described herein. Data processing system 130 or its components may include, or be associated with, one or more components or functionalities of a computing device including, for example, one or more processors coupled to a memory that may store instructions, data, or commands for implementing the functionality of DPS 130 discussed herein.
[0075] In one aspect, the technical solution may include system 100, which may include one or more processors (e.g., 610) that can be coupled to a memory (e.g., 615 or 620). The memory 615 or 620 may store instructions, computer code, or data that enable one or more processors 620 to implement any functionality of DPS 130, including, for example, any functionality of ML trainer 140, instrument model 144, processing function 170, or interface 180. For example, instructions stored in memory 615 or 620 may configure or cause one or more processors to perform various operations or tasks of DPS 130.
[0076] One or more processors 610 can recognize a series of video frames 164 of a medical procedure captured by the robotic medical system 120. For example, the DPS 130 can receive a real-time stream of incoming video from an ongoing medical procedure performed by a surgeon via the RMS 120. The series of video frames 164 can include frames of video segments, which can include, for example, 30, 45, 60, 90, or 120 video frames 164 per second. The video segments can include a length sufficient to determine or identify the device prediction 152 (e.g., medical device 112) and the action or activity performed by the device prediction 152. For example, the video segments can be 1, 2, 3, 4, 5, or 10 seconds long.
[0077] One or more processors 610 can recognize an instrument model 144 trained at least in part on video frames 164 of multiple videos (e.g., videos of previously performed surgeries). Such videos can be captured for one or more (e.g., previously performed) medical procedures labeled with data (e.g., installation data 122) identifying the installation of one or more medical devices 112 of one or more robotic medical systems 120. For example, the training dataset 142 may include hundreds, thousands, or even tens of thousands of videos, which may last for many hours and include 30, 45, 60, 90, 120, or more than 120 video frames per second. The videos can be labeled with tags 148 indicating the installation data 122. The tags 148 may indicate, for example, the time (e.g., seconds, minutes, hours, and dates) when the medical device 112 is installed, attached, configured, set up, or otherwise used on the RMS 120.
[0078] One or more processors 610 use device model 144 to determine a frame label (e.g., 148) for each video frame 164 in a series of video frames 164. The frame label 148 may indicate the probability of the presence of one or more types of medical devices 112. For example, device model 144 may receive an input real-time video stream (e.g., 162) where video frames 164 are input into the model for the detection or identification of medical devices 112. Each video frame 164 may be processed or determined individually. In some implementations, multiple video frames 164 forming a video segment may have a single video frame 164 to be labeled. For example, a video segment spanning multiple video segments 164, such as 3 seconds, may include video segments 164 that can be selected as representative video frames 164 for labeling with label 148. Label 148 may indicate the probability of the presence (e.g., confidence score or level of certainty) of device prediction 152 of any particular medical device 112 among a plurality of medical devices 112 identified within a video fragment or video segment 164.
[0079] One or more processors 610 may facilitate or trigger the system 100 to display video clips or segments on a display (such as display 116). For example, the system 100 may display indication 182 via a graphical user interface (e.g., 180). Indication 182 may include or indicate the presence of a type of medical device 112. The presence of the medical device 112 may be determined at least in part based on timestamps in the video on frame labels 148 of a series of video frames 164 determined via device model 144.
[0080] One or more processors 610 may be configured to receive a set 164 of video frames 164 of multiple videos captured for one or more medical procedures. The video frames 164 may include, for example, video frames 164 from various cameras located at various locations within the medical environment 102. The video frames 164 may be captured by a medical device 112, such as, for example, an endoscope, which may include an endoscopic camera. The video frames 164 may be used together with corresponding installation data 122 (e.g., a system device installation log) for each video of events (e.g., installation, setup, configuration) of various medical devices 112, as a training dataset 142.
[0081] One or more processors 610 may be configured to identify a label 148 of the video frame set 164 based on the last video frame 164 of the video frame set 164. The label 148 may include data indicating the installation time (e.g., 122) of one or more medical devices 112 and the last timestamp of the last video frame of the video frame set 164. One or more processors may be configured to use the label 148 of the video frame set 164 to facilitate or trigger training of the device model 144.
[0082] One or more processors 610 may be configured to identify multiple labels 148 for video frames 164 of multiple videos. Each of the multiple labels 148 may have a vector of one or more values corresponding to one or more medical devices 112. For example, each value in the vector of label 148 may correspond to a probability or confidence level of the presence or identification of a particular medical device type (e.g., 112) in video frame 164 or video segment. One or more processors 610 may be configured to train a device model 144 using the multiple labels 148.
[0083] One or more processors 610 may be configured to determine multiple labels 148 for video frames 164 of a plurality of videos. Each of the multiple labels 148 may include a value indicating whether one or more medical devices 112 are installed at one or more robotic medical systems 120 at a given time in each of the corresponding video frames 164. One or more processors 610 may be configured to train a device model 144 using the multiple labels 148.
[0084] One or more processors 610 may be configured to identify one or more installation logs (e.g., 122) of one or more robotic medical systems 120. Each log in the one or more logs (e.g., 122) may indicate the installation time of one or more medical devices 112 for a corresponding video in a plurality of videos. The one or more processors 610 may be configured to assign a label 148 indicating the installation time from a corresponding log in the one or more logs corresponding to the corresponding video in the plurality of videos to each video frame 164 in the plurality of videos.
[0085] The device model 144 may include a transformer neural network model that can apply first or more weights 146 to one or more spatial dimensions of image regions within video frames 164 of a plurality of videos, and apply second or more weights to the temporal dimension of a set of video frames 164 of the plurality of videos. For example, when certain conditions are met, such as when a specific timing occurs or when a specific feature is detected, the weights 146 may emphasize the importance or priority of a region of video frame 164 relative to other regions.
[0086] One or more processors 610 may be configured to generate a heatmap 172 based on at least a plurality of video frames 164. The heatmap 172 may indicate regions within a subset of video frames 164 where the probability of presence of a medical device 112 of that type exceeds a threshold of the heatmap 172. For example, the heatmap 172 may include highlighted portions of video frames 164 where the probability of presence of medical device 112 exceeds a specific threshold (e.g., 75% or 90%). One or more processors 610 may be configured to identify second regions within a subset of video frames 164 where the probability of presence of a medical device 112 of that type exceeds a second threshold, which exceeds a first threshold, based on at least a plurality of video frames 164. The second threshold may correspond to a second heatmap 172 to be displayed. For example, the second region may be a region within a first region, and the second heatmap 172 may encompass a region within the first heatmap 172. The first or second region of the heatmap 172 may overlay on the video frame being displayed. For example, the second heatmap 172 may be highlighted with a darker or more prominent shadow than the first heatmap 172. The second region may include a deterministic or confidence level (e.g., a threshold) above a threshold of the first heatmap 172. One or more processors 610 may be configured to display a subset of frames and an overlay of the heatmap 172 via a graphical user interface, the heatmap 172 including any combination of the first and second regions.
[0087] One or more processors 610 may be configured to compare the probability of presence of each corresponding video frame 164 in a series of video frames 164 with a threshold for the presence of a medical device 112 of that type. One or more processors 610 may be configured to determine a second per-frame label 148 for each corresponding frame in the series of video frames 164, at least in part, based on this comparison. The second per-frame label 148 may indicate the presence of a medical device 112 of that type at the robotic medical system 120.
[0088] One or more processors 610 may be configured to identify, from a series of video frames 164, a subset of video frames 164 corresponding to a portion of the video captured of a medical device 112 of that type used in a medical procedure. One or more processors 610 may be configured to determine a corresponding per-video-frame 164 tag for the subset of video frames 164 based on the subset of video frames 164 input into the model.
[0089] One or more processors 610 may be configured to receive a file from the robotic medical system 120, the file including an indication of the time at which one or more medical devices 112 are installed at the robotic medical system 120. One or more processors 610 may be configured to determine a corresponding per-frame tag 148 for at least one video frame 164, based at least on the indication of the time input into the device model 144.
[0090] One or more processors 610 may be configured to generate a series of per-frame labels based at least on per-frame labels for each of a series of video frames 164. One or more processors 610 may be configured to adjust the value of a first per-frame label 148 of a first video frame 164 in the series of video frames 164 using at least a second value of a second per-frame label 148 of a second video frame 164 adjacent to the first video frame 164. One or more processors 610 may be configured to determine the per-frame labels using an instrument model 144 based at least on the installation time of one or more medical devices 112 at one or more robotic medical systems 120.
[0091] One or more processors 610 can be configured to determine the probability of presence based on a comparison of a timestamp and an installation time. One or more processors 610 can be configured to identify the medical device 112 of that type based at least on a probability of presence exceeding a threshold for that type of medical device 112. One or more processors 610 can be configured to display an indication that the medical device 112 of that type has been identified. This indication is overlaid on a subset of a series of video frames 164 displayed on a graphical user interface, the subset of which has a probability of presence exceeding the threshold for that type of medical device 112.
[0092] One aspect of the technical solution may involve a non-transitory computer-readable medium storing processor-executable instructions. The instructions, when executed by one or more processors, cause the processors to identify a series of frames from a video of a medical procedure captured by a robotic medical system. When executed by the processors, the instructions can identify a model trained at least in part on frames from a plurality of video captured for the one or more medical procedures, the frames of which are labeled with data identifying the installation of one or more instruments of the one or more robotic medical systems. When executed by the processors, the instructions can use the model to determine a per-frame label for each frame in the series of frames, the per-frame label indicating the probability of the presence of one or more types of instruments. When executed by the processors, the instructions can display an indication of the presence of that type of instrument via a graphical user interface, the indication being at least in part based on timestamps in the video on the per-frame labels of the series of frames determined by the model. The indication can overlay a subset of the series of frames displayed on the graphical user interface, the subset having an presence probability exceeding a threshold for that type of instrument.
[0093] Figure 2 An example 200 of a graphical user interface 202 is illustrated, providing instructions 182 for capturing and identifying video frames processed by an ML model of a medical device 112. The graphical user interface 202 may be an interface 180 in which one or more instructions 182 are provided or displayed to a user (e.g., a surgeon utilizing an RMS 120). The graphical user interface 202 may include a location or window for displaying one or more video frames 164, which may include frames from an input video file processed by the device model 144. The video frames 164 may be frames processed in real-time during an ongoing surgical procedure or correspond to previously recorded medical procedures.
[0094] The graphical user interface may include a heatmap (e.g., 172) indicator 182 that indicates the location of the medical device 112, as determined by the device model 144, within video frame 164. Video frame 164 may identify the manipulator arm 206 attached to the RMS 120 as the medical device 112 used in this example. The left manipulator arm 206A may be shown performing a task together with the right manipulator arm 206B. Duration 204 may include a window indicating the current time (e.g., a timestamp) of video frame 164 within the surgical video being viewed to the user. The graphical user interface may include options or buttons for user selection, such as segment 220 or the select device 222 button. Segment 220 may provide the user with additional information about the surgical segment being performed. Select device 222 may allow the user (e.g., a surgeon) to select a specific device to manipulate, hold, or move.
[0095] Indicator 182 may be shown in various formats, such as lines or bars indicating the presence or absence of a specific device type (e.g., manipulator arm 206 or any other medical device 112) relative to the duration of the video recording. Indicator 182 may indicate or identify the type of medical device 112 identified or present within a specific video frame 164 (e.g., a time portion of the video), such as through color-coded filling of the indicated lines or bars. Indicator 182 may include or identify phase 208, which may correspond to various stages of a medical procedure, allowing the user to select or scroll to the start of a specific phase. Indicator 182 may include or identify steps 210, such as specific tasks or steps in a medical procedure. Video timer 212 may include a time bar that allows the user to temporarily scroll between video frames 164 of the video, such as, for example, selecting a specific moment in the video by clicking a point along the time bar.
[0096] Figure 3 The illustration shows a system configuration 300 for generating and deploying a noise-label-tolerant ML model to detect the presence of a device 302 and a spatial context 304. The device presence 302 may include any output corresponding to the classification, discrimination, or identification of the medical device 112, whether it is the robotic manipulator arm 206 or any tool handled or manipulated by the arm. The device presence 302 may include a probabilistic output of the device's presence, or a deterministic output with or without an associated probability or confidence score. The spatial context 304 may include any information or data regarding the spatial location or orientation of the identified medical device or tool. For example, the spatial context 304 may be indicated relative to a reference point (e.g., a location in video frame 164 or medical environment 102) and may correspond to the location of the medical device 112 (e.g., device type) identified in an image or frame.
[0097] For example, a data stream 162 with input video frames 164 can be received by a data processing system 130. Data stream 162 may include video frames 164 of the input video and installation data 122 of the RMS 120, which indicates timestamped events such as the timed occurrence of instrument installation, unloading, attachment, detachment, calibration, or use. The data processing system 130 can be deployed on a server, computer, in the cloud, or across any number of devices or systems (e.g., as a distributed system). The DPS 130 can utilize or execute an instrument model 144, such as a spatiotemporal transform-based neural network model with weights implemented and applied to specific model features to provide an attention mechanism for analyzing the input data.
[0098] Using the data stream 162 input to the device model 144, the DPS 130 can provide outputs such as device prediction 152, which may include per-frame labels 148 indicating the presence of a specific device type. For example, device presence 302 may include values in the vector of per-frame labels 148 indicating whether a specific type of device exists in video frame 164. The device model 144 may also provide spatial context 304, which may include the location of the device type determined to be present. For example, spatial context 304 may include heatmap 172 indicating locations that can identify or highlight the locations where the device type is determined to be present.
[0099] Now go to Figure 4 The diagram illustrates an example flowchart of a method 400 for generating and using a noise-label-tolerant ML model to detect objects in a robot program video. Method 400 can be executed by a system having one or more processors that execute computer-readable instructions stored in memory. Method 400 can be executed, for example, by system 100 and according to... Figures 1-3 and Figures 5-6 Any features or techniques discussed may be used to perform this. For example, method 400 may be implemented by one or more processors 610 of a computing system 600 that execute non-transitory computer-readable instructions stored in memory (e.g., memory 615, 620, or 625) and uses data from data storage 160 (e.g., storage device 625).
[0100] Method 400 can be used to train an ML model using data from the device system logs of a robotic medical system to detect objects in a video recording procedure executed via the robotic medical system. At operation 405, the method can label video frames using noisy data. At operation 410, the method can train the ML model using the labeled frames. At operation 415, the method can determine a per-frame label for one or more input video frames. At operation 420, the method can determine whether the per-frame label exceeds a threshold. At operation 425, based on the operation at 420, the method can determine that a device is present. At operation 430, based on the operation at 420, the method can determine that a device is not present. At operation 435, the method can modify the presence determination according to post-processing. At operation 440, the method can display a video frame indicating the presence of a device.
[0101] At operation 405, the method may use noise data to label video frames. The noise data may include mismatches or differences between the labels of frames from system logs indicating the presence or installation of a particular instrument (e.g., instrument type) and image frames with the same labels that do not depict the specific instrument (e.g., instrument type). For example, the method may use a series of video frame labels to label video frames where the labels of frames in the labeled frames indicate the presence of an instrument of that type in that frame, and the frame does not contain an image of that type of instrument.
[0102] This method may include using data from system logs (e.g., device installation and usage logs) to tag video frames in a training dataset to train an ML model for recognizing and detecting medical devices. For example, a data repository may store one or more training datasets comprising multiple videos of various procedures performed by a robotic medical system using medical devices. The training datasets may include installation data, which may include installation files or logs of various medical devices used in any medical procedure captured on multiple videos used to train the ML model.
[0103] This method may include identifying one or more labels (e.g., multiple labels) for frames of multiple videos used to train an ML model. Each of the multiple labels may include or correspond to a vector of one or more values among the values corresponding to one or more devices. Each label may correspond to one or more video frames among the one or more video frames used to train the ML model. For example, a label may correspond to a video segment that provides multiple seconds of video and includes multiple video frames. The label may indicate the presence or absence of a medical device (e.g., missing). For example, a label may include a vector with multiple entries, each entry corresponding to a medical device among the multiple medical devices. For each medical device, the label may identify whether the medical device is present or not present in the video in its corresponding values.
[0104] This method may include identifying a tag for the frame set based on the last frame in the frame set. The tag may include data indicating the installation time of one or more devices. The tag may include the last timestamp of the last frame in the frame set. The tag may indicate events and their timing, such as the installation, unloading, configuration, deconfiguration, attachment, detachment, movement, or use of a given medical device in a robotic medical system.
[0105] For example, the method may include determining multiple labels for frames in a plurality of videos used to train an ML model. Each of the multiple labels may include a value indicating, at time of each corresponding frame in the video frames used to train the ML model, whether one or more instruments are installed, configured, attached to one or more robotic medical systems, or otherwise used by one or more robotic medical systems.
[0106] At operation 410, the method can train an ML model using labeled frames. The method may include training an ML model for detecting, recognizing, and distinguishing medical devices from video data using a training dataset of video frames and installation data (e.g., recorded events of installation, configuration, attachment, or use of a medical device). The method may include, for example, recognizing a model trained at least in part on frames from multiple videos captured for one or more medical procedures using data labeling. The data may include, for example, information recognizing the installation, configuration, attachment, calibration, or use of one or more devices in one or more robotic medical systems. The data may include noisy data, such as data from the labels of frames taken from the system logs of the robotic system. Such data may include information about mismatches or discrepancies between timestamps indicating the presence or installation of a particular device (e.g., device type) on a robotic arm and image frames in which a particular device (e.g., device type) may not be shown or displayed. For example, the method may use labels on a series of video frames to train an ML model, where some labeled frames indicate the presence of a device of that type in the frame, while those specifically labeled frames do not include images of devices of that type indicated in the labels. Such mismatches or noise in the data can be overcome using a neural network model based on a spatiotemporal transformer, where specific weights can be applied to data portions that are more important than other data.
[0107] Data processing systems can utilize any ML architecture or framework to train ML models. For example, ML models can include transformer neural network models, graph neural network models, or any other attention-based machine learning models. ML models can be configured or trained to include, utilize, or apply first or more weights to a region of an image within one of (one or more) frames of multiple videos, in a spatial dimension. ML models can be configured or trained to include, utilize second or more weights, or apply second or more weights to a set of frames across multiple videos, in a temporal dimension. Weights can be configured or set to focus on, emphasize, or otherwise implement an attention mechanism, focusing on a specific set or combination of features encountered in the input data (e.g., video frames or installation data) of the video to be processed.
[0108] This method may include identifying and using one or more logs from one or more robotic medical systems to train an ML model. Each of the one or more logs may indicate the time of installation, attachment, configuration, or use of one or more devices for a corresponding video in multiple videos. The method may assign a label from multiple tags indicating the time of installation, configuration, attachment, or use of the medical device to each frame of the multiple videos used for training the ML model, the time derived from the corresponding log in one or more logs corresponding to the corresponding video in the multiple videos. The method may include receiving a set of frames from multiple videos captured for one or more medical procedures and using the labels of the set of frames identified at operation 405 to train the model. The method may also include using multiple labels identified or determined at operation 405 to train the model.
[0109] At operation 415, the method may determine a frame label for one or more input video frames. This method may include using an ML model trained at operation 410 to determine a frame label for one or more input video frames of a medical procedure, which will be processed to use the ML model for the presence and identification of medical devices. For example, the data processing system may receive video of a medical procedure, which will be processed to identify and detect medical devices. The data processing system may receive or access video of a medical procedure stored in a data repository. The method may include identifying a series of frames from a video of a medical procedure captured by a robotic medical system. The series of frames may correspond to video of a previously performed medical procedure or live-streamed video and to an ongoing surgical procedure.
[0110] This method may include using an ML model to determine a frame-by-frame label for each frame in a series of frames. The frame-by-frame label may indicate the probability of presence of one or more types of devices. The frame-by-frame label may include a separate label for each individual frame of the input video received for processing, or it may include labels for multiple consecutive video frames of the input video. This method may use an ML model to determine the frame-by-frame label based at least on the time of installation, configuration, attachment, or use of one or more devices at one or more robotic medical systems. This method may use an ML model to determine the probability of presence based on a comparison of timestamps and installation times.
[0111] This method may include identifying a subset of frames from a series of frames in a video that correspond to a portion of the video capturing an instrument of that type used in a medical procedure. For example, the ML model can identify video segments comprising any number of consecutive video frames (e.g., 30, 45, 60, 90, 120, 150, 180, 210 video frames) that can span any number of seconds of the video (e.g., 2, 3, 4, 5 seconds or more). The method can determine a per-frame label for the video segment. The per-frame label can be identified based on the attention mechanism (e.g., weights) of the ML model across multiple video frames. The ML model can determine the corresponding per-frame label for a subset of frames based on a subset of frames input into the model.
[0112] At operation 420, the method can determine whether each frame label exceeds a threshold. The method may include comparing the value of each frame label, corresponding to a probability or confidence level of the presence of a given medical device in the video, to a threshold representing the probability of presence. The threshold can be any acceptable threshold for determining the probability of a medical device's presence in the video (e.g., 90% certainty or confidence, 95%, 99%, 99.5%, or greater than 99.5%). For example, each frame label may include a value indicating a probability of presence greater than the threshold or a value indicating a probability not greater than the threshold (e.g., less than the threshold).
[0113] This method may include comparing the probability of presence of each corresponding frame in a series of frames with a threshold indicating the presence of a given type of device. The threshold may be the same for all device types, or it may vary based on the device type. The threshold may be the same for each frame label, or it may vary for different frames. For example, a higher threshold may be used for the frame label of a video segment across multiple video frames than for the frame label of each video frame individually. This method may receive a file from a robotic medical system that includes an indication of the time at which one or more devices are installed at the robotic medical system. This method may determine the corresponding frame label of at least one frame in the series of frames based at least on the indication of the time input into the model.
[0114] At operation 425, based on the operation at 420, the method can determine the presence of a device. Based on determining at operation 420 that each frame label exceeds a threshold, the method can determine that a specific type of device exists in a video frame or video clip. For example, if the confidence score or probability of determining the presence of a medical device in a video frame or clip, represented by each frame label, is greater than a threshold, then the method can create values indicating the presence of the device for the vector of each frame label.
[0115] At operation 430, based on the operation at 420, the method can determine that the device is not present. For example, if the value in each frame label at operation 420 does not exceed a threshold, then the ML model can determine that the device is not present. For example, if at operation 420, the confidence score or probability of determining the presence of a medical device in a video frame or segment, represented by each frame label, is not greater than a threshold, then the method can create a value for the vector of each frame label indicating that the device is not present.
[0116] At 435, the method can modify the presence determination based on post-processing. The method may include adjusting or correcting the determination at operation 425 or 430 based on post-processing functions such as frame smoothing or outlier filtering. For example, the method may monitor a series of per-frame labels for a series of consecutive video frames, where a specific determination of the presence or absence of a medical device in every frame except one exceeds a specific threshold. The method may identify, within such a series of frames, a single frame that includes a determination that contradicts (e.g., is opposite to) the determinations of all surrounding frames. In response to this identification, the processing function may correct the outlier per-frame labels to make the determination consistent with the determinations of all surrounding per-frame labels.
[0117] For example, the method may include determining a second per-frame label for each corresponding frame in a series of frames, at least in part based on a comparison at operation 420, the second per-frame label indicating the presence of an instrument of that type at the robotic medical system. The method may generate a series of per-frame labels based at least on the per-frame labels for each frame in the series of frames. The method may adjust the value of the first per-frame label of the first frame in the series of frames using at least a second value of the second per-frame label of a second frame adjacent to the first frame.
[0118] This method can generate heatmaps for identifying or highlighting regions in video frames corresponding to the location where a medical device is detected. For example, the method can generate heatmaps based on at least multiple video frames, indicating regions within a subset of frames where the probability of the presence of that type of device exceeds a threshold of the heatmap. The method can also identify second regions within a subset of frames where the probability of the presence of that type of device exceeds a second threshold, which exceeds a first threshold, based on at least multiple video frames. The second region of the second heatmap can be within the region of the first heatmap, thereby indicating a higher-probability region of the video frame to be displayed (e.g., more prominent highlighting).
[0119] At 440, the method may display a video frame indicating the presence of a device. The method may include displaying an indication of the presence of this type of device via a graphical user interface, the indication being at least partially based on timestamps in the video on each frame label of a series of frames determined via a model. For example, the method may include one or more processors displaying an indication identifying a device of this type determined at operation 425. The indication may overlay a subset of the series of frames displayed on the graphical user interface, the subset having a probability of presence exceeding a threshold for this type of device. The indication may include a heatmap that can be displayed to highlight the location of the device.
[0120] Figure 5 A surgical system 500 according to some embodiments is depicted. The surgical system 500 may be an example of a medical environment 102. The surgical system 500 may include a robotic medical system 505 (e.g., robotic medical system 120), a user control system 510, and an assistive system 515 that are communicatively coupled to each other. A visualization tool 520 (e.g., visualization tool 114) may be connected to the assistive system 515, which in turn may be connected to the robotic medical system 505. Therefore, when the visualization tool 520 is connected to the assistive system 515 and the assistive system is connected to the robotic medical system 505, the visualization tool can be considered connected to the robotic medical system. In some embodiments, the visualization tool 520 may be additionally or alternatively directly connected to the robotic medical system 505.
[0121] Surgical system 500 can be used to perform computer-aided medical procedures on patient 525. In some embodiments, the surgical team may include surgeon 530A and additional medical personnel 530B-530D, such as medical assistants, nurses, and anesthesiologists, as well as other suitable team members who may assist in the surgical procedure or medical session. A medical session may include surgical procedures performed on patient 525, as well as any preoperative (e.g., which may include setting up surgical system 500, including preparing patient 525 for the procedure) and postoperative (e.g., which may include patient cleanup or aftercare) or other procedures during the medical session. Although described in the context of surgical procedures, surgical system 500 may be implemented in non-surgical procedures or in other types of medical procedures or diagnoses where the accuracy and convenience of the surgical system may be beneficial.
[0122] The robotic medical system 505 may include a plurality of manipulator arms 535A-535D, to which a plurality of medical instruments (e.g., medical instrument 112) may be coupled or mounted. Each medical instrument may be any suitable surgical instrument (e.g., an instrument with tissue interaction capabilities), imaging device (e.g., an endoscope, ultrasound instrument, etc.), sensing instrument (e.g., a force-sensing surgical instrument), diagnostic instrument, or other suitable instrument that may be used for computer-assisted surgical procedures on a patient 525 (e.g., by being at least partially inserted into and manipulated to perform computer-assisted surgical procedures on the patient). Although the robotic medical system 505 is shown as including four manipulator arms (e.g., manipulator arms 535A-535D), in other embodiments, the robotic medical system may include more or fewer than four manipulator arms. Additionally, not all manipulator arms may have a medical instrument mounted thereon at all times during a medical session. Furthermore, in some embodiments, the medical instrument mounted on the manipulator arm may be suitably replaced with another medical instrument.
[0123] One or more of the manipulator arms 535A-535D and / or the medical instruments attached to the manipulator arms may include one or more displacement transducers, orientation sensors, positioning sensors, and / or other types of sensors and devices that measure parameters and / or generate kinematic information. One or more components of the surgical system 500 may be configured to use measured parameters and / or kinematic information to track (e.g., determine the posture of the medical instruments and objects) and / or control the medical instruments and any objects attached to the medical instruments and / or manipulator arms 535A-535D.
[0124] Surgeon 530A may use user control system 510 to control (e.g., move) one or more of manipulator arms 535A-535D and / or medical instruments connected to the manipulator arms. To facilitate control of the manipulator arms 535A-535D and tracking of the progress of a medical session, user control system 510 may include a display (e.g., display 116 or 1130) that can provide surgeon 530A with images (e.g., high-resolution 3D images) of the surgical site associated with patient 525 captured by a medical instrument (e.g., medical instrument 112, which may be an endoscope) mounted to one of the manipulator arms 535A-535D. User control system 510 may include a stereoscopic viewer with two or more displays, where surgeon 530A can view stereoscopic images of the surgical site associated with patient 525 generated by a stereoscopic imaging system. In some embodiments, user control system 510 may also receive images from assistive system 515 and visualization tool 520.
[0125] Surgeon 530A can perform one or more procedures using images displayed by user control system 510 and one or more medical instruments attached to manipulator arms 535A-535D. To facilitate control of manipulator arms 535A-535D and / or the medical instruments mounted thereon, user control system 510 may include a set of controls. These controls can be manipulated by surgeon 530A to control the movement of manipulator arms 535A-535D and / or the medical instruments mounted thereon. The controls can be configured to detect various hand, wrist, and finger movements of surgeon 530A to allow surgeon to visually perform procedures on patient 525 using one or more medical instruments mounted to manipulator arms 535A-535D.
[0126] The assistive system 515 may include one or more computing devices configured to perform processing operations within the surgical system 500. For example, one or more computing devices may control and / or coordinate operations performed by various other components of the surgical system 500 (e.g., the robotic medical system 505, the user control system 510). Computing devices included in the user control system 510 may transmit instructions to the robotic medical system 505 via one or more computing devices of the assistive system 515. The assistive system 515 may receive and process image data representing images captured by one or more imaging devices (e.g., medical tools) attached to the robotic medical system 505, as well as other data stream sources received from visualization tools. For example, one or more image capture devices (e.g., image capture device 110) may be located within the surgical system 500. These image capture devices may capture images from various viewpoints within the surgical system 500. These images (e.g., video streams) may be transmitted to the visualization tool 520, which may then transfer those images as a single combined data stream to the assistive system 515. The assistive system 515 can then transmit a single video stream (including any data stream received from one or more medical tools of the robotic medical system 505) to be displayed on the display of the user control system 510 (e.g., display 116).
[0127] In some embodiments, the assistive system 515 may be configured to present visual content (e.g., a single combined data stream) to other team members (e.g., medical personnel 530B-530D) who may not have access to the user control system 510. Therefore, the assistive system 515 may include a display 540 configured to display one or more user interfaces, such as images of the surgical site, information associated with the patient 525 and / or the surgical procedure, and / or any other visual content (e.g., a single combined data stream). In some embodiments, the display 540 may be a touchscreen display and / or include other features that allow medical personnel 530A-530D to interact with the assistive system 515.
[0128] The robotic medical system 505, user control system 510, and auxiliary system 515 can be communicatively coupled to each other in any suitable manner. For example, in some embodiments, the robotic medical system 505, user control system 510, and auxiliary system 515 can be communicatively coupled via control line 545, which can represent any wired or wireless communication link that can serve a particular implementation. Therefore, the robotic medical system 505, user control system 510, and auxiliary system 515 can each include one or more wired or wireless communication interfaces, such as one or more local area network interfaces, Wi-Fi network interfaces, cellular interfaces, etc. It should be understood that the surgical system 500 may include other or additional components or elements that may be required or are considered desirable for medical sessions using the surgical system.
[0129] Figure 6 Example block diagrams of an example computer system 600 illustrated according to some embodiments are depicted. The computer system 600 can be any computing device as used herein and can include or be used to implement a data processing system or components thereof. The computer system 600 includes at least one bus 605 or other communication component or interface for transferring information between various elements of the computer system. The computer system also includes at least one processor 610 or processing circuitry coupled to the bus 605 for processing information. The computer system 600 also includes at least one main memory 615, such as random access memory (RAM) or other dynamic storage device, coupled to the bus 605 for storing information and instructions to be executed by the processor 610. The main memory 615 can be used to store information during the execution of instructions by the processor 610. The computer system 600 may also include at least one read-only memory (ROM) 620 or other static storage device coupled to the bus 605 for storing static information and instructions for the processor 610. Storage device 625 (such as a solid-state device, disk, or optical disk) can be coupled to the bus 605 to persistently store information and instructions.
[0130] Computer system 600 may be coupled to display 630, such as a liquid crystal display or an active matrix display, via bus 605 for displaying information. Input device 635, such as a keyboard or voice interface, may be coupled to bus 605 for transmitting information and commands to processor 610. Input device 635 may include a touchscreen display (e.g., display 630). Input device 635 may also include cursor controls, such as a mouse, trackball, or arrow keys, for transmitting directional information and command selection to processor 610 and for controlling cursor movement on display 630.
[0131] The processes, systems, and methods described herein can be implemented by a computer system 600 in response to a processor 610 executing instructions contained in main memory 615. Such instructions may be read into main memory 615 from another computer-readable medium, such as storage device 625. The arrangement for executing the instructions contained in main memory 615 causes the computer system 600 to perform the illustrative processes described herein. One or more processors in a multiprocessor arrangement may also be used to execute the instructions contained in main memory 615. Hardwired circuitry may be used in place of or in combination with software instructions with the systems and methods described herein. The systems and methods described herein are not limited to any specific combination of hardware circuitry and software.
[0132] Despite Figure 6 Example computing systems have been described, but the subject matter including the operations described herein can be implemented in other types of digital electronic circuit systems, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations thereof.
[0133] The topics described herein sometimes illustrate different components contained within or connected to different other components. It should be understood that the architectures depicted are illustrative, and many other architectures can indeed be implemented to achieve the same functionality. Conceptually, any arrangement of components achieving the same functionality is effectively “associated” to achieve the desired functionality. Therefore, any two components combined herein to achieve a particular functionality can be considered “associated” with each other to achieve the desired functionality, regardless of the architecture or intermediate components. Similarly, any two components so associated can also be considered “operably connected” or “operably coupled” to each other to achieve the desired functionality, and any two components that can be so associated can also be considered “operably coupled” to each other to achieve the desired functionality. Specific examples of operably coupled components include, but are not limited to, physically matable or physically interactive components, wirelessly interactive or wirelessly interactive components, or logically interactive or logically interactive components.
[0134] Regarding the use of plural or singular terms in this document, those skilled in the art may appropriately convert from plural to singular or from singular to plural depending on the context or application. For clarity, various singular / plural transformations may be explicitly described herein.
[0135] Those skilled in the art will understand that, in general, the terms used herein, and especially those used in the appended claims (e.g., the body of the appended claims), are generally intended to be “open-ended” terms (e.g., the term “comprising” should be interpreted as “including but not limited to”, the term “having” should be interpreted as “at least having”, the term “including” should be interpreted as “including but not limited to”, etc.).
[0136] Although the accompanying drawings and detailed embodiments may illustrate a specific order of method steps, the order of such steps may differ from the order depicted and described unless otherwise specified above. Furthermore, unless otherwise specified above, two or more steps may be performed simultaneously or partially simultaneously. For example, such variations may depend on the chosen software and hardware system and the designer's choices. All such variations are within the scope of this disclosure. Similarly, software implementations of the described methods can be accomplished using standard programming techniques with rule-based logic and other logic to perform various connection steps, processing steps, comparison steps, and decision steps.
[0137] Those skilled in the art will further understand that if the introduced claim statements are intended to have a specific quantity, then such intent will be explicitly stated in the claims, and where no such statement is made, such intent does not exist. For example, to aid understanding, the appended claims may contain the use of introductory phrases “at least one” and “one or more” to introduce claim statements. However, the use of such phrases should not be construed as implying that introducing claim statements with the indefinite article “a” or “an” limits any particular claim containing such introduced claim statements to an invention containing only one such statement, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” or “an” should generally be interpreted as meaning “at least one” or “one or more”); the same applies to the use of definite articles used to introduce claim statements. Furthermore, even if a specific quantity of the introduced claim statements is explicitly stated, those skilled in the art will recognize that such statements should generally be interpreted as meaning at least the quantity stated (e.g., in the absence of other modifiers, simply stating “two statements” generally means at least two statements, or two or more statements).
[0138] Furthermore, in cases where conventions such as "at least one of A, B, and C" are used, such constructions are generally intended to be understood by a person skilled in the art in the meaning of the convention (e.g., "a system having at least one of A, B, and C" will include, but is not limited to, systems having only A, only B, only C, A and B, A and C, B and C, or A, B, and C, etc.). In cases where conventions such as "at least one of A, B, or C" are used, such constructions are generally intended to be understood by a person skilled in the art in the meaning of the convention (e.g., "a system having at least one of A, B, or C" will include, but is not limited to, systems having only A, only B, only C, A and B, A and C, B and C, or A, B, and C, etc.). A person skilled in the art will further understand that, in fact, almost any transition word or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to consider the possibility of including one, any, or both of the terms. For example, the phrase "A or B" will be understood to include the possibility of "A" or "B" or "A and B".
[0139] In addition, unless otherwise stated, the words “approximately,” “about,” “around,” “basically,” etc., are used to mean plus or minus ten percent.
[0140] For purposes of illustration and description, the foregoing description of illustrative embodiments has been presented. It is not intended to be exhaustive or limiting of the precise forms disclosed, and modifications and variations can be made in accordance with the foregoing teachings, or may be derived from practice of the disclosed embodiments. The scope of the invention is intended to be defined by the appended claims and their equivalents.
Claims
1. A system comprising: One or more processors coupled to memory, in order to: Identify a series of frames from a video of a medical procedure captured by a robotic medical system; The identification is based at least in part on a model trained on frames from multiple videos captured for one or more medical procedures, the frames of which are labeled with data identifying the installation of one or more instruments in one or more robotic medical systems; The model is used to determine a frame label for each frame in the series of frames, the frame label indicating the probability of the presence of one or more types of devices; as well as An indication of the presence of a type of device is displayed via a graphical user interface, the indication being at least in part based on timestamps in the video on the frame labels of the series of frames determined via the model.
2. The system of claim 1, wherein the one or more processors are further configured to: Receive a set of frames from the plurality of videos captured for one or more medical procedures; A tag for the frame set is identified based on the last frame in the frame set, the tag including data indicating the installation time of the one or more devices and the last timestamp of the last frame; and The model is trained using the labels of the frame set.
3. The system of claim 1, wherein the one or more processors are further configured to: For each frame of the plurality of videos, multiple tags are identified, each of the multiple tags having a vector of one or more values corresponding to one or more instruments; and The model is trained using the multiple labels.
4. The system of claim 1, wherein the one or more processors are further configured to: For each frame of the plurality of videos, a plurality of tags are determined for each frame, each tag having a value indicating whether the one or more instruments are installed at the one or more robotic medical systems at a given time in each corresponding frame of the frame; and The model is trained using the multiple labels.
5. The system of claim 1, wherein the one or more processors are further configured to: Identify one or more logs from the one or more robotic medical systems, each of the one or more logs being prepared to indicate the installation time of the one or more devices for a corresponding video among the plurality of videos; and For each frame of the plurality of videos, a tag indicating the installation time is assigned from the plurality of tags, the installation time being derived from the corresponding log in one or more logs corresponding to the corresponding video of the plurality of videos.
6. The system of claim 1, wherein the model comprises a spatial-temporal neural network model, the spatial-temporal neural network model applying a first one or more weights to one or more spatial dimensions of an image region within a frame of the plurality of videos, and applying a second one or more weights to a temporal dimension of a set of frames across the plurality of videos.
7. The system of claim 1, wherein the one or more processors are further configured to: A heatmap is generated based at least on the model, the heatmap indicating regions within a subset of the frame where the probability of the presence of an instrument of the type exceeds a threshold of the heatmap.
8. The system of claim 7, wherein the one or more processors are further configured to: At least based on the model, a second region is identified within the subset of the frame where the probability of the presence of an instrument of that type exceeds a second threshold, wherein the second threshold exceeds the threshold; and The subset of frames covered by the heatmap, which is displayed via the graphical user interface, having at least one of the first and second regions.
9. The system of claim 1, wherein the one or more processors are further configured to: The probability of presence of each corresponding frame in the series of frames is compared with a threshold for the presence of the type of device; and A second per-frame label for each corresponding frame in the series of frames is determined, at least in part, based on the comparison, the second per-frame label indicating the presence of the type of device at the robotic medical system.
10. The system of claim 1, wherein the one or more processors are further configured to: Identify and capture a subset of frames from the series of frames of the video corresponding to a portion of the video of the device of the type used in the medical procedure; and The corresponding frame label for the subset of frames is determined based on the subset of frames input into the model.
11. The system of claim 1, wherein the one or more processors are further configured to: Receive a document from the robotic medical system, the document including an indication of the time for installing the one or more instruments at the robotic medical system; and The corresponding frame label for at least one frame in the frames is determined based at least on the indication of the time input into the model.
12. The system of claim 1, wherein the one or more processors are further configured to: A series of frame-by-frame labels are generated based at least on the frame-by-frame labels for each frame in the series of frames; and The value of the first frame label of the first frame in the series of frames is adjusted using at least a second value of the second frame label of the second frame adjacent to the first frame.
13. The system of claim 1, wherein the one or more processors are further configured to: The model is used to determine the per-frame label based at least on the installation time of the one or more devices at the one or more robotic medical systems; and The probability of existence is determined based on a comparison between the timestamp and the installation time.
14. The system of claim 1, wherein the one or more processors are further configured to: The device of that type is identified at least based on a probability of presence exceeding a threshold for that type of device; and The instruction that displays the type of device is indicated.
15. The system of claim 1, wherein the indication is overlaid on a subset of the series of frames displayed on the graphical user interface, the probability of the presence of the subset of the series exceeding a threshold for the type of device.
16. A method comprising: One or more processors coupled to memory use data that identifies the installation of one or more types of instruments in one or more robotic medical systems to label frames of multiple videos captured for one or more medical procedures; The model is trained using labeled frames by one or more processors; The model is used to determine a frame label for each frame in the series of frames, the frame label indicating the probability of the presence of the one or more types of devices; as well as The presence of a certain type of device is indicated via a graphical user interface.
17. The method of claim 16, further comprising: The presence of the type of device is determined by the one or more processors based at least in part on timestamps in the video on the frame labels of the series of frames determined via the model, wherein the frame labels in the labeled frames indicate the presence of the type of device, and the frames do not include images of the type of device.
18. The method of claim 16, comprising: The processor receives a set of frames from the plurality of videos captured for one or more medical procedures. The one or more processors identify a tag for the frame set based on the last frame in the frame set, the tag including data indicating the installation time of the one or more devices and the last timestamp of the last frame; and The model is trained by the one or more processors using the labels of the frame set.
19. A non-transitory computer-readable medium storing processor-executable instructions, wherein when executed by one or more processors, the instructions cause the one or more processors to: Identify a series of frames from a video of a medical procedure captured by a robotic medical system; The identification is based at least in part on a model trained on frames from multiple videos captured for one or more medical procedures, the frames of which are labeled with data identifying the installation of one or more instruments in one or more robotic medical systems; The model is used to determine a frame label for each frame in the series of frames, the frame label indicating the probability of the presence of one or more types of devices; as well as An indication of the presence of a type of device is displayed via a graphical user interface, the indication being at least in part based on timestamps in the video on the frame labels of the series of frames determined via the model.
20. The non-transitory computer-readable medium of claim 19, wherein the indication is overlaid on a subset of the series of frames displayed on the graphical user interface, the probability of the presence of the subset of the series exceeding a threshold for the type of device.