Machine learning system for robotic medical procedure complexity
A machine learning system for robotic medical procedures automatically assesses complexity by analyzing intra-operative data, enhancing performance evaluation and reducing resource consumption through efficient video recommendations and expert review.
Patent Information
- Application Number
- PCT/US2025/039636
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-30
- Filing Date
- 2025-07-29
- Publication Date
- 2026-02-05
AI Technical Summary
Determining the complexity of medical procedures performed by robotic systems is challenging due to variability in patient anatomy, disease state, and surgical approach, leading to inaccurate evaluation of performance and resource consumption.
A machine learning system that automatically identifies video segments to predict case complexity, using models trained on intra-operative data to determine feature vectors and performance indicators, enabling scalable and efficient complexity assessment.
Improves the accuracy and efficiency of evaluating robotic medical procedure performance by providing context for surgeon skill and patient outcomes, reducing resource consumption through selective expert review and model re-training.
Smart Images

Figure US2025039636_05022026_PF_FP_ABST
Abstract
Description
MACHINE LEARNING SYSTEM FOR ROBOTIC MEDICAL PROCEDURECOMPLEXITYCROSS-REFERENCES TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority under 35 U.S.C. § 119 to U.S. Provisional Patent Application No. 63 / 677,313, filed July 30, 2024, which is hereby incorporated by reference herein in its entirety.BACKGROUND
[0002] Medical procedures can be performed using a robotic medical system. Due to the large number of procedures performed, and the variability seen in patient anatomy, disease state, and surgical approach, it can be challenging and resource-consuming to determine certain attributes of the medical procedures, such as the level of complexity of the medical procedure. Without accurately determining such attributes of the medical procedure, it can be challenging to reliably, efficiently, and accurately evaluate performance of the medical procedure.SUMMARY
[0003] Technical solutions disclosed herein can include a machine learning system to determine the level of complexity of a medical procedure that can be performed using a robotic medical system. The computing system can use machine learning to determine the level of complexity of the medical procedure, which can facilitate accurately evaluating the performance of the medical procedure and improving the accuracy, reliability or efficiency with which robotic medical systems can perform the medical procedure. By using machine learning, the computing system can implement a scalable solution that may not need manual identification of keyframes to use to identify complexity.
[0004] Indeed, aspects of the technical solutions disclosed herein can leverage the improved interpretation of surgeon skill and patient outcomes to generate more effective video recommendations, more efficient and effective video searches, a scalable, and automatic technical solution for computing complexity.
[0005] To do so, the computing system can automatically identify a video segment using machine learning and use the identified video segment to predict case complexity, thusenabling automation and scalability. The computing system can implement a quality control technique that routes video or other case information to a remote computing device (e.g., associated with an expert) for evaluation and input. The computing system can select cases for expert review based on historic model performance or model confidence scores (on a particular case) associated with determining the level of complexity of a medical procedure. Thus, by selectively routing cases to remote computing devices for expert review, the computing system can reduce the number of cases for which expert input is requested, thereby reducing delays, latency, or resource consumption associated with expert review and model re-training. The measure of complexity determined by the computing system can improve the interpretation of operator or surgeon skill. For example, by determining a complexity level for a medical procedure, context can be provided to understand the surgeon’s performance and trends in performance across cases.
[0006] At least one aspect of the present disclosure is a system. The system can include one or more processors, coupled with memory, to receive a video of a medical procedure performed on a robotic medical system. The one or more processors can select segment(s) of the video used to determine a level of complexity of the medical procedure. The one or more processors can execute, using the segment of the video, one or more models trained by machine learning to determine feature data. The one or more processors can determine, using the feature data and the one or more processors, the level of complexity of the medical procedure.
[0007] The one or more processors can execute, using the video segment as input, a trained machine learning model to predict case complexity.
[0008] The one or more processors can, alternatively, execute a process to sample the video segment into a set of single images or video clips. The one or more processors can execute, using a single image (or single clip) as input, a trained machine learning model to predict case complexity from this single image (or single clip). The one or more processors can execute, using the set of case complexity predictions and timestamps, an aggregator that aggregates these predictions into a single case complexity prediction through voting methods.
[0009] In some implementations, the one or more processors can execute a stacking of machine learning models. For example, the one or more processors can execute a process to sample the video segment into a set of single images. The one or more processors canexecute, using a single image as input, a pre-trained machine learning model to extract a feature vector that represents low-dimensional spatial features of the image. The one or more processors can execute, using feature vectors as input, a second machine learning model trained to predict case complexity on the video segment.
[0010] The one or more processors can, alternatively, execute a stacking of machine learning models. For example, the one or more processors execute a process to sample the video segment into a set of single images. The one or more processors can execute, using a single image as input, a machine learning model trained to predict case complexity. Instead of extracting the prediction, the processor can extract a feature vector that represents lowdimensional indicators of the complexity. The one or more processors can optionally execute, using feature vectors as input, a second machine learning model trained to predict case complexity on a video clip. The one or more processors can execute an aggregator to aggregate predictions across clips into a single case complexity prediction.
[0011] The one or more processors can receive a label of the type of the medical procedure. The one or more processors can select, using the label, the one or more models from different models trained by machine learning for different procedure types.
[0012] The one or more processors can receive system data of the medical procedure from the robotic medical system. The one or more processors can select performance indicators from the system data. The one or more processors can stack the feature data with the selected performance indicators. The one or more processors can train, using the stacked feature data including the selected performance indicators, the one or more models using machine learning.
[0013] The one or more processors can receive a search query for videos of medical procedures, the search query including an indication of the type of the medical procedure and the level of complexity. The one or more processors can select, responsive to the search query, the video of the medical procedure. The one or more processors can provide, for display via a graphical user interface, a graphical user interface element including an indication of the video.
[0014] The one or more processors can determine to submit the case for review based on historic model performance for a particular procedure type and / or complexity level, model confidence score on a particular case, or by random selection. The one or more processorscan generate data to cause a graphical user interface to display the video and the machinelearning prediction of the level of complexity for cases submitted by system for review. The one or more processors can receive, from the graphical user interface, an update to the level of complexity. The one or more processors can re-train the one or more models using the update to the level of complexity.
[0015] The one or more processors can generate a confusion matrix for historic model performance on one or more models. The one or more processors can use this historic model performance to determine what cases to submit for review.
[0016] At least one aspect of the present disclosure is a method including receiving, by one or more processors coupled with memory, a video of a medical procedure performed on a robotic medical system. The method can include selecting, by the one or more processors, based on a type of the medical procedure, a segment of the video to determine a level of complexity of the medical procedure with. The method can include executing, by the one or more processors, using the segment of the video, one or more models trained by machine learning to determine feature data. The method can include determining, by the one or more processors, using the feature data and the one or more models, a level of complexity of the medical procedure.
[0017] The method can include executing, by the one or more processors, using at least a frame of the segment, a first model trained by machine learning to determine at least a feature vector for the frame. The method can include executing, by the one or more processors, a second model trained by machine learning to determine, using the feature vector for the frame, the level of complexity of the medical procedure.
[0018] The method can include receiving, by the one or more processors, a label of the type of the medical procedure. The method can include selecting, by the one or more processors, using the label, the one or more models from different models trained by machine learning for different procedure types.
[0019] The method can include receiving, by the one or more processors, system data of the medical procedure from the robotic medical system. The method can include selecting, by the one or more processors, performance indicators from the system data. The method can include stacking, by the one or more processors, the feature data (or images or clips of the video) with the selected performance indicators. The method can include training, by the oneor more processors, using the stacked feature data including the selected performance indicators, the one or more models using machine learning.
[0020] The method can include receiving, by the one or more processors, a search query for videos of medical procedures, the search query including an indication of the type of the medical procedure and the level of complexity. The method can include selecting, by the one or more processors, responsive to the search query, the video of the medical procedure. The method can include providing, by the one or more processors, for display via a graphical user interface, a graphical user interface element comprising an indication of the video.
[0021] The method can include determining, by the one or more processors, to submit the video for review based on the level of complexity exceeding a first threshold and an accuracy level of the one or more models being less than a second threshold. The method can include generating, by the one or more processors, data to cause a graphical user interface to display the video responsive to the level of complexity exceeding the first threshold and the accuracy level being less than the second threshold. The method can include receiving, by the one or more processors, from the graphical user interface, an update to the level of complexity. The method can include re-training, by the one or more processors, the one or more models using the update to the level of complexity.
[0022] The method can include generating, by the one or more processors, a confusion matrix from performance of the one or more models. The method can include using, by the one or more processors, the confusion matrix to determine to submit the video for review.
[0023] The method can include generating, by the one or more processors, a feature vector using a first model of the one or more models. The method can include executing, by the one or more processors, using the feature vector, a second model trained by machine learning to determine the level of complexity of the medical procedure.
[0024] At least one aspect of the present disclosure is a non-transitory computer-readable medium storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to receive a video of a medical procedure performed by a robotic medical system. The one or more processors can select, such as based on a type of the medical procedure, a segment of the video to determine a level of complexity of the medical procedure with. The one or more processors can execute, using the segment of the video, one or more models trained by machine learning to determine feature data. Theone or more processors can determine, using the feature data and the one or more models, the level of complexity of the medical procedure.
[0025] The one or more processors can execute, using at least a frame of the segment, a first model trained by machine learning to determine at least a feature vector for the frame. The one or more processors can execute, using the feature vector for the frame, a second model trained by machine learning to determine the level of complexity of the medical procedure.
[0026] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of various aspects and implementations, and provide an overview or framework for understanding the nature and character of the claimed aspects and implementations. The drawings provide illustration and a further understanding of the various aspects and implementations, and are incorporated in and constitute a part of this specification. The foregoing information and the following detailed description and drawings include illustrative examples and should not be considered as limiting.BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings are not intended to be drawn to scale. Like reference numbers and designations in the various drawings indicate like elements. For purposes of clarity, not every component may be labeled in every drawing. In the drawings:
[0028] FIG. 1 A depicts an example computing system to determine a complexity of a medical procedure using a model.
[0029] FIG. IB depicts an example computing system including a model to determine complexity of multiple frames and an aggregator to aggregate the frame complexities into a case complexity.
[0030] FIG. 1C depicts an example computing system including a first model to determine video frame features and a second model to determine video clip complexity.
[0031] FIGS. 2A-2B depicts an example computing system to surface complexity determinations for a medical procedure to a user and re-train a machine learning model.
[0032] FIG. 3 depicts a confusion matrix for a machine learning model trained to determine complexity of a medical procedure.
[0033] FIG. 4 depicts example charts of procedure indicators for medical procedures plotted against medical procedure complexity.
[0034] FIG. 5 depicts an example method of determining complexity of a medical procedure.
[0035] FIG. 6 depicts an example architecture of a computing system.DETAILED DESCRIPTION
[0036] Following below are more detailed descriptions of various concepts related to, and implementations of, methods, apparatuses, and systems for a machine learning system for robotic medical procedure complexity. The various concepts introduced above and discussed in greater detail below may be implemented in any of numerous ways.
[0037] This disclosure is generally directed to machine learning systems to determine complexity of a medical procedure performed by a robotic medical system. A computing system can determine the performance of the robotic medical system based at least in part on the performance of an operator, user, or surgeon that controls the robotic medical system to perform the medical procedure on a patient. The performance can be an indicator or set of indicators that quantify the operator’s performance. However, the indicators may not accurately identify the performance of an operator relative to other operators or across multiple medical procedures. For example, operators that perform more complex, difficult, or challenging medical procedures may have lower performance indicators because of the difficult nature of the medical procedures they perform, not in view of their actual skill. In this regard, the indicators used to determine operator performance may be inaccurate or misleading. Some technologies can capture aspects of case complexity, such as relevant hospital records (e.g., patient BMI, number or previous abdominal surgeries, anatomical inflammation at the time of surgery, location and density of adhesions at the time of surgery, and anatomical aberrations, etc.). Some technologies can determine case complexity using intra-operative videos, but these technologies may need manual identification of key frames that inform about complexity limiting automation and scalability.
[0038] To solve these, and other technical problems, technical solutions of this disclosure can include a machine learning system to determine robotic medical procedure complexity. The computing system can use machine learning to determine the level of complexity of a medical procedure. By using machine learning, the computing system can implement a scalable solution that may not need manual identification of keyframes to use to identify complexity. The computing system can automatically identify a video segment using machine learning and use the identified video segment to predict case complexity with, thus enabling automation and scalability. The computing system can train models to determine case complexity from the intra-operative medical procedure data. For example, the computing system can receive or collect training data of various medical procedures labeled with different levels of complexity. The computing system can generate a tag for case complexity based on an intra-operative video of the patient. The intra-operative video can capture aspects of complexity otherwise difficult to capture from solely analyzing preoperative hospital records. Furthermore, the computing system can determine the level of complexity based on the type of the medical procedure and objective performance indicators (OPIs) of the medical procedure.
[0039] The computing system can implement one or more models to determine the complexity level. For example, the computing system can implement a single-image model, such as a two-dimensional convolutional neural network, to generate a feature vector for a frame of the medical procedure video. The second model can be a temporal model, such as a sequence model that processes a sequence of frames. The temporal model can be a long- short term memory model. The second model can introduce an attention mechanism to learn the weightage of a frame in a sequence of frames. The attention mechanism can allow the computing system to consider past events of the medical procedure which can influence complexity. For example, if bleeding begins early in the procedure, the second model can use attention and memory features to use the bleeding as context to inform complexity determinations for the case. The second model can output a second feature vector, which can be stacked with selected OPIs of medical procedures to determine a case complexity.
[0040] The computing system can implement quality control to select cases for expert review. The computing system can select cases for expert review based on historic model performance and confidence associated with determining the level of complexity of a medical procedure. The quality control selector can compare the results of the models with aconfusion matrix to determine the likelihood that the result is accurate, and trigger generation of the prompt or other action based on the comparison. For example, the quality control selector can select medical procedure videos with a high complexity level but a low accuracy for user review and modification. The computing system and retrain the models using the user input. The computing system can trigger re-training responsive to new labeled training data being available. The new labeled training data can be generated from the review and input provided by the expert or user.
[0041] The measure of complexity determined by the computing system can improve the interpretation of operator or surgeon skill. For example, by determining a complexity level for a medical procedure, context can be provided to understand the surgeon’s skill level and interpreted performance that quantify the surgeon’s skill level. For example, determinations of case complexity can be incorporated into trends of a surgeon’s performance indicators to help contextualize the performance of the surgeon. Furthermore, by quantifying the complexity of different medical procedures, context for a patient outcome can be provided. This improved interpretation of patient outcomes can include using case complexity to predict patient outcomes in order to separate confounding variables from the influences of surgical techniques and surgeon skill. The complexity can allow for patient outcomes to be better interpreted by patients, staff, or medical practitioners.
[0042] Furthermore, the computing system can search videos or generate video recommendations of medical procedures using the complexity level of the medical procedure. For example, a database of videos can store videos tagged with complexity levels. This can allow for an efficient, effective, and scalable video search. For example, the computing system can make effective video recommendations by recommending videos of cases that match the complexity of the case a user is currently reviewing. For example, the computing system can implement efficient and effective video search, and enable case search filters by complexity.
[0043] Referring now to FIG. 1 A-1C (and FIG. 1C in particular), among others, an example system 100 including a computing system 100 to determine complexity of a medical procedure is shown. The system 100 can include at least one computing system 105. The computing system 105 can be a data processing system, a computer, a desktop computer, a control system, a console system, an embedded system, a cloud computing system, or any other type of computing system. The computing system 105 can be an on-premise computingsystem or an off-premises computing system. The computing system 105 can be disposed on-premises within a medical facility. The medical facility can be a hospital, an outpatient center, or any other facility.
[0044] The system 100 can include at least one robotic medical system 110. The robotic medical system 110 can be a robotic system, apparatus, or assembly including at least one instrument. For example, the instrument can include an end or tip, such as a scalpel, a scissors, a monopolar curved scissors (MCS), a cautery hook tip, a cautery spatula tip, a needle driver, forceps, a tooth retractor, a drill, or a clip applier. The instrument can be, include, or be included on a robotic arm, a robotic appendage, a robotic snake, or any other motor controlled member that can be articulated by the robotic medical system. The instrument can include at least one actuator, such as a motor, servo, or other device. The instrument can be manipulated by motors, servos, actuators, or other devices to perform a medical procedure. The robotic medical system 110 can perform a medical session, surgery, therapy, or medical procedure. For example, the robotic medical system 110 can articulate the instrument to perform surgery, therapy, or a medical evaluation with the instrument. A medical practitioner, such as a surgeon, technician, nurse, or other operator can provide input via a user device or input apparatus to manipulate the instrument to perform a medical procedure.
[0045] The robotic medical system 110 can be disposed on-premises within a medical facility. The medical facility can be a hospital, an outpatient center, or any other facility. The robotic medical system 110 can perform any type of medical procedure, such as a surgery, a therapy, or a medical evaluation. The robotic medical system 110 can be disposed at the same facility as the computing system 105. In some implementations, the robotic medical system 110 can be integrated with the computing system 105. For example, the computing system 105 can be a component of the robotic medical system 110.
[0046] The robotic medical system 110 can include at least one camera, in some implementations. The camera can be or include an endoscope. For example, the camera can be an instrument that is manipulated by the medical practitioner and controlled via a motor, servo, or other input device of the robotic medical system 110. The robotic medical system 110 can produce image frames. The image frames can be frames of a procedure video 115 captured by the camera or images taken by the camera. The procedure video 115 can be captured by the camera of the robotic medical system 110 can track the medical procedureperformed by the robotic medical system 110. The procedure video 115 can capture instruments, anatomical structures (e.g., organs, muscles, bones, or skin), or the patient in the field of view of the camera. The robotic medical system 110 can send, transmit, provide, or push the procedure video 115 to the computing system 105. In some implementations, the procedure video 115 can be a video stream that the robotic medical system 110 streams to the computing system 105.
[0047] The robotic medical system 110 can generate a procedure type 120. The procedure type 120 can be a label, tag, or value that uniquely identifies the medical procedure performed by the robotic medical system 110. The procedure type 120 can identify that a medical procedure was a coloscopy, an appendicitis, a cholecystectomy, etc. The procedure type 120 can be a tag applied to the procedure video 115. The robotic medical system 110 can automatically generate the procedure type 120, or can alternatively generate the procedure type 120 based on user input. The procedure type 120 can be input on a console of the robotic medical system 110 via a console. The robotic medical system 110 can provide the procedure type 120 to the computing system 105. In some implementations, the computing system 105 can execute at least one model trained by machine learning to automatically determine the procedure type 120 from the procedure video 115 and / or the system data 185. In some implementations, the label identifying procedure type 120 can be optional depending on the type of the complexity being determined. For example, one complexity type can be agnostic to procedure type 120.
[0048] The robotic medical system 110 can generate system data 185. The system data 185 can include performance indicators 170. In some implementations, the robotic medical system 110 can stream the performance indicators 170 to the computing system 105. In some implementations, the computing system 105 can generate the performance indicators 170 from the system data 185 received from the robotic medical system 110. The various performance indicator types can be an amount of energy or power consumed by the robotic medical system 110. The performance indicators 170 can indicate an amount of energy or power consumed by the robotic medical system 110 to operate an individual instrument or endoscope. The various performance indicators 170 can indicate a total duration of a segment of the medical procedure. For example, the performance indicators 170 can be a length of time of the particular segment of the procedure video 115, such as a length of time of the step, the phase or the entire medical procedure. The performance indicators 170 canindicate a total linear distance of the instrument of the robotic medical system 110 during the segment. For example, the performance indicators 170 can indicate a total distance that the instrument traveled during the steps, the phases, or the entire video. The performance indicators 170 can indicate a total angular distance of the instrument of the robotic medical system 110 during the segment. The angular distance can indicate an amount or distance that the instrument is rotated during manipulations. The performance indicators 170 can indicate a total angular distance that the instrument traveled during the steps, the phases, or the entire video 115. The performance indicators 170 can indicate a total number of operations of a clutch or brake. For example, the performance indicators 170 can indicate a total number of actuators or activations of a clutch or brake for the instrument 195 or the endoscope of the robotic medical system 110. The clutch or brake can be operated to float joints of the instrument or endoscope. The performance indicators 170 can indicate a total number of operations or a brake or clutch during the steps, the phases, or the entire video.
[0049] The computing system 105 can include at least one model selector 125. The model selector 125 can select a first model 130 or additionally select a second model 135 using the procedure type 120. The model selector 125 can receive a label of the procedure type 120 for a particular medical procedure in question that the computing system 105 is determining the complexity of. The model selector 125 can receive the label of procedure type 120 from the robotic medical system 110. The model selector 125 can select a first model 130 from a group or set of other first models 130. The set of first models 130 can be models trained by machine learning, each for a different medical procedure type. For example, one of the first models 130 can be trained for a first procedure type, while another of the first models 130 can be trained for a second procedure type. The model selector 125 can use the procedure type 120 to look-up and retrieve the first model 130 from a database, data repository, or storage system.
[0050] The computing system 105 can include at least one segment selector 140. The segment selector 140 can receive the label of procedure type 120 and the output of the model selector 125 as an input to use in making the selection of the video 115 (although the segment selector 140 can make the selection without using the label of the procedure type 120). The segment selector 140 can select a video clip or segment 145 of the procedure video 115. For example, the computing system 105 can select the segment 145 using the model selector 125 the procedure type 120 of the medical procedure of the procedure video 115. The segmentselector 140 can identify a segment 145 of interest from the procedure video 115 for the models 130 or 135 to make a complexity prediction with. The segment selector 140 can select the segment 145 of the procedure video 115 to use to determine the level of the complexity of the medical procedure. For example, the segment selector 140 can select the segment 145 of the procedure video 115 that includes frames that have a high level of influence on determining the complexity. For example, some frames or portions of the procedure video 115 include information that is not relevant or has a low level of relevance to determine the complexity. For example, the segment selector 140 can select the segment 145 of the procedure video 115 that includes all frames of the procedure, and the model 135 can determine the relevant frames using attention mechanism to determine complexity. For example, the computing system 105 can select the first fifteen minutes of a procedure video 115 to use to determine complexity. The computing system 105 can select from different segment selectors 140 (or configure one segment selector 140) to generate input as frames, clips, etc. and / or select the video segment 145 from which to select the frames, clips, samples, etc. The computing system 105 can select different segment selectors 140 using the procedure type 120 and model selector 125. Furthermore, the computing system 105 can select which trained models 130 and / or 135 should be used based on procedure type 120 or for the desired complexity grading scale. Similarly, the computing system 105 can select the prediction aggregator 195 using the procedure type 120 and model selector 125, e.g., the prediction aggregator 195 can be paired with the models 130 and / or 135, and can be selected along with the models 130 and / or 135.
[0051] For example, portions of the video 115 such as configuration of the robotic medical system 110, insertion of instruments or an endoscope into a patient, etc. may not be relevant or have a low level of influence on determining the complexity. However, other portions of the procedure video 115 may be relevant to determining complexity. For example, a portion of a video where an anatomical structure is transected or removed from a patient may be highly relevant to determining complexity. Furthermore, portions of a video such as insertion of an instrument into a difficult to reach area of a patient may be highly relevant to determining complexity. The segment selector 140 can select the segment 145 or segments 145 such that the most pertinent segments are analyzed by the first model 130 and the second model 135. The segment selector 140 can select segments 145 of procedure videos 115 that are the most relevant to complexity determinations based on data analysis and / or clinical input data. For example, the segment selector 140 can select a fixedproportion of the procedure video 115 such as the first 10% of the procedure video 115 (e.g., without using procedure type to make the selection). The segment selector 140 can select up to the first dissection step (e.g., using procedure type input or procedure agnostic phase segmentations). The segment selector 140 can select a particular clinical step as the video segment (e.g., using procedure type and trained models as inputs).
[0052] In some implementations, the computing system 105 can execute the first model 130 and the second model 135 using the segment 145. In some implementation, the computing system 105 may not segment the procedure video 115, and can instead execute on the entire procedure video 115. In some implementations, the model selector 140 can select the entire procedure 115 as the segment 145, and the model 130 can attend to pertinent segments of the video. For example, image frames 150 of the entire procedure video 115 can be input to the first model 130, and the entire sequence of frames (e.g., where each frame is represented as a feature vector 175 that is extracted by the model 130) of the entire procedure video 115 can be input to the second model 135. The second model 135 can be trained using an attention mechanism to selectively focus on, or weight, the frames most relevant for determining complexity. In some implementations, indicators 170 computed from the system data 185 can be input into the first model 130 and / or the second model 135.
[0053] The segment selector 140 can be a model trained by machine learning. The segment selector 140 can receive a procedure video 115 and a procedure type 120 for the procedure video 115, and output an indication of a segment of the procedure video 115. The segment selector 140 can be a sequence based neural network or a recurrent neural network. For example, the segment selector 140 can be a long-short term memory (LSTM) neural network or a gated recurrent unit (GRU). In some implementations, the segment selector 140 can select phases, tasks, or steps of the procedure video 115 that are most pertinent to determine complexity. The segment selector 140 can execute at least one model trained by machine learning on the procedure video 115 (and / or the procedure type 120) to identify the phases, tasks, or steps and select video segments corresponding to the pertinent phases, tasks, or steps to use in determining complexity.
[0054] The computing system 105 can execute one or models trained by machine learning to predict case complexity. For example, the computing system 105 can execute models with model architectures that include a spatio-temporal modeling using twodimensional (2D) convolutional neural network (CNN), LSTM, and / or transformer based architectures
[0055] For example, the computing system 105 can execute the first model 130 or second model 135 to determine feature data from the segment 145 and determine the level of complexity 160 of the medical procedure. The computing system 105 can include at least one first model 130. The computing system 105 can execute using at least one frame 150 of the segment 145. The computing system 105 can execute the first model 130 to generate or determine a feature vector 175. The first model 130 can be sequentially executed with each frame 150 of the segment 145 as an input to extract a feature vector 175 for each frame 150 of the segment 145. The first model 130 can output the feature data 175. The feature data 140 can be a feature vector or embedding of the image 150. The first model 130 can receive an image frame 150 of the segment 145. The first model 130 can be a single image model 130 that outputs the feature vector 175 based on a single image frame input to the first model 130. The first model 130 can execute on one image at a time, and may not store any memory of past images input to the first model 130. The first model 130 can be a CNN. The first model 130 can be a 2D CNN, and may receive a two dimensional matrix as an input. The first model 130 can be a self-distillation with no labels (DINO) network, a masked Siamese network (MSN), an auto encoder (AE), a decoder-encoder models, a masked auto-encoder (MAE), etc. The first model 130 can be trained by machine learning with a supervised or self-supervised learning technique. The first model 130 can, in some cases, be an off-the- shelf pre-trained model. The features 175 of each frame 150 can be fed into the second model 135. The features 175 can be spatial features in tissue, inflammation of a gallbladder, dense adhesions, how tensely distended a gallbladder is, or color of gallbladder, or other feature.
[0056] The second model 135 can be a LSTM, GRU, or transformer based architecture. The second model 135 can be a temporal model. The second model 135 can produce complexity predictions 173 (or features 175 that capture temporal information for execution by a third model that generates the complexity prediction 173 from the features). The second model 135 can be a sequence based model that weighs data of past inputs or uses past inputs as contextual data to determine outputs for current inputs. The second model 135 can include or implement an attention mechanism to address variability in when (and / or for how long) information relevant to case complexity appears in procedure video 115. The second model 135 can implement scaled dot-product attention, multi -head attention, or any other attentionmechanism. The second model 135 can be trained by machine learning with a supervised or self-supervised learning technique.
[0057] The second model 135 can be a temporal model that considers context from past frames in a segment 145. The temporal model 135 can reduce noise and provide more context in terms of interactions between instruments and organs over the course of the procedure. For example, the context could be salient bleeding that is present over the course of the procedure. Furthermore, complexity determined based on a single image could result in uncertainty between whether the complexity is 3 or 4. However, by analyzing the segment 145 with the temporal model, over time, the model can determine that the complexity is really 4.
[0058] The computing system 105 can train the first model 130 and the second model 135 using a training dataset. The training dataset can be a historical dataset or include historical data. The training dataset can include complexity grades (e.g., numerical scores, letter grades, etc.) for various cases or procedure videos 115. The complexity grades can be standardized across different cases, procedure types, etc. In some implementations, different procedure types 120 can have different types of complexity grades with different ranges or values. In some implementations, a user can review historical procedure videos 115 captured by an endoscope of the robotic medical system 110 and identify the corresponding complexity grade to create the training data. The computing system 105 can determine the sample size of the training data for each complexity grade. The computing system 105 can trigger training if a predefined number of samples of each complexity grade exist in the training dataset.
[0059] The first model 130 can output the feature vector 175 based on the image 150. The computing system 105 can combine, concatenate, or stack the segment 145 with the feature vector 175. The feature vector 175 (and in some cases, the feature vector 175 together with the segment 145) can be input to the second model 135. The second model 135 can execute using the feature vector 175 (and in some cases, the feature vector 175 together with the segment 145, although the segment 145 is an optional input) to determine the case complexity prediction 160. The second model 135 can output feature data 175. The second model 135 can output video clip complexity predictions 173 for video clip 145 (e.g., the entire video clip 145) or individual frames 187. The output complexity predictions 173 can be or can be used to generate the complexity 160 of a particular case. For example, theprediction aggregator 195 can receive the complexity predictions 173 from the model 135, and either aggregate predictions across frames, or directly use the prediction to generate 160, depending on whether the output is frame predictions 187, or a video clip 145.
[0060] The computing system 105 can include at least one indicator selector 180. The indicator selector 180 can select indicators 170 to be used with the feature vector 175 to generate a complexity tag 160. The indicator selector 180 can receive the label of procedure type 120 as an input. The indicator selector 180 can, in some implementations, identify and select the indicators 170 from system data 185 received from the robotic medical system 110. In some implementations, the robotic medical system 110 can produce indicators 170, and send the indicators 170 to the indicator selector 180 as part of the system data 185. In some implementations, the robotic medical system 110 can itself generate the indicators 170, while in other implementations, the system data 185 can include all the necessary information (e.g., kinematics data, a history of movements of instruments, a number of clutch counts, power consumption data, etc.) needed to generate the indicators 170. In some implementations, the indicators 170 are fed into at least one of the first model 130 or the second model 135. In this regard, the indicator selector 180 can identify indicators 170 that are influenced by the complexity for a procedure type 120, and feed the selected indicators 170 into the first model 130 or the second model 135.
[0061] The computing system 105 can receive system data 185, such as data streams of kinematics data, system events, system parameters, etc. The computing system 105 can be input directly into the first model 130 or the second model 135, in some implementations. In some implementations, the system data 185 can be optional depending on the types of the models 130 and 135 selected to execute.
[0062] In some examples, the computing system 105 can stack the selected indicators 170 with the feature vector 175 to generate stacked features. The computing system 105 can train the model 135 using stacked indicators 170 and feature vectors 175. The computing system 105 can execute the model 135, using the stacked feature data including the selected performance indicators 170, to determine the level of complexity 173. In some examples, the second model 135 may output a feature vector based on the stacked information. In some implementations, the output of the second model 135 can be stacked with further data. In such implementations, a third model can be used to determine the complexity prediction 173 or the complexity prediction 160 using the stacked information. The stacked features can bea concatenation or combination of the selected indicators 170 and the feature vector 175. The computing system 105 can determine or predict the complexity tag 160 from the stacked features. In some implementations, the stacked features can be or include a data structure that plots the complexity values of the feature vector 175 against the selected indicators 170. The computing system 105 can generate a complexity tag 160, and associate the complexity tag to the case 115. The computing system 105 can associate the complexity tag 160 to the case 115 stored in a database.
[0063] The complexity tag 160 can include a complexity level or grade, such as a numerical value (e.g., 0-5, 0-100, etc.) or a letter grade. The grades can be Parkland grades (e.g., a grading scale of five grades) which quantify the complexity of a cholecystectomy procedure. In some implementations, the computing system 105 can grade a procedure video 115 multiple times for different complexity scores ranges or score types. In some implementations, the computing system 105 can include a different set of models for each of the score types.
[0064] In some implementations, the computing system 105 can execute at least one model to classify case complexity that utilize features from existing step segmentation models. For example, the computing system 105 can include a model that segments procedure videos based on surgical steps, surgical phases, etc. The identified steps or phases can be used by other models (e.g., the first model 130 and the second model 135) to determine case complexity. For example, a model can be trained to predict surgical steps such as dissection of Calot’s triangle in cholecystectomy could be used to assess the state of the gallbladder. For example, the computing system 105 can implement at least one machine learning model to determine steps or phases of a medical procedure (e.g., a step segmentation model) and the computing system 105 can determine the complexity from the identified steps or phases.
[0065] In some implementations, the computing system 105 can execute at least one model to identify or recognize anatomy. For example, the computing system 105 can identify anatomy or states of the anatomy in the procedure video 115. In some implementations, the computing system 105 can classify case complexity using features from anatomy recognition models. For example, the computing system 105 can localize frames where the gallbladder is present and use the localized frame to predict case complexity. The computing system 105can determine the complexity from anatomy detections determined by at least one anatomy recognition model.
[0066] In some implementations, the computing system 105 can execute in real-time or intra-operatively as the medical procedure is being performed. The computing system 105 can be implemented in real-time an output a complexity tag or level to an operator of the robotic medical system 110. In some implementations, the computing system 105 can execute after the medical procedure is completed.
[0067] Referring now more particularly to FIG. 1 A, among others, an example computing system 105 to determine a complexity 160 of a medical procedure using a model 130 is shown. The computing system 105 can deactivate, or may not run, the segment selector 140 to segment the procedure video 115 into clips or segments. Instead, the model 130 can execute using the entire procedure video 115 to output a case complexity prediction 160 for the video 115. The case complexity prediction 160 can be one prediction quantifying complexity of the entire video 115. The model 130 can receive a single input (the procedure video 115 or a portion of the procedure video 115) and directly produce or output the case complexity prediction 160. In some implementations, the model 130 executes on a given segment 145, the model 130 can output a complexity 160 for the entire video 115. For example, using at least a portion of the video 115, the model 130 can infer or determine a complexity 160 for the entire video 115.
[0068] In another example, the computing system 105 can run the segment selector 140 to segment the procedure video 115 into clips, segments or the entire procedure video, which can collectively be referred to as video clip 145 or segment 145, as depicted in FIG. 1C. The model 130 can execute using the segment 145 to output a case complexity prediction 160 for the video 115. The case complexity prediction 160 can be a prediction quantifying complexity of the entire video 115. The model 130 can receive a single input (e.g., the procedure video 115 or a portion of the procedure video 115) and directly produce or output the case complexity prediction 160. In some implementations, the model 130 executes on a given segment (e.g., video clip 145), the model 130 can output a complexity 160 for the entire video 115. For example, using at least a portion of the video 115, the model 130 can infer or determine a complexity 160 for the entire video 115.
[0069] Referring now more particularly to FIG. IB, among others, an example computing system 105, including a model 130 to determine complexity 187 of multiple frames of a procedure video 115 and an aggregator 195 to aggregate the frame complexities 187 into a case complexity 160 is shown. The computing system 105 can sample the procedure video 115 to divide the video 115 into individual images or frames 150 or individual clips of the video 115. The computing system 105 can produce a sequence of frames 150 or a sequence of clips. The model 130 can receive individual frames of the video 115. The model 130 can receive individual frames of the entire video 115, or of a segment or clip of the video 115. Furthermore, the model 130 can receive individual video segments or clips of the longer video 115 as inputs.
[0070] The model 130 can receive the video frame 150 as a single frame input. The model 130 can output a complexity prediction 187 for the input frame 150. The model 130 can output a single complexity prediction 187 for each of the input frame 150. The model 130 can be executed by the computing system 105 multiple times. For example, the computing system 105 can execute the model 130 once for each video frame 150 of the video 115. The result can be a sequence of video frame complexity predictions 187, one for each video frame 150.
[0071] The computing system 105 can include at least one prediction aggregator 195. The prediction aggregator 195 can aggregate multiple video frame complexity predictions 187 for multiple video frames 150 together to produce a case complexity prediction 160. For example, the prediction aggregator 195 can combine multiple video frame complexity predictions 187 by determining a mean, determining a median, executing a voting algorithm, or executing a scoring algorithm using the multiple video frame complexity predictions 187. The case complexity prediction 160 can be the mean, median, or score determined from the video frame complexity predictions. The prediction aggregator 195 can aggregate all the frame-wise predictions 187 into a single case complexity prediction 160. The predictions 187 can be the same as or similar to the prediction 160, but generated on the video frame level instead of the entire case level.
[0072] Referring now more particularly to FIG. 1C, among others, an example computing system 105 including a first model 130 to determine video frame features 175 and a second model 135 to determine video clip complexity 173 is shown. In FIG. ID, the model 130 and the model 135 can be stacked. The computing system 105 can stack multiple models togetherto determine the case complexity prediction 160, e.g., one, two, three, or any number of models. In FIG. ID, the first model 130 can operate on spatial data, e.g., can be a single image model, while the second model 135 can be a temporal model or include temporal components. The segment selector 140 can segment or sample the video 115 into a series of segments or video clips 150. Furthermore, the segment selector 140 can segment or sample the video clip 145 into individual frames 150.
[0073] In some implementations, the segment selector 140 can execute a stacking of machine learning models. For example, the segment selector 140 can execute a process to sample the video segment into a set of single images. The segment selector 140 can execute, using a single image as input, a pre-trained machine learning model to extract a feature vector that represents low-dimensional spatial features of the image. The segment selector 140 can execute, using feature vectors as input, a second machine learning model trained to predict case complexity on the video segment.
[0074] The first model 130 can output video frame features 175 for each video frame 150. The model 130 can execute multiple times, once for each frame 150 of each video clip 145. For example, the segment selector 140 can segment the procedure video 115 into a clip 145 of multiple frames 150. Each frame 150 can be individually or sequentially input to the model 130, and the model 130 can output video frame features 175 (e.g., a feature vector) for each input video frame 150. Each feature vector 175 can provide a compressed representation or low-dimensional representation of each video frame 150. The feature vectors 175 can advantageously be smaller than the respective video frames 150, but still capture the information of the video frames 150 relevant to determine the complexity 160.
[0075] Instead of taking the complexity prediction 187 as output from the first model 130, like in FIG. IB, the video frame features 175 can provide intermediate results for the second model 135 to operate on. The computing system 105 can stack feature vectors 175 for each respective video clip 145 to generate video clip features 183. For example, the computing system 105 can generate the video clip features 183 by organizing the feature vectors 175 into a single data element, data structure, or set of storage addresses that correspond to one video clip 145. For example, the computing system 105 can stack the feature vectors 175 of frames 150 corresponding to one video clip 145 into a data structure 183 for each individual video clip 145.
[0076] The computing system 105 can execute the second model 135 for each respective video clip 145. The computing system 105 can execute the second model 135 with each video clip feature set 183 as a single respective input to produce a complexity prediction 173 of each video clip 145. The video clip complexity predictions 173 can be provided to the prediction aggregator 195, that can aggregate the predictions 173 together to generate the case complexity prediction 160 for the entire procedure vide 115. The predictions 173 can be the same as or similar to the prediction 160, but generated on the video clip level instead of the entire case level.
[0077] Referring now to FIGS. 2A-2B, among others, an example computing system 105 to surface complexity determinations for a medical procedure to a user and to re-train a machine learning model is shown. The computing system 105 can include or be integrated a variety of components, systems, devices, or elements. For example, the computing system 105 can include or be integrated with at least one network storage system 205. The network storage system 205 can be a cloud storage system. For example, the computing system 105 can store procedure videos 115 or system data 185 collected or received from the robotic medical system 110 in the network storage system 205. The computing system 105 can receive, retrieve, or access the procedure videos 115 or the system data 185 via a network connection, such as a large area network, a wide area network, the Internet, etc. The computing system 105 can determine case complexity and model re-training with a low or minimal amount of user input.
[0078] The computing system 105 can include or be integrated with at least one database 210. The database 210 can be a relational database management system (RDMS), a not only RDMS (noSQL) database, a vector database, a graph database, a key -value database, etc. The database 210 can store procedure types 120 for various procedure videos 115 and procedure indicators 170 for the various procedure videos 115. For example, for a particular case, the database 210 can store the procedure type 120 of the case, and the procedure indicators 170 of the case. Furthermore, the network storage 205 can store the procedure video 115 for the particular case.
[0079] The computing system 105 can include or be integrated with at least one machine learning engine 215. The machine learning engine 215 can be a piece of software or can be hardware that implements training or executing of the first model 130, the second model 135, or any other machine learning model. The machine learning engine 215 can be a graphicsprocessing unit (GPU) based computing system, a central processing united (CPU) based computing system, a neural processing unit (NPU) based computing system, a tensor processing unit (TPU) based computing system, or any other type of computing system or artificial intelligence (Al) accelerator.
[0080] The machine learning engine 215 can include or run at least one complexity predictor 230. The complexity predictor 230 can be a piece of software, a software module, or a set of instructions that executes to generate or predict a complexity for a particular medical procedure case. The complexity predictor 230 can include the model selector 125, the segment selector 140, the indicator selector 180, the first model 130, the second model 135, etc. The complexity predictor 230 can output a complexity grade 210 for a given case based on the procedure video 115 for the case, the procedure type 120 of the case, and the procedure indicators 170 for the case.
[0081] The computing system 105 can include at least one quality control selector 235. The quality control selector 235 can operate based on the results of model evaluation. The quality control selector 235 can select cases for quality control with a higher probability for complexity grades with lower accuracy or higher confusion with other grades. The quality control selector 235 can be based on confusion matrices, model confidence, or other metrics of model performance. The quality control selector 235 can determine whether a complexity grade prediction 240 should be surfaced for user or expert evaluation. The quality control selector 235 can determine or receive an accuracy or confusion level for a case. If the complexity grade prediction 240 is high, e.g., greater than a first threshold, and the accuracy or confusion levels is low (e.g., less than a second threshold), the quality control selector 235 can determine that the case should be surfaced for user review. The quality control selector 235 can analyze historic model performance for a particular procedure type, complexity level, and / or model confidence score on a particular case to make selections for cases to surface for expert review. The quality control selector 235 can select cases for expert review via a random or pseudo-random selection.
[0082] The quality control selector 235 can determine a confusion matrix for the performance of the complexity predictor 230. The confusion matrix can be a matrix that identifies a confidence level or accuracy level for each predicted class and the actual class. The quality control selector 235 can use the confusion matrix to determine whether to submit the video for review.
[0083] The computing system 105 can include at least one graphic user interface (GUI) system 220. The GUI system 220 can generate data to cause a GUI 250 displayed on a client device 245 to display at least a portion of the procedure video 115 (such as the segment 145) (and / or optionally displaying the complexity grade prediction 240) to a user. The user can review the procedure video 115 and confirm whether the complexity grade prediction 240 is correct, or alternatively modify, edit, or change the grade for the procedure video 115. The computing system 105 can generate a finalized complexity grade 260 based on the input received via the GUI 250.
[0084] The computing system 105 can include at least one re-trainer 255. The re-trainer 255 can update and re-evaluate models based on the addition of manual quality control data. The re-trainer 255 can receive and collect finalized complexity grades 260. The re-trainer 255 can implement a continuous re-training loop for the complexity predictor 230. The retrainer 255 can collect expert or user input, and generate a training data set. The training data set can store, for a particular case, a procedure type 120, a procedure video 115, procedure indicators 170, and the expert or user defined finalized complexity grade 260. The re-trainer 255 can execute at least one training algorithm, such as back propagation, gradient descent, etc. to update the weights or values of the models of the complexity predictor 230, e.g., the first model 130 and the second model 135. The re-trainer 255 can wait until a predefined number of samples have been accumulated in the training data set, and trigger re-training responsive to the predefined number of samples satisfying a threshold. The re-trainer 255 can automate model re-training and re-evaluation.
[0085] The computing system 105 can generate a complexity tag 160 for the procedure video 115. For example, the computing system 105 can set the complexity tag 160 for the video to be the complexity grade prediction 240 if the quality control selector 235 determines not to raise the complexity grade prediction 240 for expert or user review. The computing system 105 can set the complexity tag 160 to be a modified version of the complexity grade prediction 240 if the quality control selector 235 raises the complexity grade prediction 240 for expert or user review and the expert or user modifies the complexity grade prediction 240. The computing system 105 can set the complexity tag 160 to be the complexity grade prediction 240 if the quality control selector 235 raises the complexity grade prediction 240 for expert or user review and the expert or user modifies the complexity grade prediction 240. The computing system 105 can set the complexity tag 160 to be the complexity gradeprediction 240 if the quality control selector 235 raises the complexity grade prediction 240 for expert or user review and the expert or user confirms the complexity grade prediction 240 is correct.
[0086] The computing system 105 can include at least one front-end system 225. The front-end system 225 can implement one or multiple software applications, control systems, software services, etc. The front-end system 225 can provide front-end or user interface services for an expert or user on the client device 245. The client device 245 can be a smartphone, a laptop, a tablet, a console, a desktop computer, a data processing system, etc. The front-end system 225 can include various services 265-290 that can operate to display data on the client device 245 or interact with the client device 245. The services 265-290 can be surgeon performance tools, such as trends, workflow trends, recommendations, or search functions that incorporate case complexity. The services 265-290 can incorporate case complexity for a more accurate interpretation of performance trends, workflow trends, and more effective and efficient video recommendations or video searches.
[0087] For example, a trend service 265 can generate data to display various trends of procedure indicators 170. The trends generated by the trend service 265 can be generated based on the complexity tags 160 for various cases. For example, the performance indicators 170 can be normalized against the complexity tags 160, such that for a particular surgeon, hospital, or entity, the performance indicators 170 can be better understood. The front-end system 225 include a workflow service 270 that trends workflows based on the complexity tag 160. The workflow service 270 can recommend changes to a surgical workflow, e.g., different phases, segments, or steps that reduce complexity of a medical procedure. For example, the workflow service 270 can execute on the procedure videos 115 and their corresponding complexity tags 160 and determine, based on the historical data, that changing one step to another step, the complexity of the medical procedure can be reduced. The frontend system 225 can include an analysis service 280 that analyzes performance indi ctors 170 based on complexity tags 160.
[0088] The front-end system 225 can include a recommendation service 285 that can generate recommendations based on complexity. For example, based on the complexity of a case, the front-end system 225 can recommend that a surgeon perform a specific act or step during surgery based on the complexity tag 160. The front-end system 225 can include ahighlight service 290 that can generate data to display highlights of surgeon or hospital cases based on their complexity tags 160.
[0089] The search service 275 can execute video searching based on complexity tags 160. For example, a client device 245 can provide data to the front-end system 225 that defines a search query. In some implementations, the client device 245 can provide the front-end system 225 the search query. The search query can be a prompt that requests a specific procedure video 115 or a specific set of procedure videos 115 that meet at least one criteria. The criteria can include an indication of a type of the medical procedure (e.g., a coloscopy, an appendicitis, a cholecystectomy, etc.). The criteria can include an indication of a complexity level. For example, the criterial can be a complexity grade, such as a numerical value 1-5, a numerical value 1-100, a letter grade, etc. The search service 275 can search through the network storage 205 or database 210 for procedure videos 115 that satisfy the criteria, or are sufficiently similar to the criteria. The search service 275 can select a set of videos 115 that are of the identified medical procedure type and have the corresponding complexity level. The search service 275 can return the identified procedure videos 115 to the user. The search service 275 can cause the GUI 250 to display the returned videos 115 to the user on the client device 245.
[0090] Referring now to FIG. 3, among others, a confusion matrix 300 for a machine learning model trained to determine complexity of a medical procedure is shown. The computing system 105 can generate the confusion matrix 300. The computing system 105 can generate the confusion matrix 300 from a trained first model 130 and / or second model 135. The confusion matrix 300 can indicate historical performance of the first model 130 and / or the second model 135. For example, the computing system 105 can train the first model 130 and the second model 135 together, or separately. The computing system 105 can generate or receive a training dataset that includes procedure types 120, procedure videos 115, system data 185 (or indicators 170), and labeled complexity tags 160 for the videos 115. For example, for each procedure video 115, the training data set can include a complexity tag 160 for the procedure video 115, a procedure type 120 of the procedure video 115, and system data 185 or indicators 170 for the procedure video 115. The training dataset can to train the first model 130 and / or the second model 135 and use the testing set to generate the confusion matrix 300. The training dataset can be based on human-annotated labels of case complexity, in order to select the best performing model. The training dataset can begenerated with a tagging process based on human annotations, generated from an internal annotator or an external product user.
[0091] Part of training the first model 130 and / or the second model 135 can include testing or validating the performance of the first model 130 and / or the second model 135. This can include generating the confusion matrix 300. The confusion matrix 300 can indicate the type and frequency of prediction errors. The confusion matrix 300 can identify true positives, false positives, true positive rates, false positive rates, Matthews correlation coefficients (MCC), Fowlkes-Mallows indexes (FMs), etc. The confusion matrix 300 can identify, for data samples that should be classified to an actual complexity level, the rate or probability that the first model 130 and second model 135 will predict each of the available classes. The confusion matrix 300 can be a confusion matrix 300 for a first model 130 and a second model 135 trained to determine the complexity of a cholecystectomy. The grades can be percentages, numerical values (e.g., 0-5, 0-10 0-100, etc.) or letter grades. The grades can be Parkland grades (e.g., a grading scale of five grades) which quantify the complexity of a cholecystectomy procedure.
[0092] Referring now to FIG. 4, among others, example charts 405-420 of performance indicators for medical procedures plotted against medical procedure complexity are shown. The charts 405-420 can be generated by the front-end system 225. The charts 405-420 can plot the procedure indicators 170 against Parkland grades. The charts 405-420 can be grades (e.g., Parkland grades) and procedure indicators 170 for a cholecystectomy. The charts 405- 420 can indicate the influence which complexity of medical procedures has on procedure indicators 170, thereby providing an interpretation of performance trends.
[0093] For example, the computing system 105 can generate a first chart 405. The first chart 405 can plot procedure indicators 170 (e.g., duration) for various procedures against parkland grades. The first chart 405 can indicate a duration of the procedure, a duration of a step of the procedure, or a duration of a phase of the procedure in seconds, minutes, or hours against parkland grades for the various procedures. For example, the computing system 105 can generate a second chart 410. The second chart 410 can plot procedure indicators 170 (e.g., path length) for various procedures against parkland grades. The second chart 410 can indicate a path length that an instrument or a combination of instruments moved during the procedure, during a step of the procedure, or during a phase of the procedure against parkland grades for the various procedures.
[0094] For example, the computing system 105 can generate a third chart 415. The third chart 415 can plot procedure indicators 170 (e.g., camera clutches) for various procedures against parkland grades. The third chart 415 can indicate a number of clutch operations for a camera or endoscope during an entire procedure, a step of the procedure, or a phase of the procedure against parkland grades for the various procedures. For example, the computing system 105 can generate a fourth chart 420. The fourth chart 420 can plot procedure indicators 170 (e.g., energy consumed by the robotic medical system 110) for various procedures against parkland grades. The fourth chart 420 can indicate energy consumed by the robotic medical system 110 during the procedure, duration of a step of the procedure, or a phase of the procedure against parkland grades for the various procedures.
[0095] Referring now to FIG. 5, among others, an example method 500 of determining complexity of a medical procedure is shown, according to an exemplary embodiment. At least a portion of the method 500 can be performed by the computing system 105, the robotic medical system 110, a data processing system, or any other kind of computing system, device, or apparatus. At least a portion of the method 500 can be performed by the network storage 205, the database 210, the machine learning engine 215, the GUI system 220, or the front-end system 225. At least a portion of the method 500 can be performed by the model selector 125, the segment selector 140, the indicator selector 180, the first model 130, or the second model 135. The method 500 can include an ACT 505 of receiving a video of a medical procedure. The method 500 can include an ACT 510 of selecting a segment of a video based on a medical procedure type. The method 500 can include an ACT 515 of executing at least one model to determine a level of complexity.
[0096] At ACT 505, the method 500 can include receiving, by the computing system 105, a video 115 of a medical procedure. The method 500 can include receiving the procedure video 115 from the robotic medical system 110. The method 500 can include receiving the procedure video 115 intraoperatively, e.g., in real-time in order for the computing system 105 to execute in real-time. The method 500 can include receiving the procedure video 115 after the robotic medical system 110 performs the medical procedure. The method 500 can include receiving the medical procedure video 115 from a camera or endoscope of the robotic medical system 110 that recorded the video from outside or from within a patient.
[0097] At ACT 510, the method 500 can include selecting, by the computing system 105, a segment 145 of a video 115 based on a medical procedure type 120. The method 500 caninclude selecting a portion of the medical procedure video 115 by implementing one or more rules, one or more models trained by machine learning, or one or more selection algorithms. For example, the method 500 can execute a model that identifies a starting time and an ending time in the entire video 114 to use as the segment 145. The model can be trained on procedure type 120, such that model selects the segment 145 to identify relevant portions of the medical procedure pertinent to the identified type of procedure. In some implementations, the computing system 105 can store different selection models for different types of medical procedures. The method 500 can include selecting a segment selection model from a set of segment selection models all trained for different medical procedure types 120.
[0098] At ACT 515, the method 500 can include executing, by the computing system 105, at least one model to determine a level of complexity. The method 500 can include executing one or a set of models using the procedure type 120, the procedure video 115, and / or the indicators 170 to determine a complexity tag 160 for the procedure video 115. The method 500 an include executing a first model 130 on a frame 150 of the segment 145. The method 500 can include executing the first model 130 on each individual frame of the segment 145. The method 500 can include generating a feature vector 175. The method 500 can include generating a feature vector 175 for each frame 150 of the segment 145. The method 500 can include combining at least one feature vector 175 with the segment 145. The method 500 can include applying the feature vector 175 (or the feature vector 175 combined with the segment 145) to the second model 135. The second model 135 can output a complexity prediction 173 (or a second feature vector if a third model is used to execute on the second feature vector). The method 500 can include aggregating video clip complexity predictions 173 (e.g., complexity predictions 173 for individual video clips 145) together to determine a case complexity prediction 160 for the entire video 115 or the case for which the video 115 is captured.
[0099] The method 500 can include stacking the feature vector 175 with selected indicators 170. The stacked features can be applied as a complexity tag 160 to the procedure video 115. In some implementations, another model can execute on the stacked features to determine a complexity level, which can be applied as a complexity tag 160 to the procedure video 115. The method 500 can collect a database of multiple procedure videos 115 with complexity tags 160. The method 500 can include executing various applications or serviceswith the complexity tags 160 and the procedure videos 115, e.g., recommending procedure videos 115 based on complexity tags 160, querying procedure videos 115 using complexity, etc.
[0100] Referring now to FIG. 6, among others, an example block diagram of a computing system 105 is shown. The computing system 105 can include or be used to implement a data processing system or its components. The architecture described in FIG. 6 can be used to implement the computing system 105, the robotic medical system 110, or the client device 245. The computing system 105 can include at least one bus 625 or other communication component for communicating information and at least one processor 630 or processing circuit coupled to the bus 625 for processing information. The computing system 105 can include one or more processors 630 or processing circuits coupled to the bus 625 for processing information. The computing system 105 can include at least one main memory 610, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 625 for storing information, and instructions to be executed by the processor 630. The main memory 610 can be used for storing information during execution of instructions by the processor 630. The computing system 105 can further include at least one read only memory (ROM) 615 or other static storage device coupled to the bus 625 for storing static information and instructions for the processor 630. A storage device 620, such as a solid state device, magnetic disk or optical disk, can be coupled to the bus 625 to persistently store information and instructions.
[0101] The computing system 105 can be coupled via the bus 625 to a display 600, such as a liquid crystal display, or active matrix display. The display 600 can display information to a user. An input device 605, such as a keyboard or voice interface can be coupled to the bus 625 for communicating information and commands to the processor 630. The input device 605 can include a touch screen of the display 600. The input device 605 can include a cursor control, such as a mouse, a trackball, or cursor direction keys, for communicating direction information and command selections to the processor 630 and for controlling cursor movement on the display 600.
[0102] The processes, systems and methods described herein can be implemented by the computing system 105 in response to the processor 630 executing an arrangement of instructions contained in main memory 610. Such instructions can be read into main memory 610 from another computer-readable medium, such as the storage device 620. Execution ofthe arrangement of instructions contained in main memory 610 causes the computing system 105 to perform the illustrative processes described herein. One or more processors in a multiprocessing arrangement can be employed to execute the instructions contained in main memory 610. Hard-wired circuitry can be used in place of or in combination with software instructions together with the systems and methods described herein. Systems and methods described herein are not limited to any specific combination of hardware circuitry and software.
[0103] Although an example computing system has been described in FIG. 6, the subject matter including the operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.
[0104] Some of the description herein emphasizes the structural independence of the aspects of the system components or groupings of operations and responsibilities of these system components. Other groupings that execute similar overall operations are within the scope of the present application. Modules can be implemented in hardware or as computer instructions on a non-transient computer readable storage medium, and modules can be distributed across various hardware or computer based components.
[0105] The systems described above can provide multiple ones of any or each of those components and these components can be provided on either a standalone system or on multiple instantiations in a distributed system. In addition, the systems and methods described above can be provided as one or more computer-readable programs or executable instructions embodied on or in one or more articles of manufacture. The article of manufacture can be cloud storage, a hard disk, a CD-ROM, a flash memory card, a PROM, a RAM, a ROM, or a magnetic tape. In general, the computer-readable programs can be implemented in any programming language, such as LISP, PERL, C, C++, C#, PROLOG, Python, or in any byte code language such as JAVA. The software programs or executable instructions can be stored on or in one or more articles of manufacture as object code.
[0106] Example and non-limiting module implementation elements include sensors providing any value determined herein, sensors providing any value that is a precursor to a value determined herein, datalink or network hardware including communication chips,oscillating crystals, communication links, cables, twisted pair wiring, coaxial wiring, shielded wiring, transmitters, receivers, or transceivers, logic circuits, hard-wired logic circuits, reconfigurable logic circuits in a particular non-transient state configured according to the module specification, any actuator including at least an electrical, hydraulic, or pneumatic actuator, a solenoid, an op-amp, analog control elements (springs, filters, integrators, adders, dividers, gain elements), or digital control elements.
[0107] The subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. The subject matter described in this specification can be implemented as one or more computer programs, e.g., one or more circuits of computer program instructions, encoded on one or more computer storage media for execution by, or to control the operation of, data processing apparatuses. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. While a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be, or be included in, one or more separate components or media (e.g., multiple CDs, disks, or other storage devices including cloud storage). The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0108] The terms “computing device”, “component” or “data processing apparatus” or the like encompass various apparatuses, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that createsan execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.
[0109] A computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program can correspond to a file in a file system. A computer program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0110] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatuses can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Devices suitable for storing computer program instructions and data can include non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0111] The subject matter described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a clientcomputer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification, or a combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0112] While operations are depicted in the drawings in a particular order, such operations are not required to be performed in the particular order shown or in sequential order, and all illustrated operations are not required to be performed. Actions described herein can be performed in a different order.
[0113] Having now described some illustrative implementations, it is apparent that the foregoing is illustrative and not limiting, having been presented by way of example. In particular, although many of the examples presented herein involve specific combinations of method acts or system elements, those acts and those elements may be combined in other ways to accomplish the same objectives. ACTs, elements and features discussed in connection with one implementation are not intended to be excluded from a similar role in other implementations.
[0114] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of “including” “comprising” “having” “containing” “involving” “characterized by” “characterized in that” and variations thereof herein, is meant to encompass the items listed thereafter, equivalents thereof, and additional items, as well as alternate implementations consisting of the items listed thereafter exclusively. In one implementation, the systems and methods described herein consist of one, each combination of more than one, or all of the described elements, acts, or components.
[0115] Any references to implementations or elements or acts of the systems and methods herein referred to in the singular may also embrace implementations including a plurality of these elements, and any references in plural to any implementation or element or act herein may also embrace implementations including only a single element. References in the singular or plural form are not intended to limit the presently disclosed systems or methods,their components, acts, or elements to single or plural configurations. References to any ACT or element being based on any information, act or element may include implementations where the act or element is based at least in part on any information, act, or element.
[0116] Any implementation disclosed herein may be combined with any other implementation or example, and references to “an implementation,” “some implementations,” “one implementation” or the like are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic described in connection with the implementation may be included in at least one implementation or example. Such terms as used herein are not necessarily all referring to the same implementation. Any implementation may be combined with any other implementation, inclusively or exclusively, in any manner consistent with the aspects and implementations disclosed herein.
[0117] References to “or” may be construed as inclusive so that any terms described using “or” may indicate any of a single, more than one, and all of the described terms. References to at least one of a conjunctive list of terms may be construed as an inclusive OR to indicate any of a single, more than one, and all of the described terms. For example, a reference to “at least one of ‘A’ and ‘B’” can include only ‘A’, only ‘B’, as well as both ‘A’ and ‘B’. Such references used in conjunction with “comprising” or other open terminology can include additional items.
[0118] Where technical features in the drawings, detailed description or any claim are followed by reference signs, the reference signs have been included to increase the intelligibility of the drawings, detailed description, and claims. Accordingly, neither the reference signs nor their absence have any limiting effect on the scope of any claim elements.
[0119] Modifications of described elements and acts such as variations in sizes, dimensions, structures, shapes and proportions of the various elements, values of parameters, mounting arrangements, use of materials, colors, orientations can occur without materially departing from the teachings and advantages of the subject matter disclosed herein. For example, elements shown as integrally formed can be constructed of multiple parts or elements, the position of elements can be reversed or otherwise varied, and the nature or number of discrete elements or positions can be altered or varied. Other substitutions, modifications, changes and omissions can also be made in the design, operating conditionsand arrangement of the disclosed elements and operations without departing from the scope of the present disclosure.
Claims
CLAIMSWhat is claimed is:
1. A system, comprising: one or more processors, coupled with memory, to: receive a video of a medical procedure performed by a robotic medical system; select, based on a type of the medical procedure, a segment of the video to determine a level of complexity of the medical procedure with; execute, using the segment of the video, one or more models trained by machine learning to determine feature data; and determine, using the feature data and the one or more models, the level of complexity of the medical procedure.
2. The system of claim 1, comprising the one or more processors to: execute, using at least a frame of the segment, a first model trained by machine learning to determine a feature vector for the frame; and execute, using the feature vector of the frame, a second model trained by machine learning to determine the level of complexity of the medical procedure.
3. The system of claim 1, comprising the one or more processors to: receive a label of the type of the medical procedure; and select, using the label, the one or more models from a plurality of different models trained by machine learning for a plurality of different procedure types.
4. The system of claim 1, comprising the one or more processors to: receive system data of the medical procedure from the robotic medical system; select performance indicators from the system data; stack the feature data with the selected performance indicators; and train, using the stacked feature data including the selected performance indicators, the one or more models using machine learning.
5. The system of claim 1, comprising the one or more processors to: receive a search query for videos of medical procedures, the search query comprising an indication of the type of the medical procedure and the level of complexity; select, responsive to the search query, the video of the medical procedure; and provide, for display via a graphical user interface, a graphical user interface element comprising an indication of the video.
6. The system of claim 1, comprising the one or more processors to: determine to submit the video for review based on the level of complexity exceeding a first threshold and an accuracy level of the one or more models being less than a second threshold.
7. The system of claim 6, comprising the one or more processors to: generate a confusion matrix from performance of the one or more models; and use the confusion matrix to determine to submit the video for review.
8. The system of claim 1, comprising the one or more processors to: transmit, based on the level of complexity, the video for review; and generate data to cause a graphical user interface to display the video and the level of complexity responsive to the level of complexity exceeding a first threshold and the accuracy level being less than a second threshold.
9. The system of claim 8, comprising the one or more processors to: receive, from the graphical user interface, an update to the level of complexity; and re-train the one or more models using the update to the level of complexity.
10. The system of claim 1, comprising the one or more processors to: generate a feature vector using a first model of the one or more models; and generate, using the feature vector, the level of complexity using a second model of the one or more models.
11. The system of claim 10, wherein: the first model is a single image model; and the second model is a temporal model.
12. A method, comprising: receiving, by one or more processors coupled with memory, a video of a medical procedure performed by a robotic medical system; selecting, by the one or more processors, based on a type of the medical procedure, a segment of the video to determine a level of complexity of the medical procedure with; executing, by the one or more processors, using the segment of the video, one or more models trained by machine learning to determine feature data; and determining, by the one or more processors, using the feature data and the one or more models, the level of complexity of the medical procedure.
13. The method of claim 12, comprising: executing, by the one or more processors, using at least a frame of the segment, a first model trained by machine learning to determine a feature vector for the frame; and executing, by the one or more processors, using the feature vector of the frame, a second model trained by machine learning to determine the level of complexity of the medical procedure.
14. The method of claim 12, comprising: receiving, by the one or more processors, a label of the type of the medical procedure; and selecting, by the one or more processors, using the label, the one or more models from a plurality of different models trained by machine learning for a plurality of different procedure types.
15. The method of claim 12, comprising: receiving, by the one or more processors, system data of the medical procedure from the robotic medical system;selecting, by the one or more processors, performance indicators from the system data; stacking, by the one or more processors, the feature data with the selected performance indicators; and training, by the one or more processors, using the stacked feature data including the selected performance indicators, the one or more models using machine learning.
16. The method of claim 12, comprising: receiving, by the one or more processors, a search query for videos of medical procedures, the search query comprising an indication of the type of the medical procedure and the level of complexity; selecting, by the one or more processors, responsive to the search query, the video of the medical procedure; and providing, by the one or more processors, for display via a graphical user interface, a graphical user interface element comprising an indication of the video.
17. The method of claim 12, comprising: determining, by the one or more processors, to submit the video for review based on a type and frequency of prediction errors of the one or more models; generating, by the one or more processors, data to cause a graphical user interface to display the video ; receiving, by the one or more processors, from the graphical user interface, an update to the level of complexity; and re-training, by the one or more processors, the one or more models using the update to the level of complexity.
18. The method of claim 12, comprising: generating, by the one or more processors, a feature vector using a first model of the one or more models; and generating, using the feature vector by the one or more processors, the level of complexity using a second model of the one or more models.
19. A non-transitory computer-readable medium storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to: receive a video of a medical procedure performed by a robotic medical system; select, based on a type of the medical procedure, a segment of the video to determine a level of complexity of the medical procedure with; execute, using the segment of the video, one or more models trained by machine learning to determine feature data; and determine, using the feature data and the one or more models, the level of complexity of the medical procedure.
20. The non-transitory computer-readable medium of claim 19, wherein the processorexecutable instructions further include instructions to cause the one or more processors to: execute, using at least a frame of the segment, a first model trained by machine learning to determine at least a feature vector for the frame; and execute, using the feature vector of the frame, a second model trained by machine learning to determine the level of complexity of the medical procedure.
Citation Information
Patent Citations
Complexity analysis and cataloging of surgical footage
US20200273561A1