Deep Learning-Based Real-Time Remaining Surgical Duration (RSD) Estimation

The system addresses the inefficiencies of existing RSD prediction methods by randomly sampling frames from endoscopic videos and using a trained model to achieve precise, real-time RSD estimation, improving OR efficiency and patient flow.

JP7708354B2Active Publication Date: 2025-07-15AURIS HEALTH INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023558189
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-22
Filing Date
2022-01-18
Publication Date
2025-07-15
Estimated Expiration
2042-01-18

AI Technical Summary

Technical Problem

Existing methods for predicting remaining surgical duration (RSD) are labor-intensive due to manual frame labeling and struggle with accurately representing long surgical videos using recurrent neural networks, leading to insufficient prediction accuracy.

Method used

A system that randomly samples N-1 additional frames from an endoscopic video feed, combines them with the current frame, and uses a trained machine learning model to predict RSD, incorporating a low-pass filter for smoothing predictions and improving accuracy.

Benefits of technology

Enables continuous and accurate real-time RSD prediction, enhancing OR management by reducing resource inefficiencies and minimizing patient waiting times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708354000001
    Figure 0007708354000001
  • Figure 0007708354000002
    Figure 0007708354000002
  • Figure 0007708354000003
    Figure 0007708354000003
Patent Text Reader

Abstract

The embodiments described herein provide a surgical time estimation system for continuously predicting a real-time remaining surgical time (RSD) of a live surgical session of a given surgical procedure based on a real-time endoscopic video of the live surgical session. In one aspect, the process receives a current frame of the endoscopic video at a current time of the live surgical session, the current time being in a sequence of prediction time points for making a continuous RSD prediction during the live surgical session. The process then randomly samples N-1 additional frames of the endoscopic video corresponding to a progress portion of the live surgical session between a start of the endoscopic video corresponding to a start of the live surgical session and a current frame corresponding to a current time. The process then combines the N-1 randomly sampled frames and the current frame in a time order to obtain a set of N frames. The process then feeds the set of N frames to a trained model for the given surgical procedure. The process then outputs a current RSD prediction based on the set of N frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to constructing machine learning-based surgical procedure analysis tools, and more specifically, to systems, devices, and techniques for performing deep learning-based real-time remaining surgical time (RSD) estimation during live surgical sessions of surgical procedures based on endoscopic video feeds.

Background Art

[0002] Operating room (OR) costs are one of the highest healthcare and healthcare-related costs. As healthcare expenditures soar, OR cost management aimed at reducing OR costs and improving OR efficiency has become an increasingly important research topic. OR costs are often measured based on the cost structure per minute. For example, a 2005 study showed that OR costs ranged from $22 to $133 per minute, with an average cost of $62 per minute. In this per-minute cost structure, the OR cost of a given surgical procedure is directly proportional to the time / length of the surgical procedure. Therefore, accurate surgical time estimation plays an important role in constructing an efficient OR management system. Note that if the OR team overestimates the surgical time, the utilization of expensive OR resources will be insufficient. On the other hand, if the surgical time is underestimated, the waiting time of other OR teams and patients will increase. However, accurately predicting the surgical time is very difficult due to patient diversity, surgeon skills, and other unpredictable factors.

[0003] One solution to the above problem is to use machine learning to automatically estimate the remaining surgery duration (RSD) from laparoscopic video feeds. For example, existing RSD estimation techniques manually label each frame of a training dataset with a predefined surgical phase. A supervised machine learning model is then trained based on that training dataset to estimate the surgical phase at each timestamp within the training dataset. Next, by utilizing the statistics of each surgical phase across the training dataset, the remaining time to complete the current surgical phase can be estimated. The RSD is estimated using an estimation combined with an estimation of which phase has been completed at the current timestamp. Unfortunately, this technique requires manual labeling of each of the frames within the training set, which is labor-intensive and expensive.

[0004] Another existing RSD estimation technique does not rely on annotation of surgical phases. In this approach, the input to the machine learning model is a single frame. However, it is extremely difficult for a machine learning model to predict what happened before a single frame from that single frame alone. To address this problem, different variants of recurrent neural networks are utilized in an unsupervised approach to implicitly encapsulate previous frames into hidden states. Unfortunately, surgical videos are usually very long, and it is not easy to teach a machine learning model to represent thousands of frames as multiple hidden states, so the RSD prediction accuracy of this approach remains insufficient. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0005] Some embodiments described herein provide various examples of a surgical time estimation system for continuously and in real-time predicting the remaining surgical time (RSD) of a live surgical session of a given surgical procedure, based on a real-time endoscopic video of the live surgical session. In certain embodiments, the disclosed RSD prediction system receives the current frame of the endoscopic video at the current time of the live surgical session, where the current time is within a sequence of prediction time points for performing continuous RSD prediction during the live surgical session. The RSD prediction system then randomly samples N-1 additional frames of the endoscopic video corresponding to the elapsed portion of the live surgical session between the start of the endoscopic video corresponding to the start of the live surgical session and the current frame corresponding to the current time. The RSD prediction system then combines the N-1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames. Next, the system supplies the set of N frames to a trained machine learning model for a given surgical procedure. Thereafter, the RSD prediction system outputs a current RSD prediction based on the set of N frames.

[0006] In some embodiments, N is selected to be large enough such that the N-1 randomly sampled frames provide a sufficiently accurate snapshot of the various events that occurred during the elapsed portion of the live surgical session.

[0007] In some embodiments, randomly sampling the elapsed portion of the live surgical session enables a given frame within the endoscopic video to be sampled more than once at different prediction time points while continuous RSD prediction is being performed.

[0008] In some embodiments, the RSD prediction system also generates a prediction of the completion rate of the live surgical session using the trained RSD ML model, based on the set of N frames.

[0009] In some embodiments, the RSD prediction system improves the current RSD prediction at the current time by repeatedly performing the following steps multiple times to generate a current set of RSD predictions: (1) randomly sampling N-1 additional frames of the endoscopic video corresponding to the elapsed portion of the live surgery session between the start of the endoscopic video corresponding to the start of the live surgery session and the current frame corresponding to the current time; combining the N-1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; supplying the set of N frames to a trained RSD ML model; and generating a current RSD prediction from the trained RSD ML model based on the set of N frames. Next, the RSD prediction system calculates the mean and variance values of the current set of RSD predictions. Thereafter, the RSD prediction system improves the current RSD prediction by using the calculated mean and variance values as the current RSD prediction.

[0010] In some embodiments, the RSD prediction system further improves the RSD prediction by generating a continuous sequence of real-time RSD predictions corresponding to a sequence of prediction time points within the endoscopic video and applying a low-pass filter to the sequence of real-time RSD predictions to smooth the RSD prediction by removing high-frequency jitter in the sequence of real-time RSD predictions.

[0011] In some embodiments, the RSD prediction system first generates a training dataset by receiving a set of surgical procedure training videos, where each video in the set of training videos corresponds to the execution of a surgical procedure performed by a surgeon skilled in the surgical procedure. Next, for each training video in the set of training videos, the RSD prediction system constructs a set of labeled training data by executing a sequence of training data generation steps at equally spaced time points throughout the training video according to a predetermined time interval. More specifically, each training data generation step in the sequence of training data generation steps at the corresponding time point in the sequence of time points includes: (1) receiving the current frame of the training video at the corresponding time point; (2) randomly sampling N - 1 additional frames of the training video corresponding to the elapsed portion of the surgical session between the start of the training video and the current frame; (3) combining the N - 1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; and (4) labeling the set of N frames with the label associated with the current frame. Finally, the RSD prediction system outputs a plurality of sets of labeled training data associated with the set of training videos.

[0012] In some embodiments, the RSD prediction system establishes a trained RSD ML model by: (1) receiving a convolutional neural network (CNN) model; (2) training the CNN model using a training dataset that includes a plurality of sets of labeled training data; and (3) obtaining a trained RSD ML model based on the trained CNN model.

[0013] In some embodiments, the RSD prediction system is further configured to label each training video in a set of training videos by automatically determining, for each video frame in a training video, the remaining surgical time from the video frame to the end of the training video before generating a training data set, and automatically annotating the video frame with the determined remaining surgical time as the label of the video frame.

[0014] In some embodiments, the label associated with the current frame includes the associated remaining surgical time in minutes.

[0015] In some embodiments, the CNN model includes an action recognition network architecture (I3d) configured to receive a sequence of video frames as a single input.

[0016] In some embodiments, training the CNN model using a training data set includes evaluating the CNN model against a validation data set.

[0017] Some embodiments described herein also provide various examples of an RSD prediction model training system for constructing a trained RSD prediction model for performing real-time RSD prediction. In certain embodiments, the disclosed RSD prediction model training system receives a set of training videos of a target surgical procedure performed by a number of surgeons who can similarly perform the target surgical procedure. Next, the disclosed model training system randomly selects a subset of the training videos from the received set of training videos based on computational resource limitations. The disclosed model training system then initiates an iterative model adjustment procedure based on the subset of training videos. Specifically, in each given iteration, the disclosed model training system selects a timestamp between the start and end of a given training video from each of the subset of training videos. The disclosed model training system then extracts video frames in each of the subset of training videos based on the randomly selected set of timestamps.

[0018] Next, for each video frame randomly selected in each subset of the training videos, the disclosed model training system randomly samples N-1 additional frames of the training video between the start of the training video and the randomly selected video frame, and by combining the N-1 randomly sampled frames and the randomly selected video frame in chronological order, constructs a set of N frames for a given video frame within the corresponding training video. In this way, the disclosed model training system generates a batch of training data that includes a set of N frames extracted from a subset of the training videos. Then, the disclosed model training system updates the model parameters of the RSD prediction model using the batch of training data. Next, the disclosed model training system evaluates the updated RSD prediction model against a validation dataset to determine whether another iteration of the model training process is needed. Next, different embodiments of the real-time RSD prediction system and the RSD prediction model training system are described in more detail below.

[0019] In some embodiments, the RSD prediction model training system is configured to train an RSD prediction model using a plurality of sets of labeled training data corresponding to a set of training videos. More specifically, the RSD prediction model is configured to train the RSD prediction model by randomly selecting one labeled training data from each set of labeled training data within the plurality of sets of labeled training data, combining the randomly selected sets of labeled training data to form a batch of training data, and training the RSD prediction model using the batch of training data to update the RSD prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The structure and operation of the present disclosure will be understood by considering the following detailed description and the accompanying drawings in which like reference numerals refer to like parts.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

[0021] The detailed description provided below is intended as an explanation of various configurations of the subject technology and is not intended to represent the only configuration in which the subject technology may be implemented. The accompanying drawings are incorporated herein and constitute a part of the detailed description. The detailed description includes specific details aimed at providing a thorough understanding of the subject technology. However, the subject technology is not limited to the specific details described herein and may be implemented without these specific details. In some cases, structures and components are shown in block diagram form to avoid obscuring the concepts of the subject technology.

[0022] Some embodiments described herein provide various examples of a surgical time estimation system for continuously and in real-time predicting the remaining surgical time (RSD) of a live surgery session for a given surgical procedure based on a real-time endoscopic video of the live surgery session. In certain embodiments, the disclosed RSD prediction system receives the current frame of the endoscopic video at the current time of the live surgery session, where the current time is within a sequence of prediction time points for performing continuous RSD prediction during the live surgery session. The RSD prediction system then randomly samples N - 1 additional frames of the endoscopic video corresponding to the elapsed portion of the live surgery session between the start of the endoscopic video corresponding to the start of the live surgery session and the current frame corresponding to the current time. The RSD prediction system then combines the N - 1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames. Next, the system supplies the set of N frames to a trained machine learning model for a given surgical procedure. Thereafter, the RSD prediction system outputs a current RSD prediction based on the set of N frames.

[0023] Some embodiments described herein also provide various examples of an RSD prediction model training system for constructing a trained RSD prediction model for performing real-time RSD prediction. In certain embodiments, the disclosed RSD prediction model training system receives a set of training videos of a target surgical procedure performed by a number of surgeons who can similarly perform the target surgical procedure. Next, the disclosed model training system randomly selects a subset of training videos from the received set of training videos based on computational resource limitations. The disclosed model training system then initiates an iterative model adjustment procedure based on the subset of training videos. Specifically, in each given iteration, the disclosed model training system selects a time stamp between the start and end of a given training video from each of the subset of training videos. Thereafter, the disclosed model training system extracts video frames in each of the subset of training videos based on the randomly selected set of time stamps.

[0024] Next, for each video frame randomly selected in each subset of the training videos, the disclosed model training system randomly samples N-1 additional frames of the training video between the start of the training video and the randomly selected video frame, and N-1 randomly sampled frames and the randomly selected video frame are combined in chronological order to construct a set of N frames for a given video frame within the corresponding training video. In this way, the disclosed model training system generates a batch of training data including a set of N frames extracted from a subset of the training videos. Thereafter, the disclosed model training system updates the model parameters of the RSD prediction model using the batch of training data. Next, the disclosed model training system evaluates the updated RSD prediction model against the validation dataset to determine whether another iteration of the model training process is required. Next, different embodiments of the real-time RSD prediction system and the RSD prediction model training system are described in more detail below.

[0025] FIG. 1 shows a diagram showing an exemplary operating room (OR) environment 110 equipped with a robotic surgical system 100 according to some embodiments described herein. As shown in FIG. 1, the robotic surgical system 100 includes a surgeon console 120, a control tower 130, and one or more surgical robotic arms 112 located on a robotic surgical platform 116 (e.g., a table or bed, etc.), and a surgical tool having an end effector is attached to the distal end of the robotic arm 112 for performing a surgical procedure. The robotic arm 112 is shown as a table-mounted system, but in other configurations, the robotic arm may be mounted on a cart, ceiling or sidewall, or another suitable support surface. The robotic surgical system 100 can include any currently existing or future developed robotic-assisted surgical system for performing robotic-assisted surgery.

[0026] Generally, a user / operator 140, such as a surgeon or other operator, can use the user console 120 to remotely operate (e.g., teleoperate) the robotic arm 112 and / or the surgical instrument. The user console 120 can be located in the same operating room as the robotic surgery system 100, as shown in FIG. 1. In other environments, the user console 120 can be located in an adjacent or nearby room, or can be remotely operated from a remote location in another building, city, or country. The user console 120 can include a seat 132, a foot-operated controller 134, one or more handheld user interface devices (UIDs) 136, and at least one user display 138 configured to display, for example, a view of the surgical site within a patient. As shown in the exemplary user console 120, a surgeon located on the seat 132 and looking at the user display 138 can operate the foot-operated controller 134 and / or the UID 136 to remotely control the robotic arm 122 and / or the surgical instrument attached to the distal end of the arm.

[0027] In some variations, the user may also operate the robotic surgery system 100 in an "over the bed" (OTB) mode, in which the user is at the patient's side and simultaneously operates the robotic drive tool / end effector attached thereto and a manual laparoscopic tool (e.g., using a handheld user interface device (UID) 136 held in one hand). For example, the user's left hand may operate the handheld UID 136 to control the robotic surgery components, while the user's right hand may operate the manual laparoscopic tool. Thus, in these variations, the user may perform both robotic-assisted (minimally invasive surgery) MIS and manual laparoscopic surgery on the patient.

[0028] During an exemplary procedure or surgery, the patient is prepared and draped in a sterile manner to receive anesthesia. The initial access to the surgical site can be performed manually with the robotic surgical system 100 in a stowed or retracted configuration to facilitate access to the surgical site. Once access is complete, initial positioning or setup of the robotic system can be performed. During the procedure, the surgeon at the user console 120 can operate various surgical tools / end effectors and / or the imaging system using the foot pedal control device 134 and / or the UID 136 to perform the surgery. Also, manual assistance may be provided at the treatment table by a person wearing a sterile gown, who may perform tasks including, but not limited to, retracting tissue or performing a manual repositioning or tool exchange involving one or more robotic arms 112. Also, there may be a non-sterile person present to assist the surgeon at the user console 120. Once the procedure or surgery is complete, the robotic surgical system 100 and / or the user console 120 can be configured or set up to facilitate one or more post-operative procedures including, but not limited to, cleaning and / or sterilization of the robotic surgical system 100 and / or input or printing of medical records, whether electronic or hard copy, via, for example, the user console 120.

[0029] In some embodiments, communication between the robotic surgery platform 116 and the user console 120 may be via a control tower 130 that can convert user commands from the user console 120 into robotic control commands and transmit the robotic control commands to the robotic surgery platform 116. The control tower 130 may also transmit status and feedback from the robotic surgery platform 116 to the user console 120. The connections between the robotic surgery platform 116, the user console 120, and the control tower 130 may be via wired and / or wireless connections, may be proprietary, and / or may be implemented using any of a variety of data communication protocols. Any wired connection may optionally be built into the floor and / or walls or ceiling of the operating room. The robotic surgery system 100 may provide video output to one or more displays, including a display within the operating room and a remote display accessible via the Internet or other network. The video output or feed may also be encrypted to ensure privacy, and all or part of the video output may be stored on a server or electronic medical record system.

[0030] FIG. 2 shows a block diagram of a remaining surgical duration (RSD) prediction system 200 for performing real-time RSD prediction during a surgical procedure based on a treatment video feed, according to some embodiments described herein. As shown in FIG. 2, the RSD prediction system 200 can include an endoscopic video reception module 202, an N-frame generation module 204, and a trained RSD prediction machine-learning (ML) model 206, which are coupled as shown. However, other embodiments of the disclosed RSD prediction system may include additional processing modules between the endoscopic video reception module 202 and the trained RSD prediction ML model 206 (or hereinafter "RSD prediction model 206") not shown in FIG. 2.

[0031] In some embodiments, the RSD prediction model 206 is specifically constructed and trained for a particular surgical procedure, such as a Roux-en-Y gastric bypass procedure or a sleeve gastrectomy procedure. Note that this particular surgical procedure typically includes a set of predetermined phases / steps that are characteristic of the particular surgical procedure (each phase / step may further include sub-phases / sub-steps). In some embodiments, the endoscopic video receiving module 202 within the RSD prediction system 200 receives a real-time / live endoscopic video feed 208 of a surgical procedure, such as a gastric bypass procedure being performed by a surgeon. In some embodiments, to provide a complete and continuous RSD prediction during a surgical procedure, the endoscopic video receiving module 202 can receive the live endoscopic video feed 208 of the surgical procedure (or "video feed 208" hereinafter) in its entirety, i.e., from the moment the surgical procedure starts to the end of the surgical procedure. Note that the video feed 208 typically consists mostly of raw video images captured from inside the patient's body without processing.

[0032] One of ordinary skill in the art will understand that multiple trained RSD prediction models can be constructed / build before performing real-time RSD prediction, such that each RSD prediction model among the multiple trained RSD prediction models is constructed for a particular surgical procedure among multiple different surgical procedures. Thus, based on the particular surgical procedure captured by the video feed 208, the RSD prediction system 200 can select the corresponding RSD prediction model 206 from among the multiple constructed RSD prediction models used by the RSD prediction system 200.

[0033] In some embodiments, the endoscopic video receiving module 202 further includes a frame buffer 210 of a predetermined buffer size. At the start of a surgical procedure when the video feed 208 has just begun to be received, it should be noted that the frame buffer 210 is essentially empty so that frames within the received video feed 208 can be buffered into the frame buffer 210 either for each frame or every other frame. However, the predetermined buffer size of the frame buffer 210 is often smaller than the space required to store each frame of the entire recorded video feed or even every other frame. Thus, in some embodiments, the frame buffer 210 can be configured as a rolling buffer. In these embodiments, when the frame buffer 210 becomes full at a particular point during a live surgical procedure, the endoscopic video receiving module 202 can receive a new video frame from the endoscopic feed 208 and store it in the frame buffer 210, while at the same time, removing an older video frame, for example, the oldest frame, from the frame buffer 210. The management of the frame buffer 210 will be described in more detail below.

[0034] The RSD prediction system 200 further includes an N-frame generation module 204 coupled to the endoscopic video reception module 202. In some embodiments, the N-frame generation module 204 is configured to operate at each time point corresponding to each newly received video frame. This means that if the video feed 208 is captured at 60 fps, the N-frame generation module 204 is activated 60 times per second for each received frame. However, this approach can be very computationally intensive and impractical. In some other embodiments, the N-frame generation module 204 can be synchronized with a timer such that the N-frame generation module 204 is triggered / activated only at a series of predetermined time points according to a predetermined time interval, for example, every 1 second, every 2 seconds, or every 5 seconds, etc. This means that if the video feed 208 is captured at 60 fps, the N-frame generation module 204 is activated only once every 60 frames, once every 120 frames, or once every 300 frames, etc., rather than for each individual frame.

[0035] Since the RSD prediction system 200 is configured to generate continuous and real-time RSD prediction / updating during a live surgical procedure, regardless of whether the RSD prediction is performed on a frame-by-frame basis or based on a predetermined time interval, the time corresponding to each RSD prediction point is referred to as the "current time", "current time T", or "current type point" of the live surgical procedure. However, if the RSD prediction is performed based on a predetermined time interval, for example, every 1 second or every 2 seconds, the current type point for performing the current RSD prediction is within a sequence of predetermined time points for performing continuous RSD prediction / updating throughout the surgical procedure. Further, the video frame of the video feed 208 corresponding to the current type point is referred to as the "current frame" of the live surgical procedure, which is also the most recent video frame within the video feed 208.

[0036] In some embodiments, in preparation for generating real-time RSD predictions, the N-frame generation module 204 obtains the current frame of the endoscopic feed 208 at the current time of the surgical procedure. Note that the N-frame generation module 204 can obtain the current frame from the frame buffer 210 within the endoscopic video reception module 202 or directly from the video feed 208. Further, the N-frame generation module 204 is configured to randomly sample N-1 additional frames from the previously received and stored video frames of the video feed 208 corresponding to the past portion of the surgical procedure, i.e., from the start of the endoscopic video corresponding to the start of the surgical procedure to the current frame of the video feed 208 corresponding to the current type point. Note that N here is a predetermined integer, which will be described in more detail below.

[0037] In some embodiments, the integer N is selected as a trade-off between the computational constraints for the RSD prediction system 200 to process a set of N frames (defining the upper limit of N) and a set of N-1 frames randomly sampled from the past portion of the surgical procedure of sufficient size to provide a representative or sufficiently accurate snapshot of the set of events that occurred during the past portion of the surgical procedure (defining the lower limit of N). In this way, the downstream RSD prediction model 206 can predict at which time point and phase / step within the entire surgical procedure the current frame is located based on analyzing the set of images captured by the N-1 randomly sampled frames and the current frame (i.e., a total of N frames). On the other hand, the upper limit of the number N indicates that a set of N frames can be processed in real time within a predetermined time interval to perform real-time RSD predictions by the downstream RSD-ML model 206.

[0038] Note that the N-frame generation module 204 can randomly sample N-1 frames from the video frames buffered in the frame buffer 210. For example, at the current time T, if a total of K frames (including the current frame) are buffered / stored in the frame buffer 210, N-1 frames can be randomly sampled from among the K-1 frames (excluding the current frame). For practical reasons, it is desirable that (N-1)≦(K-1). In some embodiments, N is a number between [4,20]. Thus, after the first few seconds of receiving and buffering video frames from the endoscope feed 208, the relationship N-1<<(K-1) can be easily satisfied and maintained throughout the surgical procedure.

[0039] In some embodiments, each video frame in the frame buffer 210 is associated with a corresponding sequence number s that represents the order in which the video frame is received by the reception module 202. For example, if a total of K frames have been received at the current time, each frame stored in the frame buffer 210 will have a corresponding sequence number s from 1 to K (where K is the current frame). Thus, in order to randomly sample N - 1 frames from the K - 1 buffered frames, the N - frame generation module 204 can generate a first random number R1 from 1 to K, and then select the buffered frame with sequence number s = R1 as the first frame among the N - 1 frames. The N - frame generation module 204 can then repeat this procedure by generating a second random number R2. If R2 ≠ R1, the N - frame generation module 204 can select the buffered frame with sequence number s = R2 as the second sampled frame among the N - 1 frames. However, if R2 = R1, the N - frame generation module 204 is configured to generate a new random number R2 to replace the previous random number R2. These random number generation and comparison steps are repeated until a random number R2 that is not equal to R1 is obtained. At this point, the N - frame generation module 204 is configured to select the buffered frame with sequence number s = R2 ≠ R1 as the second sampled frame among the N - 1 frames. Further, the N - frame generation module 204 is configured to generate a unique random number and repeat the above procedure of selecting the buffered frame with a sequence number s equal to the unique random number if less than N - 1 randomly sampled frames are obtained. However, this procedure can end when N - 1 randomly sampled frames generated based on N - 1 unique random numbers are selected from among the K - 1 buffered frames.

[0040] After N - 1 randomly sampled frames of the endoscopic video 208 are acquired, the N - frame generation module 204 is configured to combine the N - 1 randomly sampled frames with the current frame K to obtain a set of N frames, and the set of N frames is arranged in chronological order that matches the corresponding sequence number s of the set of N frames. In other words, the set of N frames is ordered in ascending order of the corresponding set of sequence numbers s such that the original chronological order of the set of N frames within the endoscopic feed 208 is maintained.

[0041] At the start of a surgical procedure, it should be noted that the set of N frames generated by the N - frame generation module 204 typically contains similar images since the N frames are in a dense space. As the surgical procedure progresses, especially as a real - time procedure moves towards the end of the surgical procedure, the individual frames within the set of N frames become increasingly spread out and the set of N frames becomes increasingly different from each other. As a result, the set of N frames continues to function as a proxy for the events that occurred during the course of the surgical procedure.

[0042] Also, note that the set of N frames generated by the above N-frame generation procedure is used to generate a single RSD prediction at the current time T, which is itself a single decision time point among a sequence of predetermined prediction time points throughout the entire surgical procedure. Therefore, the above N-frame generation procedure is continuously executed by the N-frame generation module 204 in the sequence of RSD prediction time points over the live surgical procedure to generate a sequence 214 of N randomly sampled frames (note that the "sequence 214 of N randomly sampled frames" herein means a plurality of sets of N frames generated in the sequence of RSD prediction time points and arranged in chronological order). One skilled in the art will understand that as time progresses over the live surgical procedure, the set of buffered video frames K in the frame buffer 210 will continue to grow (up to the maximum allowable value if it exists). As a result, further, the set of buffered video frames is designated as K(T), which is a function of the current time T. Therefore, at each new RSD prediction time point T, a new set of N frames is generated based on the new set of buffered video frames K(T) in the frame buffer 210.

[0043] When comparing two consecutive sets {F1} and {F2} of N frames generated at two consecutive prediction time points T1 and T2 > T1, note that two observations can be made. First, each randomly sampled frame in the set {F2} is selected from K(T2), while each randomly sampled frame in the set {F1} is selected from K(T1), and K(T2) > K(T1) (if a maximum exists, the maximum value K maxBefore reaching (), it represents a slightly longer period of the surgical procedure progress. Second, a given randomly sampled frame within the set {F2} can also be in the set {F1}. In other words, due to the randomness in the random sampling operation, the same buffered frame in the frame buffer 210 can be selected more than once in the sequence of predetermined prediction time points during a live surgical procedure. This characteristic of the disclosed N-frame generation module 204 enables the RSD prediction system 200 to revisit / reprocess some previously processed video frames at a later time, which has the advantage of gradually improving the stability and consistency of RSD prediction.

[0044] Returning to FIG. 2, it should be noted that the RSD prediction system 200 further includes a trained RSD prediction model 206 coupled to the N-frame generation module 204. As will be described in more detail below, the trained RSD prediction model 206 can be constructed / trained using the same N-frame generation procedure described above, based on one or more training data sets generated from one or more training videos of the same surgical procedure. When performing real-time RSD prediction within the RSD prediction system 200, the RSD prediction model 206 receives a single set of N frames at the current prediction time point T from among a sequence 214 of N frames randomly sampled from the N-frame generation module 204, and then is configured to generate a new current RSD prediction based on the processing of the unique set of N frames. In some other embodiments, the RSD prediction model 206 is configured to process a set of N frames based on the corresponding order of the set of frames, but does not need to know the exact time stamps associated with the set of frames.

[0045] In some embodiments, instead of generating a single randomly selected set of N frames at the current type point T and calculating a single RSD prediction, a plurality of randomly selected sets of N frames at the current type point T can be generated, and then the RSD prediction model 206 can be used to generate a plurality of RSD predictions based on the processing of the plurality of randomly selected sets of N frames corresponding to the same current type point T. Next, the variance value and the mean value of the RSD predictions for the current type point T can be calculated based on the plurality of RSD predictions corresponding to the current time point T. Then, the current RSD prediction can be replaced with the calculated mean value and variance value, which represents a more reliable RSD prediction at the current type point T than a single RSD prediction technique.

[0046] Note that, for example, in addition to performing RSD predictions as a per-minute quantity, the RSD prediction model 206 can also be configured to generate a completion rate prediction indicating what percentage of the surgical procedure has been completed at the current type point T. Note that the completion rate prediction provides another useful piece of information to both the live session surgical crew and the surgical crew waiting in line for the next surgical session.

[0047] In some embodiments, the RSD prediction ML model 206 can be implemented using various convolutional neural network (CNN / ConvNet) architectures. For example, a particular embodiment of the RSD prediction model 206 is based on using the Inflated 3D ConvNet (I3d) as the backbone of a model configured to receive a sequence of video frames as a single input. However, other implementations of the RSD prediction model 206 can also use recurrent neural network (RNN) architectures such as the long short-term memory (LSTM) network architecture. During a live surgical procedure, the trained RSD prediction model 206 is configured to generate successive real-time RSD predictions 216 (e.g., as how many minutes remaining) at a sequence of predetermined prediction time points based on a randomly sampled sequence 214 of N frames, whereby surgeons and surgical staff both inside and outside the operating room can always know the progress and remaining surgical time. In some embodiments, when new RSD predictions are generated continuously, a Butterworth filter can be used to "smooth" the RSD output, for example, by removing some high-frequency jitter, whereby the most recent set of predictions can be used as an indicator of the direction of the next RSD prediction (e.g., decreasing or increasing over time).

[0048] The disclosed RSD prediction system 200 samples buffered video frames representing portions of the surgical procedure at each prediction time point in small time steps (e.g., 1 second or 2 seconds) and makes corresponding RSD predictions at each prediction time point, taking note that a population of predictions is generated that is separated by small sampling intervals, e.g., 1 second or 2 second intervals. Using such small sampling time steps, the disclosed RSD prediction system 200 is configured to randomly and substantially sample the same set of buffered video frames multiple times in order to make a series of similar RSD predictions. Thus, consistency in a series of RSD predictions is necessary to show the effectiveness of the RSD prediction model 206 in making such RSD predictions.

[0049] As described above, when the size of the frame buffer 210 has a limit, more video frames are added, so buffer management is required. In the naive approach, when the buffer size limit is reached, the new video frame added to the frame buffer 210 will be accompanied by an older frame, for example, the oldest frame removed from the frame buffer 210. However, this approach is not desirable because it will continue to remove the early part of the surgical procedure. The disclosed random sampling technique is intended to sample the entire elapsed portion of the surgical time for each predicted time point. In some embodiments, instead of dropping the oldest video frame from the frame buffer 210, older frames in the buffer can be removed more strategically throughout the buffer. For example, the frame removal strategy can include removing every other frame in the buffer, thereby enabling the surgical procedure information of the earlier part of the video to always be saved. To keep more video frames of the early part of the video, importance weights can also be assigned to the video frames. Thus, older frames receive higher weights while newer frames receive lower weights. This technique, combined with the technique of removing every other frame described above, can make it possible to keep more video frames from the start part of the surgical video in the frame buffer 210 over the surgical procedure.

[0050] As described above, the predetermined number N should be selected as a trade-off between the computational constraints (defining the upper limit of N) for the RSD prediction model 206 to process a set of N frames and a set of N - 1 frames randomly sampled from the past portion of the surgical procedure, of sufficient size, to enable the RSD prediction model 206 to "observe" and thus estimate the progress of the surgical procedure up to the current time / frame based on the current set of N frames. In one particular embodiment, N = 8, i.e., 7 randomly sampled frames + the current frame for each RSD prediction has been found to provide an optimal balance between computational complexity and RSD prediction accuracy over the entire surgical procedure.

[0051] In some embodiments, instead of keeping the number N constant over a given surgical procedure, it is possible to add more frames by the N-frame generation module 204, such that N becomes a variable so that more buffered frames can be sampled as the live surgical procedure progresses. However, due to the consistency of image processing by the RSD prediction model 206, the disclosed RSD prediction system implementing an increased number of sampled frames needs to start with a set of P frames including the set of N frames, followed by a set of M "black frames", each black frame being used as a dummy frame that does not contribute to the decision-making. However, as time progresses, for example, at a set of predetermined time points in the surgical procedure, one or more frames within the set of M black frames are added onto the set of N frames in order to actually sample real buffered frames. By combining with the original set of N frames, the newly added frames from the set of M black frames enable the RSD prediction model 206 to sample and process more buffered frames for real-time RSD prediction.

[0052] In some embodiments, the RSD prediction system 200 may select the RSD prediction model 206 based on a particular surgical procedure. In other words, a number of RSD prediction models may be constructed for various specific surgical procedures. For example, if the surgical procedure being performed is Roux-en-Y gastric bypass, the RSD prediction system 200 is configured to select an RSD prediction model 206 constructed and trained specifically for Roux-en-Y gastric bypass. However, if the surgical procedure being performed is sleeve gastrectomy, the RSD prediction system 200 may select a different RSD prediction model 206 constructed and trained for sleeve gastrectomy in the RSD prediction system 200. Thus, before using the RSD prediction model 206 in a live surgical procedure for real-time RSD prediction, the RSD prediction model 206 needs to be trained using training videos, for example, recorded videos of the same surgical procedure performed by a surgeon implementing the gold standard or by a number of surgeons performing the same surgical procedure in substantially the same manner.

[0053] FIG. 3 shows a block diagram of an RSD prediction model training system 300 for constructing an RSD prediction model 206 in an RSD prediction system 200 according to some embodiments described herein. As shown in FIG. 3, the RSD prediction model training system 300 can include a training video receiving module 302, a training data generating module 304, and an RSD prediction model adjusting module 306, which are coupled as shown. It should be noted that the disclosed RSD prediction model training system 300 (or hereinafter "model training system 300") is configured to construct / train an RSD prediction model 206 for use in an RSD prediction system 200 for a particular surgical procedure that includes a predetermined set of phases / steps (each phase / step can further include a predetermined sub-phase / sub-step). Then, during a live surgery session for a particular surgical procedure, such as a Roux-en-Y gastric bypass procedure or a sleeve gastrectomy procedure, the trained RSD prediction model 206, which is the output of the model training system 300, can be used for real-time RSD prediction in the RSD prediction system 200.

[0054] In some embodiments, the training video receiving module 302 within the model training system 300 receives a set of recorded training videos 308 of the same surgical procedure as the RSD prediction model 206, such as a gastric bypass procedure. For example, the set of training videos 308 can include a training video A depicting a surgical procedure being performed by a surgeon performing the gold standard or, alternatively, by a surgeon skilled in performing a given surgical procedure in a standard manner. Thus, when training video A is used to train the RSD prediction model 206, training video A can be used to establish a standard for performing a given surgical procedure. A single training video of a surgeon skilled in a surgical procedure enables the RSD prediction model 206 to learn the characteristics of the surgical procedure depicted in that video, but may not be sufficient to teach the model to recognize variations to the standard performance of the surgical procedure, such as changes in the order of execution of a set of steps in the surgical procedure. Further, a single training video may also not be sufficient to teach the model to recognize different types of complications, such as adhesions in a patient, and events that are unusual but can occur during a surgical procedure. However, when multiple training videos covering various scenarios and variations of the same surgical procedure are used to train the RSD prediction model 206, the trained model becomes more robust by its ability to recognize (1) changes in the order of surgical steps, (2) patient complications, (3) unusual events, and (4) other variations.

[0055] In some embodiments, the set of training videos 308 can similarly include recorded surgical videos of the same surgical procedure performed by a number of surgeons who can perform the surgical procedure by, for example, performing the same set of standard surgical steps associated with the surgical procedure. However, the set of training videos 308 can also include variations in performing the surgical procedure. In some embodiments, the set of training videos 308 can include a first subset of recorded videos that capture variations in the order in which the set of surgical steps are performed. The set of training videos 308 can also include a second subset of recorded videos that capture known patient comorbidities and known types of events that can occur during the same surgical procedure that are different from normal. In some embodiments, the set of training videos 308 can also include the same surgical procedure captured by different camera angels. For example, two training videos can be generated by two endoscopic cameras placed at two opposing angels to capture additional timing information of the surgical procedure that cannot be obtained with a single camera angel.

[0056] In some embodiments, the training video receiving module 302 can include a storage device 312, and the training video receiving module 302 is configured to preload a set of training videos 308 into the storage device 312 prior to the actual model training process. Note that unlike the frame buffer 210 described in connection with the endoscopic video receiving module 202, the storage device 312 can store the entire set of training videos 308 without size limitations. In some embodiments, the training video receiving module 302 further includes a labeling sub-module 314 configured to label each frame within each received training video (hereinafter referred to as "a given training video 308") within the set of training videos 308 and associate the frames with their respective timing information. In some implementations, the labeling sub-module 314 is configured to label each frame within a given training video 308 with sequential frame numbers, such as 0, 1, 2, etc., based on the order of the frames within the given training video 308. Note that this frame number label can indicate the relative timing of the associated frames in the associated surgical procedure. Note that based on the frame number label of a given frame, the specific timestamp of the given frame and the RSD value from the given frame to the end of the surgical procedure in the associated surgical procedure can be automatically determined. Conversely, when a timestamp (e.g., 23 minutes and 45 seconds) for a given training video 308 is provided, the frame number label associated with the timestamp can be automatically determined, and then the corresponding video frame can be selected.

[0057] In some embodiments, labeling a given training video 308 within a set of training videos may additionally and optionally include providing labels for surgical phases / steps and / or sub - phases / sub - steps with respect to video frames within the given training video 308. For example, if a given surgical phase / step of a surgical procedure in a given training video 308 starts at 20 minutes 15 seconds and ends at 36 minutes 37 seconds, each frame between these two time stamps must be labeled as the same surgical phase / step. Note that this phase / step label for each frame is added to the above - mentioned frame number label of the given frame. In some implementations, these additional and optional phase / step labels for a given training video 308 can be generated by a human operator / labeler by manually identifying the start time point and end time point of each phase / step and then annotating the frames between the start time point and the end time point with the associated phase / step label. Thus, in subsequent parameter adjustment steps, the frame number label of a given frame can be used to determine the specific time stamp of the given frame, and the phase / step label of the given frame can be used to determine in which phase / step of the associated surgical procedure in a given training video 308 the given frame is associated.

[0058] In some other embodiments, additional and optional phase / step labels for a given training video 308 may be automatically generated by the labeling submodule 314. For example, the labeling submodule 314 may include a separate deep learning neural network trained for a particular surgical procedure associated with the given training video 308 for phase / step recognition. Thus, using the labeling submodule 314 that includes such a trained deep learning neural network, the start and end of each surgical phase / step within a given training video 308 can be automatically identified, and then the frames between the two identified boundaries of a given phase / step can be labeled with the corresponding phase / step label. Note that the training video receiving module 302 outputs a set of labeled training videos 318 corresponding to the set of received raw training videos 308.

[0059] The training video receiving module 302 is coupled to a training data generation module 304 within the model training system 300. In some embodiments, the training data generation module 304 is configured to process a set of labeled training videos 318 to generate a training data set 330. In some embodiments, a single training data point within the training data set 330 can be generated using each video frame between the start and end of a labeled training video (hereinafter referred to as a "given labeled training video 318") within the set of labeled training videos 318. In some embodiments, the training data generation module 304 is configured to randomly select a set of timestamps within a given labeled training video 318 and then construct a set of training data points corresponding to the randomly selected set of timestamps to generate a set of training data points from the given labeled training video 318. In this way, the training data generation module 304 can generate the training data set 330 by combining multiple sets of training data points generated from the set of labeled training videos 318.

[0060] In some embodiments, instead of generating training data points based on a randomly selected set of timestamps, the training data generation module 304 generates a set of training data points from a given labeled training video 318 at a set of time points based on a predetermined time interval and then constructs a set of training data points corresponding to the set of time points. For example, the training data generation module 304 can generate a set of training data points from a given labeled training video 318 at a set of time points at one-minute intervals. In some embodiments, the training data generation module 304 generates a training data set 330 as a training data sequence according to the progress of a surgical procedure in the labeled training video 318. More specifically, the training data generation module 304 can generate the training data set 330 by progressively outputting training data points at each time point within the labeled training video 318 based on a predetermined time interval.

[0061] Similar to the RSD prediction process described above, each selected time point within a given labeled training video 318 for generating a training data point within the training data set 330 corresponds to a labeled video frame (hereinafter referred to as a "selected frame") within the given labeled training video 318. Note that the training data generation module 304 also includes an N-frame generation module 324 configured to operate in the same manner as the N-frame generation module 204 within the RSD prediction system 200. More specifically, to generate the corresponding training data point at the selected time point, the N-frame generation module 324 is configured to obtain the selected frame of the given labeled training video 318 at the selected time point. Next, the N-frame generation module 324 is configured to generate N - 1 additional frames by randomly sampling N - 1 "previous" frames from a portion of the labeled training video 318 during a time period prior to the selected frame and less than the selected time point, where N is a predetermined integer herein. In other words, the N - 1 additional frames can be obtained from the entire labeled training video 318 from the start of the training video to the selected time point. Due to the randomness in sampling the set of N - 1 frames to form a given training data point, even if they are ordered in a temporal sequence, the time intervals between these N frames may be different, for example, 5 minutes between the first two frames and 10 minutes between the last two frames. In some embodiments, these N - 1 frames may not include any video frames associated with any extracorporeal events within the given labeled training video 318.

[0062] As described above, the predetermined number N is selected as a trade-off between computational constraints (defining an upper limit for N) for training the RSD prediction model 206 using the training data set 330 composed of a set of N generated frames, and a sufficiently large number of sampled frames before the selected time point / frame (defining a lower limit for N) so as to sufficiently represent the progress of the surgical procedure up to the selected time point / frame. In some embodiments, the predetermined number N used by the N-frame generation module 324 within the model training system 300 is the same as the predetermined number N used by the N-frame generation module 204 within the RSD prediction system 200. In one particular embodiment, N = 8, i.e., the N-frame generation module 324 is configured to generate an 8-frame training data point by obtaining the frame selected at the selected time point and randomly sampling 7 additional frames across the portion of the labeled training video 318 prior to the selected time point. As described above, the randomness in selecting the N - 1 additional frames allows the same frame within the labeled training video 318 to be selected more than once when generating the training data set 330 during the RSD prediction model training process.

[0063] After processing the set of labeled training videos 318, the training data generation module 304 generates a training data set 330 as an output. Referring again to FIG. 3, the training data generation module 304 is coupled to an RSD prediction model adjustment module 306 (or hereinafter "model adjustment module 306") configured to adjust a set of neural network parameters within the RSD prediction model 206 based on a received training data set 330 that includes a plurality of sets of training data points generated from a plurality of labeled training videos 318, where it should be noted that each training data point is further composed of a set of N frames ordered in a temporal sequence by their respective time / frame number labels within a given labeled training video 318.

[0064] In some implementations, the RSD prediction model training process includes collectively using the training data generation module 304 and the model adjustment module 306 to train a model based on a set of labeled training videos 318 as follows. First, a subset of M training videos is randomly selected from the set of labeled training videos 318. This step can be performed by the training data generation module 304. Note that the set of labeled training videos 318 may contain a large number (e.g., hundreds) of videos that cannot all be used for model training at the same time. Specifically, the number M can be determined based on the computational resource limitations for processing a single training data point at a given time. In some embodiments, the number M is determined based on the memory limitations of one or more processors (e.g., one or more graphic processing units (GPUs)) used by the model adjustment module 306 to process a set of video frames. Note that if the number M is found to be greater than or equal to the number of videos in the set of labeled training videos 318, the entire set of labeled training videos 318 can be selected.

[0065] Next, for each of the M selected training videos, using the training data generation module 304, a timestamp is randomly selected between the start and end of a given training video. Thereafter, using the training data generation module 304, a video frame within the given training video is selected based on the randomly selected timestamp. Next, for the randomly selected video frame, the N-frame generation module 324 is used to randomly select N - 1 additional frames using the N-frame generation technique described above, and these N - 1 additional frames are combined with the randomly selected video frame to form a set of N frames for a given training video. It should be noted that the above steps are repeated for the entire set of M selected training videos corresponding to M randomly selected frames from the M selected training videos. As a result, the training data generation module 304 generates M sets of N frames corresponding to the M selected training videos (e.g., M = 16). These M sets of N frames are referred to as a "batch" of training data, which can be considered as a single training data point within the training data set 330. It should be noted that a modification to the above steps for generating a single batch of training data is that instead of randomly selecting one frame from each of the M selected training videos, it is possible to randomly select each frame within the M frames from the entire set of M selected training videos. In other words, it is possible to select more than two frames from a given video within the set of M selected training videos, but it is also possible that a given video within the set of M selected training videos is not selected at all.

[0066] After generating a batch of training data as described above, the model adjustment module 306 is used to process the batch of training data in a process called "iteration" to optimize the model. Specifically, during the iteration, the batch passes through the neural network of the RSD prediction model, an error is estimated, and is used to update parameters such as the weights and biases of the RSD prediction model using an optimization technique such as gradient descent.

[0067] Note that the above process of generating a single batch of training data and using that batch to update the RSD prediction model represents a single iteration of the model training process. Thus, the RSD prediction model training process includes many such iterations, and the RSD prediction model is gradually optimized through each iteration in many iterations. In some embodiments, during the model training process, the RSD prediction model is evaluated at the end of each given iteration with respect to a validation data set. If the RSD prediction error from the trained model with respect to the validation data set is stable or at a plateau (e.g., within a predetermined error margin), the RSD prediction model training process can be terminated. Otherwise, another new iteration is added to the model training process.

[0068] Note that the above-described RSD prediction model training process is based on randomly selecting training data points within a set of labeled training videos 318. In some other embodiments, the disclosed RSD prediction model training process based on the model training system 300 can be a progressive training process that follows the progression of each labeled training video 318 within the set of labeled training videos 318. In this progressive model training process, the training data set 330 can be gradually generated one data point at a time based on a predetermined time interval (e.g., every 1 minute) from the start of each labeled training video 318 to the end of the labeled training video 318.

[0069] More specifically, the training data generation module 304 continues to generate a training data set 330 as a sequence of training data points based on a predetermined time interval (e.g., at 1-second or 1-minute intervals) from the start to the end of a given labeled training video 318. Similarly, for the RSD prediction process described above, the time associated with the current training data point generated in the progressive model training process can be referred to as the "current time" of the model training process. Thus, at the current time, the current frame within the given labeled training video 318 corresponding to the current time is selected. Next, for the selected current frame, the N-frame generation module 324 is used to randomly select N - 1 additional frames using the N-frame generation technique described above, and these N - 1 additional frames are combined with the current frame to form a set of N frames, i.e., the current training data point within the training data set 330.

[0070] Similar to the above, the model adjustment module 306 continues to adjust / update a set of neural network parameters based on a sequence of RSD values (i.e., the true RSD values within a given labeled training video 318) associated with a sequence of training data points using newly generated training data points within the training data set 330. Specifically, adjusting a set of neural network parameters based on a given generated training data point and the corresponding RSD value includes predicting the RSD value using an RSD prediction model trained to match the corresponding true RSD value. In some embodiments, adjusting a set of neural network parameters using the training data set 330 includes performing stochastic gradient descent to minimize the RSD prediction error. After the entire training data set 330 generated by the training data generation module 304 has been processed by the model adjustment module 306, the model training system 300 finally outputs the trained RSD prediction model 206.

[0071] In some embodiments, the disclosed progressive model training process samples the labeled training video 318 based on small time steps (e.g., a few seconds). Since the sequence of training data points within the training data set 330 is only separated by this small time step, the progressive model training process essentially samples the substantially same set of video frames repeatedly over a relatively short time period (e.g., 1 minute) and performs a series of model adjustment procedures during this short time period based on the substantially same target RSD. Thus, the disclosed model training process also progressively improves the consistency and accuracy of the predictions of the trained RSD prediction model 206.

[0072] In some embodiments, instead of generating a single training data point at each current type point, a plurality of training data points may be generated using the N-frame generation module 324 at the same current type point. Since randomness was involved in generating each set of N frames, each of the plurality of training data points generated at the same current type point is likely to be composed of a different set of N - 1 randomly sampled frames, and thus, most likely, a different set of N frames. However, since these plurality of training data points are also associated with the same target RSD value, as a result of using them to adjust / train a set of neural network parameters at the associated current type point, the convergence time can be faster than using a single training data point at a given time point.

[0073] In some embodiments, the model adjustment module 306 is configured to train the RSD prediction model using a set of labeled training videos 318. For example, the model adjustment module 306 may be configured to sequentially process each training video in the set of labeled training videos 318 based on the above-described progressive training process for a single labeled training video 318 to adjust a set of neural network parameters until the entire set of labeled training videos 318 has been processed.

[0074] In some embodiments, the model adjustment module 306 is also configured to train the RSD prediction model 206 to perform a completion rate prediction indicating what percentage of the surgical procedure has been completed at the current prediction time point using a set of labeled training videos 318. Since the completion rate value at each prediction time point is known, the model adjustment module 306 is configured to adjust / train a set of neural network parameters to cause the completion rate prediction by the RSD prediction model 206 to match the actual completion rate value at a given prediction time point.

[0075] Note that the progressive model training process described above represents one epoch of training, i.e., a training data point selected from the set of labeled training videos 318 is used only once in a single pass. In some embodiments, to ensure convergence of the model training process, the set of labeled training videos 318 is repeatedly used over multiple epochs / passes.

[0076] More specifically, a plurality of sets of training data points can be first generated from a set of labeled training videos 318 based on a predetermined time interval. Next, the RSD prediction model is trained through a plurality of epochs using the same set of training data points. Specifically, in each epoch of training, the same training step is executed by the model adjustment module 306 using the same set of training data points, whereby the RSD prediction model approaches convergence by one step. In practice, 20 to 50 epochs of training can be executed based on the same set of training data points. This means using the training data generation module 304 20 to 50 times to execute the same random sampling procedure 20 to 50 times for each data point within a plurality of sets of training data points for a given data point. Thus, in each epoch of the model training process, a unique training data set 330 is generated based on the same set of training data points. This is because due to the random nature of the N-frame generation module 324, different epochs of the model training process can use different sets of previous frames / images for each data point within the same set of training data points.

[0077] It should be noted that when comparing with the uniform sampling of previous frames at each time point T, performing RSD prediction using either the uniform sampling technique or the disclosed random sampling technique may generate relatively similar real-time RSD prediction results. However, using the disclosed random sampling technique for RSD prediction model training can generally achieve significantly better model training results, such as a faster convergence rate, than using the uniform sampling technique, at least because the uniform sampling technique cannot generate a new training data set when multiple training data points are generated at the same time point T or in different epochs. Further, as described above, the randomness in selecting N - 1 additional frames using the disclosed random sampling technique also has the advantage of allowing the same previous frame within a given training video to be selected more than once within the generated training data set, thereby enabling the prediction stability of the trained RSD prediction model to be gradually but more effectively improved.

[0078] It should be noted that the disclosed random sampling technique also provides a mechanism for testing the model prediction reliability of a trained RSD prediction model during real-time RSD prediction. For example, after performing a first RSD prediction at a given time point T using the trained RSD prediction model 206, the N-frame generation technique can be reapplied to randomly sample earlier frames again, and a second prediction can be made using the trained RSD prediction model 206. Then, the N-frame generation technique can be used again to randomly sample earlier frames once more, and a third prediction can be made using the trained RSD prediction model 206. Next, multiple RSD predictions can be compared for prediction reliability. If all three predictions are substantially the same (e.g., the difference is within a few seconds), the reliability of the RSD prediction result can be high. However, if multiple RSD predictions at time point T produce significantly different results (e.g., additional RSD predictions bounce up or down around the first RSD prediction), it may indicate that the trained model cannot learn what has occurred from the start of the surgical procedure to time point T. In such a scenario, manual intervention and / or postoperative analysis may be required to understand what has occurred during the surgical procedure.

[0079] Figure 4 presents a flowchart showing an exemplary process 400 for performing real-time RSD prediction during a live surgery session based on a treatment video feed, according to some embodiments described herein. In one or more embodiments, one or more of the steps in FIG. 4 may be omitted, repeated, and / or performed in a different order. Accordingly, the specific arrangement of steps shown in FIG. 4 should not be construed as limiting the scope of the technique.

[0080] Process 400 begins by receiving a real-time / live endoscopic video feed of a live surgery session of a particular surgical procedure being performed by a surgeon (step 402). In some embodiments, the surgical procedure is a Roux-en-Y gastric bypass procedure or a sleeve gastrectomy procedure. Next, process 400 obtains the current frame of the endoscopic feed at the current time of the live surgery session (step 404). Process 400 also randomly samples N-1 additional frames from the stored / buffered video frames of the video feed corresponding to the elapsed portion of the surgery session (step 406). Various embodiments of randomly sampling N-1 additional frames were described above in connection with the RSD prediction system 200 and FIG. 2. Note that the N-1 frames randomly sampled from the elapsed portion of the surgery session provide a representative snapshot of the set of events that occurred during the elapsed portion of the surgery session. Thereafter, process 400 combines the N-1 randomly sampled frames with the current frame to generate a set of N frames arranged in their original temporal order (step 408). Next, process 400 processes the set of N frames using a trained RSD prediction model for the surgical procedure to generate a real-time RSD prediction for the live surgery session (step 410). Next, process 400 determines whether the end of the live surgery session has been reached (step 412). If not, process 400 subsequently returns to step 404 and continues to process the live video feed at the next prediction time point to generate the next real-time RSD prediction. When the end of the live surgery session is reached, the real-time RSD prediction process 400 ends.

[0081] Figure 5 presents a flowchart showing an exemplary process 500 for building a trained RSD prediction model in the disclosed RSD prediction system according to some embodiments described herein. In one or more embodiments, one or more of the steps in FIG. 5 may be omitted, repeated, and / or performed in a different order. Accordingly, the specific arrangement of steps shown in FIG. 5 should not be construed as limiting the scope of the technique.

[0082] Process 500 begins by receiving a recorded training video of a target surgical procedure, e.g., a gastric bypass procedure (step 502). In some embodiments, the target surgical procedure is a Roux-en-Y gastric bypass procedure or a sleeve gastrectomy procedure. In some embodiments, the target surgical procedure depicted in the training video is performed by a surgeon performing a gold standard or, alternatively, a surgeon skilled in performing a given surgical procedure in a standard manner. Next, process 500 labels each frame in the received training video and associates the frame with respective timing information (step 504). In some embodiments, the timing information is a sequential frame number based on the order of frames in the training video. In some implementations, process 500 may additionally and optionally label a set of frames in the training video with surgical phase / step labels.

[0083] Next, process 500 proceeds to an incremental semi-supervised learning process according to the progress of the labeled training video. Specifically, process 500 generates a training data point within the training data set at the current time T in the labeled training video (step 506). In some embodiments, process 500 generates a training data point at the current time T by first obtaining the current frame of the labeled training video at the current time T. Next, process 500 generates N-1 additional frames by randomly sampling N-1 "previous" frames from the portion of the labeled training video prior to the current time T. Process 500 then combines the N-1 randomly sampled frames with the current frame arranged in the original temporal order to obtain the corresponding training data point at the current time T.

[0084] Next, process 500 uses the generated training data point and the corresponding target RSD value at the current time to adjust a set of neural network parameters in the RSD prediction model (step 508). In some embodiments, adjusting a set of neural network parameters based on a given generated training data point and the corresponding target RSD includes predicting an RSD value using an RSD prediction model that has been trained to match the corresponding target RSD value. Next, process 500 determines whether the end of the labeled training video has been reached (step 510). If not, process 500 then returns to step 506 and continues the incremental training process by generating the next training data point within the training data set at the next current time T in the labeled training video based on a predetermined time interval. When the end of the labeled training video is reached, the incremental training process 500 ends.

[0085] Although process 500 has been described based on the use of a single training video, it should be noted that process 500 can be easily modified to include multiple training videos according to several embodiments described in conjunction with the model training system 300 of FIG. 3. Further, although the model training process in process 500 is described in a single-epoch manner, process 500 can be easily modified to a multi-epoch model training process according to several embodiments described in conjunction with the model training system 300 of FIG. 3.

[0086] FIG. 6 presents a flowchart showing another exemplary process 600 for training an RSD prediction model in the disclosed RSD prediction system according to several embodiments described herein. In one or more embodiments, one or more of the steps of FIG. 6 may be omitted, repeated, and / or performed in a different order. Accordingly, the specific arrangement of steps shown in FIG. 6 should not be construed as limiting the scope of the present technique.

[0087] Process 600 begins by receiving a labeled training video of a target surgical procedure, such as a gastric bypass procedure (step 602). In some embodiments, the set of labeled training videos can likewise include recorded surgical videos of the target surgical procedure performed by a number of surgeons who were able to perform the target surgical procedure, for example, by performing the same set of standard surgical steps. However, the set of labeled training videos can also include variations in performing the target surgical procedure, such as variations in the order of performing the set of surgical steps. In some embodiments, the set of labeled training videos is obtained from the corresponding set of raw training videos using the training video labeling technique described above. Specifically, each labeled training video includes a time stamp of the associated video frames. In some embodiments, the time stamps of the video frames are represented by a set of frame number labels.

[0088] Next, process 600 randomly selects a subset of M training videos from the set of labeled training videos (step 604). In some embodiments, the number M can be determined based on computational resource limitations for processing a plurality of training data points at a given time. Specifically, the number M can be determined based on the memory limitations of one or more processors (e.g., one or more graphics processing units (GPUs)) used by the system to train the RSD prediction model. Process 600 then initiates an iterative model adjustment procedure based on the M selected training videos.

[0089] Specifically, in a given iteration of the model adjustment procedure, process 600 randomly selects time stamps between the start and end of a given training video for each of the M training videos (step 606). Note that since the frame number labels and corresponding actual time stamps in the labeled training videos are uniquely related to each other, the random time stamps in step 606 can be provided in the form of either frame numbers or actual times. Thereafter, process 600 extracts the video frames of each of the M training videos based on the M randomly selected time stamps associated with the M training videos (step 608). Next, for each of the M randomly selected video frames of each of the M training videos, process 600 constructs a set of N frames for a given video frame in the corresponding training video using the N-frame generation technique described above (step 610). As a result, process 600 generates a batch of training data that includes M sets of N frames extracted from the M training videos.

[0090] Thereafter, process 600 uses the batch of training data to update model parameters such as the weights and biases of the RSD prediction model (step 612). For example, process 600 can use an optimization technique such as gradient descent when training the model parameters based on the batch of training data. Next, process 600 evaluates the updated RSD prediction model against the validation data set (step 614). Process 600 then determines whether the RSD prediction error from the trained model has reached a plateau or has reached within an acceptable error margin (step 616). If not, process 600 returns to step 606 and starts a new iteration of the RSD prediction model training process. If so, the RSD prediction model training process 600 ends.

[0091] Further benefits and applications For the purpose of OR schedule setting, in current surgical procedures, typically, an OR time scheduled based on statistical average time is assigned, and it should be noted that after the current surgical procedure, the next scheduled surgical procedure follows. Since previous RSD prediction techniques are not very different from the scheduled OR time, conventional RSD prediction techniques are typically more accurate at the start of a surgical procedure. However, in RSD prediction, the influence of various complication factors and events different from normal becomes increasingly prominent towards the end of a surgical procedure, so the RSD prediction error typically increases towards the end of a surgical procedure.

[0092] In contrast, the real-time RSD prediction generated by the disclosed RSD prediction model 206 becomes increasingly accurate towards the end of a surgical procedure, and the accurate RSD prediction near the end of a surgical procedure can be used by the next surgical crew for preparation. For example, in the case of a surgical procedure with a length of 2 hours, the RSD prediction accuracy continues to increase in the last 30 minutes of the surgical procedure towards the end. This characteristic of the disclosed RSD prediction system and techniques provides a highly reliable RSD prediction to the surgeon and surgical crew of the next scheduled surgical procedure near the end of the ongoing surgical procedure, for example, when the RSD prediction is less than 30 minutes. Therefore, the surgeon of the next scheduled surgical procedure can accurately know when the current surgical procedure will end, and thereby this surgeon can plan the time buffer for preparation and arrival at the OR accordingly.

[0093] Alternatively, the surgical team can start preparing for the next scheduled surgical procedure when the real-time RSD prediction equals a predetermined time buffer for preoperative preparation. In other words, accurate RSD prediction enables preoperative preparation for the next scheduled surgical procedure to be performed while the current surgical procedure is still in progress. For example, when the RSD prediction reaches the 20-minute mark, the surgical team can start preparing the OR for the next scheduled surgical procedure and does not have to wait for the last few minutes of the current surgical procedure. This allows for a seamless transition from the current surgical procedure to the next scheduled surgical procedure with a very short or no gap in between.

[0094] Note that the ideal RSD prediction curve as a function of time appears to be a line with a negative slope where the y-value decreases linearly with time, i.e., the RSD value decreases. In many situations, the actual RSD prediction often follows the ideal prediction curve. However, in some situations, various abnormal events can cause the actual RSD prediction to deviate significantly from the ideal RSD prediction curve. One type of abnormal event during a surgical procedure is due to the occurrence of complications associated with difficult anatomical structures. Another type of abnormal event is due to the occurrence of events that are different from normal, such as bleeding or occlusion of the camera view (e.g., due to fogging or being covered by blood). The occurrence of abnormal events usually adds extra time / delay to the surgical procedure.

[0095] In some embodiments, the disclosed RSD prediction technique is configured to predict delays caused by anomalous events that are reflected in the RSD prediction output / curve such that the RSD prediction output / curve deviates (e.g., jumps) suddenly from a standard RSD prediction curve (e.g., generated by a surgeon performing a gold standard). The disclosed RSD prediction technique facilitates the automatic and instantaneous identification of such complication events and other abnormal / unusual events early or at the onset of such events based on the real-time RSD prediction output. For example, the real-time RSD prediction output may cause a sudden change in the slope of the real-time RSD prediction curve indicating the likelihood of a complication. As another example, using the RSD prediction curve, a surgeon can be identified as switching the order of a surgical procedure by noticing that the prediction output suddenly rises and then falls back down following the general slope of the RSD prediction curve. Note that the ability to instantaneously identify such anomalous events in real-time enables the identification of potential problems during a surgical procedure that require an external assistant.

[0096] Another advantage is comparing two surgeons performing the same surgical procedure. A model can be trained with a surgeon performing a gold standard who is highly skilled at the surgical procedure. The trained model can then be applied to another surgeon in training to become like the surgeon performing the gold standard. Then, comparing the entire RSD curve of the surgeon performing the gold standard (also referred to as the "gold standard RSD curve") with the entire RSD curve from the surgeon in training, locations within the training RSD curve where the surgeon in training is faster or slower than the gold standard curve can be identified as the training RSD curve accelerates and decelerates throughout the surgical procedure. Such comparison enables the surgeon in training to perform a postoperative review of the RSD prediction output.

[0097] FIG. 7 conceptually shows a computer system that can implement some embodiments of the subject technology. The computer system 700 can be a client, server, computer, smartphone, PDA, laptop, or tablet computer with one or more processors embedded or coupled thereto, or any other type of computing device. Such a computer system includes various types of computer-readable media and interfaces for various other types of computer-readable media. The computer system 700 includes a bus 702, processing unit(s) 712, system memory 704, read-only memory (ROM) 710, persistent storage device 708, input device interface 714, output device interface 706, and network interface 716. In some embodiments, the computer system 700 is part of a robotic surgery system.

[0098] Bus 702 collectively represents all system buses, peripheral buses, and chipset buses that communicatively connect many internal devices of computer system 700. For example, bus 702 communicatively connects processing unit(s) 712 to ROM 710, system memory 704, and persistent storage device 708.

[0099] From these various memory units, the processing unit(s) 712 retrieves the instructions to be executed and the data to be processed in order to execute the various processes described in the disclosure of this patent, including the various real-time RSD prediction procedures and the various RSD prediction model training procedures described in connection with FIGS. 2-6. The processing unit(s) 712 can include any type of processor, including, but not limited to, a microprocessor, a graphics processing unit (GPU), a tensor processing unit (TPU), an intelligent processor unit (IPU), a digital signal processor (DSP), a field programmable gate array (FPGA), and an application specific integrated circuit (ASIC). In different implementations, the processing unit(s) 712 can be a single processor or a multi-core processor.

[0100] The ROM 710 stores the static data and instructions required by the processing unit(s) 712 and other modules of the computer system. On the other hand, the persistent storage device 708 is a read / write memory device. This device is a non-volatile memory unit that stores instructions and data even when the computer system 700 is off. Some implementations of the subject disclosure use a mass storage device (such as a magnetic disk or an optical disk, and its corresponding disk drive, etc.) as the persistent storage device 708.

[0101] Other implementations use a removable storage device (such as a floppy disk, flash drive, and its corresponding disk drive) as the persistent device 708. Similar to the persistent storage device 708, the system memory 704 is a read / write memory device. However, unlike the storage device 708, the system memory 704 is a volatile read / write memory device such as random access memory. The system memory 704 stores some of the instructions and data that the processor needs during execution. In some implementations, various processes described in connection with FIGS. 2-6, including various real-time RSD prediction procedures and various RSD prediction model training procedures, are stored in the system memory 704, the persistent storage device 708, and / or the ROM 710. From these various memory units, the processing unit(s) 712 retrieves the instructions to be executed and the data to be processed in order to execute the processes of some implementations.

[0102] The bus 702 is also connected to an input device interface 714 and an output device interface 706. The input device interface 714 enables a user to communicate information to and select commands for the computer system. Input devices used with the input device interface 714 include, for example, an alphanumeric keyboard and a pointing device (also referred to as a "cursor control device"). The output device interface 706 enables, for example, the display of images generated by the computer system 700. Output devices used with the output device interface 706 include, for example, a printer and a display device such as a cathode ray tube (CRT) or a liquid crystal display (LCD). Some implementations include a device such as a touch screen that functions as both an input device and an output device.

[0103] Finally, as shown in FIG. 7, bus 702 also couples computer system 700 to a network (not shown) via network interface 716. In this way, a computer can be part of a network of computers, such as a local area network (“LAN”), a wide area network (“WAN”), an intranet, or a network of networks such as the Internet. Any or all of the components of computer system 700 can be used in conjunction with the subject disclosure.

[0104] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed in this patent disclosure can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0105] The hardware used to implement the various exemplary logics, logic blocks, modules, and circuits described in connection with the aspects disclosed in this specification can be implemented or executed by a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate logic or discrete transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of receiver devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with DSP cores, or any other such configuration. Alternatively, some steps or methods may be performed by circuitry specific to a given function.

[0106] In one or more exemplary aspects, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The steps of a method or algorithm disclosed herein may be embodied in processor-executable instructions that may reside on a non-transitory computer-readable storage medium or a processor-readable storage medium. A non-transitory computer-readable storage medium or a processor-readable storage medium may be any storage medium that can be accessed by a computer or a processor. By way of example and not limitation, such non-transitory computer-readable or processor-readable storage media may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, the terms “disk” and “disc” include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy discs, and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. The above combinations are also included within the scope of non-transitory computer-readable media and non-transitory processor-readable media. Further, the operations of a method or algorithm may reside as one or any combination, or set of code and / or instructions, on a non-transitory processor-readable storage medium and / or a non-transitory computer-readable storage medium that may be incorporated into a computer program product.

[0107] This patent document contains many details, but these should be construed not as limitations on the scope of any disclosed technology or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. Specific features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable partial combination. Further, although some features are described above as functioning in a particular combination, even if initially claimed as such, one or more features from the claimed combination may optionally be excluded from the combination, and the claimed combination may be for the purpose of a partial or variant of a partial combination.

[0108] Similarly, although operations are shown in the drawings in a particular order, this should not be understood as requiring that such operations be performed in that particular order or sequentially, or that all illustrated operations be performed to achieve the desired result. Further, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0109] Only some implementations and examples are described, and other implementations, extensions, and variations may be made based on what is described and illustrated in this patent document.

[0110] 〔Embodiment〕 (1) A computer-implemented method for continuously and real-time predicting the remaining surgical time (RSD) of a live surgical session of a surgical procedure based on a real-time endoscopic video of the live surgical session, comprising: Receiving the current frame of the endoscopic video at the current time of the live surgery session, wherein the current time is within a sequence of prediction time points for performing continuous RSD prediction during the live surgery session; Randomly sampling N-1 additional frames of the endoscopic video corresponding to the elapsed portion of the live surgery session between the start of the endoscopic video corresponding to the start of the live surgery session and the current frame corresponding to the current time; Combining the N-1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; Feeding the set of N frames to a trained RSD machine learning (ML) model for the surgical procedure; Outputting a current RSD prediction from the trained RSD ML model based on the set of N frames; A computer-implemented method comprising the above steps. (2) The computer-implemented method according to embodiment 1, wherein N is selected to be large enough such that the N-1 randomly sampled frames provide a sufficiently accurate snapshot of various events occurring during the elapsed portion of the live surgery session. (3) The computer-implemented method according to embodiment 1, wherein randomly sampling the elapsed portion of the live surgery session enables sampling a given frame within the endoscopic video more than once at different prediction time points while performing continuous RSD prediction. (4) The computer-implemented method according to embodiment 1, further comprising generating a prediction of the completion rate of the live surgery session using the trained RSD ML model based on the set of N frames. (5) The method is for generating a set of current RSD predictions Randomly sampling N-1 additional frames of the endoscopic video corresponding to the elapsed portion of the live surgery session between the start of the endoscopic video corresponding to the start of the live surgery session and the current frame corresponding to the current time; Combining the N-1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; Supplying the set of N frames to a trained RSD ML model; Generating a current RSD prediction from the trained RSD ML model based on the set of N frames; Repeating the above steps multiple times; Calculating the mean and variance of the set of current RSD predictions; Improving the current RSD prediction by using the calculated mean and variance as the current RSD prediction; Further comprising improving the current RSD prediction at the current time, according to the computer-implemented method described in Embodiment 1.

[0111] (6) The method further comprises: Generating a continuous sequence of real-time RSD predictions corresponding to the sequence of prediction time points in the endoscopic video; Applying a low-pass filter to the sequence of real-time RSD predictions to smooth the RSD prediction by removing high-frequency jitter in the sequence of real-time RSD predictions; Further comprising improving the RSD prediction, according to the computer-implemented method described in Embodiment 1. (7) The method further comprises: Receiving a set of training videos of the surgical procedure, wherein each video in the set of training videos corresponds to the execution of the surgical procedure performed by a surgeon skilled in the surgical procedure; For each training video in the set of training videos, constructing a set of labeled training data by executing a sequence of training data generation steps in a sequence of equidistant time points across the entire training video according to a predetermined time interval, wherein each training data generation step in the sequence of training data generation steps at the corresponding time point in the sequence of time points: receiving the current frame of the training video at the corresponding time point; randomly sampling N-1 additional frames of the training video corresponding to the elapsed portion of the surgical session between the start of the training video and the current frame; combining the N-1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; labeling the set of N frames with the label associated with the current frame; including; outputting a plurality of sets of labeled training data associated with the set of training videos; thereby generating a training data set; The computer-implemented method according to Embodiment 1, further comprising. (8) The method includes receiving a convolutional neural network (CNN) model; training the CNN model using the training data set including the plurality of sets of labeled training data; obtaining the trained RSD ML model based on the trained CNN model; The computer-implemented method according to Embodiment 7, further comprising establishing the trained RSD ML model thereby. (9) The method further includes, before generating the training data set, for each video frame in the training video, automatically determining the remaining surgical time from the video frame to the end of the training video, automatically annotating the video frame with the determined remaining surgical time as the label of the video frame, and thereby labeling each training video in the set of training videos, as further included in Embodiment 7, the computer-implemented method described in Embodiment 7. (10) The label associated with the current frame includes the associated remaining surgical time in minutes, the computer-implemented method described in Embodiment 7.

[0112] (11) The CNN model includes an action recognition network architecture (I3d) configured to receive a sequence of video frames as a single input, the computer-implemented method described in Embodiment 8. (12) Training the CNN model using the training data set includes evaluating the CNN model against a validation data set, the computer-implemented method described in Embodiment 8. (13) A system for continuously and real-time predicting the remaining surgical time (RSD) of a live surgery session of a surgical procedure based on a real-time endoscopic video of the live surgery session, one or more processors, a memory coupled to the one or more processors, the memory, when executed by the one or more processors, causes the system to receive the current frame of the endoscopic video at the current time of the live surgery session, the current time being within a sequence of prediction time points for performing continuous RSD prediction during the live surgery session, Randomly sampling N-1 additional frames of the endoscopic video corresponding to the elapsed portion of the live surgery session between the start of the endoscopic video corresponding to the start of the live surgery session and the current frame corresponding to the current time; Combining the N-1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; Feeding the set of N frames to a trained RSD machine learning (ML) model for the surgical procedure; Outputting a current RSD prediction from the trained RSD ML model based on the set of N frames; A memory storing instructions to cause the above to be performed; A system comprising the above. (14) The system according to embodiment 13, wherein N is selected to be sufficiently large such that the N-1 randomly sampled frames provide a sufficiently accurate snapshot of various events occurring during the elapsed portion of the live surgery session. (15) When executed by the one or more processors, the memory causes the system to To generate a set of current RSD predictions, Randomly sampling N-1 additional frames of the endoscopic video corresponding to the elapsed portion of the live surgery session between the start of the endoscopic video corresponding to the start of the live surgery session and the current frame corresponding to the current time; Combining the N-1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; Feeding the set of N frames to a trained RSD ML model; Generating a current RSD prediction from the trained RSD ML model based on the set of N frames; Repeating the above multiple times; Calculating the mean and variance values of the current set of RSD predictions; Improving the current RSD prediction by using the calculated mean and variance values as the current RSD prediction; The system according to embodiment 13, further storing instructions for improving the current RSD prediction at the current time.

[0113] (16) When the memory is executed by the one or more processors, the system is caused to Generate a continuous sequence of real-time RSD predictions corresponding to the sequence of predicted time points in the endoscopic video; Apply a low-pass filter to the sequence of real-time RSD predictions to smooth the RSD prediction by removing high-frequency jitter in the sequence of real-time RSD predictions; The system according to embodiment 13, further storing instructions for improving the RSD prediction. (17) When the memory is executed by the one or more processors, the system is caused to Receive a set of training videos of the surgical procedure, wherein each video in the set of training videos corresponds to the execution of the surgical procedure performed by a surgeon skilled in the surgical procedure; For each training video in the set of training videos, construct a set of labeled training data by executing a sequence of training data generation steps in a sequence of equally spaced time points over the entire training video according to a predetermined time interval, wherein each training data generation step in the sequence of training data generation steps at the corresponding time point in the sequence of time points Receives the current frame of the training video at the corresponding time point; Randomly sampling N-1 additional frames of the training video corresponding to the elapsed portion of the surgical procedure between the start of the training video and the current frame; Combining the N-1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; Labeling the set of N frames with the label associated with the current frame; Constructing, including; Outputting a plurality of sets of labeled training data associated with the set of training videos; The system according to embodiment 13, further storing instructions for generating a training data set. (18) When the memory is executed by the one or more processors, the system is caused to Receive a convolutional neural network (CNN) model; Training the CNN model using the training data set including the plurality of sets of labeled training data corresponding to the set of training videos; Obtaining the trained RSD ML model based on the trained CNN model; The system according to embodiment 17, further storing instructions for establishing the trained RSD ML model. (19) When the memory is executed by the one or more processors, the system is caused to Randomly selecting one labeled training data from each set of labeled training data within the plurality of sets of labeled training data; Combining the randomly selected sets of labeled training data to form a batch of training data; Training the CNN model using the batch of training data to update the CNN model; The system according to embodiment 18, further storing instructions for training the CNN model using the plurality of sets of labeled training data. (20) The system according to embodiment 18, wherein the CNN model includes an action recognition network architecture (I3d) configured to receive a sequence of video frames as a single input.

Claims

1. A computer-implemented method for continuously and real-time predicting the remaining surgical duration (RSD) of a live surgical session of a surgical procedure based on a real-time endoscopic video of the live surgical session, comprising: Receiving a current frame of the endoscopic video at the current time of the live surgical session, wherein the current time is within a sequence of prediction time points for performing continuous RSD prediction during the live surgical session; Randomly sampling N-1 additional frames of the endoscopic video corresponding to the elapsed portion of the live surgical session between the start of the endoscopic video corresponding to the start of the live surgical session and the current frame corresponding to the current time; Combining the N-1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; Feeding the set of N frames to a trained RSD machine learning (ML) model for the surgical procedure; Outputting a current RSD prediction from the trained RSD ML model based on the set of N frames. A computer-implemented method as described above.

2. The computer-implemented method according to claim 1, wherein N is selected to be large enough such that the N-1 randomly sampled frames provide sufficiently accurate snapshots of various events that occurred during the elapsed portion of the live surgical session.

3. The computer-implemented method according to claim 1, wherein randomly sampling the elapsed portion of the live surgical session enables a given frame within the endoscopic video to be sampled more than once at different prediction time points while continuous RSD prediction is being performed.

4. The computer-implemented method according to claim 1, further comprising generating a prediction of the completion rate of the live surgical session using the trained RSD ML model based on the set of N frames.

5. The method further comprises, For generating a set of current RSD predictions Randomly sampling N-1 additional frames of the endoscopic video corresponding to the elapsed portion of the live surgery session between the start of the endoscopic video corresponding to the start of the live surgery session and the current frame corresponding to the current time; Combining the N-1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; Supplying the set of N frames to a trained RSD ML model; Generating a current RSD prediction from the trained RSD ML model based on the set of N frames; Repeating the above steps multiple times; Calculating the mean and variance of the set of current RSD predictions; Improving the current RSD prediction by using the calculated mean and variance as the current RSD prediction; Further comprising improving the current RSD prediction at the current time, the computer-implemented method according to claim 1.

6. The method further comprises: Generating a continuous sequence of real-time RSD predictions corresponding to the sequence of prediction time points in the endoscopic video; Applying a low-pass filter to the sequence of real-time RSD predictions to smooth the RSD prediction by removing high-frequency jitter in the sequence of real-time RSD predictions; Thereby improving the RSD prediction, the computer-implemented method according to claim 1.

7. The method further comprises: Receiving a set of training videos of the surgical procedure, wherein each video in the set of training videos corresponds to the execution of the surgical procedure performed by a surgeon skilled in the surgical procedure; For each training video in the set of training videos, constructing a set of labeled training data by executing a sequence of training data generation steps in a sequence of equally spaced time points over the entire training video according to a predetermined time interval, wherein each training data generation step in the sequence of training data generation steps at the corresponding time point in the sequence of time points is Receiving the current frame of the training video at the corresponding time point; Randomly sampling N - 1 additional frames of the training video corresponding to the elapsed portion of the surgical session between the start of the training video and the current frame; Combining the N - 1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; Labeling the set of N frames with the label associated with the current frame; Including constructing; Outputting a plurality of sets of labeled training data associated with the set of training videos; Generating a training data set by; Further including the computer - implemented method according to claim 1.

8. The method includes: Receiving a convolutional neural network (CNN) model; Training the CNN model using the training data set including the plurality of sets of labeled training data; Obtaining the trained RSD ML model based on the trained CNN model; Further including establishing the trained RSD ML model by the computer - implemented method according to claim 7.

9. Before generating the training data set, the method includes: For each video frame in the training video, Automatically determining the remaining surgical time from the video frame to the end of the training video; Automatically annotating the video frame with the determined remaining surgical time as the label of the video frame; Further including labeling each training video in the set of training videos by the computer - implemented method according to claim 7.

10. The label associated with the current frame includes the associated remaining surgical time in minutes, according to the computer - implemented method of claim 7.

11. The CNN model includes an action recognition network architecture (I3d) configured to receive a sequence of video frames as a single input, according to the computer - implemented method of claim 8.

12. The computer-implemented method of claim 8, wherein training the CNN model using the training data set includes evaluating the CNN model against a validation data set.

13. A system for continuously and real-time predicting the remaining surgical duration (RSD) of a live surgical session of a surgical procedure based on a real-time endoscopic video of the live surgical session, one or more processors; a memory coupled to the one or more processors, the memory, when executed by the one or more processors, causes the system to receive a current frame of the endoscopic video at a current time of the live surgical session, the current time being within a sequence of prediction time points for performing continuous RSD prediction during the live surgical session; randomly sample N - 1 additional frames of the endoscopic video corresponding to an elapsed portion of the live surgical session between a start of the endoscopic video corresponding to a start of the live surgical session and the current frame corresponding to the current time; combine the N - 1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; supply the set of N frames to a trained RSD machine learning (ML) model for the surgical procedure; output a current RSD prediction from the trained RSD ML model based on the set of N frames; a memory storing instructions that cause the system to comprising a system.

14. The system of claim 13, wherein N is selected to be large enough such that the N - 1 randomly sampled frames provide a sufficiently accurate snapshot of various events that occurred during the elapsed portion of the live surgical session.

15. The memory, when executed by the one or more processors, causes the system to generate a set of current RSD predictions by randomly sampling N - 1 additional frames of the endoscopic video corresponding to an elapsed portion of the live surgical session between a start of the endoscopic video corresponding to a start of the live surgical session and the current frame corresponding to the current time; Combining the N - 1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; Supplying the set of N frames to a trained RSD ML model; Generating a current RSD prediction from the trained RSD ML model based on the set of N frames; Repeating the above steps multiple times; Calculating the mean value and variance value of the set of current RSD predictions; Improving the current RSD prediction by using the calculated mean value and variance value as the current RSD prediction; The system according to claim 13, further storing instructions for improving the current RSD prediction at the current time.

16. When the memory is executed by the one or more processors, the system is caused to Generate a continuous sequence of real - time RSD predictions corresponding to the sequence of prediction time points in the endoscopic video; Apply a low - pass filter to the sequence of real - time RSD predictions to smooth the RSD prediction by removing high - frequency jitter in the sequence of real - time RSD predictions; The system according to claim 13, further storing instructions for improving the RSD prediction.

17. When the memory is executed by the one or more processors, the system is caused to Receive a set of training videos of the surgical procedure, wherein each video in the set of training videos corresponds to the execution of the surgical procedure performed by a surgeon skilled in the surgical procedure; For each training video in the set of training videos, construct a set of labeled training data by executing a sequence of training data generation steps in a sequence of equally - spaced time points over the entire training video according to a predetermined time interval, wherein each training data generation step in the sequence of training data generation steps at the corresponding time point in the sequence of time points Receives the current frame of the training video at the corresponding time point; Randomly sampling N-1 additional frames of the training video corresponding to the elapsed portion of the surgical procedure between the start of the training video and the current frame; Combining the N-1 randomly sampled frames and the current frame in chronological order to obtain a set of N frames; Labeling the set of N frames with the label associated with the current frame; Constructing, including; Outputting a plurality of sets of labeled training data associated with the set of training videos; The system according to claim 13, further storing instructions for generating a training data set by.

18. When the memory is executed by the one or more processors, the system Receiving a convolutional neural network (CNN) model; Training the CNN model using the training data set including the plurality of sets of labeled training data corresponding to the set of training videos; Obtaining the trained RSD ML model based on the trained CNN model; The system according to claim 17, further storing instructions for establishing the trained RSD ML model by.

19. When the memory is executed by the one or more processors, the system Randomly selecting one labeled training data from each set of labeled training data in the plurality of sets of labeled training data; Combining the randomly selected sets of labeled training data to form a batch of training data; Training the CNN model using the batch of training data to update the CNN model; The system according to claim 18, further storing instructions for training the CNN model using the plurality of sets of labeled training data by.

20. The system according to claim 18, wherein the CNN model includes an action recognition network architecture (I3d) configured to receive a sequence of video frames as a single input.

Citation Information

Patent Citations

  • Online, incremental, real-time learning for tagging and labeling data streams for deep neural networks and neural network applications

    JP2020511723A

  • Surgical decision support using decision-theoretic models

    JP2020537205A

  • Step-based system for providing surgical intraoperative cues

    US20190223961A1

  • Post discharge risk prediction

    US20200273581A1