Automated quality assurance of machine-learning model output
A quality assurance decision engine for surgical video annotations addresses the inefficiencies and errors in existing systems by using confidence thresholds to automate the review process, improving the speed and accuracy of surgical video annotation.
Patent Information
- Application Number
- PCT/EP2025/065114
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-03
- Filing Date
- 2025-06-02
- Publication Date
- 2025-12-11
AI Technical Summary
The process of annotating large volumes of surgical video data for computer-assisted surgery systems is highly subjective, time-intensive, and prone to errors due to the numerous factors involved, such as patient condition and annotator training, which can lead to delays and inefficiencies.
Implementing a quality assurance decision engine that assesses machine-learning model outputs for surgical video annotations using a confidence threshold system, triggering secondary reviews for annotations with low confidence and preventing publication until errors are corrected, thereby reducing the need for manual intervention and optimizing processing resources.
This approach enhances the efficiency and accuracy of surgical video annotation by selectively flagging annotations for review, reducing processing time and resource usage while ensuring high-quality output.
Smart Images

Figure EP2025065114_11122025_PF_FP_ABST
Abstract
Description
AUTOMATED QUALITY ASSURANCE OF MACHINE-LEARNING MODELOUTPUTCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 655,246, filed June 3, 2024, the entire content of which is incorporated herein by reference.BACKGROUND
[0002] The present disclosure relates in general to computing technology and relates more particularly to computing technology for automated quality assurance of machinelearning model output.
[0003] Computer-assisted systems, particularly computer-assisted surgery systems (CASs), rely on video data digitally captured during a surgery. Such video data can be stored and / or streamed. In some cases, the video data can be used to augment a person’s physical sensing, perception, and reaction capabilities. For example, such systems can effectively provide the information corresponding to an expanded field of vision, both temporal and spatial, that enables a person to adjust current and future actions based on the part of an environment not included in his or her physical field of view.
[0004] Alternatively, or in addition, the video data can be stored and / or transmitted for several purposes such as archival, training, post-surgery analysis, and / or patient consultation. The process of analyzing a large amount of video data from surgical procedures and generating annotations can be highly subjective, time intensive, and error- prone due, for example, to the volume of data and the numerous factors (e.g., patient condition, annotator training, etc.).SUMMARY
[0005] According to an aspect, a system includes a memory device and one or more processors coupled with the memory device. The one or more processors are configured to apply a machine-learning model to a video of a surgical procedure to produce a plurality of annotations and machine-learning model confidence values associated with the annotations, where the annotations are phases of the surgical procedure. The one or more processors can also be configured to determine a quality assurance confidence value based on the annotations and the machine-learning model confidence values, publish the videowith the annotations based on determining that the quality assurance confidence value is above a first confidence threshold, and trigger a secondary quality assurance review and prevent publishing of the annotations until completion of the secondary quality assurance review based on determining that the quality assurance confidence value is below a second confidence threshold. The one or more processors can also be configured to publish the video with the annotations and trigger the secondary quality assurance review based on determining that the quality assurance confidence value is below the first confidence threshold and above the second confidence threshold.
[0006] According to another aspect, a computer-implemented method for automated quality assurance of machine-learning model output can include applying a machinelearning model to a video of a surgical procedure to produce a plurality of annotations and machine-learning model confidence values associated with the annotations, determining a quality assurance confidence value based on the annotations and the machine-learning model confidence values, publishing the video with the annotations based on determining that the quality assurance confidence value is above a first confidence threshold, triggering a secondary quality assurance review and prevent publishing of the annotations until completion of the secondary quality assurance review based on determining that the quality assurance confidence value is below a second confidence threshold, and publishing the video with the annotations and trigger the secondary quality assurance review based on determining that the quality assurance confidence value is below the first confidence threshold and above the second confidence threshold.
[0007] According to a further aspect, a computer program product includes a memory device with computer readable instructions stored thereon, where executing the computer readable instructions by one or more processing units causes the one or more processing units to perform a plurality of operations. The operations include can include applying a machine-learning model to a video of a surgical procedure to produce a plurality of annotations, determining a quality assurance confidence value based on the annotations, publishing the video with the annotations based on determining that the quality assurance confidence value is above a first confidence threshold, triggering a secondary quality assurance review and prevent publishing of the annotations until completion of the secondary quality assurance review based on determining that the quality assurance confidence value is below a second confidence threshold, and publishing the video withthe annotations and trigger the secondary quality assurance review based on determining that the quality assurance confidence value is below the first confidence threshold and above the second confidence threshold.
[0008] Additional technical features and benefits are realized through the techniques of the present invention. Aspects of the invention are described in detail herein and are considered a part of the claimed subject matter. For a better understanding, refer to the detailed description and to the drawings.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The specifics of the exclusive rights described herein are particularly pointed out and distinctly claimed in the claims at the conclusion of the specification. The foregoing and other features and advantages of the aspects of the invention are apparent from the following detailed description taken in conjunction with the accompanying drawings in which:
[0010] FIG. 1 depicts a computer-assisted surgery (CAS) system according to one or more aspects;
[0011] FIG. 2 depicts a surgical procedure system according to one or more aspects;
[0012] FIG. 3 depicts a system for analyzing video captured by a video recording system according to one or more aspects;
[0013] FIG. 4 depicts a video processing pipeline according to one or more aspects;
[0014] FIG. 5 depicts a video processing pipeline according to one or more aspects;
[0015] FIG. 6 depicts components of a quality assurance decision engine according to one or more aspects;
[0016] FIG. 7 depicts an autoencoder of a workflow outlier detector according to one or more aspects;
[0017] FIG. 8 depicts a local outlier factor determination of a workflow outlier detector according to one or more aspects;
[0018] FIG. 9 depicts a process of automated quality assurance of machine-learning model output according to one or more aspects; and
[0019] FIG. 10 depicts a computer system according to one or more aspects.
[0020] The diagrams depicted herein are illustrative. There can be many variations to the diagrams and / or the operations described herein without departing from the spirit ofthe invention. For instance, the actions can be performed in a differing order, or actions can be added, deleted, or modified. Also, the term “coupled” and variations thereof describe having a communications path between two elements and do not imply a direct connection between the elements with no intervening elements / connections between them. All of these variations are considered a part of the specification.DETAILED DESCRIPTION
[0021] Aspects of the technical solutions described herein relate to automated quality assurance of machine-learning model output. Exemplary aspects of the technical solutions described herein include systems and methods for machine-learning generated annotations of surgical video to be selectively flagged for quality assurance review through an automated process using a quality assurance decision engine.
[0022] In exemplary aspects of the technical solutions described herein, surgical data that is captured by a computer-assisted surgical (CAS) system can be annotated with surgical phases to establish surgical workflows and perform further analysis operations. Phase annotation can be used to segment surgical videos of surgical procedures. Using machine-learning for annotation can speed the process of video review. However, if the annotations include errors, it can result in further processing and time delays to fix the errors. To reduce time delays and apply a uniform approach to surgical video annotation, machine-learning models that provide annotations can be assessed using a quality assurance decision engine to identify which annotations need review and which can be published via an automated pipeline, reducing the need for manual intervention. The quality assurance decision engine can identify annotations with lower confidence that may need further more detailed review. By only selectively using detailed reviews for a subset of annotations, processing resource usage and time can be reduced.
[0023] Turning now to FIG. 1, an example computer-assisted system (CAS) system 100 is generally shown in accordance with one or more aspects. The CAS system 100 includes at least a computing system 102, a video recording system 104, and a surgical instrumentation system 106. As illustrated in FIG. 1, an actor 112 can be medical personnel that uses the CAS system 100 to perform a surgical procedure on a patient 110. Medical personnel can be a surgeon, assistant, nurse, administrator, or any other actor that interacts with the CAS system 100 in a surgical environment. The surgical procedure canbe any type of surgery. In other examples, actor 112 can be a technician, an administrator, an engineer, or any other such personnel that interacts with the CAS system 100. For example, actor 112 can record data from the CAS system 100, configure / update one or more attributes of the CAS system 100, review past performance of the CAS system 100, repair the CAS system 100, and / or the like including combinations and / or multiples thereof.
[0024] A surgical procedure can include multiple phases, and each phase can include one or more surgical actions. A “surgical action” can include an incision, a compression, a stapling, a clipping, a suturing, a cauterization, a sealing, or any other such actions performed to complete a phase in the surgical procedure. A “phase” represents a surgical event that is composed of a series of steps (e.g., closure). A “step” refers to the completion of a named surgical objective (e.g., hemostasis). During each step, certain surgical instruments 108 (e.g., forceps) are used to achieve a specific objective by performing one or more surgical actions. In addition, a particular anatomical structure of the patient may be the target of the surgical action(s).
[0025] The video recording system 104 includes one or more cameras 105, such as operating room cameras, endoscopic cameras, and / or the like including combinations and / or multiples thereof. The cameras 105 capture video data of the surgical procedure being performed. The video recording system 104 includes one or more video capture devices that can include cameras 105 placed in the surgical room to capture events surrounding (i.e., outside) the patient being operated upon. The video recording system 104 further includes cameras 105 that are passed inside (e.g., endoscopic cameras) the patient 110 to capture endoscopic data. The endoscopic data provides video and images of the surgical procedure.
[0026] The computing system 102 includes one or more memory devices, one or more processors, a user interface device, among other components. All or a portion of the computing system 102 shown in FIG. 1 can be implemented for example, by all or a portion of computer system 1200 of FIG. 10. Computing system 102 can execute one or more computer-executable instructions. The execution of the instructions facilitates the computing system 102 to perform one or more methods, including those described herein. The computing system 102 can communicate with other computing systems via a wired and / or a wireless network.
[0027] A data collection system 150 can be employed to store the surgical data, including the video(s) captured during the surgical procedures. The data collection system 150 includes one or more storage devices 152. The data collection system 150 can be a local storage system, a cloud-based storage system, or a combination thereof. Further, the data collection system 150 can use any type of cloud-based storage architecture, for example, public cloud, private cloud, hybrid cloud, and / or the like including combinations and / or multiples thereof. In some examples, the data collection system can use a distributed storage, i.e., the storage devices 152 are located at different geographic locations. The storage devices 152 can include any type of electronic data storage media used for recording machine-readable data, such as semiconductor-based, magnetic-based, optical-based storage media, and / or the like including combinations and / or multiples thereof. For example, the data storage media can include flash-based solid-state drives (SSDs), magnetic-based hard disk drives, magnetic tape, optical discs, and / or the like including combinations and / or multiples thereof.
[0028] In one or more examples, the data collection system 150 can be part of the video recording system 104, or vice-versa. In some examples, the data collection system 150, the video recording system 104, and the computing system 102, can communicate with each other via a communication network, which can be wired, wireless, or a combination thereof. The communication between the systems can include the transfer of data (e.g., video data, instrumentation data, and / or the like including combinations and / or multiples thereof), data manipulation commands (e.g., browse, copy, paste, move, delete, create, compress, and / or the like including combinations and / or multiples thereof), data manipulation results, and / or the like including combinations and / or multiples thereof. In one or more examples, the computing system 102 can manipulate the data already stored / being stored in the data collection system 150. Alternatively, or in addition, the computing system 102 can manipulate the data already stored / being stored in the data collection system 150 based on information from the surgical instrumentation system 106.
[0029] In one or more examples, the video captured by the video recording system 104 is stored on the data collection system 150. In some examples, the computing system 102 curates parts of the video data being stored on the data collection system 150. In some examples, the computing system 102 filters the video captured by the video recording system 104 before it is stored on the data collection system 150. Alternatively, or inaddition, the computing system 102 filters the video captured by the video recording system 104 after it is stored on the data collection system 150. Instrument data (e.g., robotic logs, electrosurgical instrument logs, etc.) can also be stored in the data collection system 150.
[0030] A surgical data management system 160 can provide access to portions of data captured in the data collection system 150, as well as data and records stored in other systems. The surgical data management system 160 can establish user access permissions to patient and surgical data. The surgical data management system 160 can also control access through an interface based on the user access permissions. Access to the surgical data management system 160 can be provided through one or more applications or secure web pages. The surgical data management system 160 can be a stand-alone application, module, and / or an extension of another system. Additional aspects of the surgical data management system 160 can include accessing artificial intelligence (Al)-powered surgical video and analytics. Further aspects of the surgical data management system 160 can include accessing simulation materials that can assist surgeons to prepare, practice, and teach surgical procedures. Further aspects of the surgical data management system 160 can include integrating aspects of equipment in an operating room, surgery planning, rating surgeon performance, and other such features. The surgical data management system 160 can also provide access to technical specifications and information relating to the use of surgical instruments, for example.
[0031] Turning now to FIG. 2, a surgical procedure system 200 is generally shown according to one or more aspects. The example of FIG. 2 depicts a surgical procedure support system 202 that can include or may be coupled to the CAS system 100 of FIG. 1. The surgical procedure support system 202 can acquire image or video data using one or more cameras 204. The surgical procedure support system 202 can also interface with one or more sensors 206 and / or one or more effectors 208. The sensors 206 may be associated with surgical support equipment and / or patient monitoring. The effectors 208 can be robotic components or other equipment controllable through the surgical procedure support system 202. The surgical procedure support system 202 can also interact with one or more user interfaces 210, such as various input and / or output devices. The surgical procedure support system 202 can store, access, and / or update surgical data 214 associated with a training dataset and / or live data as a surgical procedure is being performed onpatient 110 of FIG. 1. The surgical procedure support system 202 can store, access, and / or update surgical objectives 216 to assist in training and guidance for one or more surgical procedures. User configurations 218 can track and store user preferences.
[0032] The surgical procedure support system 202 can also communicate with other systems through a network 230. For example, the surgical procedure support system 202 can communicate with a quality assurance decision engine 240 and a surgical data postprocessing system 250 through the network 230. Other types of devices, such as a computing device 234 (e.g., a mobile phone, laptop, personal computer, or tablet computer), can communicate directly with the surgical procedure support system 202 or through the network 230. As one example, user interfaces 210 may be connected to or integrated with the surgical procedure support system 202 by a wired connection while the computing device 234 connects to the surgical procedure support system 202 via a wireless connection. In some aspects, the computing device 234 can execute or link to another computer system that executes the surgical data management system 160 of FIG. 1 to access various data sources through the network 230.
[0033] The surgical data post-processing system 250 can receive surgical data and associated data generated by the surgical procedure support system 202 and may be separately stored and secured through other data storage. Access to specific data or portions of data through the surgical data post-processing system 250 may be limited by associated permissions. The surgical data post-processing system 250 may include features such as video viewing, video sharing, data analytics, and selective data extraction.
[0034] The quality assurance decision engine 240 can assess the results of machinelearning models that may be executed during a surgical procedure or as part of post- surgical processing by the surgical data post-processing system 250. For example, artificial intelligence generated annotations related to phase, anatomy, surgical instruments, and / or other such aspects can be assessed by the quality assurance decision engine 240 to determine whether confidence was sufficiently low to trigger additional quality assurance analysis before publishing the annotations for use. The surgical data post-processing system 250 may generate surgical performance metrics and comparison data across data sets collected at multiple locations, making data available to the quality assurance decision engine 240. The quality assurance decision engine 240 can perform analysis when a video of a surgical procedure is initially annotated. In some aspects, thequality assurance decision engine 240 can determine whether annotation issues are likely occurring such that a video with incorrect annotations can be adjusted prior to publishing. Further, in tracking performance characteristics of a machine-learning model that generates annotations, the maturity of the machine-learning model can be tracked and adjusted through retraining or redesign as needed to reduce annotation errors. In some aspects, the quality assurance decision engine 240 and / or surgical data post-processing system 250 can be components of the surgical data management system 160 of FIG. 1.
[0035] One or more computing device 264 (e.g., a mobile phone, laptop, personal computer, or tablet computer), can execute the surgical data management system 160 of FIG. 1 to access various data sources through a network 260. The network 230 may be within a facility or multiple facilities maintained within a private network. The network 260 may be a wider area network, such as the internet. Accordingly, the networks 230 and 260 may have access to different files and data sets along with shared access to select files and data sets. In some aspects, networks 230 and 260 can be combined.
[0036] Turning now to FIG. 3, a system 300 for analyzing video and data is generally shown according to one or more aspects. In accordance with aspects, the video and data is captured from video recording system 104 of FIG. 1. The analysis can result in predicting features that include surgical phases and structures (e.g., instruments, anatomical structures, etc.) in the video data using machine learning. System 300 can be the computing system 102 of FIG. 1, or a part thereof in one or more examples. System 300 uses data streams in the surgical data to identify procedural states according to some aspects.
[0037] System 300 includes a data reception system 305 that collects surgical data, including the video data and surgical instrumentation data. The data reception system 305 can include one or more devices (e.g., one or more user devices and / or servers) located within and / or associated with a surgical operating room and / or control center. The data reception system 305 can receive surgical data in real-time, i.e., as the surgical procedure is being performed. Alternatively, or in addition, the data reception system 305 can receive or access surgical data in an offline manner, for example, by accessing data that is stored in the data collection system 150 of FIG. 1.
[0038] System 300 further includes a machine learning processing system 310 that processes the surgical data using one or more machine learning models to identify one ormore features, such as surgical phase, instrument, anatomical structure, etc., in the surgical data. It will be appreciated that machine learning processing system 310 can include one or more devices (e.g., one or more servers), each of which can be configured to include part or all of one or more of the depicted components of the machine learning processing system 310. In some instances, a part or all of the machine learning processing system 310 is in the cloud and / or remote from an operating room and / or physical location corresponding to a part or all of data reception system 305. It will be appreciated that several components of the machine learning processing system 310 are depicted and described herein. However, the components are just one example structure of the machine learning processing system 310, and that in other examples, the machine learning processing system 310 can be structured using a different combination of the components. Such variations in the combination of the components are encompassed by the technical solutions described herein.
[0039] The machine learning processing system 310 includes a machine learning training system 325, which can be a separate device (e.g., server) that stores its output as one or more trained machine learning models 330. The machine learning models 330 are accessible by a machine learning execution system 340. The machine learning execution system 340 can be separate from the machine learning training system 325 in some examples. In other words, in some aspects, devices that “train” the models are separate from devices that “infer,” i.e., perform real-time processing of surgical data using the trained machine learning models 330.
[0040] Machine learning processing system 310, in some examples, further includes a data generator 315 to generate simulated surgical data, such as a set of virtual images, or record the video data from the video recording system 104, to train the machine learning models 330. Data generator 315 can access (read / write) a data store 320 to record data, including multiple images and / or multiple videos. The images and / or videos can include images and / or videos collected during one or more procedures (e.g., one or more surgical procedures). For example, the images and / or video may have been collected by a user device worn by the actor 112 of FIG. 1 (e.g., surgeon, surgical nurse, anesthesiologist, etc.) during the surgery, anon-wearable imaging device located within an operating room, or an endoscopic camera inserted inside the patient 110 of FIG. 1. The data store 320 isseparate from the data collection system 150 of FIG. 1 in some examples. In other examples, the data store 320 is part of the data collection system 150.
[0041] Each of the images and / or videos recorded in the data store 320 for training the machine learning models 330 can be defined as a base image and can be associated with other data that characterizes an associated procedure and / or rendering specifications. For example, the other data can identify a type of procedure, a location of a procedure, one or more people involved in performing the procedure, surgical objectives, and / or an outcome of the procedure. Alternatively, or in addition, the other data can indicate a stage of the procedure with which the image or video corresponds, rendering specification with which the image or video corresponds and / or a type of imaging device that captured the image or video (e.g., and / or, if the device is a wearable device, a role of a particular person wearing the device, etc.). Further, the other data can include image-segmentation data that identifies and / or characterizes one or more objects (e.g., tools, anatomical objects, etc.) that are depicted in the image or video. The characterization can indicate the position, orientation, or pose of the object in the image. For example, the characterization can indicate a set of pixels that correspond to the object and / or a state of the object resulting from a past or current user handling. Localization can be performed using a variety of techniques for identifying objects in one or more coordinate systems.
[0042] The machine learning training system 325 uses the recorded data in the data store 320, which can include the simulated surgical data (e.g., set of virtual images) and actual surgical data to train the machine learning models 330. The machine learning model 330 can be defined based on a type of model and a set of hyperparameters (e.g., defined based on input from a client device). The machine learning models 330 can be configured based on a set of parameters that can be dynamically defined based on (e.g., continuous or repeated) training (i.e., learning, parameter tuning). Machine learning training system 325 can use one or more optimization algorithms to define the set of parameters to minimize or maximize one or more loss functions. The set of (learned) parameters can be stored as part of a trained machine learning model 330 using a specific data structure for that trained machine learning model 330. The data structure can also include one or more non-leamable variables (e.g., hyperparameters and / or model definitions).
[0043] Machine learning execution system 340 can access the data structure(s) of the machine learning models 330 and accordingly configure the machine learning models 330 for inference (i.e., prediction). The machine learning models 330 can include, for example, a fully convolutional network adaptation, an adversarial network model, an encoder, a decoder, or other types of machine learning models. The type of the machine learning models 330 can be indicated in the corresponding data structures. The machine learning model 330 can be configured in accordance with one or more hyperparameters and the set of learned parameters.
[0044] The machine learning models 330, during execution, receive, as input, surgical data to be processed and subsequently generate one or more inferences according to the training. For example, the video data captured by the video recording system 104 of FIG.1 can include data streams (e.g., an array of intensity, depth, and / or RGB values) for a single image or for each of a set of frames (e.g., including multiple images or an image with sequencing data) representing a temporal window of fixed or variable length in a video. The video data that is captured by the video recording system 104 can be received by the data reception system 305, which can include one or more devices located within an operating room where the surgical procedure is being performed. Alternatively, the data reception system 305 can include devices that are located remotely, to which the captured video data is streamed live during the performance of the surgical procedure. Alternatively, or in addition, the data reception system 305 accesses the data in an offline manner from the data collection system 150 or from any other data source (e.g., local or remote storage device).
[0045] The data reception system 305 can process the video and / or data received. The processing can include decoding when a video stream is received in an encoded format such that data for a sequence of images can be extracted and processed. The data reception system 305 can also process other types of data included in the input surgical data. For example, the surgical data can include additional data streams, such as audio data, RFID data, textual data, measurements from one or more surgical instruments / sensors, etc., that can represent stimuli / procedural states from the operating room. The data reception system 305 synchronizes the different inputs from the different devices / sensors before inputting them in the machine learning processing system 310.
[0046] The machine learning models 330, once trained, can analyze the input surgical data, and in one or more aspects, predict and / or characterize features (e.g., structures) included in the video data included with the surgical data. The video data can include sequential images and / or encoded video data (e.g., using digital video file / stream formats and / or codecs, such as MP4, MOV, AVI, WEBM, AVCHD, OGG, etc.). The prediction and / or characterization of the features can include segmenting the video data or predicting the localization of the structures with a probabilistic heatmap. In some instances, the one or more machine learning models include or are associated with a preprocessing or augmentation (e.g., intensity normalization, resizing, cropping, etc.) that is performed prior to segmenting the video data. An output of the one or more machine learning models can include image-segmentation or probabilistic heatmap data that indicates which (if any) of a defined set of structures are predicted within the video data, a location and / or position and / or pose of the structure(s) within the video data, and / or state of the structure(s). The location can be a set of coordinates in an image / frame in the video data. For example, the coordinates can provide a bounding box. The coordinates can provide boundaries that surround the structure(s) being predicted. The machine learning models 330, in one or more examples, are trained to perform higher-level predictions and tracking, such as predicting a phase of a surgical procedure and tracking one or more surgical instruments used in the surgical procedure.
[0047] While some techniques for predicting a surgical phase (“phase”) in the surgical procedure are described herein, it should be understood that any other technique for prediction can be used without affecting the aspects of the technical solutions described herein. In some examples, the machine learning processing system 310 includes a detector 350 that uses the machine learning models to, for instance, identify a phase within the surgical procedure (“procedure”). The detector 350 uses a particular procedural tracking data structure 355 from a list of procedural tracking data structures. The detector 350 selects the procedural tracking data structure 355 based on the type of surgical procedure that is being performed. In one or more examples, the type of surgical procedure is predetermined or input by actor 112. The procedural tracking data structure 355 identifies a set of potential phases that can correspond to a part of the specific type of procedure.
[0048] In some examples, the procedural tracking data structure 355 can be a graph that includes a set of nodes and a set of edges, with each node corresponding to a potentialphase. The edges can provide directional connections between nodes that indicate (via the direction) an expected order during which the phases will be encountered throughout an iteration of the procedure. The procedural tracking data structure 355 may include one or more branching nodes that feed to multiple next nodes and / or can include one or more points of divergence and / or convergence between the nodes. In some instances, a phase indicates a procedural action (e.g., surgical action) that is being performed or has been performed and / or indicates a combination of actions that have been performed. In some instances, a phase relates to a biological state of a patient undergoing a surgical procedure. For example, the biological state can indicate a complication (e.g., blood clots, clogged arteries / veins, etc.), pre-condition (e.g., lesions, polyps, etc.). In some examples, the machine learning models 330 are trained to detect an “abnormal condition,” such as hemorrhaging, arrhythmias, blood vessel abnormality, etc.
[0049] Each node within the procedural tracking data structure 355 can identify one or more characteristics of the phase corresponding to that node. The characteristics can include visual characteristics. In some instances, the node identifies one or more tools that are typically in use or availed for use (e.g., on a tool tray) during the phase. The node also identifies one or more roles of people who are typically performing a surgical task, a typical type of movement (e.g., of a hand or tool), etc. Thus, the detector 350 can use the segmented data generated by machine learning execution system 340 that indicates the presence and / or characteristics of particular objects within a field of view to identify an estimated node to which the real image data corresponds. Identification of the node (i.e., phase) can further be based upon previously detected phases for a given procedural iteration and / or other detected input (e.g., verbal audio data that includes person-to-person requests or comments, explicit identifications of a current or past phase, information requests, etc.).
[0050] The detector 350 can output a prediction associated with a portion of the video data that is analyzed by the machine learning processing system 310. For instance, a phase prediction can be associated with a portion of video data by identifying a start time and an end time of the portion of the video that is analyzed by the machine learning execution system 340. The prediction that is output can include segments of the video where each segment corresponds to and includes an identity of a surgical phase or other aspect as detected by the detector 350 based on the output of the machine learningexecution system 340. Further, the prediction, in one or more examples, can include additional data dimensions such as, but not limited to, identities of the structures (e.g., instrument, anatomy, etc.) that are identified by the machine learning execution system 340 in the portion of the video that is analyzed. The prediction can also include a confidence score of the prediction. Other examples can include various other types of information in the prediction that is output.
[0051] It should be noted that although some of the drawings depict endoscopic videos being analyzed, the technical solutions described herein can be applied to analyze video and image data captured by cameras that are not endoscopic (i.e. , cameras external to the patient’s body) when performing open surgeries (i.e., not laparoscopic surgeries). For example, the video and image data can be captured by cameras that are mounted on one or more personnel in the operating room, e.g., surgeon. Alternatively, or in addition, the cameras can be mounted on surgical instruments, walls, or other locations in the operating room.
[0052] Turning now to FIG. 4, a video processing pipeline 400 is depicted according to one or more aspects. In the example of FIG. 4, a video 402 can be processed 404, for instance, to select portions for further annotation and analysis. As one example, preliminary processing can remove portions of video that are not part of a surgical procedure, such as when an endoscopic or laparoscopic camera is outside of a patient's body. A machine-learning model 406 can perform annotations of the video 402 after the video 402 is processed 404 to mark portions of the video 402 that are associated with classification or segmentation criteria. For example, the machine-learning model 406 may identify a surgical phase on a frame-by-frame basis. The machine-learning model 406 can also determine confidence values associated with the annotations. For instance, the confidence values can be normalized as an average or standard deviation of pixels or patches within a frame. Further, the annotation and confidence can be determined relative to a sequence of frames. Resulting annotations and machine-learning confidence values can be provided to quality assurance decision engine 240 to determine a quality assurance confidence level. For example, a quality assurance confidence value that is considered high can result in the publication 410 of the video 402 and annotations without a need for further quality assurance review. A quality assurance confidence value that is considered medium or mid-level can result in the publication 412 of the video 402 and annotations asa provisional action before performing a secondary quality assurance review 414 to adjust any issues that may be minor in nature. Upon making adjustments during the secondary quality assurance review 414, the result can be output as publication 410 of the video 402 and annotations. If the quality assurance confidence value is considered low, publication of the video 402 and annotations can be prevented or blocked until a secondary quality assurance review 416 is performed to adjust the annotations in error. The secondary quality assurance review 414, 416 can be performed by a human using information provided by the quality assurance decision engine 240. For example, the quality assurance decision engine 240 can provide a summary of the detected annotations that led to a reduced quality assurance confidence value and / or other detected conditions that resulted in the quality assurance confidence value being below a threshold level associated with a high quality assurance confidence value. After the secondary quality assurance review 416 is complete, the result can be output as publication 410 of the video 402 and annotations.
[0053] In some aspects, annotations generated by the machine-learning model 406 can be used by one or more secondary machine-learning models 418. For example, phase determination inferences from the machine-learning model 406 can be used to infer instrument and / or anatomy with greater accuracy. Secondary annotations generated by the one or more secondary machine-learning models 418 can be output as publication 410 of the video 402 and annotations with secondary annotations. If there are problems with the annotations produced by the machine-learning model 406, which are then fed to the one or more secondary machine-learning models 418, the annotations or predictions of the secondary machine-learning models 418 may be incorrect. In some aspects, if the secondary quality assurance review 414, 416 detects a problem with the annotations of the machine-learning model 406, then output of the secondary machine-learning models 418 can also be flagged as potentially in need of review. An accuracy monitor 420 can periodically check the publication 410 to determine whether annotations by the machinelearning model 406 or secondary annotations by the secondary machine-learning models 418 are accurate. Further, the accuracy monitor 420 may also identify whether the secondary quality assurance review 414, 416 likely corrected issues with annotations of the machine-learning model 406.
[0054] Values of high confidence, medium confidence, and low confidence can be configurable, for instance, as a first confidence threshold, where values above the first confidence threshold are considered high confidence. Values below a second confidence threshold can be considered low confidence. Values between the first confidence threshold and the second confidence threshold can be considered medium confidence.
[0055] FIG. 5 depicts a video processing pipeline 500 according to one or more aspects. The video processing pipeline 500 includes similar elements as the video processing pipeline 400 of FIG. 4. One difference is that the secondary machine-learning models 418 of FIG. 5 can directly make inferences off of the video 402 as processed 404 without relying upon annotations from the machine-learning model 406. Results of inferences by the secondary machine-learning models 418 can also be subject to the quality assurance decision engine 240.
[0056] FIG. 6 depicts components 600 of a quality assurance decision engine, such as quality assurance decision engine 240, according to one or more aspects. In the example of FIG. 6, the quality assurance decision engine 240 can include rule-based logic 602, a workflow outlier detector 604, a machine-learning confidence verifier 606, and a quality assurance confidence value generator 608. In some aspects, the rule-based logic 602 can be logical rules based on domain knowledge to identify illogical or unusual annotations in the annotations 610 generated by the machine-learning model 406 and / or the secondary machine-learning model 418. The annotation rules can include one or more of: a missing phase, a phase limit break, an atypical phase transition, an atypical phase sequence, a toggle breach limit, a minimum duration not met, and a maximum duration exceeded. In some aspects, a missing phase can be determined based on an expected phase not being present. The phase limit breach can be determined where a phase occurs more than a maximum number of times. An unusual transition from one phase to another, e.g., phase A directly after B, is an example of an unusual phase transition. An unusual ordering of phases within an overall workflow can be, for example, a first occurrence of phase B occurring before a first occurrence of phase A. In some aspects, a toggle limit breach can be determined based on alternation between two phases occurring more than the maximum number of times. As another example, a minimum duration not met can be where a case or phase duration is less than a predefined minimum time for an associated case type or phase type. A maximum duration exceeded is another type of rule, where a case or phaseduration is greater than a maximum duration for an associated case type or phase type. Rules can be defined with a severity level to distinguish between conditions that are possible with a lower likelihood of occurrence and conditions that should not occur. For example, rules can have severity levels established to distinguish between high, medium, and low levels.
[0057] As an example of rules for a specific type of case, such as a prostatectomy, a missing phase could be identified if a bladder neck transection phase is not annotated. A phase limit breach could be detected if 4 occurrences of vas and seminal vesicles phase occur in one case. An unusual phase transition could be detected if specimen retrieval is followed by port insertion. An unusual phase sequence could be detected if a bladder neck transection occurs before the first occurrence of releasing the bladder. A toggle limit breach could be detected if alternation between vas and seminal vesicles and transection of prostatic pedicles occurs 3 times. A minimum duration not met could be detected if the case duration is less than 1 hour. A maximum duration exceeded could be detected if the case duration is greater than 5 hours.
[0058] The workflow outlier detector 604 can use one or more of: an autoencoder reconstruction accuracy or variational autoencoder reconstruction accuracy to identify unusual phase sequences, a local outlier factor to identify outlier phases, and a statistical model to compare phase transitions and durations to historical data. The autoencoder reconstruction accuracy can be determined using an autoencoder, such as autoencoder 700 of FIG. 7. The autoencoder 700 can receive input 702, perform encoding 704 that compresses layers down to a layer of compressed data 706 and expand the compressed data 706 through expanding layers to perform decoding 708 and produce an output 710. The autoencoder 700 can be trained on historical annotations to learn structure data. The annotations 610 can be passed as input 702 to the autoencoder 700 to compress the data into a set of encodings as compressed data 706 and perform decoding 708 to create a reconstruction of the annotations 610 as the output 710. For annotations 610 that are unusual compared to training data, the output 710 can exhibit poor reconstruction with low reconstruction accuracy. This can be determined, for example, by comparing the output 710 to the input 702 for differences.
[0059] The local outlier factor can be determined, for example, by a local outlier factor determination 800 of the workflow outlier detector 604. The local outlier factordetermination 800 can include an autoencoder 802 that is trained on historical annotations to leam structure data, similar to the encoding 704 portion of the autoencoder 700 of FIG. 7. The local outlier factor determination 800 can also include an outlier model 804 that can be generated based on encodings of the training data. The encoding process can convert a whole video of annotations down to a one-dimensional array. The onedimensional array of the outlier model 804 can support using clustering methods to identify local outliers. The annotations 610 of FIG. 6 can be passed to the autoencoder 802 to generate encodings that are passed to the outlier model 804, and outliers can be identified using a local outlier factor 806.
[0060] In some aspects, the workflow outlier detector 604 of FIG. 6 can use a statistical model to identify outliers. As an example, historical annotations can be used to create a transition metric that calculates a likelihood of transitioning from one phase to another. This data can also be used to determine the distribution of phase durations. The annotations 610 can be used to generate a workflow indicating a sequence of transitions between phases. Likelihoods of transitions of the workflow can be determined using the transition matrix, and likelihood of the durations of each phase can be calculated based on the historical data. Likelihood scores can be determined for videos based on workflow and phase durations. Videos with likelihood scores below a likelihood threshold can be identified as outliers.
[0061] With reference to FIG. 6, the machine-learning confidence verifier 606 can receive a machine-learning confidence value 612 from the machine-learning model 406 and / or from the secondary machine-learning model 418. The machine-learning confidence verifier 606 can compare the machine-learning confidence value 612 to one or more thresholds to convert numerical values into a high, medium, or low confidence level designation.
[0062] A machine-learning model maturity value 614 can be associated with the machine-learning model 406 and / or the secondary machine-learning model 418 indicating an accuracy level with respect to time. For example, during initial testing, the machinelearning model 406 may be deemed to have a machine-learning model maturity value 614 that is low until an accuracy target (e.g., 90%) is maintained for a first maturity period before advancing to a medium level. After maintaining the accuracy target for a second maturity period, the machine-learning model maturity value 614 can be deemed as high.If the accuracy dips below a minimum threshold, model retraining can be triggered. The machine-learning model maturity value 614 can be adjusted based on an accuracy of the machine-learning model over a period of time. Accuracy monitor 420 can be used to monitor the accuracy of the machine-learning model 406 and / or the secondary machinelearning model 418. The accuracy monitor 420 can include random reviews to gauge performance. The accuracy monitor 420 can also check results of the secondary quality assurance review 414, 416. Changes made by the secondary quality assurance review 414, 416 can also impact the accuracy scores and result in changes to the machine-learning model maturity value 614.
[0063] The quality assurance confidence value generator 608 can use a combination of outputs from the rule-based logic 602, workflow outlier detector 604, machine-learning confidence verifier 606, and machine-learning model maturity value 614 to determine a quality assurance confidence value 616 as output of the quality assurance decision engine 240. As one example, the quality assurance confidence value 616 can be set as high confidence if the machine-learning model maturity value 614 has a high maturity value and no issues are detected by the rule-based logic 602, workflow outlier detector 604, and machine-learning confidence verifier 606. The quality assurance confidence value 616 can be set as medium confidence if the machine-learning model maturity value 614 has a medium maturity value or an issue was detected by workflow outlier detector 604 while the machine-learning model maturity value 614 has a high maturity value. The quality assurance confidence value 616 can be set as low confidence if the machine-learning model maturity value 614 has a low maturity value or an issue was detected by the rulebased logic 602, or an issue was detected by the workflow outlier detector 604 while the machine-learning model maturity value 614 has a medium maturity value.
[0064] As another example, the quality assurance confidence value generator 608 can determine the quality assurance confidence value 616 using various rule combinations. For instance, cases that break at least one high severity rule, at least 2 medium severity rules or at least 4 low severity rules can be sent to secondary quality assurance review 416 before publishing. Cases that don’t break rules but are identified as an outlier can be sent to publication 412 and secondary quality assurance review 414. Incomplete cases, e.g., cases that are missing required phases, may not be passed through the outlier model, butinstead can be flagged for secondary quality assurance review 416 based on relevant annotation rules.
[0065] Cases can be classified as needing a major edit if a phase is completely missed or has moved over a threshold amount of the video length (e.g., 50%). Major edit cases can be sent to secondary quality assurance review 416. Cases can be classified as needing a minor edit if a phase repetition or phase toggling is missed, an extra phase was added, or a phase moved over a lower threshold amount of the video length (e.g., 25%). Threshold values can be set according to different options to establish expected quality assurance review targets, for instance, based on an option table. Performance of the quality assurance decision engine 240 and resulting reviews can be monitored over time along with the accuracy and maturity of machine-learning models to determine whether the overall performance aligns with a target option. Subsequent adjustments can be made to thresholds and / or logic of the quality assurance confidence value generator 608.
[0066] Turning now to FIG. 9, a flowchart of a method 1100 for automated quality assurance of machine-learning model output is generally shown in accordance with one or more aspects. All or a portion of method 1100 can be implemented, for example, by all or a portion of CAS system 100 of FIG. 1, the system 200 of FIG. 2, and / or computer system 1200 of FIG. 10, for instance through execution of the surgical data management system 160, quality assurance decision engine 240, and / or surgical data post-processing system 250.
[0067] At block 1102, a machine-learning model 406 can be applied to a video 402 of a surgical procedure to produce a plurality of annotations 610 and machine-learning model confidence values 612 associated with the annotations 610. At block 1104, a quality assurance confidence value 616 can be determined based on the annotations 610 and the machine-learning model confidence values 612. At block 1106, the video 402 with the annotations 610 can be published based on determining that the quality assurance confidence value 616 is above a first confidence threshold. At block 1108, a secondary quality assurance review 416 can be triggered and publishing of the annotations 610 prevented until completion of the secondary quality assurance review 416 based on determining that the quality assurance confidence value 616 is below a second confidence threshold. At block 1110, the video 402 with the annotations 610 can be published and the secondary quality assurance review 414 can be triggered based on determining that thequality assurance confidence value 616 is below the first confidence threshold and above the second confidence threshold.
[0068] In some aspects, the annotations 610 can be phases of the surgical procedure.
[0069] In some aspects, quality assurance confidence value 616 can be determined by a quality assurance decision engine 240 using rule-based logic 602 to check the annotations 610. In some aspects, the rule-based logic 602 can be configured to check for one or more of: a missing phase, a phase limit break, an atypical phase transition, an atypical phase sequence, a toggle breach limit, a minimum duration not met, and a maximum duration exceeded. In some aspects, one or more rules can be defined with a severity level to distinguish between conditions that are possible with a lower likelihood of occurrence and conditions that should not occur.
[0070] In some aspects, the quality assurance confidence value 616 can be determined by a quality assurance decision engine 240 using a workflow outlier detector 604 to check the annotations 610. In some aspects, the workflow outlier detector 604 can use one or more of: an autoencoder reconstruction accuracy to identify unusual phase sequences, a local outlier factor to identify outlier phases, and a statistical model to compare phase transitions and durations to historical data.
[0071] In some aspects, the quality assurance confidence value 616 can be determined by a quality assurance decision engine 240 using a machine-learning confidence verifier 606 to check the machine-learning model confidence values 612. In some aspects, the machine-learning confidence verifier 606 can be configured to check the machine-learning model confidence values 612 on a per-frame basis relative to one or more level thresholds.
[0072] In some aspects, the quality assurance confidence value 616 can be determined by a quality assurance decision engine 240 using a machine-learning model maturity value 614 that is adjusted based on an accuracy of the machine-learning model 406 over a period of time.
[0073] In some aspects, one or more secondary machine-learning models 418 can be applied to the video 402. The one or more secondary machine-learning models 418 can use the annotations 610 generated by the machine-learning model 406 to generate additional annotations published with the video 402.
[0074] In some aspects, the one or more secondary machine-learning models 418 can generate additional annotations published with the video 402, where a secondary qualityassurance confidence value is determined by a quality assurance decision engine 240 that checks both the annotations and the additional annotations. In some aspects, the one or more secondary machine-learning models 418 can produce secondary machine-learning model confidence values associated with the additional annotations, and the quality assurance decision engine 240 can be configured to check the secondary machine-learning model confidence values to determine the secondary quality assurance confidence value.
[0075] The processing shown in FIG. 9 is not intended to indicate that the operations are to be executed in any particular order or that all of the operations shown in FIG. 9 are to be included in every case. Additionally, the processing shown in FIG. 9 can include any suitable number of additional operations.
[0076] Turning now to FIG. 10, a computer system 1200 is generally shown in accordance with an aspect. The computer system 1200 can be an electronic computer framework comprising and / or employing any number and combination of computing devices and networks utilizing various communication technologies, as described herein. The computer system 1200 can be easily scalable, extensible, and modular, with the ability to change to different services or reconfigure some features independently of others. The computer system 1200 may be, for example, a server, desktop computer, laptop computer, tablet computer, or smartphone. In some examples, computer system 1200 may be a cloud computing node. Computer system 1200 may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, and so on that perform particular tasks or implement particular abstract data types. Computer system 1200 may be practiced in distributed cloud computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.
[0077] As shown in FIG. 10, the computer system 1200 has one or more central processing units (CPU(s)) 1201a, 1201b, 1201c, etc. (collectively or generically referred to as processor(s) 1201). The processors 1201 can be a single-core processor, multi-core processor, computing cluster, or any number of other configurations. The processors 1201 can be any type of circuitry capable of executing instructions. The processors 1201, alsoreferred to as processing circuits, are coupled via a system bus 1202 to a system memory 1203 and various other components. The system memory 1203 can include one or more memory devices, such as read-only memory (ROM) 1204 and a random-access memory (RAM) 1205. The ROM 1204 is coupled to the system bus 1202 and may include a basic input / output system (BIOS), which controls certain basic functions of the computer system 1200. The RAM is read- write memory coupled to the system bus 1202 for use by the processors 1201. The system memory 1203 provides temporary memory space for operations of said instructions during operation. The system memory 1203 can include random access memory (RAM), read-only memory, flash memory, or any other suitable memory systems.
[0078] The computer system 1200 comprises an input / output (I / O) adapter 1206 and a communications adapter 1207 coupled to the system bus 1202. The I / O adapter 1206 may be a small computer system interface (SCSI) adapter that communicates with a hard disk 1208 and / or any other similar component. The I / O adapter 1206 and the hard disk 1208 are collectively referred to herein as a mass storage 1210.
[0079] Software 1211 for execution on the computer system 1200 may be stored in the mass storage 1210. The mass storage 1210 is an example of a tangible storage medium readable by the processors 1201, where the software 1211 is stored as instructions for execution by the processors 1201 to cause the computer system 1200 to operate, such as is described hereinbelow with respect to the various Figures. Examples of computer program product and the execution of such instruction is discussed herein in more detail. The communications adapter 1207 interconnects the system bus 1202 with a network 1212, which may be an outside network, enabling the computer system 1200 to communicate with other such systems. In one aspect, a portion of the system memory 1203 and the mass storage 1210 collectively store an operating system, which may be any appropriate operating system to coordinate the functions of the various components shown in FIG. 10.
[0080] Additional input / output devices are shown as connected to the system bus 1202 via a display adapter 1215 and an interface adapter 1216 and. In one aspect, the adapters 1206, 1207, 1215, and 1216 may be connected to one or more I / O buses that are connected to the system bus 1202 via an intermediate bus bridge (not shown). A display 1219 (e.g., a screen or a display monitor) is connected to the system bus 1202 by a display adapter1215, which may include a graphics controller to improve the performance of graphicsintensive applications and a video controller. A keyboard, a mouse, a touchscreen, one or more buttons, a speaker, etc., can be interconnected to the system bus 1202 via the interface adapter 1216, which may include, for example, a Super I / O chip integrating multiple device adapters into a single integrated circuit. Suitable I / O buses for connecting peripheral devices such as hard disk controllers, network adapters, and graphics adapters typically include common protocols, such as the Peripheral Component Interconnect (PCI). Thus, as configured in FIG. 10, the computer system 1200 includes processing capability in the form of the processors 1201, and storage capability including the system memory 1203 and the mass storage 1210, input means such as the buttons, touchscreen, and output capability including the speaker 1223 and the display 1219.
[0081] In some aspects, the communications adapter 1207 can transmit data using any suitable interface or protocol, such as the internet small computer system interface, among others. The network 1212 may be a cellular network, a radio network, a wide area network (WAN), a local area network (LAN), or the Internet, among others. An external computing device may connect to the computer system 1200 through the network 1212. In some examples, an external computing device may be an external web server or a cloud computing node.
[0082] It is to be understood that the block diagram of FIG. 10 is not intended to indicate that the computer system 1200 is to include all of the components shown in FIG. 10. Rather, the computer system 1200 can include any appropriate fewer or additional components not illustrated in FIG. 10 (e.g., additional memory components, embedded controllers, modules, additional network interfaces, etc.). Further, the aspects described herein with respect to computer system 1200 may be implemented with any appropriate logic, wherein the logic, as referred to herein, can include any suitable hardware (e.g., a processor, an embedded controller, or an application-specific integrated circuit, among others), software (e.g., an application, among others), firmware, or any suitable combination of hardware, software, and firmware, in various aspects. Various aspects can be combined to include two or more of the aspects described herein.
[0083] In some aspects, a computer-implemented method for automated quality assurance of machine-learning model output can include applying a machine-learning model to a video of a surgical procedure to produce a plurality of annotations and machine-learningmodel confidence values associated with the annotations, determining a quality assurance confidence value based on the annotations and the machine-learning model confidence values, publishing the video with the annotations based on determining that the quality assurance confidence value is above a first confidence threshold, triggering a secondary quality assurance review and prevent publishing of the annotations until completion of the secondary quality assurance review based on determining that the quality assurance confidence value is below a second confidence threshold, and publishing the video with the annotations and trigger the secondary quality assurance review based on determining that the quality assurance confidence value is below the first confidence threshold and above the second confidence threshold.
[0084] In some aspects, the computer-implemented method can include where the quality assurance confidence value is determined by a quality assurance decision engine using rule-based logic to check the annotations and a workflow outlier detector to check the annotations.
[0085] In some aspects, the quality assurance confidence value can be determined by the quality assurance decision engine on a per-frame basis relative to one or more level thresholds, and the quality assurance confidence value is determined based on a machinelearning model maturity value that is adjusted based on an accuracy of the machinelearning model over a period of time.
[0086] In some aspects, the annotations and machine-learning model confidence values are produced by the machine-learning model during the surgical procedure, post- operatively, or a combination thereof.
[0087] In some aspects, a computer program product includes a memory device with computer readable instructions stored thereon, wherein executing the computer readable instructions by one or more processing units causes the one or more processing units to perform a plurality of operations. The operations can include applying a machine-learning model to a video of a surgical procedure to produce a plurality of annotations, determining a quality assurance confidence value based on the annotations, publishing the video with the annotations based on determining that the quality assurance confidence value is above a first confidence threshold, triggering a secondary quality assurance review and prevent publishing of the annotations until completion of the secondary quality assurance review based on determining that the quality assurance confidence value is below a secondconfidence threshold, and publishing the video with the annotations and trigger the secondary quality assurance review based on determining that the quality assurance confidence value is below the first confidence threshold and above the second confidence threshold.
[0088] In some aspects, the quality assurance confidence value can be determined by a quality assurance decision engine using rule-based logic to check the annotations, wherein one or more rules are defined with a severity level to distinguish between conditions that are possible with a lower likelihood of occurrence and conditions that should not occur.
[0089] In some aspects, the quality assurance confidence value can be determined by the quality assurance decision engine using a workflow outlier detector to check the annotations, wherein the workflow outlier detector uses one or more of: an autoencoder reconstruction accuracy, a local outlier factor, and a statistical model.
[0090] In some aspects, the quality assurance confidence value can be determined by the quality assurance decision engine using a machine-learning model maturity value that is adjusted based on an accuracy of the machine-learning model over a period of time.
[0091] In some aspects, the operations can include other aspects as previously described.
[0092] The present invention may be a system, a method, and / or a computer program product at any possible technical detail level of integration. The computer program product may include a computer-readable storage medium (or media) having computer- readable program instructions thereon for causing a processor to carry out aspects of the present invention.
[0093] The computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non- exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in agroove having instructions recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0094] Computer-readable program instructions described herein can be downloaded to respective computing / processing devices from a computer-readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network, and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0095] Computer-readable program instructions for carrying out operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuitry, or either source-code or object code written in any combination of one or more programming languages, including an object-oriented programming language such as Smalltalk, C++, or the like, and procedural programming languages, such as the “C” programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user’s computer, partly on the user’s computer, as a stand-alone software package, partly on the user’s computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some aspects, electronic circuitry including, for example, programmable logic circuitry, field- programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute thecomputer-readable program instruction by utilizing state information of the computer- readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present invention.
[0096] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to aspects of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer- readable program instructions.
[0097] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.
[0098] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0099] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various aspects of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, thefunctions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
[0100] The descriptions of the various aspects of the present invention have been presented for purposes of illustration but are not intended to be exhaustive or limited to the aspects disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described aspects. The terminology used herein was chosen to best explain the principles of the aspects, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the aspects described herein.
[0101] Various aspects of the invention are described herein with reference to the related drawings. Alternative aspects of the invention can be devised without departing from the scope of this invention. Various connections and positional relationships (e.g., over, below, adjacent, etc.) are set forth between elements in the following description and in the drawings. These connections and / or positional relationships, unless specified otherwise, can be direct or indirect, and the present invention is not intended to be limiting in this respect. Accordingly, a coupling of entities can refer to either a direct or an indirect coupling, and a positional relationship between entities can be a direct or indirect positional relationship. Moreover, the various tasks and process steps described herein can be incorporated into a more comprehensive procedure or process having additional steps or functionality not described in detail herein.
[0102] The following definitions and abbreviations are to be used for the interpretation of the claims and the specification. As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” “contains,” or “containing,” or any other variation thereof are intended to cover a non-exclusive inclusion. For example, a composition, a mixture, process, method, article, or apparatus that comprises a list ofelements is not necessarily limited to only those elements but can include other elements not expressly listed or inherent to such composition, mixture, process, method, article, or apparatus.
[0103] Additionally, the term “exemplary” is used herein to mean “serving as an example, instance or illustration.” Any aspect or design described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects or designs. The terms “at least one” and “one or more” may be understood to include any integer number greater than or equal to one, i.e., one, two, three, four, etc. The terms “a plurality” may be understood to include any integer number greater than or equal to two, i.e., two, three, four, five, etc. The term “connection” may include both an indirect “connection” and a direct “connection.”
[0104] The terms “about,” “substantially,” “approximately,” and variations thereof are intended to include the degree of error associated with measurement of the particular quantity based upon the equipment available at the time of filing the application. For example, “about” can include a range of ± 8% or 5%, or 2% of a given value.
[0105] For the sake of brevity, conventional techniques related to making and using aspects of the invention may or may not be described in detail herein. In particular, various aspects of computing systems and specific computer programs to implement the various technical features described herein are well known. Accordingly, in the interest of brevity, many conventional implementation details are only mentioned briefly herein or are omitted entirely without providing the well-known system and / or process details.
[0106] It should be understood that various aspects disclosed herein may be combined in different combinations than the combinations specifically presented in the description and accompanying drawings. It should also be understood that, depending on the example, certain acts or events of any of the processes or methods described herein may be performed in a different sequence, may be added, merged, or left out altogether (e.g., all described acts or events may not be necessary to carry out the techniques). In addition, while certain aspects of this disclosure are described as being performed by a single module or unit for purposes of clarity, it should be understood that the techniques of this disclosure may be performed by a combination of units or modules associated with, for example, a medical device.
[0107] In one or more examples, the described techniques may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include non-transitory computer-readable media, which corresponds to a tangible medium such as data storage media (e.g., RAM, ROM, EEPROM, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer).
[0108] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term “processor” as used herein may refer to any of the foregoing structure or any other physical structure suitable for implementation of the described techniques. Also, the techniques could be fully implemented in one or more circuits or logic elements.
[0109] The following examples are illustrative of the techniques described herein.
[0110] Example 1. A system comprising: a memory device; and one or more processors coupled with the memory device, the one or more processors configured to: apply a machine-learning model to a video of a surgical procedure to produce a plurality of annotations and machine-learning model confidence values associated with the annotations, wherein the annotations are phases of the surgical procedure; determine a quality assurance confidence value based on the annotations and the machine-learning model confidence values; publish the video with the annotations based on determining that the quality assurance confidence value is above a first confidence threshold; trigger a secondary quality assurance review and prevent publishing of the annotations until completion of the secondary quality assurance review based on determining that the quality assurance confidence value is below a second confidence threshold; and publish the video with the annotations and trigger the secondary quality assurance review based on determining that the quality assurance confidence value is below the first confidence threshold and above the second confidence threshold.
[0111] Example 2. The system of Example 1, wherein the quality assurance confidence value is determined by a quality assurance decision engine using rule-based logic to check the annotations.
[0112] Example 3. The system of Example 2, wherein the rule-based logic is configured to check for one or more of: a missing phase, a phase limit break, an atypical phase transition, an atypical phase sequence, a toggle breach limit, a minimum duration not met, and a maximum duration exceeded.
[0113] Example 4. The system of Example 3, wherein one or more rules are defined with a severity level to distinguish between conditions that are possible with a lower likelihood of occurrence and conditions that should not occur.
[0114] Example 5. The system of Example 1, wherein the quality assurance confidence value is determined by a quality assurance decision engine using a workflow outlier detector to check the annotations.
[0115] Example 6. The system of Example 5, wherein the workflow outlier detector uses one or more of: an autoencoder reconstruction accuracy to identify unusual phase sequences, a local outlier factor to identify outlier phases, and a statistical model to compare phase transitions and durations to historical data.
[0116] Example 7. The system of Example 1, wherein the quality assurance confidence value is determined by a quality assurance decision engine using a machine-learning confidence verifier to check the machine-learning model confidence values.
[0117] Example 8. The system of Example 7, wherein the machine-learning confidence verifier is configured to check the machine-learning model confidence values on a perframe basis relative to one or more level thresholds.
[0118] Example 9. The system of Example 1, wherein the quality assurance confidence value is determined by a quality assurance decision engine using a machine-learning model maturity value that is adjusted based on an accuracy of the machine-learning model over a period of time.
[0119] Example 10. The system of Example 1, wherein the one or more processors are configured to: apply one or more secondary machine-learning models to the video, wherein the one or more secondary machine-learning models use the annotations generated by the machine-learning model to generate additional annotations published with the video.
[0120] Example 11. The system of Example 1, wherein the one or more processors are configured to: apply one or more secondary machine-learning models to the video, wherein the one or more secondary machine-learning models generate additional annotations published with the video, wherein a secondary quality assurance confidence value is determined by a quality assurance decision engine that checks both the annotations and the additional annotations.
[0121] Example 12. The system of Example 11, wherein the one or more secondary machine-learning models produce secondary machine-learning model confidence values associated with the additional annotations, and the quality assurance decision engine is configured to check the secondary machine-learning model confidence values to determine the secondary quality assurance confidence value.
[0122] Example 13. A computer-implemented method for automated quality assurance of machine-learning model output, the method comprising: applying a machine-learning model to a video of a surgical procedure to produce a plurality of annotations and machinelearning model confidence values associated with the annotations; determining a quality assurance confidence value based on the annotations and the machine-learning model confidence values; publishing the video with the annotations based on determining that the quality assurance confidence value is above a first confidence threshold; triggering a secondary quality assurance review and prevent publishing of the annotations until completion of the secondary quality assurance review based on determining that the quality assurance confidence value is below a second confidence threshold; and publishing the video with the annotations and trigger the secondary quality assurance review based on determining that the quality assurance confidence value is below the first confidence threshold and above the second confidence threshold.
[0123] Example 14. The computer-implemented method of Example 13, wherein the quality assurance confidence value is determined by a quality assurance decision engine using rule-based logic to check the annotations and a workflow outlier detector to check the annotations.
[0124] Example 15. The computer-implemented method of Example 14, wherein the quality assurance confidence value is determined by the quality assurance decision engine on a per-frame basis relative to one or more level thresholds, and the quality assuranceconfidence value is determined based on a machine-learning model maturity value that is adjusted based on an accuracy of the machine-learning model over a period of time.
[0125] Example 16. The computer-implemented method of Example 13, wherein the annotations and machine-learning model confidence values are produced by the machinelearning model during the surgical procedure, post-operatively, or a combination thereof.
[0126] Example 17. A computer program product comprising a memory device with computer readable instructions stored thereon, wherein executing the computer readable instructions by one or more processing units causes the one or more processing units to perform a plurality of operations comprising: applying a machine-learning model to a video of a surgical procedure to produce a plurality of annotations; determining a quality assurance confidence value based on the annotations; publishing the video with the annotations based on determining that the quality assurance confidence value is above a first confidence threshold; triggering a secondary quality assurance review and prevent publishing of the annotations until completion of the secondary quality assurance review based on determining that the quality assurance confidence value is below a second confidence threshold; and publishing the video with the annotations and trigger the secondary quality assurance review based on determining that the quality assurance confidence value is below the first confidence threshold and above the second confidence threshold.
[0127] Example 18. The computer program product of Example 17, wherein the quality assurance confidence value is determined by a quality assurance decision engine using rule-based logic to check the annotations, wherein one or more rules are defined with a severity level to distinguish between conditions that are possible with a lower likelihood of occurrence and conditions that should not occur.
[0128] Example 19. The computer program product of Example 18, wherein the quality assurance confidence value is determined by the quality assurance decision engine using a workflow outlier detector to check the annotations, wherein the workflow outlier detector uses one or more of: an autoencoder reconstruction accuracy, a local outlier factor, and a statistical model.
[0129] Example 20. The computer program product of Example 19, wherein the quality assurance confidence value is determined by the quality assurance decision engine using a machine-learning model maturity value that is adjusted based on an accuracy of the machine-learning model over a period of time.
Claims
CLAIMSWhat is claimed is:
1. A system (100, 200, 1200) comprising: a memory device; and one or more processors coupled with the memory device, the one or more processors configured to: apply a machine-learning model (406) to a video (402) of a surgical procedure to produce a plurality of annotations and machine-learning model confidence values associated with the annotations, wherein the annotations are phases of the surgical procedure; determine a quality assurance confidence value based on the annotations and the machine-learning model confidence values; publish the video (402) with the annotations based on determining that the quality assurance confidence value is above a first confidence threshold; trigger a secondary quality assurance review and prevent publishing of the annotations until completion of the secondary quality assurance review based on determining that the quality assurance confidence value is below a second confidence threshold; and publish the video (402) with the annotations and trigger the secondary quality assurance review based on determining that the quality assurance confidence value is below the first confidence threshold and above the second confidence threshold.
2. The system (100, 200, 1200) of claim 1, wherein the quality assurance confidence value is determined by a quality assurance decision engine using rule-based logic to check the annotations.
3. The system (100, 200, 1200) of claim 2, wherein the rule-based logic is configured to check for one or more of: a missing phase, a phase limit break, an atypical phase transition, an atypical phase sequence, a toggle breach limit, a minimum duration not met, and a maximum duration exceeded.
4. The system (100, 200, 1200) of claim 3, wherein one or more rules are defined with a severity level to distinguish between conditions that are possible with a lower likelihood of occurrence and conditions that should not occur.
5. The system (100, 200, 1200) of any of claims 1 to 4, wherein the quality assurance confidence value is determined by a quality assurance decision engine using a workflow outlier detector to check the annotations, and optionally wherein the workflow outlier detector uses one or more of: an autoencoder reconstruction accuracy to identify unusual phase sequences, a local outlier factor to identify outlier phases, and a statistical model to compare phase transitions and durations to historical data.
6. The system (100, 200, 1200) of any of claims 1 to 5, wherein the quality assurance confidence value is determined by a quality assurance decision engine using a machinelearning confidence verifier to check the machine-learning model confidence values, and optionally wherein the machine-learning confidence verifier is configured to check the machine-learning model confidence values on a per-frame basis relative to one or more level thresholds.
7. The system (100, 200, 1200) of any of claims 1 to 6, wherein the quality assurance confidence value is determined by a quality assurance decision engine using a machinelearning model maturity value that is adjusted based on an accuracy of the machinelearning model over a period of time.
8. The system (100, 200, 1200) of any of claims 1 to 7, wherein the one or more processors are configured to: apply one or more secondary machine-learning models (418) to the video (402), wherein the one or more secondary machine-learning models (418) use the annotations generated by the machine-learning model (406) to generate additional annotations published with the video (402).
9. The system (100, 200, 1200) of any of claims 1 to 8, wherein the one or more processors are configured to: apply one or more secondary machine-learning models (418) to the video (402), wherein the one or more secondary machine-learning models (418) generate additional annotations published with the video (402), wherein a secondary quality assurance confidence value is determined by a quality assurance decision engine that checks both the annotations and the additional annotations, and optionally wherein the one or more secondary machine-learning models (418) produce secondary machine-learning model confidence values associated with the additional annotations, and the quality assurance decision engine is configured to check the secondary machine-learning model confidence values to determine the secondary quality assurance confidence value.
10. A computer-implemented method for automated quality assurance of machine-learning model output, the method comprising: applying a machine-learning model (406) to a video (402) of a surgical procedure to produce a plurality of annotations and machine-learning model confidence values associated with the annotations; determining a quality assurance confidence value based on the annotations and the machine-learning model confidence values; publishing the video (402) with the annotations based on determining that the quality assurance confidence value is above a first confidence threshold; triggering a secondary quality assurance review and prevent publishing of the annotations until completion of the secondary quality assurance review based on determining that the quality assurance confidence value is below a second confidence threshold; and publishing the video (402) with the annotations and trigger the secondary quality assurance review based on determining that the quality assurance confidence value is below the first confidence threshold and above the second confidence threshold.
11. The computer-implemented method of claim 10, wherein the quality assurance confidence value is determined by a quality assurance decision engine using rule-based logic to check the annotations and a workflow outlier detector to check the annotations,and optionally wherein the quality assurance confidence value is determined by the quality assurance decision engine on a per-frame basis relative to one or more level thresholds, and the quality assurance confidence value is determined based on a machine-learning model maturity value that is adjusted based on an accuracy of the machine-learning model over a period of time.
12. The computer-implemented method of claim 10 or 11, wherein the annotations and machine-learning model confidence values are produced by the machine-learning model (406) during the surgical procedure, post-operatively, or a combination thereof.
13. A computer program product comprising a memory device with computer readable instructions stored thereon, wherein executing the computer readable instructions by one or more processing units causes the one or more processing units to perform a plurality of operations comprising: applying a machine-learning model to a video (402) of a surgical procedure to produce a plurality of annotations; determining a quality assurance confidence value based on the annotations; publishing the video (402) with the annotations based on determining that the quality assurance confidence value is above a first confidence threshold; triggering a secondary quality assurance review and prevent publishing of the annotations until completion of the secondary quality assurance review based on determining that the quality assurance confidence value is below a second confidence threshold; and publishing the video (402) with the annotations and trigger the secondary quality assurance review based on determining that the quality assurance confidence value is below the first confidence threshold and above the second confidence threshold.
14. The computer program product of claim 13, wherein the quality assurance confidence value is determined by a quality assurance decision engine using rule-based logic to check the annotations, wherein one or more rules are defined with a severity level to distinguish between conditions that are possible with a lower likelihood of occurrence and conditions that should not occur, and optionally wherein the quality assuranceconfidence value is determined by the quality assurance decision engine using a workflow outlier detector to check the annotations, wherein the workflow outlier detector uses one or more of: an autoencoder reconstruction accuracy, a local outlier factor, and a statistical model.
15. The computer program product of claim 14, wherein the quality assurance confidence value is determined by the quality assurance decision engine using a machinelearning model maturity value that is adjusted based on an accuracy of the machinelearning model (406) over a period of time.
Citation Information
Patent Citations
Computer vision-based surgical workflow recognition system using natural language processing techniques
US20230017202A1
Surgical workflow and activity detection based on surgical videos
US20230165643A1