Anatomy state tracking based on machine learning

CN122847743APending Publication Date: 2026-09-29INTUITIVE SURGICAL OPERATIONS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580017008.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-18
Filing Date
2025-03-14
Publication Date
2026-09-29

Smart Images

  • Figure CN122847743A_ABST
    Figure CN122847743A_ABST
Patent Text Reader

Abstract

Anatomical state tracking based on machine learning is described. A system can include one or more processors coupled with a memory to receive data of a medical procedure performed on a subject with a robotic medical system. The one or more processors can utilize one or more models trained with machine learning to identify an anatomical structure based on the data. The one or more processors can utilize the one or more models and based on the identified anatomical structure to detect a state of the anatomical structure. The one or more processors can provide an indication of a performance of the medical procedure based at least in part on the state of the anatomical structure.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing of related patent applications Pursuant to 35 USC § 119, this application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 566,531, filed March 18, 2024, the entire contents of which are incorporated herein by reference. Background Technology

[0002] Robotic medical systems may include instruments for conducting medical sessions or procedures. For example, the instrument may be used to perform surgery, treatment, or medical evaluation. Robotic medical systems may include endoscopes that capture video of medical procedures. Summary of the Invention

[0003] The technical solutions disclosed herein may include machine learning-based anatomical state tracking. The computational system may include a multi-layered machine learning model to detect and evaluate anatomical states. This model may include layers or branches built upon a baseline model or a baseline or base branch. The base branch may use deep learning for anatomical body segmentation to identify and track anatomical structures such as organs, tissues, or bones. Other branches may use the identified anatomical structures from the base branch to determine the anatomical state of the structure. For example, the computational system may implement a machine learning model on the output of a base anatomical body segmentation model to determine anatomical states such as size, color, bleeding, burns, or amputation. The computational system may implement a machine learning model to segment anatomical structures into different regions, for example, identifying areas of bruising, charring, burns, adipose tissue, muscle tissue, or connective tissue. Through this hierarchical approach, the model can be trained, stored, or executed to detect states with less computational resources (e.g., processor resources, memory resources, power resources). Because multiple layers of a model can be combined into a single model or set of models and trained together, the processing resources required to train a model can be reduced compared to multiple individual models that might be trained and executed separately. Furthermore, the trained model can be stored in less memory and consume fewer processing resources to execute. Additionally, since models or branches can be executed based on the underlying base model or branch, higher-level model branches can be lighter or smaller models that require less training data or a greater number of dimensional inputs.

[0004] One aspect of this disclosure relates to a system. The system may include one or more processors coupled to memory to receive data from medical procedures performed on a subject using a robotic medical system. The one or more processors may utilize one or more models trained with machine learning to identify anatomical structures based on the data. The one or more processors may utilize the one or more models and detect the state of the anatomical structures based on the identified anatomical structures. The one or more processors may provide indications of the performance of the medical procedure, at least in part, based on the state of the anatomical structures.

[0005] One or more processors can execute one or more models on frames of a video of a medical protocol. One or more processors can use identified anatomical structures to generate a first mask for frames representing anatomical structures. One or more processors can use detected states to generate a second mask for frames representing states.

[0006] One or more processors can execute one or more models trained using machine learning. The one or more models may include a first branch, which includes at least one first model, for recognizing anatomical structures. The one or more models may include a second branch, which includes at least one second model, for detecting the state of anatomical structures using the anatomical structures recognized by the first branch.

[0007] One or more processors can receive video frames of medical procedures and identifiers of the medical procedures from the robotic medical system. One or more processors can generate encodings of the frames. One or more processors can generate embeddings of the medical procedure identifiers. One or more processors can execute a first branch of one or more models on the frame encoding and identifier embeddings to identify anatomical structures. One or more processors can execute a second branch of one or more models on the frame encoding, identifier embeddings, and identified anatomical structures to detect the state of the anatomical structures.

[0008] One or more processors can use the identified anatomical structures to generate a first mask for frames representing the anatomical structures. One or more processors can use the detected states to generate a second mask for frames representing the states. One or more processors can determine the overlap level between the first and second masks. One or more processors can train one or more models to maximize the overlap level.

[0009] One or more processors can determine a first loss for a first branch of one or more models, the first branch being used to identify anatomical structures. One or more processors can determine a second loss for a second branch of one or more models, the second branch being used to detect the state of the anatomical structures. One or more processors can train the first and second branches in parallel using the first and second losses.

[0010] One or more processors can receive video frames of a medical procedure, hyperspectral data of the frames, and an identifier of the medical procedure from the robotic medical system. One or more processors can execute a first branch of one or more models on the frames and the identifier of the medical procedure to identify anatomical structures. One or more processors can execute a second branch of one or more models on the frames, the identifier of the medical procedure, and the identified anatomical structures to detect the state of the anatomical structures. One or more processors can execute a third branch of one or more models on the hyperspectral data and at least one of the identified anatomical structures or the state of the anatomical structures to detect the oxygen saturation of the anatomical structures.

[0011] One or more processors may use the identified anatomical structures and detected states to execute a first model of one or more models to generate a first frame including a first mask indicating the state. One or more processors may use the identified anatomical structures to execute a second model of one or more models to generate a second frame including a second mask indicating the anatomical structures.

[0012] One or more processors can utilize one or more models and, based on the identified anatomical structures, use frames from a video of a medical protocol to detect the state of the anatomical structures at points in time. One or more processors can use the state detected from the frames to track the state over time. One or more processors can use the changes in state over time to generate indications of the performance of the medical protocol.

[0013] One or more processors can use the identified anatomical structures and their states to search for videos of medical procedures to identify a subset of the videos. One or more processors can generate data to enable a graphical user interface to display indications of the subset of videos.

[0014] One or more processors can receive data during medical procedures. One or more processors can generate alerts during medical procedures using the state of anatomical structures detected by one or more models. One or more processors can generate data to display the alerts in a graphical user interface.

[0015] One or more processors can generate a video of a medical procedure, including a mask, to indicate the state of anatomical structures identified and detected by one or more models. One or more processors can generate data to enable a graphical user interface to display the masked video.

[0016] At least one aspect of this disclosure relates to a method. The method may include receiving data from a medical procedure performed on a subject using a robotic medical system by one or more processors coupled to memory. The method may include identifying anatomical structures based on the data by one or more processors using one or more models trained with machine learning. The method may include detecting the state of anatomical structures by one or more processors using one or more models and based on the identified anatomical structures. The method may include providing indications of the performance of the medical procedure by one or more processors, at least in part, based on the state of the anatomical structures.

[0017] The method may include one or more models trained with machine learning, executed by one or more processors. The one or more models may include a first branch comprising at least one first model for identifying anatomical structures. The one or more models may include a second branch comprising at least one second model for detecting the state of anatomical structures using the anatomical structures identified by the first branch.

[0018] The method may include receiving video frames of a medical procedure and identifiers of the medical procedure from a robotic medical system by one or more processors. The method may include generating encodings of the frames by one or more processors. The method may include generating embeddings of the identifiers of the medical procedure by one or more processors. The method may include executing a first branch of one or more models by one or more processors on the encoded frames and the embedded identifiers to identify anatomical structures. The method may include executing a second branch of one or more models by one or more processors on the encoded frames, the embedded identifiers, and the identified anatomical structures to detect the state of the anatomical structures.

[0019] The method may include generating a first mask for frames representing anatomical structures using identified anatomical structures by one or more processors. The method may include generating a second mask for frames representing states using detected states by one or more processors. The method may include determining the overlap level between the first and second masks by one or more processors. The method may include training one or more models by one or more processors to maximize the overlap level.

[0020] The method may include a first loss determined by one or more processors for a first branch of one or more models, the first branch being used to identify anatomical structures. The method may also include a second loss determined by one or more processors for a second branch of one or more models, the second branch being used to detect the state of the anatomical structures. Furthermore, the method may include training the first and second branches in parallel using the first and second losses by one or more processors.

[0021] At least one aspect of this disclosure relates to a non-transient computer-readable medium. This medium can store processor-executable instructions that, when executed by one or more processors, cause one or more processors to receive data from a medical procedure performed on a subject using a robotic medical system. The instructions can cause one or more processors to identify anatomical structures based on the data using one or more models trained with machine learning. The instructions can cause one or more processors to detect the state of anatomical structures based on the identified anatomical structures using one or more models. The instructions can cause one or more processors to provide instructions on the performance of the medical procedure, at least in part, based on the state of the anatomical structures.

[0022] Instructions can cause one or more processors to receive video frames of medical procedures and identifiers of the medical procedures from the robotic medical system. Instructions can cause one or more processors to generate encodings of the frames. Instructions can cause one or more processors to generate embeddings of the medical procedure identifiers. Instructions can cause one or more processors to execute a first branch of one or more models based on the frame encodings and identifier embeddings to identify anatomical structures. Instructions can cause one or more processors to execute a second branch of one or more models based on the frame encodings, identifier embeddings, and identified anatomical structures to detect the state of the anatomical structures.

[0023] The instructions can cause one or more processors to generate a first mask for frames representing the anatomical structures using the identified structures. The instructions can also cause one or more processors to generate a second mask for frames representing the states using the detected states. The instructions can further cause one or more processors to determine the overlap level between the first and second masks. Finally, the instructions can cause one or more processors to train one or more models to maximize the overlap level.

[0024] These and other aspects and implementations are discussed in detail below. The above information and the following detailed description include illustrative examples of various aspects and implementations, and provide an overview or framework for understanding the nature and characteristics of the claimed aspects and implementations. The accompanying drawings provide illustration and further understanding of the various aspects and implementations, and are incorporated in and constitute a part of this specification. The above information and the following detailed description and accompanying drawings include illustrative examples and should not be considered limiting. Attached Figure Description

[0025] The accompanying drawings are not to scale. The same reference numerals and names in the various drawings indicate the same elements. For clarity, not every component can be labeled in every drawing. In the drawings: Figure 1 An example computational system for identifying anatomical structures and detecting their state is described.

[0026] Figure 2An example model is depicted for identifying anatomical structures and detecting their state.

[0027] Figure 3 An example method for identifying anatomical structures and detecting their state is described.

[0028] Figure 4 An example method for training a model using combined branch loss is described.

[0029] Figure 5 An example computing architecture for a computing system is described. Detailed Implementation

[0030] The various concepts and implementations related to methods, apparatuses, and systems for machine learning-based anatomical state tracking are described in more detail below. The various concepts introduced above and discussed in more detail below can be implemented in any of a variety of ways.

[0031] This disclosure generally relates to a machine learning system for identifying anatomical structures and their states in frames of video of medical procedures performed by a robotic medical system. Performance metrics can track the performance of medical procedures such as surgery. Performance metrics can include manipulator travel distance, energy consumption, number of clutches, etc. During surgery, surgeons can quickly assess the state of a patient's organs and take action to ensure effective and efficient treatment of the condition. The actions taken by the surgeon can be linked to performance metrics or can be used to generate performance metrics. However, simply using performance metrics to track surgeon performance may only quantify the actions taken by the surgeon and may not provide contextual insights into why the surgeon took the actions or the consequences of those actions on the patient. Reliance on performance metrics can confuse surgeons, as they may not be able to articulate why performance metrics decrease or increase during or after surgery. For example, a performance metric for manipulator travel distance may be low due to an enlarged or swollen liver (which makes navigation of the robotic manipulator more difficult during surgery). Similarly, performance metrics such as energy usage may increase due to significant bleeding or thickened connective tissue (which requires additional surgical work). However, because the context of the anatomical structures being manipulated by the surgeon is not factored into the calculation of performance metrics, surgeons may not clearly understand why their performance metrics increase or decrease as the manipulation progresses. Without knowledge of the state of the anatomical structures being manipulated, surgeons may find it difficult to quantify their performance during surgical procedures.

[0032] To assess anatomical status during and after surgery, surgeons can manually record and report the status of various organs or other anatomical structures. Additionally, surgeons may collect pathological or biopsy samples or perform additional imaging during surgery to determine the status of anatomical structures. Sample collection or additional imaging may result in longer and more invasive or traumatic surgeries, potentially increasing the surgeon's workload, prolonging surgical duration, and increasing the power consumption of the robotic medical system. The extended duration of surgeries may lead to further wear and tear on the medical robotic system, and even make routine surgeries take longer, resulting in faster wear and tear on the medical robotic system and requiring more maintenance. Furthermore, the accuracy of the surgeon's reports may depend on the accurate information provided by the surgeon. If the surgeon makes an error and incorrectly indicates which area the sample was taken from, the reported status of various anatomical structures may be inaccurate. Moreover, if the surgeon reviews the video after surgery, the quantification of the anatomical status provided by the surgeon may be subjective, and different surgeons may draw different conclusions about the anatomical structures' status based on the same video.

[0033] Models can be trained to detect the state of anatomical structures from input frames. However, since anatomical states can appear differently for different anatomical structures, model training may require large training datasets of various types of anatomical structures, which may not be available. Training a single model to detect the state of anatomical structures from frames with large datasets can lead to increased training time, increased model size, increased memory usage, or increased processing resource usage. This can make model training difficult and resource-intensive (e.g., consuming significant processing, memory, or storage resources), and the large model size can make real-time or intraoperative implementation difficult. For postoperative implementation, the model may consume significant processing, memory, or power resources.

[0034] Therefore, it is necessary to effectively assess the anatomical state of structures in an objective and quantifiable manner. Objective assessment of anatomical state allows surgeons to understand the success of surgical procedures and the final anatomical state, thereby enabling them to predict whether a patient will heal successfully post-surgery. Objectively determining anatomical state can help surgeons identify the difficulty of a case and respond to sudden changes or indications of damage to anatomical structures, such as bleeding, burns, or tears, which may require surgeon attention or a change in approach.

[0035] To address these and other technical issues, the technical solutions of this disclosure may include multi-layer machine learning models for detecting and evaluating anatomical states. The computational system can implement multi-layer machine learning models that incorporate multiple data sources to provide evaluations of different categories of anatomical states in an efficient manner. Each evaluation can be a separate branch built upon a baseline model or a baseline branch. The base branch can use machine learning or deep learning to segment anatomical bodies to identify and track anatomical structures such as organs, tissues, or bones. Other branches can use the identified anatomical structures from the base branch to determine the anatomical state or part of a structure. For example, the computational system can implement machine learning models or deterministic computations on the output of the base anatomical body segmentation model to determine anatomical states such as size, color, burns, bleeding, or amputation. The computational system can implement machine learning models to segment anatomical structures into different parts, such as identifying areas of bruising, charring, burns, adipose tissue, muscle tissue, connective tissue, etc. Furthermore, the computational system can implement machine learning models that incorporate additional data sources; for example, for states such as oxygen (O2) saturation, a model incorporating hyperspectral image data can be implemented. Other branches can be used to identify surgical equipment or instruments, such as sutures, grafts, or clips.

[0036] By employing a hierarchical approach, the model can be trained, stored, or executed to detect states with less computational resources (e.g., processor, memory, power). Because multiple layers of the model can be combined into a single model or set of models and trained together, the processing resources required to train the model can be reduced compared to multiple individual models that might be trained and executed separately. Furthermore, the trained model can be stored in less memory and consume less processing resources to execute. Additionally, since models or branches can be executed based on the underlying base model or branch, higher-level model branches can be lighter or smaller models that require less training data or a greater number of dimensional inputs. At least one piece of data input to the model system can be medical robotic system data, such as medical procedure types, medical procedure stages, or medical procedure steps. Medical robotic system data can provide latent states that can be centralized and accelerate the training of the model system. Moreover, using a hierarchical model approach, the model system can be flexible and adaptable to new data modalities that become available to the computational system. For example, if hyperspectral image data becomes available to the computational system, additional branches or models can be implemented on top of the basic anatomical segmentation model. In addition to anatomical body detection using the basic anatomical body segmentation model, additional models can also consume hyperspectral image data to detect states, such as tissue oxygen saturation.

[0037] Using a multi-layered modeling approach, this system can accurately describe anatomical structures and states within video footage. These states can be tracked over time to quantify changes in the state of anatomical structures. These states can be outputs that can be utilized by other systems or applications, such as surgical quality quantification and assessment systems. Furthermore, by detecting the state of anatomical structures, video review software can be improved by allowing surgeons to quickly navigate to segments of the procedural video where the state of the anatomical structures requiring review or where the state has drastically changed (such as transection or hemostasis). Additionally, historical states and procedural outcomes can be tracked and used to make postoperative recommendations. Similarly, computational systems can use this state to make intraoperative recommendations based on real-time state assessments and comparisons.

[0038] Now for reference Figure 1 Among other things, an example system 100 is shown, which includes a computing system 110 for identifying anatomical structures and detecting the state of the anatomical structures. System 100 may include at least one computing system 110. The computing system 110 may be a data processing system, computing system, computer system, computer, desktop computer, laptop computer, tablet computer, control system, console system, embedded system, cloud computing system, server system, or any other type of computing system. The computing system 110 may be a local system or a non-local system. The computing system 110 may be a hybrid system in which some components of the computing system 110 are located locally and some components of the computing system 110 are located non-locally.

[0039] System 100 may include at least one robotic medical system 105. Robotic medical system 105 may be a robotic system, device, or component that includes at least one instrument. The instrument may be or include a tip or end effector. The tip or end effector may be mounted or attached to the instrument. The tip may be removable or may be a permanent component of the instrument or robotic medical system 105. For example, the tip may be a scalpel, scissors, single-pole curved scissors (MCS), cauterization hook tip, cauterization spatula tip, needle actuator, forceps, tooth retractor, drill, or clamp applicator. The instrument may be or include a robotic arm, robotic appendage, robotic snake arm, or any other motor-controlled component that can be articulated by the robotic medical system. The instrument may include at least one actuator, such as a motor, servo mechanism, or other actuator device. The instrument may be operated by a motor, servo mechanism, actuator, or other device to perform medical procedures. Robotic medical system 105 may perform medical sessions or medical procedures. For example, robotic medical system 105 may articulate an instrument to utilize the instrument for surgical procedures, treatments, or medical evaluations. Medical procedures can be performed on subjects (e.g., humans, adults, children, or animals). Medical practitioners (such as surgeons, technicians, nurses, or other operators) can provide input via user equipment or input devices (e.g., joysticks, buttons, touchpads, keyboards, steering mechanisms, etc.) to manipulate instruments to perform medical procedures. In some embodiments, the robotic medical system 105 may include an endoscope. The endoscope may be an instrument manipulated by a medical practitioner and controlled via motors, servo mechanisms, or other input devices.

[0040] The computing system 110 can receive data from medical procedures performed on a subject using the robotic medical system 105. The computing system 110 can receive at least one image frame 115 from the robotic medical system 105. The computing system 110 can receive system data 120 from the robotic medical system 105. The image frame 115 can be an image of at least one anatomical structure of a patient (e.g., a person, animal, or biological material) captured during the medical procedure. The system data 120 can provide an indication, name, or identifier specifying the type of medical procedure being performed or having been performed. The system data 120 can be linked to or labeled to the image frame 115 to indicate that the image frame 115 was captured for a specific type of medical procedure. The system data 120 can indicate that the medical procedure is a polyp removal, cataract surgery, cesarean section, appendectomy, or any other type of medical or surgical procedure.

[0041] The robotic medical system 105 can generate, store, or produce at least one image frame 115. At least one endoscope of the robotic medical system 105 can capture at least one image frame 115 and provide the image frame 115 to the computing system 110. The image frame 115 may be a portion of video captured by the endoscope of the robotic medical system 105. The image frame 115 may include an image, picture, or pixel representing at least one anatomical structure of a subject or patient. The robotic medical system 105 can transmit the image frame 115 to the computing system 110. The computing system 110 can receive the image frame 115 from the robotic medical system 105. The robotic medical system 105 can use system data 120 to tag each image frame 115 or a group of image frames 115 to indicate for what type of medical procedure the image frame 115 was captured.

[0042] During a medical procedure, image frame 115 or system data 120 may be streamed to computing system 110. For example, when robotic medical system 105 performs a medical procedure, it may stream image frame 115 or system data 120 to computing system 110. For example, robotic medical system 105 may transmit image frame 115 or system data 120 in real time or as image frame 115 or system data 120 is generated, captured, or stored. Robotic medical system 105 may provide image frame 115 or system data 120 intraoperatively (e.g., during a medical procedure).

[0043] The computing system 110 may implement a machine learning-based model 125. The computing system 110 may execute one or more models 125 trained via machine learning to detect anatomical structures 145. Anatomical structures 145 may be organs (e.g., brain, liver, pancreas, stomach, intestine), tissues (e.g., muscle tissue, connective tissue, nervous tissue), or bones (e.g., ribs, tibia, or skull). In some embodiments, the computing system 110 may simultaneously identify multiple different anatomical structures 145 within a single image frame 115. In some embodiments, even when multiple anatomical structures are included within a single image frame, the computing system 110 may execute model 125 to identify a single anatomical structure of interest, such as the anatomical structure being operated on by the robotic medical system 105, an anatomical structure located at the center of the image frame 115, or an anatomical structure that is the focal point of the image frame 115.

[0044] The computing system 110 can execute one or more models 125 trained through machine learning to detect the state of an anatomical structure 145 or to detect the location of an anatomical structure. The location of the anatomical structure 145 can be a fragment, segment, or portion of an image associated with a specific state or condition. The model 125 can be a multi-level segmentation model 125. For example, the computing system 110 can execute a first branch 135 including at least one first model to identify the anatomical structure 145. The computing system 110 can execute a second branch 140 including at least one second model to identify the anatomical structure 145. The second branch 140 can use the anatomical structure 145 identified by the first branch 135 to detect the state of the anatomical structure 145.

[0045] Model 125 may be or include at least one neural network, such as a transformer network, embedding network, encoding network, convolutional neural network, feedforward neural network, or any other neural network topology. Model 125 may include multiple branches. Model 125 may include a first branch 135 and a second branch 140. Branches 135 and 140 may each include at least one model or multiple models. The first branch 135 may be a first part, a first model, or a first set of models of model 125. The second branch 140 may be a second part, a second model, or a second set of models of model 125. The computing system 110 can identify anatomical structure 145 or the state of anatomical structure 145 by executing model 125 based on data. The data may be information, images, videos, data packets, time-series information, datasets, or data structures. For example, the computing system 110 can identify anatomical structure 145 and the state of anatomical structure based on image frame 115 or system data 120.

[0046] The computing system 110 may include a machine learning engine 130. The machine learning engine 130 may train a model 125. The machine learning engine 130 may train model 125 to learn to identify anatomical bodies of interest. Model 125 may include at least one base branch or base level 135 to identify anatomical bodies. Model 125 may further segment the state of the anatomical structure into parts or states (e.g., burns, bleeding, or other treated sites within the anatomical body). Model 125 may include at least one part-level or state branch 140 to segment the state of the anatomical structure. Anatomical structures 145 detected by branch 135 may be encoded as features used as context to help the second branch 140 segment the parts or states of the anatomical structure 145. This can help improve the accuracy of state or part detection by the second branch 140 and reduce training time, processing or memory resources used for training, or the size of training data 150. Branches 135 and 140 may be trained together by the machine learning engine 130. Because multiple branches of model 125 can be trained together, this reduces the total amount of training data 150 and processing resources (e.g., processors, memory, or power) required by the machine learning engine 130 for segmenting anatomical structures or states. Furthermore, the computational system 110 can use system data 120, such as procedure type or step labels, as input to model 125. System data 120 can be a time series, which can be encoded as embedded features associated with or linked to visual features of image frame 115. The embedded features generated from system data 120 can provide latent labels to narrow down the scope and accelerate learning to detect states.

[0047] The computing system 110 can utilize at least one model 125 trained with machine learning to identify anatomical structures 145. The computing system 110 can store at least one model 125. The computing system 110 can retrieve a model 125 from storage. The computing system 110 can execute the retrieved model 125. The computing system 110 can execute at least one model 125 on image frames 115 or system data 120 to identify anatomical structures 145. The computing system 110 can execute one or more models 125 on frames 115 of a video of a medical procedure.

[0048] The computing system 110 can use image frame 115 and system data 120 to execute a first branch 135 to determine anatomical structure 145. The computing system 110 can execute model 125 to encode or embed image frame 115 or system data 120. The computing system 110 can execute the first branch 135 on the encoding of frame 115 and the embedding of system data 120 to identify anatomical structure 145. Executing the first branch 135 can generate anatomical structure detection 145. The computing system 110 can generate a first mask 175 for image frame 115 for anatomical structure 145. The first branch 135 can output the first mask 175. The first mask 175 can identify pixels in image frame 115 that correspond to anatomical structure 145. The first mask 175 can identify pixels in image frame 115 that do not correspond to anatomical structure 145. For example, the first mask 175 may include a label, identifier, or number indicating that a pixel is associated with anatomical structure 145, and a second label, identifier, or number indicating that a pixel is not associated with anatomical structure 145. The first branch 135 can generate the first mask 175 based on the detection or recognition of the anatomical structure 145.

[0049] The computational system 110 can detect the state or location of the identified anatomical structure 145 based on the identified anatomical structure 145. The computational system 110 can execute model 125 based on the identified anatomical structure 145. For example, the state or location detected by model 125 can be specific to the anatomical structure 145. For example, if the state or location is burn tissue, burn tissue may look different on the skin than on an internal organ such as the liver. In this respect, the state or location detected by model 125 can be at least partially based on the anatomical structure 145. The computational system 110 can execute a second branch 140 on the identified anatomical structure 145 to identify the location or state of the anatomical structure 145. For example, the identified anatomical structure 145 can be the input to the second branch 140. The computing system 110 can perform a second branch 140 on frame 115 (e.g., the encoding of frame 115), system data 120 (e.g., the embedding of system data 120), and identified anatomical structure 145 (e.g., the hidden or latent state of model 125 representing anatomical structure 145) to identify the state or location of anatomical structure 145.

[0050] The state of the anatomical structure 145 determined by the second branch 140 can be a native state. A native state can be the state of the anatomical structure 145 prior to a medical procedure performed on or near the anatomical structure 145. A native state may be inflammatory staining, perfusion, or deformity. The state of the anatomical structure 145 determined by the second branch 140 can be an altered state. An altered state of the anatomical structure 145 can be the state of the anatomical structure after a medical procedure performed on or near the anatomical structure 145. An altered state can be removal, detachment, exposure, stretching, or deformation. The state of the anatomical structure 145 determined by the second branch 140 can be a reconstructed state. For example, a reconstructed state can be closure, transplantation, or anastomosis. The state of the anatomical structure 145 determined by the second branch 140 can be a damaged state. A damaged state can include tearing, burns, bruising, or cutting. The state of the anatomical structure 145 determined by the second branch 140 can be a final state. For example, a final state can be the final output of a medical procedure that includes the anatomical structure 145, or the state of the last frame 115. The final state can provide an indication of the condition of the anatomy 145 or the patient at the end or completion of the medical procedure. The state can be a Boolean state, for example, indicating the presence or absence of a symptom or condition. The state can be a variable, or it can be the degree of a symptom or condition on a certain scale or within a certain range.

[0051] In some implementations, states can be decomposed or combined into other states. Furthermore, as new branches or new data become available to the computational system 110, the model 125 can be updated to generate new or different states, thereby providing more detail and definition of the anatomy 145. The nature of some states may be complex, with some states being interconnected and interdependent or alternatively independent. To account for this in the output, each category and subcategory of state can be measured independently and quantified as a percentage of its relation to the base branch 135.

[0052] The computing system 110 can use the detected state or location to generate a second mask 180 for the image frame 115. A second branch 140 can generate the second mask 180. The second mask 180 can identify the location or state of the anatomical structure 145. The second mask 180 can identify pixels in the image frame 115 that correspond to the location or state of the anatomical structure 145. The second mask 180 can identify pixels in the image frame 115 that do not correspond to the location or state of the anatomical structure 145. For example, the second mask 180 may include a label, identifier, or number indicating that a pixel is associated with the state of the anatomical structure 145, and a second label, identifier, or number indicating that a pixel is not associated with the state of the anatomical structure 145. The second branch 140 can generate the second mask 180 based on the detection or identification of the anatomical structure 145 or the identified location or state of the anatomical structure 145.

[0053] In some implementations, the computing system 110 can determine states with various levels of granularity. In some implementations, a user can provide input via client device 199 to identify the granularity level of the state determined by the second branch 140. The granularity level can be a Boolean indication, such as true or false. The granularity level can be a range of values, such as a percentage value or a value between 0 and 10. Model 125 can adjust the granularity of the state based on user input. For example, for a Boolean granularity level, the state could indicate whether the anatomical structure 145 is burned or not. The level can be a sequential level, such as the degree of bruising.

[0054] In some implementations, model 125 may include two or more branches. In some implementations, model 125 may be adjustable, and new branches may be added to model 125 over time. Branches may operate in a hierarchy, where the determination, detection, or identification of one branch is used by another branch to produce more determinations, detections, or identifications. For example, a base branch (e.g., first branch 135) may detect anatomical structure 145, while higher-level branches operating on or on the base branch may detect information about anatomical structure 145, or information about the state or location of anatomical structure 145. For example, as new data or a new data stream becomes available to computing system 110, computing system 110 may load or train new branches to perform on the new data stream.

[0055] For example, if computing system 110 is coupled to robotic medical system 105 that generates hyperspectral data, or if a new hyperspectral data sensor is installed in existing robotic medical system 105, the new hyperspectral data can be provided to a new branch of model 125. The architecture of computing system 110 and model 125 can support new products, robotic medical system 105, or updates to robotic medical system 105. For example, hyperspectral capabilities can be added to robotic medical system 105, and computing system 110 can be updated to utilize hyperspectral capabilities.

[0056] The computing system 110 can be updated via software updates or by updating model 125, adding new branches to allow for real-time or near-real-time assessment and tracking using new data sources during surgical procedures or medical protocols, and to allow for postoperative review of new imaging techniques through artificial intelligence suite products. As new or future imaging techniques (such as perfusion imaging, white light firefly imaging, and other advanced imaging protocols) are implemented in the robotic medical system 105, model 125 can be updated to operate on the new data, allowing for the analysis of more anatomical states, or allowing for alternative methods of analyzing anatomical states. The computing system 110 can provide identified anatomical structures 145 and their status during surgical procedures and can be applied to intraoperative or postoperative surgical feeds. This provides a sustainable system that can evolve as available imaging technologies change.

[0057] For example, computing system 110 may receive hyperspectral data of image frame 115. The hyperspectral data may correspond to image frame 115. Computing system 110 may execute a first branch 135 of one or more models 125 on image frame 115 and system data 120 including medical protocol identifiers to identify anatomical structures 145. Computing system 110 may execute a second branch 140 on image frame 115, system data 120 including medical protocol identifiers, and the identified anatomical structures 145 to detect the location or state of the anatomical structures. Computing system 110 may execute a third branch of one or more models 125 based on the hyperspectral data and at least one of the identified anatomical structures 145 or the state or location of the anatomical structures 145. The third branch may detect or identify the oxygen saturation of the anatomical structures 145. The third branch may generate a mask indicating the oxygen saturation level in the anatomical structures 145.

[0058] Model 125 may include additional branches directly built on the base branch or the first branch 135. The model may include at least one additional branch built on the second branch 140. These additional branches may operate on additional input data, such as hyperspectral data. Furthermore, model 125 may include additional branches for post-processing. For example, the post-processing branch may determine the color gradient of the segmented region identified as anatomical structure 145 across image frame 115. Model 125 may include additional branches for direct processing of segmentation masks (e.g., a first mask 175 or a second mask 180). These direct processing branches may detect whether the anatomical body has been severed or cut based on the segmentation of the anatomical structure 145.

[0059] The computing system 110 may include at least one machine learning engine 130. Machine learning engine 130 can train model 125. For example, machine learning engine 130 can train model 125 using training data 150. Machine learning engine 130 can train model 125 based on training data 150. Machine learning engine 130 may include at least one of losses 155-170. Machine learning engine 130 can generate, compute, or determine a first branch loss 155 indicating the loss of the first branch 135. Machine learning engine 130 can train the first branch 135 to minimize, reduce, or lower the first branch loss 155. Machine learning engine 130 can generate, compute, or determine a second branch loss 160 indicating the loss of the second branch 140. Machine learning engine 130 can train the second branch 140 to minimize, reduce, or lower the second branch loss 160.

[0060] Machine learning engine 130 can determine the overlap level between first mask 175 and second mask 180. Machine learning engine 130 can train one or more models 125 to maximize the overlap level. For example, computation system 110 can determine a combination loss 170. Machine learning engine 130 can utilize the combination loss 170 to improve model performance to maintain consistency in segmentation results. The combination loss 170 can be a loss that considers the loss of both first branch 135 and second branch 140 (e.g., multiple branches together). Machine learning engine 130 can determine the combination loss 170 based on the overlap or intersection between regions of image frames 115 identified as anatomical structures 145 and regions of image frames 115 identified as sites or states of anatomical structures 145. For example, computation system 110 can indicate the overlap level between first mask 175 and second mask 180. Machine learning engine 130 can minimize or reduce the combination loss 170 to maximize or increase the overlap between first mask 175 and second mask 180. For example, since the second mask 180 can represent the location or state of the anatomical structure represented by the first mask 175, the second mask 180 should be completely included or contained within the first mask 175. More specifically, the first mask 175 should completely enclose the second mask 180.

[0061] Machine learning engine 130 can train multiple models, branches, or levels of model 125 in parallel, simultaneously, or concurrently. For example, machine learning engine 130 can train a first branch 135 and a second branch 140 simultaneously. Machine learning engine 130 can train the first branch 135 and the second branch 140 in parallel using at least one of a first branch loss 155, a second branch loss 160, and a combined loss 170. For example, machine learning engine 130 can train the first branch 135 and the second branch 140 in parallel using a first loss 155 and a second loss 160. For example, machine learning engine 130 can train the first branch 135 and the second branch 140 in parallel using a first loss 155, a second loss 160, and a combined loss 170. For example, training can generate, identify, or determine the parameters or configuration of model 125 that minimize or reduce losses 155-170. Machine learning engine 130 can perform backpropagation to train model 125. For example, machine learning engine 130 can execute machine learning algorithms, such as gradient descent with a loss of 155–170 or stochastic gradient descent with a loss of 155–170 relative to the parameters of model 125. The machine learning algorithms can implement second-order gradient descent, Newton's method, conjugate gradient method, quasi-Newton method, or Levenberg-Marquardt algorithm to train model 125.

[0062] The machine learning engine 130 can train a model 125 using training data 150. The training data 150 may include multiple image frames 115, corresponding system data 120, and corresponding masks. For example, for each image frame 115, the training data 150 may include system data 120 indicating the type of medical procedure for which the captured image frame 115 is targeted. Furthermore, for each image frame 115, the training data 150 may identify a first mask 175 that identifies the anatomical structures 145 depicted in the image frame 115. Additionally, for each image frame 115, the training data 150 may identify a second mask 180 that identifies the location or state of the anatomical structures 145 depicted in the image frame 115.

[0063] The computing system 110 may include at least one interface manager 190. The interface manager 190 may generate at least one graphical user interface 197. The interface manager 190 may generate data to cause at least one user or client device 199 to display the graphical user interface 197. The graphical user interface 197 may be displayed on at least one display of the client device 199. The client device 199 may be a smartphone, laptop, desktop computer, console, tablet, or any other computing system or device. The client device 199 may be integrated with the computing system 110, or it may be a separate device or system. The client device 199 may be a mobile device or a fixed device. The client device 199 may be communicatively coupled to the computing system 110. The client device 199 may communicate with the computing system 110 via at least one network, such as the Internet, a local area network (LAN), a wide area network (WAN), or a Wi-Fi network.

[0064] Interface manager 190 or computing system 110 can generate or provide indications of the performance of robotic medical system 105 or medical procedures. For example, interface manager 190 can provide indications of the performance of the medical procedure based at least in part on the state of anatomical structure 145. For example, the performance of the medical procedure can be an indication of a state or a characteristic of a state determined by model 125. For example, if the state is bleeding, interface manager 190 can make graphical user interface 197 display the amount or level of bleeding in anatomical structure 145.

[0065] Interface manager 190 can generate at least one metric indicating the performance of a medical protocol. Interface manager 190 can generate the metric at least in part based on masked frame 115, first mask 175, second mask 180, detected anatomical structure 145, or the state of detected anatomical structure 145. This metric can quantify the state of anatomical structure 145. The metric can be an objective measurement determined by tracking the state over time and determining changes in the state. The metric can be determined by analyzing changes in patterns, such as by optical flow mapping, e.g., assigning directionality and amplitude to pixel changes. The metric can be determined based on pixel area ratios, e.g., the ratio of the area associated with the state to the entire visible area. The metric can be paired with kinematic data, such as energy data from an energy application. The metric can allow the incorporation of additional image data, such as the hyperspectral data output of a robotic medical system 105.

[0066] Metrics or performance indicators generated by the computing system 110 can be used for research or evaluation. For example, metrics or performance indicators can be used to improve clinical outcomes and can allow for improved outcome assessment, or enable new healthcare practitioner bases or groups to use metrics or performance indicators to care for or treat patients, such as pathologists or postoperative care teams. Metrics or performance indicators can provide information about surgical skills; for example, metrics can be linked to the results or consequences of medical procedures. Metrics linked to clinical outcomes can reflect a surgeon's skills in handling different situations or types of surgery. Metrics can anatomically illustrate how different surgical skills affect anatomical states. Recorded outputs can be organized to be added to postoperative reports to interpret the patient's condition after surgery.

[0067] In some implementations, the computing system 110 can track the state of the anatomical structure 145 determined by the model 125. For example, the computing system 110 can record the state of the anatomical structure 145 at multiple time points (e.g., for multiple different image frames 115). For example, the computing system 110 can utilize one or more models 125 and, based on the identified anatomical structure 145, detect the state of the anatomical structure at multiple time points using multiple frames 115 of the video of the medical protocol. The computing system 110 can track the state throughout the medical protocol. For example, the computing system 110 can generate a trend or a series of values ​​for the state of the anatomical structure 145 over time, such as a first level of state at a first time point, a second level of state at a second time point, and a third level of state at a third time point. For example, a change in state can indicate that bleeding has begun, increased to a certain amount, and then subsided. This can provide the computing system or the user with an understanding of how the state originates, changes, alleviates, resolves, or remains unresolved during the medical protocol. The interface manager 190 can use the trend to generate indications of the medical protocol's performance.

[0068] For example, a decreasing trend in the bleeding level of anatomical structure 145 over time can indicate that anatomical structure 145 is healing. A decreasing trend in the bleeding level of anatomical structure 145 over time can indicate that anatomical structure 145 is deteriorating. The calculation system 110 can generate performance indicators of medical procedures based on changes in the state of anatomical structure 145 over time. The interface manager 190 can display an indication of the rate of change of the state, or can indicate whether the state is increasing or decreasing within a time window. The calculation system 110 can generate a measure of the trend in the rate of change of the state of the anatomical structure. The rate of change can refer to the amount of change, deterioration, degradation, or other types of change of the anatomical structure over a period of time. For example, a positive rate of change of the state can correspond to the healing of the anatomical structure, while a negative rate of change of the state can correspond to the deterioration, damage, injury, or degradation of the anatomical structure.

[0069] In some implementations, model 125 may be executed in real time or intraoperatively to determine the state of anatomical structure 145 during medical procedures. Interface manager 190 may generate alerts, alarms, or messages in response to detecting a specific state or a state associated with a level that meets a threshold (e.g., a level greater than or less than a threshold). Interface manager 190 may generate data to cause graphical user interface 197 to display alerts, alarms, or messages. Alerts may include indications of anatomical structure 145, indications of the state of anatomical structure 145, indications of the level of the state, or indications of the rate of change of the level of the state.

[0070] The computing system 110 (e.g., via interface manager 190) can prevent damage to anatomical structures by providing alerts, or otherwise mitigate or reduce the amount or likelihood of undesirable changes in the state of the anatomical structure. For example, the computing system 110 can improve safety by providing alerts based on tracking the state of the anatomical structure (e.g., tracking the amount of bleeding or burns over time), thereby preventing damage to the anatomical structure based on analyzing changes in state or on a sequence or trend of the state of the anatomical structure. The computing system 110 can automatically provide alerts based on the amount of damage (such as bleeding or burns exceeding a threshold amount (e.g., percentage, absolute value, ratio relative to the size on the anatomical structure)). In some cases, when a threshold for the burn state of the anatomical structure (e.g., the amount of burn) is exceeded, the computing system 110 can automatically reduce the amount of energy supplied by the instruments or tools of the robotic medical system 105 to reduce, mitigate, or stop the burn to the anatomical structure, thereby slowing or stopping the rate of change of the anatomical structure. In some cases, the computing system 110 detects slow bleeding that has not stopped and provides an alarm or automatically controls the instruments of the robotic medical system 105 to reduce the amount of bleeding or provides feedback or indication that the bleeding has not stopped. Slow bleeding may indicate incomplete anastomosis, and the computing system 110 may automatically perform or prompt actions to complete anastomosis or cross-connection between the subject's anatomical structures or components.

[0071] Interface manager 190 can generate video based on masked frames 115. For example, interface manager 190 can generate a video of a medical procedure. The video can include various masks 175 or 180 to indicate the anatomical structure 145 recognized by one or more models 125 and the state of the anatomical structure 145 detected by one or more models 125. For example, masked frames 115 can include a first mask 175 and a second mask 180. Interface manager 190 can combine multiple masked frames 115 of a medical procedure into a single video of the medical procedure. Interface manager 190 can store the video of the masked frames 115 as a video file (e.g., WMV, AV1, FLV, or any other type of video file). Interface manager 190 can generate data to display the video, segmenting the anatomical structure 145 and its state, in a graphical user interface 197. Interface manager 190 can overlay state information onto the video. For example, the interface manager 190 can overlay an indication of the damage condition onto the anatomical structure. This indication could be a symbol, color, measurement, numerical value, or other indication corresponding to the amount of damage (e.g., burn amount). The user, via the client device 199, can play, stop, pause, view, scan forward, scan backward, or move to a specific state in the video.

[0072] In some implementations, the interface manager 190 may use identified anatomical structures 145 or identified anatomical states to navigate a medical protocol video. For example, the interface manager 190 may provide state-based navigation, where a user can navigate the medical protocol video based on states such as state importance. A user may select a specific state or request to view the most important state or the most important anatomical structure via a client device 199, and the interface manager 190 may navigate the video to the frame associated with the selected state, the most important state, or the most important anatomical structure. For example, the graphical user interface 197 may include menus, optional elements, sliders, or drop-down menus that allow a user to select a specific type of anatomical structure 145 or a specific state of anatomical structure 145. The interface manager 190 may display a portion of the video or navigate to a specific image frame 115 of the video based on the user's selection. For example, the interface manager 190 may navigate to an image frame 115 that includes the selected anatomical structure 145 or a selected anatomical structure 145 containing a specific selected state.

[0073] In some implementations, the interface manager 190 may generate at least one timeline for a video of a medical protocol. The timeline may be interactive and allow a user to view or navigate to specific segments or frames of the video. The timeline may identify anatomical structures. For example, the timeline may indicate the names of anatomical structures displayed in different parts of the video, or it may display colors to identify specific anatomical structures. The timeline may indicate the start time when an anatomical structure becomes visible in the video, and the end time when the anatomical structure becomes no longer visible. When an anatomical structure is visible, the timeline may indicate multiple segments, for example, a specific anatomical structure being visible during two or more different time segments of the video. The timeline may include similar indications of states, such as when a state begins, ends, or its level changes. For example, the timeline may indicate when anatomical structure 145 begins to bleed, the level of bleeding changes over time, and when the bleeding stops. The interface manager 190 may include a first timeline and another timeline, the first timeline showing when different anatomical structures are visible in the video, and the other timeline showing when the state of an anatomical structure begins or stops. The interface manager 190 may populate the timeline with major transitions in state. For example, a timeline can be marked with major state transitions, when a burn occurs, when anatomical structures are cut, when anatomical structures are sutured, when bleeding begins in the anatomical structures, and when bleeding stops. A timeline can include elements such as rectangular blocks, arrows, or segments spanning a length of time, which include different color gradients. Color gradients can indicate states or levels to provide a visual representation of how states change over time.

[0074] The computing system 110 can search video collections, sets, or databases of medical procedures using anatomical structures 145 or the segmentation of anatomical structure states. The computing system 110 can navigate through recorded surgical or medical procedure videos by using state changes of key anatomical structures in the procedures as targets or filters to find certain state values. The computing system 110 can implement software or artificial intelligence programs to group data for study or viewing based on the determination of model 125.

[0075] For example, the database can store videos segmented based on anatomical structure 145 and the state of the anatomical structure. The computing system 110 can search the video database using a mask 175 or 180, or the type of anatomical structure 145, or the type of state identified for anatomical structure 145. The computing system 110 can search multiple videos of a medical procedure using the identified anatomical structure 145 and the state of the anatomical structure 145 to identify subsets of multiple videos. For example, the computing system 110 can identify which videos include the same type of anatomical structure identified by model 125. The computing system 110 can then determine whether the identified anatomical structure includes the state identified by model 125. If the video includes both the type and state of the anatomical structure detected by model 125, the computing system 110 can generate a result set of videos including the same anatomical structure and state. The interface manager 190 can display the videos of the result set to a user via a graphical user interface 197 on a client device 199. The user can play or interact with various videos in the returned video set. Because model 125 includes modular branches, the exact architecture of the model used for segmentation can vary.

[0076] The computing system 110 may include at least one controller 195. The controller 195 may control the robotic medical system 105 using the detection of anatomical structures 145 by model 125 and the detection of the state of the anatomical structures by model 125. The controller 195 may implement rule-based control algorithms, plant-based control algorithms, or any other control algorithms that can move or manipulate instruments or endoscopes of the robotic medical system 105. The controller 195 may detect cuts or burns to anatomical structures caused by unintended instrument movement based on the state detection and anatomical structure detection of model 125. The controller 195 may manipulate the instruments to avoid further cuts or burns. The controller 195 may automatically control the instruments to perform medical procedures based on the state or anatomical structures detected by model 125; for example, state detection may guide the controller 195 to determine whether the robotic medical system 105 has correctly severed an organ or correctly re-sutured an organ. The detected state and anatomical structures may be inputs or feedback to the control algorithms of the fully autonomous or semi-autonomous robotic medical system 105. In some embodiments, the controller 195 is located within the computing system 110, while in other embodiments, the controller 195 is part of the robotic medical system 105. In some embodiments, components of the controller 195 are arranged on both the computing system 110 and the robotic medical system 105.

[0077] Now for reference Figure 2 Among other things, an example model 125 for identifying anatomical structures and detecting their state is shown. Figure 2In this model, model 125 identifies the liver and its burn state. However, model 125 can detect various types of anatomical structures and different states or locations of these structures. Model 125 can be a pipeline. Model 125 can be modularized into individual steps, segments, components, executable modules, functions, sections, or equations. Model 125 can receive image frames 115 and system data 120 as input. Model 125 can discretize the surgical video into continuous frame images. System data 120, which can provide type or step labels, can be synchronized with the images 115. In this respect, each image 115 can have a corresponding type or label.

[0078] Model 125 may include at least one component for embedding or encoding input. For example, model 125 may include at least one image encoder 205. Image encoder 205 may receive frame 115 of a video of a medical protocol. Image encoder 205 may generate an encoding of frame 115. Image encoder 205 may encode or embed image 115. Image encoder 205 may transform a high-dimensional vector or matrix into a low-dimensional vector or matrix. The encoding may be a numerical vector or matrix representation of image 115. Image encoder 205 may use image 115 to generate the hidden internal state of model 125. Image encoder 205 may be a machine learning core or neural network employing image frame 115. In some embodiments, image encoder 205 may be a visual transformer. Image encoder 205 may receive image 115 as input and transform image 115 into a compact feature vector representing the input image 115. Image encoder 205 may be a network operating on spatial and temporal correlations. Image encoder 205 may decipher and extract representative semantics from image 115 by applying an attention mechanism. The attention mechanism can be a selective weighting technique. It can apply different weights to emphasize different parts of image 115, leading to the finding of the optimal and compressed representation of image 115 that satisfies the task or objective of model 125. The attention mechanism can be trained via machine learning based on the task or objective of model 125. If needed, the attention mechanism can apply weights in both spatial and temporal dimensions.

[0079] System data embedding 210 can receive identifiers of medical procedures from the robotic medical system 105, such as system data 120. System data embedding 210 can embed system data 120. System data embedding 210 can generate embeddings using system data 120. System data embedding 210 can convert system data 120 into a numerical vector. System data 120 can be time-series data, which includes an indication of the type of medical procedure being performed and a corresponding timestamp.

[0080] Model 125 may include modules, software components, or operations 215. Operation 215 may combine the encoding of image 115 with the embedding of system data 120. Operation 215 may combine the encoding of image 115 with the embedding of system data 120. Operation 215 may link the numerical vector generated from system data 120 through system data embedding 210 with the visual features encoded from image 115 by image encoder 205. The combined encoding of image 115 and the embedding of system data 120 may be provided to part decoder 220 of part branch 140. The combined encoding of image 115 and the embedding of system data 120 may be provided to anatomical decoder 225 of base branch 135.

[0081] The base branch 135 may include at least one anatomical decoder 225. The anatomical decoder 225 may be a neural network. The anatomical decoder 225 may transform the hidden internal states generated by the image encoder 205 and the system data embedding 210 into... F a Features, identification, segmentation, or detection 145. For example, the anatomical body decoder 225 can output features indicating or representing anatomical structures 145. The anatomical body decoder 225 provides anatomical segmentation using additional transform layers. The anatomical body decoder 225 can output anatomical structure detection 145. Anatomical structure detection 145 can uniquely identify the type of anatomical structure through markers, indicators, labels, or numerical values. The base branch 135 can provide anatomical structure detection 145 to the state or location branch 140. For example, the anatomical structure detection 145 output by the anatomical body decoder 225 for encoded features can be fed into the state or location branch 140 to develop a multi-level hierarchical model 125. Because the anatomical segmentation features 145 are provided to the state or location branch 140, model learning for detecting locations or states can also be based on anatomical structure detection 145.

[0082] State branch 140 may include at least one state or part decoder 220. Decoder 220 can decode the combined image encoding and system data embedding into... F p The state or location decoder 220 can output features, indications, representations, or segments of the state or location 230 of the anatomical structure detected by the anatomical decoder 225. Detection 230 can be a label, name, numerical value, or other identification data segment that identifies the location or state of the detected anatomical structure 145. The state or location decoder 220 can generate a segmentation or mask 180 of the state (e.g., a burn area or a bleeding area) based on an integrated or learned feature representation of the image 115 and system data 120.

[0083] Model 125 may include at least one model 240. Model 125 may include at least one model 245. Models 240 and 245 may be feedforward neural networks or convolutional neural networks that generate image 115 or segments (e.g., masks 180 or 175). Computational system 110 may execute model 240 to generate frame 115 including first mask 180 or generate first mask 180. First mask 180 may include an indication or segmentation of image 115 that identifies portions of image 115 corresponding to state detection 230. Model 240 may be executed based on detections 230 and 145. Detections 230 and 145 may be combined by modules, software components, or operations 235. Operation 235 may combine or link detections 230 and 145. Operation 235 may provide the combined detections 230 and 145 as input to model 240. Model 240 may generate image 115 including mask 180 based on the combined detections 230 and 145. The computing system 110 can execute model 245 to generate frame 115 including second mask 180 or generate second mask 180. Second mask 180 may include indications or segmentations of image 115 that identify anatomical structures. Model 245 can be executed based on detection 145.

[0084] Machine learning engine 130 can train model 125 using training data 150. Machine learning engine 130 can calculate, determine, or generate loss. L p (Loss 160) L ap (Loss 170) or loss L a (Loss 155). Loss 160 can be a loss for state or part branch 140. Loss 155 can be a loss for base branch 135. Computation system 110 can execute a loss function to generate loss 160 for state or part branch 140. Each branch can have its own loss function that measures the difference between the predicted mask and the ground truth mask. Computation system 110 can execute a loss function to generate loss 155 for base branch 135. The loss function can generate loss 155 based on the difference between the mask 175 of the anatomical structure indicated by training data 150 and the ground truth mask of the anatomical structure. Computation system 110 can execute a loss function to generate loss 160. The loss function can generate loss 160 based on the difference between the mask 180 of the state or part of the anatomical structure indicated by training data 150 and the ground truth mask of the state or part of the anatomical structure.

[0085] The machine learning engine 130 may include a loss function to maintain consistency between masks 175 and 180. The machine learning engine 130 may execute the loss function to determine the consistency between masks 175 and 180. Consistency may be that the part mask 180 is contained within (e.g., partially or completely contained) the anatomical mask 175. The machine learning engine 130 may execute the loss function to determine the loss. L a (Loss 170). Loss 170 measures how much overlap there is between masks 180 and 175. For example, if part mask 180 is completely contained within anatomical mask 175, loss 170 can be zero or a low value. The more part mask 180 lies outside anatomical mask 175, the higher loss 170 can be. Machine learning engine 130 can use the amount of part mask 180 lying outside anatomical mask 175 to determine loss 170. Machine learning engine 130 can use the amount of part mask 180 lying within anatomical mask 175 to determine loss 170. Machine learning engine 130 can use the number of pixels of part mask 180 lying outside the pixel fragment or region defined by anatomical mask 175 to determine loss 170. In this respect, training with loss 170 can maximize the overlap between masks 180 and 175, for example, maximizing the amount of mask 180 within a larger segmentation region defined by mask 175.

[0086] Machine learning engine 130 can train model 125 to convergence. For example, machine learning engine 130 can execute machine learning algorithms, such as training algorithms (e.g., backpropagation), to minimize losses 160, 170, and 155. During the training phase of model 125, machine learning engine 130 can feed video frames 115 labeled with anatomical masks 175 and site or state masks 180. Training data 150 can store masks 175 and 180 for various frames 115 of medical or surgical procedures. Because the losses for branches 140 and 135 are independently defined (i.e., losses 160 and 155) and commonly defined (loss 170), training of branches 140 and 135 can be performed simultaneously. Training of all models included in model 125 can be performed in parallel, simultaneously, or concurrently.

[0087] Now for reference Figure 3Among other things, a method 300 for identifying anatomical structures and detecting their state is shown. A computing system 110, a robotic medical system 105, or a client device 199 may perform at least a portion of method 300. Any type of computing system, data processing system, processing system, or processing circuitry may execute instructions to perform at least a portion of method 300. Method 300 may be executed as instructions or as hardware circuitry or logic circuitry. Method 300 may include ACT 305 for receiving data from medical procedures. Method 300 may include ACT 310 for identifying anatomical structures. Method 300 may include ACT 315 for detecting state. Method 300 may include ACT 320 for providing performance indications.

[0088] In ACT 305, method 300 may include receiving medical procedure data by computing system 110. Method 300 may include receiving data from robotic medical system 105 by computing system 110. Method 300 may include receiving one, multiple, or a series of image frames 115 by computing system 110. Method 300 may include receiving system data 120 by computing system 110. System data 120 and image frames 115 may correspond to each other or be linked via timestamps. For example, for a particular image frame 115, system data 120 may indicate the type of medical procedure, the stage of the medical procedure, or the step of the medical procedure. In some embodiments, method 300 may include receiving advanced imaging data, such as hyperspectral image data, by computing system 110.

[0089] In ACT 310, method 300 may include identifying anatomical structures by computing system 110. Method 300 may include executing model 125 to determine anatomical structures 145. Model 125 may include at least one branch. For example, method 300 may include executing a first branch 135 of model 125 by computing system 110 to identify anatomical structures 145. Method 300 may include executing the first branch 135 of model 125 by computing system 110 to select a label, identifier, value, or marker that uniquely identifies the type of anatomical structure 145. For example, the type of structure may be liver, pancreas, muscle, a specific muscle (such as the biceps brachii), a specific bone (such as the tibia), etc.

[0090] Method 300 may include performing an image encoder 205 by computing system 110 to encode frame 115. Method 300 may include performing system data embedding 210 by computing system 110 to embed system data 120. Method 300 may include combining the encoding of image 115 and the embedding of system data 120 by computing system 110, and providing the combined features to state or location decoder 220 and anatomical decoder 225. Method 300 may include performing anatomical decoder 225 by computing system 110. Method 300 may include performing anatomical decoder 225 on the combined features by computing system 110 to detect or output anatomical structures 145. Method 300 may include performing model 245 by computing system 110 to generate a mask 175.

[0091] Method 300 may include detecting a state by computing system 110. Method 300 may include performing a state or location branch 140 using detection 145 performed by base branch 135 to generate a mask 180 for the state. For example, method 300 may include decoding encoded image 115 and embedded system data 120 into state or location detection 230 by state or location decoder 220. Method 300 may include combining detection 230 and detection 145 by operator 235. Method 300 may include generating a mask 180 by model 240 using the combined detection 230 and detection 145.

[0092] In ACT 320, method 300 may include performance indicators provided by computing system 110. Method 300 may include generating at least one metric indicating the performance of a medical procedure performed by an operator of robotic medical system 105. Method 300 may include generating a metric based on detections of anatomical structure 145 by first branch 135 and detections of the state of the anatomical structure by second branch 140. Method 300 may include generating a metric indicating the likelihood that the medical procedure will be a successful procedure. Method 300 may include generating a metric indicating how much unnecessary bleeding, burns, or cuts occurred during the medical procedure. Method 300 may include generating a metric indicating how well the anatomical structure was repaired after cuts. The metric may quantify the state determined by model 125 or the change of the state determined by model 125 over time. For example, if the bleeding state changes over time, indicating that bleeding is slowing or stopping, method 300 may include generating a metric that quantifies how the bleeding is decreasing and whether the bleeding will stop.

[0093] Now for reference Figure 4Among other things, a method 400 for training model 125 using combined branch loss 170 is shown. Computing system 110, robotic medical system 105, or client device 199 may perform at least a portion of method 400. Any type of computing system, data processing system, processing system, or processing circuitry may execute instructions to perform at least a portion of method 400. Method 400 may include at least one ACT 405. Method 400 may include at least one ACT 410 for training the model using combined branch loss. Method 400 may include at least one ACT 415 for deploying the model.

[0094] In ACT 405, method 400 may include receiving a model 125 comprising a first branch 135 and a second branch 140 by computing system 110. Method 400 may include receiving the model 125 from internal components of computing system 110, from a storage system or device of computing system 110, or from an external system. For example, model 125 may be developed or designed on an external system and then sent from the external system to computing system 110. In some embodiments, model 125 is developed or designed on computing system 110. Method 400 may include storing model 125 on at least one storage device or memory device of computing system 110. Method 400 may include receiving model 125 by machine learning engine 130. Method 400 may include storing model 125 on at least one storage device, memory device, or memory device by machine learning engine 130.

[0095] In ACT 410, method 400 may include training model 125 by computation system 110 using combined branch loss. Method 400 may include determining a combined loss 170 by computation system 110 using the determination of a first branch 135 and a second branch 140. Method 400 may include determining a combined loss 170 by computation system 110, which considers the losses in the first branch 135 and the second branch 140 simultaneously, in parallel, or concurrently. Method 400 may include determining a combined loss 170 by computation system 110. Method 400 may include executing a loss function by computation system 110 to determine a combined loss 170. Method 400 may include executing a loss function to determine a combined loss 170 based on the overlap between mask 180 and mask 175. The combined loss 170 may be higher when multiple pixels or fragments of masks 175 and 180 do not overlap. For example, the larger the number of pixels in a mutually exclusive region or a mutually exclusive region, the higher the combined loss 170 may be.

[0096] Method 400 may include training model 125 by machine learning engine 130 to minimize combinatorial loss 170. Method 400 may include training multiple branches of model 125 simultaneously, in parallel, or concurrently by machine learning engine 130. Machine learning engine 130 may train multiple branches of model 125 within the same training process, training epoch, or training segment. Machine learning engine 130 may use combinatorial loss 170 to tune, adjust, or change the values ​​of both first branch 135 and second branch 140. Machine learning engine 130 may train branches 135 and 140 using training data 150. For example, machine learning engine 130 may perform backpropagation or another training algorithm to train model 125 using combinatorial loss 170. In addition to training using combinatorial loss 170, machine learning engine 130 may also use individual branch losses to train model 125. For example, method 400 may include a first branch loss 155 determined by machine learning engine 130 for first branch 135, and a second branch loss 160 determined by machine learning engine 130 for second branch 140. Machine learning engine 130 can train model 125 using first branch loss 155 and second branch loss 160. Machine learning engine 130 can use branch loss 155 to train, tune, or adjust the parameters of first branch 135 and second branch 140. Machine learning engine 130 can use branch loss 160 to train, tune, or adjust the parameters of first branch 135 and second branch 140.

[0097] Now for reference Figure 5 Among other things, an example block diagram of computing system 110 is shown. Computing system 110 may include or be used to implement a data processing system or components thereof. Figure 5 The architecture described herein can be used to implement computing system 110, robotic medical system 105, or client device 199. Computing system 110 may include at least one bus 525 or other communication component for transmitting information, and at least one processor 530 or processing circuitry coupled to bus 525 for processing information. Computing system 110 may include one or more processors 530 or processing circuitry coupled to bus 525 for processing information. Computing system 110 may include at least one main memory 510, such as random access memory (RAM) or other dynamic storage device, coupled to bus 525 for storing information and instructions to be executed by processor 530. Main memory 510 may be used to store information during instruction execution by processor 530. Computing system 110 may further include at least one read-only memory (ROM) 515 or other static storage device coupled to bus 525 for storing static information and instructions of processor 530. Storage device 520 (such as a solid-state device, disk, or optical disk) may be coupled to bus 525 to persistently store information and instructions.

[0098] The computing system 110 can be coupled to a display 500, such as a liquid crystal display or an active matrix display, via a bus 525. The display 500 can display information to a user. An input device 505 (such as a keyboard or voice interface) can be coupled to the bus 525 for transmitting information and commands to the processor 530. The input device 505 may include a touchscreen of the display 500. The input device 505 may include cursor controls, such as a mouse, trackball, or arrow keys, for transmitting directional information and command selection to the processor 530 and for controlling cursor movement on the display 500.

[0099] The processes, systems, and methods described herein can be implemented by a computing system 110 in response to an arrangement of processor 530 executing instructions contained in main memory 510. Such instructions may be read into main memory 510 from another computer-readable medium, such as storage device 520. The arrangement of executing the instructions contained in main memory 510 causes the computing system 110 to perform the illustrative processes described herein. One or more processors in a multiprocessor arrangement may be employed to execute the instructions contained in main memory 510. Hardwired circuitry may be used in place of software instructions, or in combination with software instructions, and used in conjunction with the systems and methods described herein. The systems and methods described herein are not limited to any particular combination of hardware circuitry and software.

[0100] Despite Figure 5 An example computing system is described herein, but the subject matter including the operations described herein can be implemented in other types of digital electronic circuit systems, or in computer software, firmware, or hardware (including the structures disclosed herein and their equivalents), or in a combination of one or more of them.

[0101] Some descriptions herein emphasize the structural independence of aspects of system components or the grouping of operations and responsibilities of these system components. Other groups performing similar overall operations are within the scope of this application. Modules may be implemented in hardware or as computer instructions on a non-transient computer-readable storage medium, and modules may be distributed across various hardware or computer-based components.

[0102] The systems described above can provide any one or more of these components, and these components can be provided on a standalone system or on multiple instances of a distributed system. Furthermore, the systems and methods described above can be provided as one or more computer-readable programs or executable instructions embodied on or in one or more artifacts. The artifacts can be cloud storage devices, hard disks, CD-ROMs, flash memory cards, PROMs, RAM, ROMs, or magnetic tapes. Generally, computer-readable programs can be implemented in any programming language, such as LISP, PERL, C, C++, C#, PROLOG, Python, or in any bytecode language, such as JAVA. Software programs or executable instructions can be stored as object code on or in one or more artifacts.

[0103] Examples and non-limiting module implementation elements include sensors that provide any value determined herein, sensors that provide any value that is a precursor to the value determined herein, data link or network hardware including communication chips, oscillating crystals, communication links, cables, twisted pairs, coaxial cabling, shielded cabling, transmitters, receivers or transceivers, logic circuits, hardwired logic circuits, reconfigurable logic circuits configured in a particular non-transient state according to the module specification, any actuators including at least electrical, hydraulic or pneumatic actuators, solenoids, operational amplifiers, analog control elements (springs, filters, integrators, adders, dividers, gain elements), or digital control elements.

[0104] The subject matter and operations described in this specification can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of these. The subject matter described in this specification can be implemented as one or more computer programs encoded on one or more computer storage media, such as one or more computer program instruction circuits, for execution by or control of the operation of a data processing device. Alternatively or additionally, program instructions can be encoded on artificially generated propagating signals (e.g., machine-generated electrical, optical, or electromagnetic signals) that are generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium can be or is included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of these. While the computer storage medium is not a propagating signal, it can be a source or destination of computer program instructions encoded in artificially generated propagating signals. The computer storage medium can also be or be included in one or more separate components or media (e.g., multiple CDs, disks, or other storage devices including cloud storage). The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0105] The terms "computing device," "component," or "data processing apparatus," etc., include various means, devices, and machines for processing data, including, for example, programmable processors, computers, systems-on-a-chip, or a combination thereof. The device may include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the device may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The device and execution environment can implement various different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0106] Computer programs (also known as programs, software, software applications, applications, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, objects, or other units suitable for a computing environment. A computer program may correspond to a file in a file system. A computer program may be stored as a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on a single computer, located at a site, or distributed across multiple sites and interconnected via a communication network.

[0107] The processes and logic flows described in this specification can be implemented by one or more programmable processors, which execute one or more computer programs by manipulating input data and generating output. The processes and logic flows can also be implemented by special-purpose logic circuitry, and the apparatus can be implemented as special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). Applicable devices for storing computer program instructions and data can include non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory can be supplemented or incorporated therein by special-purpose logic circuitry.

[0108] The subject matter described herein can be implemented in a computing system that includes back-end components (e.g., data servers), middleware components (e.g., application servers), front-end components (e.g., client computers or web browsers with graphical user interfaces through which users can interact with the implementation of the subject matter described herein), or combinations of one or more such back-end, middleware, or front-end components. The components of the system can communicate digitally via any form or medium, such as a communication network. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), the Internet (e.g., the Internet) and peer-to-peer networks (e.g., self-organizing peer-to-peer networks).

[0109] Although the operations are described in a specific order in the accompanying drawings, these operations do not need to be performed in the specific order or sequential order shown, nor is it necessary to perform all of the operations shown. The actions described herein can be performed in different orders.

[0110] Some illustrative embodiments have now been described. It is clear that the foregoing is illustrative and not restrictive, and is presented merely by way of example. In particular, although many of the examples presented herein relate to specific combinations of method actions or system elements, these actions and elements can be combined in other ways to achieve the same objective. An action, element, or feature discussed in connection with one embodiment is not intended to exclude a similar role in other embodiments.

[0111] The wording and terminology used herein are for descriptive purposes and should not be considered limiting. The terms “comprising,” “including,” “having,” “containing,” “involving,” “characterized by,” “characterized by,” and variations thereof, as used herein, are intended to cover items listed below, their equivalents, and additional items, as well as alternative implementations consisting only of items listed below. In one implementation, the system and method described herein comprises one of the said elements, actions, or components; each combination of a plurality of said elements, actions, or components; or all of said elements, actions, or components.

[0112] Any reference to embodiments of systems and methods, or elements or actions, mentioned in the singular form herein may also cover embodiments that include multiple such elements, and any reference to embodiments, elements, or actions mentioned in the plural form herein may also cover embodiments that include only a single element. References in the singular or plural form are not intended to limit the currently disclosed systems or methods, their components, actions, or elements to a single or multiple configuration. References to any ACT or element based on any information, action, or element may include embodiments in which the action or element is at least partially based on any information, action, or element.

[0113] Any implementation disclosed herein may be combined with any other implementation or example, and references to "implementation," "some implementations," or "one implementation," etc., are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic associated with that implementation may be included in at least one implementation or example. These terms used herein do not necessarily all refer to the same implementation. Any implementation may be inclusively or exclusively combined with any other implementation in any manner consistent with the aspects and implementations disclosed herein.

[0114] A reference to "or" can be interpreted as inclusive, so any term described using "or" can refer to any one of the single, multiple, or all described terms. A reference to at least one of the conjunctions in a list of terms can be interpreted as an inclusive "or" to refer to any one of the single, multiple, or all described terms. For example, a reference to "at least one of A and B" can include only "A", only "B", or both "A" and "B". Such references used in conjunction with "include" or other open terms can include additional items.

[0115] If a reference numeral is used after a technical feature in the drawings, detailed descriptions, or any claims, it is included to enhance the comprehensibility of the drawings, detailed descriptions, and claims. Therefore, the reference numerals and their absence do not limit the scope of any claim element.

[0116] Modifications may be made to the described elements and actions without substantially departing from the teachings and advantages of the subject matter disclosed herein, such as changing the size, dimensions, structure, shape and proportions, parameter values, installation arrangement, material use, color, and orientation of various elements. For example, an element presented as a whole may be composed of multiple parts or elements, the positions of elements may be reversed or otherwise varied, and the nature or number of discrete elements or positions may be changed or varied. Other substitutions, modifications, alterations, and omissions may also be made in the design, operating conditions, and arrangement of the disclosed elements and operations without departing from the scope of this disclosure.

Claims

1. A system comprising: One or more processors, coupled to memory, for use in: Receive data on medical procedures performed on subjects using robotic medical systems; Using one or more models trained with machine learning, anatomical structures are identified based on the data; The state of the anatomical structure is detected using one or more of the models and based on the identified anatomical structure; as well as The indication of the performance of the medical procedure is provided at least in part based on the state of the anatomical structure.

2. The system according to claim 1, comprising: The one or more processors are used for: The one or more models are executed on frames of the video of the medical procedure; A first mask for the frame of the anatomical structure is generated using the identified anatomical structure; and A second mask for the frame is generated using the detected state.

3. The system according to claim 1, comprising: The one or more processors are used for: Execute the one or more models trained with the machine learning method, the one or more models comprising: A first branch, comprising at least one first model, the first branch being used to identify the anatomical structure; and The second branch includes at least one second model, which is used to detect the state of the anatomical structure using the anatomical structure identified by the first branch.

4. The system according to claim 1, comprising: The one or more processors are used for: Receive video frames of the medical procedure and the identifier of the medical procedure from the robotic medical system; Generate the encoding of the frame; The embedding of the identifier in the medical protocol is generated; The first branch of the one or more models is executed on the encoding of the frame and the embedding of the identifier to identify the anatomical structure; and The second branch of the one or more models is executed on the encoding of the frame, the embedding of the identifier, and the identified anatomical structure to detect the state of the anatomical structure.

5. The system according to claim 1, comprising: The one or more processors are used for: A first mask for the frame of the anatomical structure is generated using the identified anatomical structure; A second mask for the frame is generated using the detected state; Determine the overlap level between the first mask and the second mask; and Train the one or more models to maximize the overlap level.

6. The system according to claim 1, comprising: The one or more processors are used for: Determine a first loss for a first branch of the one or more models, the first branch being used to identify the anatomical structures; Determine a second loss for a second branch of the one or more models, the second branch being used to detect the state of the anatomical structure; and The first branch and the second branch are trained in parallel using the first loss and the second loss.

7. The system according to claim 1, comprising: The one or more processors are used for: The robotic medical system receives video frames of a medical procedure, hyperspectral data of the frames, and an identifier of the medical procedure. The first branch of the one or more models is executed on the frame and the identifier of the medical protocol to identify the anatomical structure; The second branch of the one or more models is executed on the frame, the identifier of the medical procedure, and the identified anatomical structure to detect the state of the anatomical structure; and The third branch of the one or more models is executed on at least one of the hyperspectral data and the identified anatomical structure or the state of the anatomical structure to detect the oxygen saturation of the anatomical structure.

8. The system according to claim 1, comprising: The one or more processors are used for: The first model of the one or more models is executed using the identified anatomical structures and detected states to generate a first frame including a first mask indicating the states; and The identified anatomical structures are used to perform a second model of the one or more models to generate a second frame that includes a second mask indicating the anatomical structures.

9. The system according to claim 1, comprising: The one or more processors are used for: Using the one or more models and based on the identified anatomical structures, the state of the anatomical structures is detected at multiple time points using multiple frames of the video of the medical protocol; The state is tracked over time using the state detected from the plurality of frames; and The indication of the performance of the medical protocol is generated by using the change of the state over time.

10. The system of claim 1, comprising: The one or more processors are used for: Using the identified anatomical structures and their states, multiple videos of medical procedures are searched to identify a subset of the videos; and Generate data to enable the graphical user interface to display instructions for the subset of the plurality of videos.

11. The system of claim 1, comprising: The one or more processors are used for: The data is received during the medical procedure. During the medical procedure, alerts are generated using the status of the anatomical structures detected by the one or more models; and Generate data to display the alert in the graphical user interface.

12. The system according to claim 1, comprising: Generate a video of the medical procedure including multiple masks to indicate the anatomical structures identified by the one or more models and the state of the anatomical structures detected by the one or more models; as well as Data is generated to enable the graphical user interface to display the video, which includes the plurality of masks.

13. A method comprising: Data on medical procedures performed on a subject using a robotic medical system are received by one or more processors coupled to memory; The one or more processors use one or more models trained with machine learning to identify anatomical structures based on the data; The state of the anatomical structure is detected by the one or more processors using the one or more models and based on the identified anatomical structure; as well as The one or more processors provide indications of the performance of the medical protocol based at least in part on the state of the anatomical structure.

14. The method of claim 13, further comprising: The one or more processors execute the one or more models trained using the machine learning method, the one or more models comprising: A first branch, comprising at least one first model, the first branch being used to identify the anatomical structure; and The second branch includes at least one second model, which is used to detect the state of the anatomical structure using the anatomical structure identified by the first branch.

15. The method of claim 13, further comprising: The one or more processors receive video frames of the medical procedure and the identifier of the medical procedure from the robotic medical system; The encoding of the frame is generated by the one or more processors; The embedding of the identifier of the medical protocol is generated by the one or more processors; The one or more processors execute a first branch of the one or more models on the encoding of the frame and the embedding of the identifier to identify the anatomical structure; as well as The one or more processors execute a second branch of the one or more models on the encoding of the frame, the embedding of the identifier, and the identified anatomical structure to detect the state of the anatomical structure.

16. The method of claim 13, further comprising: The one or more processors use the identified anatomical structure to generate a first mask for the frame of the anatomical structure; The one or more processors generate a second mask for the frame using the detected state; The overlap level between the first mask and the second mask is determined by the one or more processors; as well as The one or more models are trained by the one or more processors to maximize the overlap level.

17. The method of claim 13, further comprising: The one or more processors determine a first loss for a first branch of the one or more models, the first branch being used to identify the anatomical structures; The one or more processors determine a second loss for a second branch of the one or more models, the second branch being used to detect the state of the anatomical structure; as well as The first branch and the second branch are trained in parallel by the one or more processors using the first loss and the second loss.

18. A non-transitory computer-readable medium storing processor-executable instructions, said instructions causing the one or more processors, when executed by said one or more processors, to: Receive data on medical procedures performed on subjects using robotic medical systems; Using one or more models trained with machine learning, anatomical structures are identified based on the data; The state of the anatomical structure is detected using one or more of the models and based on the identified anatomical structure; as well as The medical protocol provides an indication of its performance, at least in part, based on the state of the anatomical structures.

19. The non-transient computer-readable medium of claim 18, wherein the instructions cause the one or more processors to: Receive video frames of the medical procedure and the identifier of the medical procedure from the robotic medical system; Generate the encoding of the frame; The embedding of the identifier in the medical protocol is generated; The first branch of the one or more models is executed on the encoding of the frame and the embedding of the identifier to identify the anatomical structure; as well as The second branch of the one or more models is executed on the encoding of the frame, the embedding of the identifier, and the identified anatomical structure to detect the state of the anatomical structure.

20. The non-transient computer-readable medium of claim 18, wherein the instructions cause the one or more processors to: A first mask for the frame of the anatomical structure is generated using the identified anatomical structure; A second mask for the frame is generated using the detected state; Determine the overlap level between the first mask and the second mask; as well as Train the one or more models to maximize the overlap level.