Automatic generation of procedure metadata using multi-modal data streams

A machine learning model with an embedding layer and transformer automatically generates procedure metadata from multi-modal data streams, addressing manual entry issues and improving real-time control and post-procedure data analysis in computer-assisted systems.

WO2026006212A1PCT designated stage Publication Date: 2026-01-02INTUITIVE SURGICAL OPERATIONS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/034882
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-25
Filing Date
2025-06-24
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Current computer-assisted systems rely on manual entry of procedure metadata, leading to errors and difficulties in accessing and integrating procedure-related data for post-procedure analysis, especially in systems like surgical robots, which hinders real-time control and integration with Electronic Health Record systems.

Method used

A system utilizing a machine learning model with an embedding layer and transformer to generate procedure metadata from multi-modal data streams, enabling automatic metadata generation without manual input, and integrating with control systems for improved real-time control and post-procedure data processing.

Benefits of technology

Automatically generates accurate procedure metadata from multi-modal data streams, reducing manual errors and enhancing the integration of computer-assisted systems with Electronic Health Record systems, facilitating improved real-time control and post-procedure data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025034882_02012026_PF_FP_ABST
    Figure US2025034882_02012026_PF_FP_ABST
Patent Text Reader

Abstract

A system may obtain post-operative medical procedure data representative of a medical procedure that was performed using a computer-assisted system. The post-operative medical procedure data comprises a plurality of data streams. A first data stream comprises time-series data relating to operation of the computer-assisted system. The system inputs each of the plurality of data streams into a respective projection layer of a first machine learning model. The projection layer for the first data stream includes an embedding layer and a transformer. The first machine learning model is configured to output procedure metadata indicators based on outputs from the projection layer, and the first machine learning model is trained using historical procedure data labeled with known metadata. The system receives the medical procedure metadata indicators as an output of the first machine learning model and associate the post-operative medical procedure data with the output medical procedure metadata indicators.
Need to check novelty before this filing date? Find Prior Art

Description

AUTOMATIC GENERATION OF PROCEDURE METADATA USING MULTI- MODAE DATA STREAMSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of the filing date of provisional U.S. Patent Application No. 63 / 663,996 entitled “AUTOMATIC GENERATION OF PROCEDURE METADATA USING MULTI-MODAL DATA STREAMS,” filed on June 25, 2024. The entire contents of the provisional application are hereby expressly incorporated herein by reference.FIELD

[0002] The present disclosure relates generally to computer-assisted systems and more particularly to training and utilizing artificial intelligence to automatically generate procedure metadata from multi-modal data streams connected with the procedure.BACKGROUND

[0003] Computer-assisted systems, including robotically assisted systems or robotic systems, may include one or more manipulators that can be operated with the assistance of an electronic controller (e.g., computer or control system) to move and control functions of one or more instruments coupled to the manipulators. A manipulator generally includes mechanical links connected by joints. An instrument is removably (or permanently) coupled to one of the links, typically a distal link of the plural links. In some embodiments, manipulator systems are used in conjunction with one or more auxiliary devices (e.g., a surgical bed, an insufflator, etc.).

[0004] Machine learning models such as Robotic Transformer (RT) models or the like can be used to control operation of the computer-assisted systems in real-time during a procedure using multi-modal data streams related to the procedure. Typically, the multi-modal data streams include data generated in relation to operation of computer-assisted systems, video, of the procedure, etc. However, automated operation of the computer-assisted systems can be improved in some case by including procedure related metadata such as an indicator of the type or the modality of the procedure being performed among the multi-modal data streams. Furthermore, the procedure related metadata can be useful in analyzing post-procedure data sets, categorizing procedure data sets for training machine learning models, evaluating facility and / or personnel at performing the procedure, etc. Particular examples of post-procedure data set analysis / processing that utilize procedure related metadata include, but are not limited to, generation of surgeon performance assessments, medical procedure workflow, medical environment (e.g., operating room) workflow assessments and analytics, and / or automatic procedure documentation tools (e.g., data that can be input to Electronic Health Record systems).

[0005] Currently procedure related metadata is only available as manual inputs through interaction with the operators of the computer-assisted system (e.g., hospital staff or surgeons for medical procedures). Reliance on manual entry can lead to error in entered metadata and / or complete omission of the metadata in relation to intra-procedure applications. Furthermore, in relation to post-procedure applications, retrieving the metadata for untagged post-procedure data sets compiled by the manufacturer of the computer-assisted system or other entities may be difficult to impossible because those entities do not have direct access to the metadata, which remains in the control of the operators after completion of the procedure.

[0006] Accordingly, there is a need for systems and methods that generate procedure related metadata without direct access to the operators’ protected metadata and / or to avoid reliance on manual entry. In particular, there is a need for systems that can generate the procedure related metadata from real-time multi-modal data streams during a procedure and / or from archived instances of the multi-modal data streams after completion of the procedure. Such systems and methods would improve the real-time control applications of the manipulator systems and enable post-procedure data processing operations without needing to rely on the system operators and thereby reduce the need for manual documentation and improve integration of manipulator systems such as surgical robots into Electronic Health Record systems and tools.SUMMARY

[0007] In some aspects, the techniques described herein relate to a computer system, the system including: one or more processors; and one or more non-transitory, computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to: obtain post-operative medical procedure data representative of a medical procedure that was performed using a computer-assisted system, the post-operative medical procedure data including a plurality of data streams from one or more data sources, wherein a first data stream of the plurality of data streams includes time- series data relating to operation of one or more repositionable structures by the computer-assisted system; input each of the plurality of data streams into a respective projection layer of a first machinelearning model, wherein: the projection layer for the first data stream includes an embedding layer and a transformer, the first machine learning model is configured to output one or more procedure metadata indicators based on outputs from the projection layer for at least a portion of the plurality of data streams, and the first machine learning model is trained using historical procedure data labeled with known metadata; receive the one or more medical procedure metadata indicators as an output of the first machine learning model; and associate the post-operative medical procedure data with the output medical procedure metadata indicators.

[0008] In some aspects, the techniques described herein relate to a computer-implemented method including: obtaining post-operative medical procedure data representative of a medical procedure that was performed using a computer-assisted system, the post-operative medical procedure data including a plurality of data streams from one or more data sources, wherein a first data stream of the plurality of data streams includes time- series data relating to operation of the computer-assisted system; inputting each of the plurality of data streams into a respective projection layer of a first machine learning model, wherein: the projection layer for the first data stream includes an embedding layer and a transformer, the first machine learning model is configured to output one or more procedure metadata indicators based on outputs from the projection layer for at least a portion of the plurality of data streams, and the first machine learning model is trained using historical procedure data labeled with known metadata; receiving the one or more medical procedure metadata indicators as an output of the first machine learning model; and associating the post-operative medical procedure data with the output medical procedure metadata indicators.

[0009] In some aspects, the techniques described herein relate to a computer-assisted system, the system including: one or more repositionable structures configured to support one or more instruments; and a control system operably coupled to the one or more repositionable structures, wherein the control system is configured to: obtain a plurality of data streams from one or more data sources, wherein a first data stream of the plurality of data streams includes time-series data relating to operation of the computer-assisted system; input each of the plurality of data streams into a respective projection layer of a first machine learning model, wherein: the projection layer for the first data stream includes an embedding layer and a transformer, the first machine learning model configured to output one or more medical procedure metadata indicators based on outputs from the projecting layer for at least a portion of the plurality of data streams, and the first machine learning model is trained using historical medical procedure data labeled with known metadata; receive the one or moremedical procedure metadata indicators as an output of the first machine learning model; and control the computer-assisted system based upon the one or more medical procedure metadata indicators.

[0010] In some aspects, the techniques described herein relate to a computer-implemented method including: obtaining a plurality of data streams from one or more data sources, wherein a first data stream of the plurality of data streams includes time- series data relating to operation of a computer-assisted system coupled to one or more repositionable structures; inputting each of the plurality of data streams into a respective projection layer of a first machine learning model, wherein: the projection layer for the first data stream includes an embedding layer and a transformer, the first machine learning model configured to output one or more medical procedure metadata indicators based on outputs from the projecting layer for at least a portion of the plurality of data streams, and the first machine learning model is trained using historical medical procedure data labeled with known metadata; receiving the one or more medical procedure metadata indicators as an output of the first machine learning model; and controlling the computer-assisted system based upon the one or more medical procedure metadata indicators.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] FIGs. 1A and IB are diagrams for a computer-assisted system in accordance with one or more embodiments.

[0012] FIG. 2 is a schematic diagram of a system for determining task specific data streams for controlling a repositionable structure.

[0013] FIG. 3 is a block diagram of an artificial intelligence model in accordance with one or more embodiments.

[0014] FIG. 4 is a block diagram of an artificial intelligence model in accordance with one or more embodiments.

[0015] FIG. 5 is a block diagram of an artificial intelligence model in accordance with one or more embodiments.

[0016] FIG. 6 is a block diagram of an artificial intelligence model in accordance with one or more embodiments.

[0017] FIG. 7 is a flow diagram of a method in accordance with one or more embodiments.

[0018] FIG. 8 is a flow diagram of a method in accordance with one or more embodiments.

[0019] Examples of the present disclosure and their advantages are best understood by referring to the detailed description that follows. It should be appreciated that like reference numerals are used to identify like elements illustrated in one or more of the figures, whereinshowings therein are for purposes of illustrating examples of the present disclosure and not for purposes of limiting the same.DETAILED DESCRIPTION

[0020] In this description, specific details are set forth describing some embodiments consistent with the present disclosure. Numerous specific details are set forth in order to provide a thorough understanding of the embodiments. It will be apparent, however, to one skilled in the art that some embodiments may be practiced without some or all of these specific details. The specific embodiments disclosed herein are meant to be illustrative but not limiting. One skilled in the art may realize other elements that, although not specifically described here, are within the scope and the spirit of this disclosure. In addition, to avoid unnecessary repetition, one or more features shown and described in association with one embodiment may be incorporated into other embodiments unless specifically described otherwise or if the one or more features would make an embodiment non-functional.

[0021] Further, the terminology in this description is not intended to limit the invention. For example, spatially relative terms-such as “beneath”, “below”, “lower”, “above”, “upper”, “proximal”, “distal”, and the like-may be used to describe the relation of one element or feature to another element or feature as illustrated in the figures. These spatially relative terms are intended to encompass different positions (z.e., locations) and orientations (z.e., rotational placements) of the elements or their operation in addition to the position and orientation shown in the figures. For example, if the content of one of the figures is turned over, elements described as “below” or “beneath” other elements or features would then be “above” or “over” the other elements or features. A device may be otherwise oriented and the spatially relative descriptors used herein interpreted accordingly. Likewise, descriptions of movement along and around various axes include various special element positions and orientations. In addition, the singular forms “a”, “an”, and “the” are intended to include the plural forms as well, unless the context indicates otherwise. Additionally, the terms “comprises”, “comprising”, “includes”, and the like specify the presence of stated features, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups. Components described as coupled may be electrically or mechanically directly coupled, or they may be indirectly coupled via one or more intermediate components.

[0022] Elements described in detail with reference to one embodiment, implementation, system, or module may, whenever practical, be included in other embodiments,implementations, systems, or modules in which they are not specifically shown or described. For example, if an element is described in detail with reference to one embodiment and is not described with reference to a second embodiment, the element may nevertheless be claimed as included in the second embodiment. Thus, to avoid unnecessary repetition in the following description, one or more elements shown and described in association with one embodiment, implementation, or application may be incorporated into other embodiments, implementations, or aspects unless specifically described otherwise, unless the one or more elements would make an embodiment or implementation non-functional, or unless two or more of the elements provide conflicting functions.

[0023] In some instances, well known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0024] This disclosure describes various devices, elements, and portions of computer- assisted systems and elements in terms of their state in three-dimensional space. As used herein, the term “position” refers to the location of an element or a portion of an element (e.g., three degrees of translational freedom in a three-dimensional space, such as along Cartesian x-, y-, and z-coordinates). As used herein, the term “orientation” refers to the rotational placement of an element or a portion of an element (e.g., three degrees of rotational freedom in three-dimensional space, such as about roll, pitch, and yaw axes, represented in angle-axis, rotation matrix, quaternion representation, and / or the like). As used herein, and for a device with a kinematic series, such as with a repositionable structure with a plurality of links coupled by one or more joints, the term “proximal” refers to a direction toward a base of the kinematic series, and “distal” refers to a direction away from the base along the kinematic series.

[0025] As used herein, the term “pose” refers to the multi-degree of freedom (DOF) spatial position and orientation of a coordinate system of interest attached to a rigid body. In general, a pose includes a pose variable for each of the DOFs in the pose. For example, a full 6-DOF pose for a rigid body in three-dimensional space would include 6 pose variables corresponding to the 3 positional DOFs (e.g., x, y, and z) and the 3 orientational DOFs (e.g., roll, pitch, and yaw). A 3-DOF position only pose would include only pose variables for the 3 positional DOFs. Similarly, a 3-DOF orientation only pose would include only pose variables for the 3 rotational DOFs. Further, a velocity of the pose captures the change in pose over time (e.g., a first derivative of the pose). For a full 6-DOF pose of a rigid body in three- dimensional space, the velocity would include 3 translational velocities and 3 rotationalvelocities. Poses with other numbers of DOFs would have a corresponding number of velocities translational and / or rotational velocities.

[0026] This disclosure occasionally refers to the disclosed techniques being applied to “patients” undergoing a “medical procedure.” It should be appreciated that these references are not intended to limit the application of the disclosed techniques to applied medicine contexts. For example, the described techniques can be applied to facilitate physician training, equipment testing and / or calibration, and / or other contexts. Accordingly, any reference to the term “patient” is done for ease of explanation and also envisions the application of the described techniques to a generic “subject.”

[0027] The word “task” is used herein to refer to a discrete portion of a procedure that may be autonomously, semi-autonomously, or manually implemented in furtherance of the procedure. For example, a task may be to move an endoscope to a particular position, to advance an instrument to a particular depth, to replace an instrument coupled to a manipulator, and so on. In some embodiments, a task is associated with component tasks to accomplish an overall goal. For example, a task to analyze a worksite may include component tasks related to moving an endoscope to view the worksite, advancing an instrument to predetermined depth, and enabling a functionality supported by the instrument. These component tasks may also be referred to as “subtasks.”

[0028] Aspects of this disclosure are described in reference to computer-assisted systems, which can include devices that are teleoperated, externally manipulated, autonomous, semiautonomous, and / or the like. Further, aspects of this disclosure are described in terms of an implementation using a teleoperated surgical system, such as the da Vinci® Surgical System commercialized by Intuitive Surgical, Inc. of Sunnyvale, California. Knowledgeable persons will understand, however, that inventive aspects disclosed herein may be embodied and implemented in various ways, including teleoperated and non-teleoperated, and medical and non-medical embodiments and implementations. Implementations on da Vinci® Surgical Systems are merely exemplary and are not to be considered as limiting the scope of the inventive aspects disclosed herein. For example, techniques described with reference to surgical instruments and surgical methods may be used in other contexts. Thus, the instruments, systems, and methods described herein may be used for humans, animals, portions of human or animal anatomy, industrial systems, general robotic, or teleoperated systems. As further examples, the instruments, systems, and methods described herein may be used for non-medical purposes including industrial uses, general robotic uses, sensing or manipulating non-tissue work pieces, cosmetic improvements, imaging of human or animalanatomy, gathering data from human or animal anatomy, setting up or taking down systems, training medical or non-medical personnel, and / or the like. Additional example applications include use for procedures on tissue removed from human or animal anatomies (with or without return to a human or animal anatomy) and for procedures on human or animal cadavers. Further, these techniques can also be used for medical treatment or diagnosis procedures that include, or do not include, surgical aspects.

[0029] FIG. 1A illustrates an embodiment of a computer-assisted system. The system can be used, for example, in surgical, diagnostic, therapeutic, biopsy, or non-medical procedures, and is generally indicated by the reference numeral 100. As shown in FIG. 1A, the computer- assisted system 100 can include one or more manipulator assemblies 102 for operating one or more medical instrument systems 104 in performing various procedures on a patient P positioned on a table T in a medical environment 101. For example, the manipulator assembly 102 can drive catheter or end effector motion, can apply treatment to target tissue, and / or can manipulate control members. The manipulator assembly 102 can be teleoperated, non-teleoperated, or a hybrid teleoperated and non-teleoperated assembly with select degrees of freedom of motion that can be motorized and / or teleoperated and select degrees of freedom of motion that can be non-motorized and / or non-teleoperated. An operator input system 106, which can be inside or outside of the medical environment 101, generally includes one or more input devices 107 (see FIG. IB) for controlling manipulator assembly 102. Manipulator assembly 102 supports medical instrument system 104 and can optionally include a plurality of actuators or motors that drive inputs on medical instrument system 104 in response to commands from a control system 112. The actuators can optionally include drive systems that when coupled to medical instrument system 104 can advance medical instrument system 104 into a naturally or surgically created anatomic orifice. Other drive systems can move the distal end of medical instrument in multiple degrees of freedom, which can include three degrees of linear motion (e.g., linear motion along the X, Y, Z Cartesian axes) and in three degrees of rotational motion (e.g., rotation about the X, Y, Z Cartesian axes). The manipulator assembly 102 can support various other systems for irrigation, treatment, or other purposes. Such systems can include fluid systems (including, for example, reservoirs, heating / cooling elements, pumps, and valves), generators, lasers, interrogators, and ablation components.

[0030] Computer-assisted system 100 also includes a display system 110 for displaying an image or representation of the surgical site and medical instrument system 104 generated by an imaging system 109 which can include an imaging system, such as an endoscopic imagingsystem. The outputs of the imaging system 109 can comprise a portion of multimodal data 202 (FIG. 2) as described in more detail below. Display system 110 and operator input system 106 can be oriented so an operator O can control medical instrument system 104 and operator input system 106 with the perception of telepresence. A graphical user interface can be displayable on the display system 110 and / or a display system of an independent planning workstation.

[0031] In some examples, the endoscopic imaging system components of the imaging system 109 can be integrally or removably coupled to medical instrument system 104. However, in some examples, a separate imaging device, such as an endoscope, attached to a separate manipulator assembly can be used with medical instrument system 104 to image the surgical site. The endoscopic imaging system 109 can be implemented as hardware, firmware, software, or a combination thereof which interact with or are otherwise executed by one or more computer processors, which can include a processor system 114 of the control system 112.

[0032] Computer-assisted system 100 can also include a sensor system 108. The sensor system 108 can include a position / location sensor system (e.g., an actuator encoder or an electromagnetic (EM) sensor system) and / or a shape sensor system (e.g., an optical fiber shape sensor) for determining the position, orientation, speed, velocity, pose, and / or shape of the medical instrument system 104. The sensor system 108 can also include temperature, pressure, force, or contact sensors or the like. The outputs of the sensor system 108 can comprise a portion of the multimodal data 202 (FIG. 2) as described in more detail below.

[0033] Computer-assisted system 100 can also include the control system 112. Control system 112 includes memory 116 and the processor system 114 for effecting control between medical instrument system 104, operator input system 106, sensor system 108, and display system 110. Control system 112 also includes programmed instructions (e.g., a non-transitory machine-readable or computer-readable mediums storing the instructions) to implement a procedure using the manipulator assembly 102 including for navigation, steering, imaging, engagement feature deployment or retraction, applying treatment to target tissue (e.g., via the application of energy), or the like.

[0034] Control system 112 can optionally further include a virtual visualization system to provide navigation assistance to operator O when controlling medical instrument system 104 during an image-guided surgical procedure. Virtual navigation using the virtual visualization system can be based upon reference to an acquired pre-operative or intra-operative dataset of anatomic passageways. The virtual visualization system processes images of the surgical siteimaged using imaging technology such as computerized tomography (CT), magnetic resonance imaging (MRI), fluoroscopy, thermography, ultrasound, optical coherence tomography (OCT), thermal imaging, impedance imaging, laser imaging, nanotube X-ray imaging, and / or the like. The control system 112 can use a pre-operative image to locate the target tissue (using vision imaging techniques and / or by receiving user input) and create a pre-operative plan, including an optimal first location for performing treatment. The preoperative plan can include, for example, a planned size to expand an expandable device, a treatment duration, a treatment temperature, and / or multiple deployment locations.

[0035] FIG. IB is another example diagram of a computer-assisted system 100, according to various embodiments. In the example of FIG. IB, the operator input system 106 comprises a workstation that includes the one or more input devices 107. The one or more input devices 107 may comprise one or more leader input devices that are designed to be contacted and manipulated by the operator O. For example, the one or more input devices 107 may be configured for use by the hands, the head, or some other body part(s) of the operator O. The one or more input devices 107, in the example of FIG. IB, are supported by the operator input system 106 and can be mechanically grounded. In some embodiments, an ergonomic support 118 (e.g., forearm rest) can be provided on which the operator O can rest his or her forearms. In some examples, the operator O can perform tasks at a worksite within a workspace near the manipulator assembly 102 during a procedure, by commanding the manipulator assembly 102 using the one or more input devices 107. In a medical example, the worksite may be a surgical worksite within the medical environment 101 associated with the patient P.

[0036] In the example of FIG. IB, the display system 110 may be configured to display images for viewing by the operator O and be configured moved in various degrees of freedom (DOFs) to accommodate the viewing position of the operator O and / or to provide control functions. In embodiments where the display system 110 provides control functions, the one or more input devices 107 may include the display system 110. The display system 110 may display images that depict a worksite at which the operator O is performing various tasks by manipulating the one or more input devices 107 and / or the display system 110. In some examples, images displayed by display system 110 may be received by the operator input system 106 from one or more imaging devices arranged at a worksite (e.g., images from the imaging system 109). In other examples, the images displayed by display system 110 may be generated by the display system 110 (or by a different connected device or system), such as for virtual representations of tools, the worksite, or for user interface components. As willbe explained below, in some embodiments the display system 110 may display one or more tasks for the operator O to perform with respect to any component of the computer-assisted system 100.

[0037] As illustrated, the computer-assisted system 100 also includes the manipulator assembly 102 acting as a follower device that can be commanded by the operator input system 106. In a medical example, the manipulator assembly 102 can be located near an operating table (e.g., the table T of FIG. 1A a bed, or other support) on which the patient P can be positioned. In some medical examples, the manipulator assembly 102 is provided on the operating table, e.g., on or in a patient, simulated patient, or model, training dummy, etc. (not shown). As illustrated, the manipulator assembly 102 may include a plurality of repositionable structures 120 (sometimes referred to as “manipulator arms” in robotic embodiments). In some embodiments, the repositionable structures 120 may include a plurality of links that are rigid members and joints that can be individually actuated as part of a kinematic series. Additionally, each of the repositionable structures 120 is configured to couple to specific instrument 122 (such as an instrument of the medical instrument system 104 of FIG. 1A). While FIG. IB illustrates the manipulator assembly 102 having four repositionable structures 120a- 120d, in other embodiments, the manipulator assembly 102 may include one, two, three, four, five, six, or additional or fewer repositionable structures 120a- 120d.

[0038] The instrument 122 of the medical instrument system 104 can include, for example, a working portion 126 and one or more structures for supporting and / or driving the working portion 126. Example working portions 126 include end effectors that physically contact or manipulate material, energy application elements that apply electrical, RF, ultrasonic, or other types of energy, sensors that detect characteristics of the workspace environment (such as temperature sensors, imaging devices, etc.), and the like. In various embodiments, examples of instruments 122 include, without limitation, a sealing instrument, a cutting instrument, a sealing-and-cutting instrument, an energy instrument for applying energy, a gripping instrument (e.g., clamps, jaws), a stapler, an imaging instrument such as one using optical, RF, or ultrasonic imaging modalities, a sensing instrument, an irrigation instrument, a suction instrument, and / or the like. In addition, the instruments 122 may include a transmission mechanism 128 that can be coupled to a drive assembly 130 of the respective repositionable structure 120a-120d. The drive assembly 130 may include a drive and / or other mechanisms controllable from the operator input system 106 that transmit forces to the transmission mechanism 128 to articular or otherwise actuate the instrument! 122.

[0039] As illustrated, each instrument 122 may be mounted to a portion of a respective repositionable structure 120a- 120d. In FIG. IB, this is shown with the drive assembly 130 physically coupled to the transmission mechanism 128. The distal portion of each repositionable structure 120a- 120d further includes a cannula mount 124 to which a cannula (not shown) is mounted. When a cannula is mounted to the cannula mount 124, a shaft of the instrument 122 passes through the cannula and into a workspace.

[0040] In various embodiments, one or more of the working portions 126 of the instruments 122 may include an imaging device for capturing images as part of the imaging system 109. The imaging devices may include any sensing technology capable of acquiring an image. Example imaging instruments include an optical endoscope, a hyperspectral camera, an ultrasonic sensor, etc. Imaging instruments may comprise monoscopic imagers, stereoscopic imagers, and / or the like. Imaging devices based on radiofrequency domains may capture images in any frequency spectrum, including visible light, infrared light, ultraviolet light, and / or the like. The imaging devices may include an illumination source to light the region being imaged. In embodiments where the working portions 126 of one or more of the instruments 122 include an imaging device, the instrument 122 may be configured to capture images of a portion of the workspace for display via the display system 110.

[0041] In some embodiments, the repositionable structures 120a- 120d and / or instruments 122 can be controlled to move the working portions 126 in response to manipulation of the one or more input devices 107 by the operator O. Accordingly, the repositionable structures 120a- 120d and / or instruments 122 may be said to “follow” the one or more input devices 107 through teleoperation. This enables the operator O to perform tasks at the worksite using the repositionable structures 120a- 120d and / or instruments 122. For a surgical example, the operator O can direct the repositionable structures 120a- 120d of the manipulator assembly 102 to move the working portions 126 as part of a surgical procedure performed at an internal surgical site that is entered via one or more minimally invasive apertures or natural orifices. It should be appreciated that, in some embodiments, the manipulator assembly 102 may include non-teleoperated components that the operator O or other medical professional must manually manipulate to a desired pose.

[0042] In some embodiments, the repositionable structure 120a of the computer-assisted system 100 may be configured to support a working portion 126a that includes an imaging device (also referred to herein as an “imaging device 126a”). For convenience, an instrument 122 that includes an imaging device is also referred to as an “imaging instrument” herein. The control system 112 may be configured to command the repositionable structure 120aand / or the imaging instrument 122 comprising the imaging device 126a to automatically position and / or orient (“pose”) the field of view (FOV) of the imaging device 126a to provide images of the workspace and / or other instruments 122.

[0043] In the illustrated embodiment, the control system 112 is communicatively coupled to the operator input system 106. In other embodiments, the control system 112 may be provided as a component of the operator input system 106 and / or the manipulator assembly 102. During teleoperation, as the operator O moves the one or more input devices 107, one or more sensors configured to detect the one or more input devices 107 generate spatial and / or orientation movement data that is provided to control system 112. The control system 112 may interpret the spatial and / or orientation information to determine and / or provide control signals to the manipulator assembly 102 to control the movement of repositionable structures 120a- 120d, instruments 122, and / or working portions 126. In addition to the components of the manipulator assembly 102, in some embodiments, the control system 112 is configured to interpret inputs received from the operator input system 106 to control operation of one or more auxiliary devices (not depicted) utilized in a procedure. For example, the operator input system 106 may be used to control a pose of a surgical bed or operation of an insufflator.

[0044] In one embodiment, the control system 112 supports one or more wired communication protocols, (e.g., Ethernet, USB, and / or the like) and / or one or more wireless communication protocols (e.g., Bluetooth, IrDA, HomeRF, IEEE 2102.11, DECT, Wireless Telemetry, and / or the like) for communications between the control system 112 and the operator input system 106 and / or the manipulator assembly 102.

[0045] In some embodiments, the control system 112 may be implemented at one or more computing systems. For example, one or more computing systems may be used to control the manipulator assembly 102. As another example, one or more computing systems may be used to control components of the operator input system 106, such as movement of display system 110.

[0046] As illustrated, the control system 112 includes the processor system 114, the memory 116, a control module 132 and an artificial intelligent (Al) assist module 134. The memory 116 may store the control module 132 and the Al assist module 134. The processor system 114 may include one or more processors having different processing architectures for processing instructions. For example, the one or more processors may be one or more cores or micro-cores of a multi-core processor, a central processing unit (CPU), a microprocessor, a field-programmable gate array (FPGA), an application- specific integrated circuit (ASIC), adigital signal processor (DSP), a graphics processing unit (GPU), a tensor processing unit (TPU), and / or the like.

[0047] In some embodiments, the processor system 114 includes circuity to support one or more communication interfaces (e.g., Bluetooth interface, infrared interface, network interface, optical interface, etc.). Additionally, a communication interface of control system 112 may include an integrated circuit for connecting the control system 112 to a network (not shown) (e.g., a local area network (LAN), a wide area network (WAN) such as the Internet, mobile network, or any other type of network) and / or to another device, such as the operator input system 106 and / or the manipulator assembly 102.

[0048] Additionally, the memory 116 may include non-persistent storage (e.g., volatile memory, such as random access memory (RAM), cache memory), persistent storage (e.g., a hard disk, an optical drive such as a compact disk (CD) drive or digital versatile disk (DVD) drive, a flash memory, a floppy disk, a flexible disk, a magnetic tape, any other magnetic medium, any other optical medium, programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a FLASH-EPROM, and / or any other memory chip or cartridge. The non-persistent storage and persistent storage are examples of non- transitory, tangible machine-readable media that can store executable code that, when run by one or more processors (e.g., processor system 114), can cause the one or more processors to perform one or more of the techniques and / or methods disclosed herein.

[0049] The Al assist module 134 may implement one or more machine learning models and / or training protocols therefor. For example, the Al assist module 134 may implement one or more neural networks, deep learning models, decision trees, support vector machines, linear regression, generative Al models, reinforced learning models, random forests, Naive Bayes models, large language models (LLMs), generative adversarial networks, foundation models, image recognition models, linear discriminant analysis models, creative applications, autoregressive models, supervised or unsupervised learning models, multimodal models, vision language models (VLMs), vision foundation models (VFMs), large multi-modal models (LMMs), Transformer models (including Robotic Transformer models), or another machine learning or Al model for performing the methods described herein. The structure of the one or more machine learning models is described in more detail below. The Al assist module 134 may include dedicated processors and memory for storing and performing Al processes, or the Al assist module 134 may utilize resources of the processor system 114 and the memory 116 to store and / or perform any processing or tasks required to perform the methods described herein.

[0050] Additionally, the control system 112 may also include one or more input devices (such as a touchscreen, keyboard, mouse, microphone, touchpad, electronic pen, or any other type of input device) and / or output devices (such as a display device, a speaker, external storage, a printer, or any other output device). In some embodiments, the control system 112 may be implemented on a particular node of a distributed computing system (e.g., a cloud computing system). As another example, different functionalities associated with the control system 112 may be implemented on different nodes of the distributed computing system. Further, one or more elements of the aforementioned control system 112 may be located at a remote location and connected to the other elements over a network.

[0051] In an endoscopic surgery example, the imaging instrument comprising the imaging device 126a may be inserted into the patient prior to the other instruments 122, including a second instrument 122b comprising a second working portion 126b. The second instrument 122b can include any appropriate working portion 126b, and can even include a second imaging device. Accordingly, the imaging device 126a may be maneuvered to positioned to identify a target to which other instruments may interact with as part of another task. The control system 112 may, for example, automatically command the corresponding repositionable structures 120a and 120b to position respective instruments 122a and 122b to perform one or more tasks in tandem, or sequentially based on the specific task, instruments, and positions of the repositionable structures 120a and 120b. In examples described further herein, the control system 112 may perform Al processes and algorithms via the Al assist module 134 to derive procedure metadata indicators from different data stream modalities either in real-time during a procedure or on complete data sets after a procedure is completed. In the real-time examples, these procedure metadata indicators may be used to control the repositionable structures 120, for example by utilizing the procedure metadata as part of a task planning operation. In the post-procedure examples, the procedure metadata may be associated with post-procedure data for use in future data processing (as described in more detail herein) and / or evaluating performance of a facility or personnel performing the procedure. Utilizing the Al assist module 134 to generate the procedure metadata indicators using only the different data stream modalities means the systems can avoid the problems associated with manual data entry by the operator O or other personnel following completion of the procedure. The Al assist module 134 may be coupled to the manipulator assembly 102 or be a stand alone device located in the medical environment 101 or other location that employs the medical environment 101. In either case, the Al assist module 134 may act as an independent node of the computer-assisted system 100 to generate the procedure metadataindicators in conjunction with procedures that utilize and do not utilize the manipulator assembly 102. In particular, operating the Al assist module 134 as a distinct computational node can enable the computer-assisted system 100 to provide real-time control applications for systems other than the manipulator assembly 102 and to accurately and automatically document procedures that do not employ the manipulator assembly 102.

[0052] FIG. 2 is a schematic diagram of a system 200 for determining task specific data streams for controlling a manipulator assembly or similar repositionable structure (such as the manipulator assembly 102). Additional details of the system 200 are shown and described in U.S. Provisional application 63 / 663,539 titled MULTI-MODAL AND TASK-ADAPTIVE SYSTEM ARCHITECTURE FOR AN INTELLIGENT SURGICAL ROBOT and filed on June 24, 2024, which is incorporated by reference herein in its entirety. The system includes a series of modules that may be executed by the control system 112 (e.g., via the Al assist module 134 of FIG. IB) to control the repositionable structure to perform one or more tasks. The system 200 includes sources of multi-modal data 202 including one or more multi-modal data streams, a task generation module 210, a task selection module 220, a data modality selection module 230, an instrument / arm select module 240, and a robotic action module 250. It should be appreciated that in other embodiments, additional or fewer modules may be implemented. Further, in other embodiments, one or more of the described modules may be combined into a single module.

[0053] The system 200 includes one or more sources of multi-modal data 202. At least some of multimodal data 202 may be outputs of the sensor system 108 and / or the imaging system 109 and additional sources of the multimodal data 202 may be other data sources or data derived from outputs of the sensor system 108, the imaging system 109, and / or other components of the computer-assisted system 100. The multimodal data 202 may include a data stream 202a indicative of force exerted upon an instrument, a data stream 202b indicative of system events generated by the repositionable structure and / or a control system thereof, a data stream 202c indicative of kinematic data associated with the repositionable structure and / or instruments or auxiliary devices associated therewith, an external video data stream 202d, a procedure video data stream 202e (such as image data generated by an endoscope, a fluorescent imaging device, a hyper- spectral imaging device, an ultrasound device, etc.), procedure metadata indicators 202f, and / or other sources of data that indicate a state of a procedure facilitated by the control system 112 (such as 3D depth data from a time- of-flight (ToF) sensor).

[0054] The system 200 may be configured to route the multimodal data 202 to a task generation module 210 configured to output one or more tasks that may be performed in furtherance of the procedure based on a procedure state represented by the input multimodal data 202. In some embodiments, the task generation module 210 determines the procedure and / or the state thereof based in part on the procedure metadata indicators 202f. The task generation module 210 may identify tasks that are to be performed in the near term (e.g., tasks that respond to conditions detected in an input set of multi-modal data) and / or tasks that are to be performed in the long term (e.g., tasks that will need to be performed later in the procedure after one or more near-term tasks are performed). By also generating long-term tasks, the system 200 is able to track preparedness for performing the long-term task and generate additional tasks to prepare the control system for performing the long-term task as the procedure advances closer to the appropriate time for performing the long-term tasks. In some embodiments, the task generation module 210 may include component models to facilitate the analysis.

[0055] The system 200 may then provide the output tasks to a task selection module 220 to select which tasks should actually be implemented by the control system and / or an operator thereof. The system 200 may then provide the selected tasks to the data modality selection module 230, instrument / arm selection module 240, and / or the robotic action module 250.

[0056] The data modality selection module 230 may be configured to analyze the one or more selected tasks to determine one or more task specific data streams needed to implement the selected tasks. As described above, each task may require a particular type of data to safely execute the task. Accordingly, the data modality selection module 230 may identify the required data modalities for performing the selected tasks. It should be appreciated that a required data modality may include a particular data stream of the multimodal data 202 (e.g., kinematic data or procedure video) or a particular state of the multimodal data (e.g., that the procedure video is required to captured image data of a particular object, such a target anatomy). It should be appreciated if a required data stream is unavailable, the data modality selection module 230 may interact with the task selection module 220 to select a different task where the required data streams are all available and / or the task generation module 210 to generate a task to make the required data stream available (e.g., by enabling a sensor or by commanding a camera to reposition a field of view). The data modality selection module 230 may then provide a selection of the available data streams to the robotic action module 250. If a task cannot be performed due to the unavailability of a specific data modality, the data modality selection 230 may generate an alert to an operator of the system 200.

[0057] Because the system 200 may be configured to process multiple tasks synchronously, the data modality selection module 230 may be configured to associate the selected available data streams with each task prior to providing the selection to the robotic action module 250. As a result, the robotic action module 250 may be able to switch data modalities based on the particular tasks being converted into robotic action.

[0058] Like the data modality selection module 230, the instrument / arm selection module 240 may be configured to receive the one or more selected tasks from the task selection module 220. In response, the instrument / arm selection module 240 may be configured to determine a particular instrument of the one or more instruments and / or a particular arm of the one or more arms of the repositionable structures for performing the selected tasks. Accordingly, the instrument / arm selection module 240 may identify which instruments and / or arms of the manipulator or repositionable structure are capable of implementing the selected tasks and assign the task to a particular instrument / arm. For example, if the system 200 is used to control multiple arms / instruments coupled to gripper end effectors, the instrument / arm selection module 240 may identify the instrument / arm corresponding to a particular gripper end effector selected to perform the task. It should be appreciated that in some scenarios, the appropriate instrument is within the procedural theater, but not coupled to the manipulator or repositionable structure. Accordingly, in this scenario, the instrument / arm selection module 240 may interface with the task generation module 210 to generate one or more tasks related to swapping an instrument coupled to arm such that the task can be performed. The instrument / arm selection module 240 then provides data indicative of specific instruments and / or arms for performing one or more of the selected tasks to the robotic action module 250.

[0059] In examples, the outputs of task selection module 220 and / or the instrument / arm selection module 240 may be provided to the operator O, such as via the operator input system 106 of FIGs. 1A and IB. The operator O may then approve of the tasks to be performed and / or the instrument / arm assignment prior to implementing the selected tasks. If the operator O disapproves the selected tasks, the task selection module 220 may provide alternative tasks that can be performed. In some embodiments, if the operator O disapproves the assignment of the tasks, the operator O may be able to manually assign the task to the preferred instrument / arm via the display system 110.

[0060] As described above, the robotic action module 250 may be configured to receive the selected tasks from the task selection module 220, task specific data modalities from the data modality selection module 230, the instrument / arm assignment information from theinstrument / arm selection module 240, and the multimodal data 202. The robotic action module 250 may then process the multimodal data 202 corresponding to the selected data modalities to generate action tokens to control the instruments / arms assigned to perform the selected tasks. The robotic action module 250 may then convert the action tokens to command one or more components of a computer-assisted robotic device (such as the manipulator assembly 102 of FIGs. 1 A and IB) to perform one or more of the selected tasks utilizing one or more of the selected instrument(s), arms(s) and auxiliary device(s).

[0061] It should be appreciated that while the system 200 includes both the data modality selection module 230 and the instrument / arm selection module 240, in some embodiments, the system 200 only includes one of the modules 230, 240. Additionally, in some embodiments, one or more of the modules 210, 220, 230, 240 may be combined into a single module, for example, by implementing chain of thought (CoT) or other recurrent prompting techniques such that a single prompt performs the described functionality corresponding to the individual modules.

[0062] FIG. 3 is a diagram of a system 300 for generating medical procedure metadata indicators from medical procedure data. These medical procedure metadata indicators can include the procedure metadata indicators 202f discussed above with respect to FIG. 2 where the procedure performed using the computer-assisted system 100 (e.g., employing the manipulator assembly 102) includes a medical procedure.

[0063] As shown in FIG. 3, the system 300 includes medical procedure data comprising a plurality of data streams 302 from one or more data sources. The medical procedure data comprising the plurality of data streams 302 may be either post-operative data from completed procedures or real-time data gathered from a procedure that is currently in process. A first data stream 302A of the plurality of data streams 302 comprises time- series data relating to operation of the computer-assisted system 100 of FIG. 1A or IB during the medical procedure. The first data stream 302A may include the data stream 202a indicative of the force exerted upon an instrument (e.g., one of instruments 122), the data stream 202b indicative of system events generated by the repositionable structures 120 and / or the control system 112, the data stream 202c indicative of kinematic data associated with the repositionable structures 120, instruments 122, and / or auxiliary devices associated therewith (e.g., the surgical bed, insufflator, etc.). A second data stream 302B of the plurality of data streams 302 may comprises image data (e.g., image data from the imaging system 109). In particular, the image data can include medical environment (e.g., operating room) image data such as external video data stream 202d and / or endoscope image data such as procedurevideo data stream 202e. The plurality of data streams 302 may include a plurality of additional data streams 302C-302N that may have data modalities the same as or different from the first data stream 302A and the second data stream 302B. For example, in some embodiments, the plurality of additional data streams 302C-302N may include text data, audio data, 3D data from sensors, time-of-flight data, etc.

[0064] The system 300 also includes a first machine learning model 303. The first machine learning model 303 comprises a set of interconnected nodes, layers, trained parameter values (e.g., multiplicative weights, additive bias, etc.), etc. The trained parameters are set via backpropagation techniques in a training process that uses historical data inputs. In particular, the historical data inputs can include historical procedure data labeled with known metadata. The known metadata can include known medical procedure indicators (e.g., known medical procedure types and / or modalities) and / or other metadata values useful in setting the parameters of the first machine learning model 303. In some embodiments, the first machine learning model 303 is configured as a multi modal procedure classification Transformer (MM-PCT) that can receive inputs of different modalities (e.g., the plurality of data streams 302) in order to generate the medical procedure metadata indicators as described herein.However, it should be appreciated that MM-PCT architecture may be substituted or supplemented with other Al architectures including, but not limited to, convolutional neural network (CNN) architectures, recurrent / recursive neural network (RNN) architectures, sorting / clustering architectures, etc.

[0065] The first machine learning model 303 may be executable by the control system 112, such as by the Al assist module 134 and / or another similar computing system connected to the control system 112. The first machine learning model 303, when executed, is configured to output one or more procedure metadata indicators based on input of the plurality of data streams 302. As described herein, execution of the first machine learning model 303 can include transforming the input plurality of data streams 302 into embedded tokens, data values, etc. to which various modification functions and the trained parameter values are applied to generate the output one or more procedure metadata indicators.

[0066] As shown in FIG. 3, the first machine learning model 303 includes a plurality of projection layers 304. The plurality of projection layers 304 are configured to receive the plurality of data streams 302 and convert the underlying data into machine readable formats suitable for processing by the first machine learning model 303. As described in more detailed herein, the machine-readable formats can include token elements representative of the data in the plurality of data streams 302. In some embodiments, the token elements arenatural language descriptions of corresponding data stream 302. The plurality of projection layers 304 may include a distinct layer for each of the plurality of data streams 302 or, alternatively, some of the plurality of projection layers 304 can be configured to receive multiple data streams of the plurality of data streams 302 such as ones having the same modality. For example, a first projection layer 304A can be configured to receive the first data stream 302A a second projection layer 304B can be configured to receive the second data stream 302B, and corresponding projection layers 304C-304N can be configured to receive the additional data streams 302C-302N. However, in cases where some of the plurality of projection layers 304 are configured to receive multiple data streams, the first projection layer 304A may be configured to receive multiple time series type data streams and the second projection layer 304B may be configured to receive multiple image type data streams (e.g., both the external video data stream 202d and the procedure video data stream 202e).

[0067] In either case, each of the plurality of data streams 302 are fed into one of the plurality of projection layers 304 and passed onto additional / hidden layers 306 before being received by an output layer 308. The output layer 308 aggregates outputs from each of the plurality of projection layers 304 and the additional / hidden layers 306 to generate the one or more procedure metadata indicators. As such, the one or more procedure metadata indicators output by the first machine learning model 303 are based on outputs from the projection layers 304 for at least a portion of the plurality of data streams 302.

[0068] The first projection layer 304A includes an embedding layer 310 and a transformer 312. The embedding layer 310 is configured to project the time-series data into machine readable tokens with embedded positional values. In some embodiments, the first projection layer 304A or the embedding layer 310 in particular is trained using a plurality of historical time series data streams having associated time stamp information such that the embedded positional values are directly derived from the time stamp information encoded with the timeseries data itself. Furthermore, the transformer 312 of the first projection layer 304A includes at least one attention layer. This attention layer is trained to modify its received inputs into refined embeddings that represent contextual relationships between the data points in the first data stream. The attention layer may be trained using ProbSparse Attention techniques or similar so that important data points are identified while avoiding quadratic complexity issues associated with longer-duration time series data. Additionally, AutoFormer, Spacetimeformer, Informer, Pyraformer, FEDformer, etc. type architectures can be employed as the classifier that accepts the time- series data as an input. In some embodiments, theinputs into the attention layer comprise outputs of the embedding layer 310. However, the first projection layer 304A may employ multiple sequential attention layers interposed between other neural network layers such as feed forward layers or the like.

[0069] In general, the second projection layer 304B is configured to transform the image data into embedded data useable by the remaining portions of the first machine learning model 303. For example, the second projection layer 304B can include a convolutional neural network or a large multimodal model (LMM) trained to process and classify image data.

[0070] The additional / hidden layers 306 may comprise various Al layers known in the art. Such layers include additional multiheaded self-attention or cross-attention layers, feed forward layers, etc. In some embodiments, the additional / hidden layers 306 may be omitted and the output layer 308 may be directly connected to the plurality of projection layers 304. However, in other embodiments, the additional / hidden layers 306 and the output layer 308 may comprise layers of a pretrained Al model such as a large language model (LLM) or similar. This pretrained Al model can include third party provided models that are either fine tuned for use with the computer-assisted system 100 or that are used as is without any additional tunning or training. When the pretrained Al model utilized is an LLM, the plurality of projection layers 304 are configured to transform the input plurality of data streams 302 into machine readable tokens representing textual descriptions of the underlying data elements that can be fed into the general word embedding layers of the pretrained LLM. As described in more detail below, the plurality of projection layers 304 may be configured to generate these textual descriptions via dedicated training processes.

[0071] Once the medical procedure metadata indicators are generated and output from the output layer 308 of the first machine learning model 303 they can be received by the control system 112 or similar computing environment that implements the system 300. Once received, the control system 112 or similar computing environment can leverage the medical procedure metadata indicators to control the repositionable structures 120 or other aspects of the computer-assisted system 100 (e.g., control other portions of the manipulator assembly 102, control the one or more auxiliary devices, present data on the display system 110, change settings of the operator input system 106, etc.). The procedure metadata indicators may also be utilized to perform various data processing tasks on the plurality of data streams 302. For example, where the medical procedure metadata indicators are utilized to control the repositionable structures 120 or other aspects of the manipulator assembly 102, the control system 112 may provide the output medical procedure metadata indicators as the procedure metadata indicators 202f for the system 200 of FIG. 2, described above. Furthermore, wherethe medical procedure metadata indicators are utilized for other data processing tasks, the control system 112, or similar computing environment, can associate the post-operative medical procedure data that forms the plurality of data streams 302 with the medical procedure metadata indicators output from the first machine learning model 303.

[0072] Furthermore, in embodiments where the output medical procedure metadata indicators are used to control the computer-assisted system 100 in real-time during a procedure, the first machine learning model 303 and / or another component of the system 300 may be configured to output a confidence score documenting the reliability of the output medical procedure metadata indicators. This confidence score may be used to dictate when the output medical procedure metadata indicators are useable for automatic performance of the control operations. For example, the confidence score may indicate when a sufficient threshold of reliability is reached, which limits situations where incorrect medical procedure metadata indicators (e.g., those generated from too little data) would be used. For example, during the initial phases of the procedure, the outputs of the first machine learning model 303 may not be usable for real-time control application because the low amount of data generated at that point in the procedure correlates with a low confidence score. However, as the first machine learning model 303 receives more data the confidence score may increase until the confidence score passes over a predefined reliability threshold. The predefined reliability threshold may include a value at which outputs from the first machine learning model 303 are deemed to be sufficiently reliable for use in real-time tasks. In some embodiments, the predefined reliability threshold may be identified from testing the first machine learning model 303 on historical data sets (e.g., post-operative medical procedure data as described herein).

[0073] The medical procedure metadata indicators that are output from the first machine learning model 303 (e.g., the output of the output layer 308) may include text or other data indicative of a medical procedure type and / or a medical procedure modality. The medical procedure type indicators may include a description of a specific procedure being performed (e.g., urology procedures such as prostatectomy and nephrectomy; general surgery procedures such as hernia, chole, and colorectal; gynecology procedures such as a hysterectomy; etc.) and the medical procedure modality indicator may include a description of the equipment employed for use in the procedure. In some cases, the modality identifier can be used to determine whether the procedure is or is not a robotic surgery procedure that employs the repositionable structures 120 or other components of the manipulator assembly102. In these embodiments, the modality identifier may be one of open, laparoscopic, robot- assisted, image-guided surgery, orthopedic, neuro-surgery, etc.

[0074] While, as described above, the first machine learning model 303 may be trained to output both the medical procedure type and / or a medical procedure modality indicators, in some alternative embodiments, the first machine learning model 303 may be configured to only to output the medical procedure type indicators. In these embodiments, the historical data inputs used to train the first machine learning model 303 can include historical procedure data labeled with known medical procedure type indicators. Furthermore, in these embodiments, a second machine learning model 403 of the system 400 shown in FIG. 4 may be used to separately generate the medical procedure modality indicators.

[0075] The second machine learning model 403, like the first machine learning model 303, comprises a set of interconnected nodes, layers, trained parameter values (e.g., multiplicative weights, additive bias, etc.), etc. The trained parameters of the second machine learning model 403 are different from those of the various embodiments of the first machine learning model 303 described herein and are set via backpropagation techniques in another training process that uses additional historical data inputs. These additional historical inputs may include medical environment data labeled with known medical procedure modality indicators. Furthermore, the second machine learning model 403 is configured to receive multi-modal medical environment data 402 as an input for use in generating the medical procedure modality indicators. The multi-modal medical environment data 402 may comprise data sources from which procedure modality can be inferred such as senor or video data of an operating or procedure room. For example, the multi-modal medical environment data 402 comprises an medical environment sensor data portion of the plurality of data streams 302. However, in some cases the multi-modal medical environment data 402 may comprise medical environment sensor data different from the plurality of data streams 302.

[0076] The structure of the second machine learning model 403 is similar to that of the first machine learning model 303 and includes one or more projections layers 404 similar to the plurality of projection layers 304, additional / hidden layers 406 like the additional / hidden layers 306, and an output layer 408 similar to the output layer 308. The output layer 408 aggregates outputs from each of the one or more projections layers 404 and the additional / hidden layers 406 to generate the medical procedure modality indicators.

[0077] Once the medical procedure modality indicators are generated and output from the output layer 408 they can be received by the control system 112 or similar computing environment that employs the system 400. The control system 112 or similar computingenvironment may combine the medical procedure modality indicators with the medical procedure type indicators as output from the first machine learning model 303 and leverage the indicators to perform control operation and data processing task in the same manner as the described above. Furthermore, as shown in FIG. 4, the control system 112 or similar computing environment can utilize the medical procedure modality indicators output from the second machine learning model 403 as an input data stream of the first machine learning model 303 for use in generating the medical procedure type indicators. In particular, the medical procedure modality indicator may be one of the additional data streams 302C-302N of the plurality of data streams 302.

[0078] Furthermore, associations of the medical procedure modality indicators and the medical procedure type indicators (whether output together from the first machine learning model 303 or from the second machine learning model 403 and the first machine learning model 303 separately) with the medical procedure data (e.g., the plurality of data streams 302) may comprise labeled post-operative medical procedure data. This labeled postoperative medical procedure data may be used to further train or refine the first machine learning model 303 and / or the second machine learning model 403. The control system 112 and / or the similar computing environment can execute a procedure- specific analysis of the labeled post-operative medical procedure data. This procedure-specific analysis may include an evaluation of a facility or personnel that performed the medical procedure.

[0079] An example process for training the first machine learning model 303 will be described with reference now to FIG. 5. First, historical time series data 502 and video data 504 are compiled and associated with ground truth data 506. The historical time series data 502 may include saved time series data streams from past procedures that includes force data, event data, and / or kinematic data as described herein. It should be appreciated that the event data may include a state of any robotic component associated with the system 200 at a particular time. The historical video data 504 may include saved external video data streams (e.g., medical environment video) and / or procedure video data streams (e.g., image data generated by an endoscope) form the same procedures that produced the historical time series data 502. The ground truth data 506 may comprise case metadata text (e.g., procedure metadata indicators) that is associated with the historical time series data 502 and the historical video data 504. In some embodiments, the ground truth data 506 may also include data such as a procedure date and time, procedure name, hospital name, surgeon name, etc.

[0080] The historical time series data 502 and the historical video data 504 may be input into an initialized version of the first machine learning model 303 (e.g., a version where theparameter values are randomized) and processed through the plurality of projection layers 304 and the additional / hidden layers 306 until the output layer 308 generates training output values. These training output values may then be compared to the ground truth data 506 to identify an error associated with the training output values. The error may then used to update each parameter value in the first machine learning model 303 using backpropagation techniques as known in the art. This process of generating training output values from the historical video data 504 and historical time series data 502, and backpropagating the resulting error to update the parameter values of the first machine learning model 303 may be repeated until a threshold condition indicating reliability of the first machine learning model 303 is achieved. This threshold can include a predetermined or minimum number of iterations through the process and / or value indicating negligible improvement to the error value form further updates to the parameter values of the first machine learning model 303.

[0081] It should be appreciated that similar training techniques may be employed to train the second machine learning model 403 as shown in FIG. 4. However, in that particular case, the historical data input may include the historical medical environment data and the ground truth data may include historical procedure modality indicators.

[0082] Furthermore, as shown in FIG. 6 these training techniques may be employed to train a third machine learning model 603 that uses only times series data (e.g., force data, event data, and / or kinematic data from the manipulator assembly 102, repositionable structures 120, etc.) as the inputs. The third machine learning model 603 includes a projection layer 604 similar to the first projection layer 304A, additional / hidden layers 606 similar to the additional / hidden layers 306, and an output layer 608 similar to the output layer 308. In some embodiments the third machine learning model 603 once trained can be integrated into the first machine learning model 303 in place of or in addition to the first projection layer 304A.

[0083] FIG. 7 shows a method 700 for generating post-operative medical procedure metadata indicators using an artificial intelligence model (such as first machine learning model 303, second machine learning model 403, third machine learning model 603, etc.). The method 700 may be executed by the control system 112, or a similar computing system having a processing unit and a memory, such as a remote server communicative coupled to the control server 112.

[0084] At block 710, the control system 112 obtains post-operative medical procedure data representative of a medical procedure that was performed using a computer-assisted system (e.g., the computer-assisted system 100 of FIGS. 1A or IB). The post-operative medical procedure data comprises a plurality of data streams from one or more data sources (e.g.,multimodal data 202 of FIG. 2 and / or the plurality of data streams 302 of FIG. 3). A first data stream of the plurality of data streams (e.g., first data stream 302A) comprises time-series data relating to operation of the computer-assisted system (e.g., operation of the repositionable structures 120, the one or more auxiliary devices, etc.). The time-series data may include one or more of event data, kinematics data, or force data of the computer- assisted system.

[0085] At block 720, the control system 112 inputs each of the plurality of data streams into a respective projection layer of a first machine learning model (e.g., the plurality of projection layers 304 of the first machine learning model 303). The projection layer for the first data stream includes an embedding layer and a transformer (e.g., embedding layer 310 and transformer 312). The first machine learning model is configured to output one or more procedure metadata indicators based on outputs from the projection layer for at least a portion of the plurality of data streams, and the first machine learning model is trained using historical procedure data labeled with known metadata. A second data stream of the plurality of data streams comprises image data. The image data may be endoscope image data or medical environment image data. Furthermore, the respective projection layer for the second data stream includes a convolutional neural network or a large multimodal model (LMM). The projection layer for the first data stream is configured to convert the time-series data into machine readable tokens with embedded positional values and is trained using a plurality of historical time series data streams having associated time stamp information. The projection layer for the first data stream may include an attention layer trained to modify inputs into the attention layer into refined embeddings that represent contextual relationships between the data points in the first data stream. The inputs into the attention layer may comprise outputs of the embedding layer.

[0086] At block 730, the control system 112 receives the one or more medical procedure metadata indicators as an output of the first machine learning model. The output medical procedure metadata indicators may include indicators indicative of a medical procedure type and / or indicators indicative of a medical procedure modality.

[0087] At block 740, the control system 112 associates the post-operative medical procedure data with the output medical procedure metadata indicators.

[0088] In some embodiments of the method 700, the first machine learning model is configured to output medical procedure metadata indicators indicative of a procedure type. In these embodiments the method 700 further comprises inputting multi-modal medical environment data into a second machine learning model (e.g., the second machine learning 1model 403) and receiving one or more medical procedure modality indicators as an output of the second machine learning model. The multi-modal medical environment data may comprise the post-operative medical procedure data; an medical environment sensor data portion of the post-operative medical procedure data; and / or medical environment sensor data different from the post-operative medical procedure data.

[0089] The second machine learning model is configured to output one or more medical procedure modality indicators indicative of a medical procedure modality based on outputs from a projection layer (e.g., one or more projections layers 404) of the second machine learning model for at least a portion of the multi-modal medical environment data. The second machine learning model is trained using historical multi-modal medical environment data labeled with known medical procedure modality indicators. In some embodiments, the method 700 includes including the output of the second machine learning model as an input data stream of the first machine learning model. Furthermore, the method 700 may include associating the post-operative medical procedure data with the output medical procedure modality indicators. The associations of the output medical procedure metadata indicators and the output medical procedure modality indicators with the post-operative medical procedure data comprise labeled post-operative medical procedure data. The method 700 may include further training the first machine learning model using the labeled post-operative medical procedure data; and / or performing a procedure- specific analysis of the labeled postoperative medical procedure data. The procedure-specific analysis includes an evaluation of a facility or personnel that performed the medical procedure.

[0090] FIG. 8 shows a method 800 for generating medical procedure metadata indicators using an artificial intelligence model (such as first machine learning model 303, second machine learning model 403, third machine learning model 603, etc.). The method 700 may be executed by the control system 112 or a similar computing system having a processing unit and a memory.

[0091] At block 810, the control system 112 obtains a plurality of data streams from one or more data sources (e.g., multimodal data 202 of FIG. 2 and / or the plurality of data streams 302 of FIG. 3). A first data stream of the plurality of data streams (e.g., first data stream 302A) comprises time-series data relating to operation of a computer-assisted system (e.g., the computer-assisted system 100). The time-series data may include one or more of event data, kinematics data, or force data of the computer-assisted system.

[0092] At block 820, the control system 112 inputs each of the plurality of data streams into a respective projection layer of a first machine learning model (e.g., the plurality ofprojection layers 304 of the first machine learning model 303). The projection layer for the first data stream includes an embedding layer and a transformer (e.g., embedding layer 310 and transformer 312). The first machine learning model is configured to output one or more medical procedure metadata indicators based on outputs from the projecting layer for at least a portion of the plurality of data streams, and the first machine learning model is trained using historical medical procedure data labeled with known metadata. A second data stream of the plurality of data streams comprises image data. The image data may be endoscope image data or medical environment image data. Furthermore, the respective projection layer for the second data stream includes a convolutional neural network or a large multimodal model (LMM). The projection layer for the first data stream is configured to convert the timeseries data into machine readable tokens with embedded positional values and is trained using a plurality of historical time series data streams having associated time stamp information. The projection layer for the first data stream may include an attention layer trained to modify inputs into the attention layer into refined embeddings that embody contextual relationships between the data points in the first data stream. The inputs into the attention layer may comprise outputs of the embedding layer.

[0093] At block 830, the control system 112 receives the one or more medical procedure metadata indicators as an output of the first machine learning model. The output medical procedure metadata indicators may include indicators indicative of a medical procedure type and / or indicators indicative of a medical procedure modality.

[0094] At block 840, the control system 112 controls the computer-assisted system based upon the one or more medical procedure metadata indicators. The control system 112 may input at least a portion of the plurality of data streams into a robotics transformer model (e.g., the modules of the system 200 of FIG. 2) to generate control outputs for controlling the computer-assisted system and control the computer-assisted system by inputting the medical procedure metadata indicators into an input layer of the robotics transformer model (e.g., task generation module 210). The control system 112 may analyze the plurality of data streams with a projection layer for the robotics transformer model to detect anomalous performance. In some embodiments, the projection layer for the robotics transformer model detects the anomalous performance based upon the one or more medical procedure metadata indicators. In some embodiments of the method 800, the control system 112 may analyze the plurality of data streams with a projection layer the robotics transformer model to predict a future action to be performed during the medical procedure. The projection layer for the roboticstransformer model may predict the future action based upon the one or more medical procedure metadata indicators.

[0095] In some embodiments of the method 800, the first machine learning model is configured to output medical procedure metadata indicators indicative of a procedure type. In these embodiments the method 800 may further comprise inputting multi-modal medical environment data into a second machine learning model (e.g., the second machine learning model 403) and receiving one or more medical procedure modality indicators as an output of the second machine learning model. The multi-modal medical environment data may comprise the plurality of data streams; an medical environment sensor data; and / or medical environment sensor data different from the plurality of data streams.

[0096] The second machine learning model is configured to output one or more medical procedure modality indicators indicative of a medical procedure modality based on outputs from a projection layer (e.g., one or more projections layers 404) of the second machine learning model for at least a portion of the multi-modal medical environment data. The second machine learning model is trained using historical multi-modal medical environment data labeled with known medical procedure modality indicators. In some embodiments, the output of the second machine learning model is an input data stream of the first machine learning model. Furthermore, the method 800 may include controlling the computer-assisted system based upon the output medical procedure modality indicators.

[0097] It is understood that the blocks of the methods 700 and 800 need not occur strictly in the order shown.

[0098] Although the systems, methods, devices, and components thereof, have been described in terms of exemplary embodiments, they are not limited thereto. The detailed description is to be construed as exemplary only and does not describe every possible embodiment of the invention because describing every possible embodiment would be impractical, if not impossible. Numerous alternative embodiments could be implemented, using either current technology or technology developed after the filing date of this patent that would still fall within the scope of the claims defining the invention.

[0099] Those skilled in the art will recognize that a wide variety of modifications, alterations, and combinations can be made with respect to the above described embodiments without departing from the scope of the invention, and that such modifications, alterations, and combinations are to be viewed as being within the ambit of the inventive concept.

Claims

What is claimed is:

1. A computer system comprising: one or more processors; and one or more non-transitory, computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to: obtain post-operative medical procedure data representative of a medical procedure that was performed using a computer-assisted system, the post-operative medical procedure data comprising a plurality of data streams from one or more data sources, wherein a first data stream of the plurality of data streams comprises timeseries data relating to operation of the computer-assisted system; input each of the plurality of data streams into a respective projection layer of a first machine learning model, wherein: the projection layer for the first data stream includes an embedding layer and a transformer, the first machine learning model is configured to output one or more procedure metadata indicators based on outputs from the projection layer for at least a portion of the plurality of data streams, and the first machine learning model is trained using historical procedure data labeled with known metadata; receive the one or more medical procedure metadata indicators as an output of the first machine learning model; and associate the post-operative medical procedure data with the output medical procedure metadata indicators.

2. The computer system of claim 1, wherein a second data stream of the plurality of data streams comprises image data.

3. The computer system of claim 2, wherein the image data is endoscope image data or medical environment image data.

4. The computer system of claim 2, wherein the respective projection layer for the second data stream includes a convolutional neural network or a large multimodal model (LMM).

5. The computer system of claim 1, wherein the projection layer for the first data stream is configured to: convert the time-series data into machine readable tokens with embedded positional values, and wherein the projection layer is trained using a plurality of historical time series data streams having associated time stamp information.

6. The computer system of claim 1, wherein the projection layer for the first data stream includes an attention layer trained to modify inputs into the attention layer into refined embeddings that represent contextual relationships between data points in the first data stream.

7. The computer system of claim 6, wherein the inputs into the attention layer comprise outputs of the embedding layer.

8. The computer system of claim 1, wherein the time-series data includes one or more of event data, kinematics data, or force data of the computer-assisted system.

9. The computer system of claim 1, wherein the output medical procedure metadata indicators include indicators indicative of a medical procedure type or indicators indicative of a medical procedure modality.

10. The computer system of claim 1, wherein: the first machine learning model is configured to output medical procedure metadata indicators indicative of a procedure type, and the instructions, when executed by the one or more processors, further cause the one or more processors to: input multi-modal medical environment data into a second machine learning model, wherein: the second machine learning model is configured to output one or more medical procedure modality indicators indicative of a medical procedure modality based on outputs from a projection layer of the second machinelearning model for at least a portion of the multi-modal medical environment data, and the second machine learning model is trained using historical multimodal medical environment data labeled with known medical procedure modality indicators; and receive the one or more medical procedure modality indicators as an output of the second machine learning model.

11. The computer system of claim 10, wherein the output of the second machine learning model is an input data stream of the first machine learning model.

12. The computer system of claim 10, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to: associate the post-operative medical procedure data with the output medical procedure modality indicators.

13. The computer system of claim 12, wherein the associations of the output medical procedure metadata indicators and the output medical procedure modality indicators with the post-operative medical procedure data comprise labeled post-operative medical procedure data.

14. The computer system of claim 13, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to: fine-tune the first machine learning model using the labeled post-operative medical procedure data.

15. The computer system of claim 13, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to: perform a procedure-specific analysis of the labeled post-operative medical procedure data.

16. The computer system of claim 15, wherein the procedure- specific analysis includes an evaluation of a facility or personnel that performed the medical procedure.

17. The computer system of claim 10, wherein the multi-modal medical environment data comprises the post-operative medical procedure data.

18. The computer system of claim 10, wherein the multi-modal medical environment data comprises an medical environment sensor data portion of the post-operative medical procedure data.

19. The computer system of claim 10 wherein the multi-modal medical environment data comprises medical environment sensor data different from the post-operative medical procedure data.

20. A computer-implemented method comprising: obtaining post-operative medical procedure data representative of a medical procedure that was performed using a computer-assisted system, the post-operative medical procedure data comprising a plurality of data streams from one or more data sources, wherein a first data stream of the plurality of data streams comprises time-series data relating to operation of the computer-assisted system; inputting each of the plurality of data streams into a respective projection layer of a first machine learning model, wherein: the projection layer for the first data stream includes an embedding layer and a transformer, the first machine learning model is configured to output one or more procedure metadata indicators based on outputs from the projection layer for at least a portion of the plurality of data streams, and the first machine learning model is trained using historical procedure data labeled with known metadata; receiving the one or more medical procedure metadata indicators as an output of the first machine learning model; and associating the post-operative medical procedure data with the output medical procedure metadata indicators.

21. The computer- implemented method of claim 20, wherein a second data stream of the plurality of data streams comprises image data.

22. The computer-implemented method of claim 21, wherein the image data is endoscope image data or medical environment image data.

23. The computer- implemented method of claim 21, wherein the respective projection layer for the second data stream includes a convolutional neural network or a large multimodal model (LMM).

24. The computer-implemented method of claim 20, wherein the projection layer for the first data stream is configured to convert the time- series data into machine readable tokens with embedded positional values, and wherein the projection layer is trained using a plurality of historical time series data streams having associated time stamp information.

25. The computer- implemented method of claim 20, wherein the projection layer for the first data stream includes an attention layer trained to modify inputs into the attention layer into refined embeddings that represent contextual relationships between data points in the first data stream.

26. The computer- implemented method of claim 25, wherein the inputs into the attention layer comprise outputs of the embedding layer.

27. The computer- implemented method of claim 20, wherein the time-series data includes one or more of event data, kinematics data, or force data of the computer-assisted system.

28. The computer- implemented method of claim 20, wherein the output medical procedure metadata indicators include indicators indicative of a medical procedure type or indicators indicative of a medical procedure modality.

29. The computer- implemented method of claim 20, wherein the first machine learning model is configured to output medical procedure metadata indicators indicative of a procedure type, and the method further comprises: inputting multi-modal medical environment data into a second machine learning model, wherein: the second machine learning model is configured to output one or more medical procedure modality indicators indicative of a medical procedure modalitybased on outputs from a projection layer of the second machine learning model for at least a portion of the multi-modal medical environment data, and the second machine learning model is trained using historical multi-modal medical environment data labeled with known medical procedure modality indicators; and receiving the one or more medical procedure modality indicators as an output of the second machine learning model.

30. The computer-implemented method of claim 29, further comprising including the output of the second machine learning model as an input data stream of the first machine learning model.

31. The computer- implemented method of claim 29, further comprising associating the post-operative medical procedure data with the output medical procedure modality indicators.

32. The computer- implemented method of claim 31, wherein the associations of the output medical procedure metadata indicators and the output medical procedure modality indicators with the post-operative medical procedure data comprise labeled post-operative medical procedure data.

33. The computer- implemented method of claim 32, further comprising further training the first machine learning model using the labeled post-operative medical procedure data.

34. The computer-implemented method of claim 32, further comprising performing a procedure- specific analysis of the labeled post-operative medical procedure data.

35. The computer- implemented method of claim 34, wherein the procedure-specific analysis includes an evaluation of a facility or personnel that performed the medical procedure.

36. The computer-implemented method of claim 29, wherein the multi-modal medical environment data comprises the post-operative medical procedure data.

37. The computer-implemented method of claim 29, wherein the multi-modal medical environment data comprises an medical environment sensor data portion of the post-operative medical procedure data.

38. The computer- implemented method of claim 29 wherein the multi-modal medical environment data comprises medical environment sensor data different from the postoperative medical procedure data.

39. A non-transitory machine-readable medium comprising a plurality of machine- readable instructions that when executed by one or more processors are adapted to cause the one or more processors to perform the method of any one of claims 20-38.

40. A computer-assisted system comprising: one or more repositionable structures configured to support one or more instruments; and a control system operably coupled to the one or more repositionable structures, wherein the control system is configured to: obtain a plurality of data streams from one or more data sources, wherein a first data stream of the plurality of data streams comprises time-series data relating to operation of the computer-assisted system; input each of the plurality of data streams into a respective projection layer of a first machine learning model, wherein: the projection layer for the first data stream includes an embedding layer and a transformer, the first machine learning model configured to output one or more medical procedure metadata indicators based on outputs from the projecting layer for at least a portion of the plurality of data streams, and the first machine learning model is trained using historical medical procedure data labeled with known metadata; receive the one or more medical procedure metadata indicators as an output of the first machine learning model; and control the computer-assisted system based upon the one or more medical procedure metadata indicators.

41. The computer-assisted system of claim 40, wherein: the control system is configured to input at least a portion of the plurality of data streams into a robotics transformer model to generate control outputs for controlling computer assisted system; and controlling the computer-assisted system comprises inputting the medical procedure metadata indicators into an input layer of the robotics transformer model.

42. The computer-assisted system of claim 41, wherein a projection layer for the robotics transformer model is trained to analyze the plurality of data streams to detect anomalous performance.

43. The computer-assisted system of claim 42, wherein a projection layer for the robotics transformer model is trained to detect the anomalous performance based upon the one or more medical procedure metadata indicators.

44. The computer-assisted system of claim 41, wherein a projection layer for the robotics transformer model is trained to analyze the plurality of data streams to predict a future action to be performed during the medical procedure.

45. The computer-assisted system of claim 44 wherein a projection layer for the robotics transformer model is trained to predict the future action based upon the one or more medical procedure metadata indicators.

46. The computer-assisted system of claim 40, wherein a second data stream of the plurality of data streams comprises image data.

47. The computer-assisted system of claim 46, wherein the image data is endoscope image data or medical environment image data.

48. The computer-assisted system of claim 46, wherein the respective projection layer for the second data stream includes a convolutional neural network or a large multimodal model (LMM).

49. The computer-assisted system of claim 40, wherein the projection layer for the first data stream is configured to convert the time-series data into machine readable tokens with embedded positional values, and wherein the projection layer is trained using a plurality of historical time series data streams having associated time stamp information.

50. The computer-assisted system of claim 40, wherein the projection layer for the first data stream includes an attention layer trained to modify inputs into the attention layer into refined embeddings that represent contextual relationships between data points in the first data stream.

51. The computer-assisted system of claim 50, wherein the inputs into the attention layer comprise outputs of the embedding layer.

52. The computer-assisted system of claim 40, wherein the time-series data includes one or more of event data, kinematics data, or force data of the computer-assisted system.

53. The computer-assisted system of claim 40, wherein the output medical procedure metadata indicators include indicators indicative of a medical procedure type or indicators indicative of a medical procedure modality.

54. The computer-assisted system of claim 40, wherein the first machine learning model is configured to output medical procedure metadata indicators indicative of a procedure type, and wherein the control system is configured to: input multi-modal medical environment data into a second machine learning model, wherein: the second machine learning model is configured to output one or more medical procedure modality indicators indicative of a medical procedure modality based on outputs from a projection layer of the second machine learning model for at least a portion of the multi-modal medical environment data, and the second machine learning model is trained using historical multi-modal medical environment data labeled with known medical procedure modality indicators; and receive the one or more medical procedure modality indicators as an output of the second machine learning model.

55. The computer-assisted system of claim 54, wherein the output of the second machine learning model is an input data stream of the first machine learning model.

56. The computer-assisted system of claim 54, wherein the control system is configured to control the computer-assisted system based upon the output medical procedure modality indicators.

57. The computer-assisted system of claim 54, wherein the multi-modal medical environment data comprises the plurality of data streams.

58. The computer-assisted system of claim 54, wherein the multi-modal medical environment data comprises medical environment sensor data.

59. The computer-assisted system of claim 54 wherein the multi-modal medical environment data comprises medical environment sensor data different from the plurality of data streams.

60. A computer-implemented method comprising: obtaining a plurality of data streams from one or more data sources, wherein a first data stream of the plurality of data streams comprises time- series data relating to operation of a computer-assisted system; inputting each of the plurality of data streams into a respective projection layer of a first machine learning model, wherein: the projection layer for the first data stream includes an embedding layer and a transformer, the first machine learning model configured to output one or more medical procedure metadata indicators based on outputs from the projecting layer for at least a portion of the plurality of data streams, and the first machine learning model is trained using historical medical procedure data labeled with known metadata; receiving the one or more medical procedure metadata indicators as an output of the first machine learning model; andcontrolling the computer-assisted system based upon the one or more medical procedure metadata indicators.

61. The computer- implemented method of claim 60, further comprising: inputting at least a portion of the plurality of data streams into a robotics transformer model to generate control outputs for controlling the computer-assisted system; and controlling the computer-assisted system comprises inputting the medical procedure metadata indicators into an input layer of the robotics transformer model.

62. The computer- implemented method of claim 61, further comprising: analyzing the plurality of data streams with a projection layer for the robotics transformer model to detect anomalous performance.

63. The computer- implemented method of claim 62, wherein the projection layer for the robotics transformer model detects the anomalous performance based upon the one or more medical procedure metadata indicators.

64. The computer- implemented method of claim 61, further comprising: analyzing the plurality of data streams with a projection layer for the robotics transformer model to predict a future action to be performed during the medical procedure.

65. The computer- implemented method of claim 64 wherein the projection layer for the robotics transformer model predicts the future action based upon the one or more medical procedure metadata indicators.

66. The computer- implemented method of claim 60, wherein a second data stream of the plurality of data streams comprises image data.

67. The computer- implemented method of claim 66, wherein the image data is endoscope image data or medical environment image data.

68. The computer- implemented method of claim 66, wherein the respective projection layer for the second data stream includes a convolutional neural network or a large multimodal model (LMM).

69. The computer-implemented method of claim 60, wherein the projection layer for the first data stream is configured to convert the time- series data into machine readable tokens with embedded positional values, and wherein the projection layer is trained using a plurality of historical time series data streams having associated time stamp information.

70. The computer-implemented method of claim 60, wherein the projection layer for the first data stream includes an attention layer trained to modify inputs into the attention layer into refined embeddings that represent contextual relationships between data points in the first data stream.

71. The computer- implemented method of claim 70, wherein the inputs into the attention layer comprise outputs of the embedding layer.

72. The computer-implemented method of claim 60, wherein the time-series data includes one or more of event data, kinematics data, or force data of the control system.

73. The computer- implemented method of claim 60, wherein the output medical procedure metadata indicators include indicators indicative of a medical procedure type or indicators indicative of a medical procedure modality.

74. The computer-implemented method of claim 60, wherein the first machine learning model is configured to output medical procedure metadata indicators indicative of a procedure type, and further comprising: inputting multi-modal medical environment data into a second machine learning model, wherein: the second machine learning model is configured to output one or more medical procedure modality indicators indicative of a medical procedure modality based on outputs from a projection layer of the second machine learning model for at least a portion of the multi-modal medical environment data, and the second machine learning model is trained using historical multimodal medical environment data labeled with known medical procedure modality indicators; andreceiving the one or more medical procedure modality indicators as an output of the second machine learning model.

75. The computer- implemented method of claim 74, further comprising including the output of the second machine learning model as an input data stream of the first machine learning model.

76. The computer-implemented method of claim 74, further comprising: controlling the computer-assisted system based upon the output medical procedure modality indicators.

77. The computer-implemented method of claim 74, wherein the multi-modal medical environment data comprises the plurality of data streams.

78. The computer- implemented method of claim 74, wherein the multi-modal medical environment data comprises medical environment sensor data.

79. The computer- implemented method of claim 74 wherein the multi-modal medical environment data comprises medical environment sensor data different from the plurality of data streams.

80. A non-transitory machine-readable medium comprising a plurality of machine- readable instructions that when executed by one or more processors are adapted to cause the one or more processors to perform the method of any one of claims 60-79.

Citation Information

Patent Citations

  • System and method for connecting garments and related articles

    US62636635P0

  • Detecting events during a surgery

    US20220129822A1