Systems and methods for automatically generating evaluation and guidance for human machine interactions

A machine learning model processes multi-modal data to automate evaluation and guidance for human-machine interactions, addressing the inefficiencies of manual expert reviews and enhancing interaction management in computer-assisted systems.

WO2026030507A1PCT designated stage Publication Date: 2026-02-05INTUITIVE SURGICAL OPERATIONS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/039990
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-01
Filing Date
2025-07-31
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

The management and monitoring of human interactions with computer-assisted systems, such as MRI machines and robotic systems, are costly and ineffective due to the reliance on manual expert reviews, which lack insight into specific human-machine interactions.

Method used

A computer system utilizing a machine learning model that processes multi-modal data streams from human-machine interactions to generate automated evaluations and guidance, incorporating a rubric and reference documents to assess and improve interaction quality.

Benefits of technology

Facilitates accurate and efficient monitoring and management of human-machine interactions by providing real-time feedback and recommendations, reducing the need for manual expert intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025039990_05022026_PF_FP_ABST
    Figure US2025039990_05022026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for systems and methods for automatically generating evaluation and guidance for human machine interactions are provided. The techniques may be performed on human-machine interaction data to evaluate a past human-machine interaction or on intra-human-machine interaction data to provide guidance during the human-machine interaction. Regardless, the techniques may involve obtaining the human-machine interaction data and inputting the data into respective projection layers of a machine learning model. The machine learning model may be configured to output the evaluation or guidance based on outputs from the projection layers, a rubric, and one or more reference documents. In the evaluation scenario, the evaluation may be associated with the human-machine interaction event. In the guidance scenario, the guidance may be provided to an operator performing the human-machine interaction.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR AUTOMATICALLY GENERATINGEVALUATION AND GUIDANCE FOR HUMAN MACHINE INTERACTIONSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of the filing date of provisional U.S. Patent Application No. 63 / 678,437 entitled “SYSTEMS AND METHODS FOR AUTOMATICALLY GENERATING EVALUATION AND GUIDANCE FOR HUMAN MACHINE INTERACTIONS,” filed on August 1, 2024. The entire contents of the provisional application are hereby expressly incorporated herein by reference.FIELD

[0002] The present disclosure relates generally to human-machine interactions and more particularly to training and utilizing artificial intelligence to automatically generate evaluation and guidance for human machine interactions from multi-modal data streams associated with the interactions.BACKGROUND

[0003] Computer-assisted systems have become more ubiquitous as the cost of computing and manufacturing has decreased. Such systems can include medical procedure equipment such as magnetic resonance imaging (MRI) machines, X-ray machines, robotically assisted systems or robotic systems, etc. The robotically assisted systems or robotic systems may include one or more manipulators that can be operated with the assistance of an electronic controller (e.g., computer or control system) to move and control functions of one or more instruments coupled to the manipulators. A manipulator generally includes mechanical links connected by joints. An instrument is removably (or permanently) coupled to one of the links, typically a distal link of the plural links. In some embodiments, manipulator systems are used in conjunction with one or more auxiliary devices (e.g., a surgical bed, an insufflator, etc.).

[0004] As the use of these computer-assisted systems has proliferated the ability to manage and monitor human interactions with such systems has become more complex and timeconsuming. Some example types of interactions include presenting and / or demonstrating capabilities of the computer-assisted system, training others to use computer-assisted systems, and / or troubleshooting issues with a computer-assisted system. The typical methods for managing and monitoring such interactions by utilizing trained experts to manually review post-interaction data to provide evaluations and / or to be present during the interactions for intra-interaction guidance have proven to be cost prohibitive and ineffectiveas the number of interactions that need monitoring and management has increased. Moreover, the experts may not have insight into the specific human-machine interactions being performed by the operator to provide appropriate evaluations or guidance.

[0005] Accordingly, there is a need for systems and methods that accurately and automatically monitor and / or manage human interactions with computer-assisted systems. In particular, there is a need for systems that can generate interaction evaluations and guidance both post-interaction and intra-interaction.SUMMARY

[0006] In some aspects, the techniques described herein relate to a computer system including: (i) one or more processors; and (ii) one or more non-transitory, computer-readable media storing instructions that, when executed by the one or more processors, cause the computer system to: (1) obtain human-machine interaction data representative of a humanmachine interaction event with respect to a computer-assisted system, wherein: (a) the human-machine interaction data includes a plurality of data streams from one or more data sources, (b) a first data stream of the plurality of data streams includes time-series data relating to operation of the computer-assisted system during the human-machine interaction event, and (c) a second data stream of the plurality of data streams is indicative of language presented as part of the human-machine interaction event; (2) input each of the plurality of data streams into a projection layer of a machine learning model, wherein the machine learning model is configured to output an evaluation based on an evaluation rubric, one or more reference documents, and outputs from the projection layers for at least a portion of the plurality of data streams; (3) receive an evaluation for the human-machine interaction event as an output of the machine learning model; and (4) associate the output evaluation with the human-machine interaction event.

[0007] In some additional aspects, the techniques described herein relate to a computer- implemented method including: (1) obtaining human-machine interaction data representative of a human-machine interaction event with respect to a computer-assisted system, wherein: (a) the human-machine interaction data includes a plurality of data streams from one or more data sources, (b) a first data stream of the plurality of data streams includes time-series data relating to operation of the computer-assisted system during the human-machine interaction event, and (c) a second data stream of the plurality of data streams is indicative of language presented as part of the human-machine interaction event; (2) inputting each of the plurality of data streams into a projection layer of a machine learning model, wherein the machinelearning model is configured to output an evaluation based on an evaluation rubric, one or more reference documents, and outputs from the projection layers for at least a portion of the plurality of data streams; (3) receiving an evaluation for the human-machine interaction event as an output of the machine learning model; and (4) associating the output evaluation with the human-machine interaction event.

[0008] In some yet further aspects, the techniques described herein relate to a computer- assisted system including: (i) a manipulator system; and (ii) a control system operably coupled to the manipulator system, wherein the control system is configured to: (1) obtain a plurality of data streams from one or more data sources, wherein: (a) a first data stream of the plurality of data streams includes time-series data relating to operation of the manipulator system or the control system as part of a human-machine interaction event for the computer- assisted system, and (b) a second data stream of the plurality of data streams is indicative of language presented as part of the human-machine interaction event; (2) input each of the plurality of data streams into a respective projection layer of a machine learning model, wherein the machine learning model is configured to output guidance based on a guidance rubric, one or more reference documents, and outputs from the projection layers for at least a portion of the plurality of data streams; (3) receive guidance as an output of the machine learning model, wherein the guidance includes recommendations to conform the humanmachine interaction event to expected baselines; and (4) present the guidance on a display device operably coupled to the control system.

[0009] In some even further aspects, the techniques described herein relate to a computer- implemented method including: (1) obtaining a plurality of data streams from one or more data sources, wherein: (a) a first data stream of the plurality of data streams includes timeseries data relating to operation of a manipulator system by a control system as part of a human-machine interaction event for a computer-assisted system, and (b) a second data stream of the plurality of data streams is indicative of language presented as part of the human-machine interaction event; (2) inputting each of the plurality of data streams into a respective projection layer of a machine learning model, wherein the machine learning model is configured to output guidance based on a guidance rubric, one or more reference documents, and outputs from the projection layers for at least a portion of the plurality of data streams; (3) receiving guidance as an output of the machine learning model, wherein the guidance includes recommendations to conform the human-machine interaction event to expected baselines; and (4) presenting the guidance on a display device operably coupled to the control system.

[0010] In some furth aspects, a non-transitory machine-readable medium comprising a plurality of machine-readable instructions that when executed by one or more processors are adapted to cause the one or more processors to perform any of the methods described herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] FIG. 1A is a block diagram of a computing system in accordance with one or more embodiments.

[0012] FIG. IB is a block diagram of an artificial intelligence model in accordance with one or more embodiments.

[0013] FIG. 1C is a block diagram of an artificial intelligence model in accordance with one or more embodiments.

[0014] FIG. 2 is a diagram for a computer-assisted system in accordance with one or more embodiments.

[0015] FIG. 3 is a flow diagram of a method in accordance with one or more embodiments.

[0016] FIG. 4 is a flow diagram of a method in accordance with one or more embodiments.

[0017] FIG. 5 is a flow diagram of a method in accordance with one or more embodiments.

[0018] FIG. 6 is a flow diagram of a method in accordance with one or more embodiments.

[0019] FIG. 7 is a flow diagram of a method in accordance with one or more embodiments.

[0020] Examples of the present disclosure and their advantages are best understood by referring to the detailed description that follows. It should be appreciated that like reference numerals are used to identify like elements illustrated in one or more of the figures, wherein showings therein are for purposes of illustrating examples of the present disclosure and not for purposes of limiting the same.DETAILED DESCRIPTION

[0021] In this description, specific details are set forth describing some embodiments consistent with the present disclosure. Numerous specific details are set forth in order to provide a thorough understanding of the embodiments. It will be apparent, however, to one skilled in the art that some embodiments may be practiced without some or all of these specific details. The specific embodiments disclosed herein are meant to be illustrative but not limiting. One skilled in the art may realize other elements that, although not specifically described here, are within the scope and the spirit of this disclosure. In addition, to avoid unnecessary repetition, one or more features shown and described in association with one embodiment may be incorporated into other embodiments unless specifically described otherwise or if the one or more features would make an embodiment non-functional.

[0022] Further, the terminology in this description is not intended to limit the invention. For example, spatially relative terms-such as “beneath”, “below”, “lower”, “above”, “upper”, “proximal”, “distal”, and the like-may be used to describe the relation of one element or feature to another element or feature as illustrated in the figures. These spatially relative terms are intended to encompass different positions (z.e., locations) and orientations (z.e., rotational placements) of the elements or their operation in addition to the position and orientation shown in the figures. For example, if the content of one of the figures is turned over, elements described as “below” or “beneath” other elements or features would then be “above” or “over” the other elements or features. A device may be otherwise oriented and the spatially relative descriptors used herein interpreted accordingly. Likewise, descriptions of movement along and around various axes include various special element positions and orientations. In addition, the singular forms “a”, “an”, and “the” are intended to include the plural forms as well, unless the context indicates otherwise. Additionally, the terms “comprises”, “comprising”, “includes”, and the like specify the presence of stated features, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups. Components described as coupled may be electrically or mechanically directly coupled, or they may be indirectly coupled via one or more intermediate components.

[0023] Elements described in detail with reference to one embodiment, implementation, system, or module may, whenever practical, be included in other embodiments, implementations, systems, or modules in which they are not specifically shown or described. For example, if an element is described in detail with reference to one embodiment and is not described with reference to a second embodiment, the element may nevertheless be claimed as included in the second embodiment. Thus, to avoid unnecessary repetition in the following description, one or more elements shown and described in association with one embodiment, implementation, or application may be incorporated into other embodiments, implementations, or aspects unless specifically described otherwise, unless the one or more elements would make an embodiment or implementation non-functional, or unless two or more of the elements provide conflicting functions.

[0024] In some instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0025] This disclosure describes various devices, elements, and portions of computer- assisted systems and elements in terms of their state in three-dimensional space. As used herein, the term “position” refers to the location of an element or a portion of an element(e.g., three degrees of translational freedom in a three-dimensional space, such as along Cartesian x-, y-, and z-coordinates). As used herein, the term “orientation” refers to the rotational placement of an element or a portion of an element (e.g., three degrees of rotational freedom in three-dimensional space, such as about roll, pitch, and yaw axes, represented in angle-axis, rotation matrix, quaternion representation, and / or the like). As used herein, and for a device with a kinematic series, such as with a repositionable structure with a plurality of links coupled by one or more joints, the term “proximal” refers to a direction toward a base of the kinematic series, and “distal” refers to a direction away from the base along the kinematic series.

[0026] As used herein, the term “pose” refers to the multi-degree of freedom (DOF) spatial position and orientation of a coordinate system of interest attached to a rigid body. In general, a pose includes a pose variable for each of the DOFs in the pose. For example, a full 6-DOF pose for a rigid body in three-dimensional space would include 6 pose variables corresponding to the 3 positional DOFs (e.g., x, y, and z) and the 3 orientational DOFs (e.g., roll, pitch, and yaw). A 3-DOF position only pose would include only pose variables for the 3 positional DOFs. Similarly, a 3-DOF orientation only pose would include only pose variables for the 3 rotational DOFs. Further, a velocity of the pose captures the change in pose over time (e.g., a first derivative of the pose). For a full 6-DOF pose of a rigid body in three- dimensional space, the velocity would include 3 translational velocities and 3 rotational velocities. Poses with other numbers of DOFs would have a corresponding number of velocities translational and / or rotational velocities.

[0027] This disclosure occasionally refers to the disclosed techniques being applied to “patients” undergoing a “medical procedure.” It should be appreciated that these references are not intended to limit the application of the disclosed techniques to applied medicine contexts. For example, the described techniques can be applied to facilitate physician training, equipment testing and / or calibration, and / or other contexts. Accordingly, any reference to the term “patient” is done for ease of explanation and also envisions the application of the described techniques to a generic “subject.”

[0028] The term “human-machine interaction” may refer to any interaction associated with a human and a computer-assisted system. In some portions of the interaction, an operator may operate or interact with the computer-assisted system (e.g., by demonstrating a capability of the computer-assisted system, performing a training task, etc.). In other portions of the interaction, an operator may verbally describe features and / or describe tasks that need to be performed with the computer-assisted system. It should be appreciated that the “human-machine interaction” may refer to an event where multiple component interactions are to occur. For example, a human-machine interaction as part of a demonstration of the capabilities of the computer-assisted system may include component interactions associated with each capability and / or facet to demonstrate.

[0029] The word “task” is used herein to refer to a discrete portion of a procedure that may be autonomously, semi-autonomously, or manually implemented in furtherance of the procedure. For example, a task may be to move an endoscope to a particular position, to advance an instrument to a particular depth, to replace an instrument coupled to a manipulator, and so on. In some embodiments, a task is associated with component tasks to accomplish an overall goal. For example, a task to analyze a worksite may include component tasks related to moving an endoscope to view the worksite, advancing an instrument to predetermined depth, and enabling a functionality supported by the instrument. These component tasks may also be referred to as “subtasks.”

[0030] Aspects of this disclosure are described in reference to computer-assisted systems, which can include devices that are teleoperated, externally manipulated, autonomous, semiautonomous, and / or the like. Further, aspects of this disclosure are described in terms of an implementation using a teleoperated surgical system, such as the da Vinci® Surgical System commercialized by Intuitive Surgical, Inc. of Sunnyvale, California. Knowledgeable persons will understand, however, that inventive aspects disclosed herein may be embodied and implemented in various ways, including teleoperated and non-teleoperated, and medical and non-medical embodiments and implementations. Implementations on da Vinci® Surgical Systems are merely exemplary and are not to be considered as limiting the scope of the inventive aspects disclosed herein. For example, techniques described with reference to surgical instruments and surgical methods may be used in other contexts. Thus, the instruments, systems, and methods described herein may be used for humans, animals, portions of human or animal anatomy, industrial systems, general robotic, or teleoperated systems. As further examples, the instruments, systems, and methods described herein may be used for non-medical purposes including industrial uses, general robotic uses, sensing or manipulating non-tissue work pieces, cosmetic improvements, imaging of human or animal anatomy, gathering data from human or animal anatomy, setting up or taking down systems, training medical or non-medical personnel, and / or the like. Additional example applications include use for procedures on tissue removed from human or animal anatomies (with or without return to a human or animal anatomy) and for procedures on human or animalcadavers. Further, these techniques can also be used for medical treatment or diagnosis procedures that include, or do not include, surgical aspects.

[0031] FIG. 1A illustrates a system 100. The system 100 includes a computing system 102 that includes a processing unit 104 and a memory unit 106. The processing unit 104 includes one or more processors, each of which may be a programmable microprocessor or the like that executes software instructions stored in memory unit 106 to execute some or all of the functions of the system 100 as described herein. Processing unit 104 may include one or more graphics processing units (GPUs) and / or one or more central processing units (CPUs), for example. Alternatively, or in addition, one or more processors in processing unit 104 may be other types of processors (e.g., application- specific integrated circuits (ASICs), field- programmable gate arrays (FPGAs), etc.), and some of the functionality of the system 100 as described herein may instead be implemented in hardware.

[0032] Memory unit 106 may include one or more volatile and / or non-volatile memories. Any suitable memory type or types may be included in memory unit 106, such as read-only memory (ROM) and / or random-access memory (RAM), flash memory, a solid-state drive (SSD), a hard disk drive (HDD), and so on. Collectively, memory unit 106 may store one or more software applications, the data received / used by those applications, and the data output / generated by those applications.

[0033] In particular, memory unit 106 stores the software that, when executed by processing unit 104, perform various functions of the system 100 related to monitoring and / or managing interactions between a human and a computer-assisted system 108. The human interacting with the computer-assisted system 108 can include an operator O of the system 100 and / or a distinct presenter for an underlying human-machine interaction event.

[0034] As shown in FIG. 1A, the computing system 102 is configured to obtain humanmachine interaction data 110 representative of a human-machine interaction event with respect to the computer-assisted system 108. The human-machine interaction data 110 may comprise post-human-machine interaction data or intra-human-machine interaction data. In either case, the human-machine interaction data 110 may comprise sensor data, video data, etc. documenting how the computer-assisted system 108 is being operated during the interaction and actions, statements, movement, etc. of the operator O or other presenter.

[0035] The computing system 102 is further configured to input the human-machine interaction data 110 into a machine learning model 112. The machine learning model 112 is configured to generate an output 114 based on a rubric 116, one or more reference documents 118 and the human-machine interaction data 110. The rubric 116 may include a set ofinstructions of context setting inputs that define how the machine learning model 112 is to generate the output 114. The one or more reference documents 118 may include one or more of product user manuals, demonstrator training documents, instrument documentation, clinical documents, operational documents, procedural guidelines, troubleshooting guidelines, or human-machine interaction instructions.

[0036] When utilized in the various manners described herein, the rubric 116 and one or more reference documents 118 may individually or together define expected baselines, standards, and / or other expectations for the human-machine interaction event, particular facets to be presented during the event, how to evaluate examples presented during the human-machine interaction event, etc. For example, the rubric 116 and one or more reference documents 118 may individually or together define metrics for quality and effectiveness of the interaction in relation to different portions of the human-machine interaction data 110 that can be measured, assessed, and analyzed (such as accuracy or truthfulness of content delivered; completeness of content delivered; flow / sequence of content delivered, and actions performed; question answering abilities; etc.). These variables can be further divided into categories such as hard metrics and soft metrics. The hard metrics may include assessments of: technical knowledge and experience about the computer-assisted system 108; descriptions of new features of the computer-assisted system 108 compared to other systems; descriptions of advantages of the computer-assisted system 108 compared to other systems; clinical or operational knowledge about procedures, workflows, etc. performed using the computer- assisted system 108; financial knowledge about the computer-assisted system 108; overall structure of the human-machine interaction event; (when to say what, what to prioritize first, what’s the right flow) an ability to answer questions about the computer-assisted system 108; etc. The soft metrics may include assessments of: speed of the interaction event (e.g., how fast or slow the delivery of the operator O or presenter is and whether the delivery is understandable); body language of the operator O or other presenter (e.g., whether the body movements indicate a trustworthy, friendly, approachable, etc. disposition); the location and amount of silences in the interaction event; how enthusiastic the operator O or other presenter appears during the interaction event; tone of voice and word choices of the operator O or other presenter; etc.

[0037] As shown in FIG. 1A, the system 100 may include a display device 119 that is operably coupled to the computing system 102. The display device 119 may be configured to receive and display the output 114 and / or display information that is based on the output 114. The display device 119 may be included in a virtuality reality (VR) headset, an augmentedreality (AR) headset, a mixed reality (MR) headset, or smart glasses. Furthermore, in some embodiments, some of the human-machine interaction data 110 may be recorded by the display device 119.

[0038] The machine learning model 112 may be executable by the computing system 102, such as by the processing unit 104 and may comprises a set of interconnected nodes, layers, trained parameter values (e.g., multiplicative weights, additive bias, etc.), etc. The trained parameters may be set via backpropagation techniques in a training process that uses historical data inputs as described in more detail in connection with FIG. 1C. In some embodiments, the machine learning model 112 is configured using a multi-modal transformer type architecture that can receive inputs of different modalities relating to the operation of the computer-assisted system 108 and actions of the operator O in order to generate the output 114 as described herein. However, it should be appreciated that this architecture may be substituted or supplemented with other Al architectures including, but not limited to, convolutional neural network (CNN) architectures, recurrent / recursive neural network (RNN) architectures, sorting / clustering architectures, etc.

[0039] The output 114 may comprise post-interaction evaluations of the human-machine interaction event and / or intra-interaction guidance as described in more detail herein. When the output 114 is intra-interaction guidance, the computing system 102 may be configured to present the guidance on the display device 119. The guidance may include a list of a list of tasks to perform and / or suggestions for how to modify a current presentation style of the operator O or other presenter to conform to expected baselines as documented in the rubric 116 and / or one or more reference documents 118. The list of tasks may include tasks to present a facet of the computer-assisted system 108 such as a diagnostic task to perform. Such a diagnostic task may be a task configured to assist in identifying possible malfunctions in the computer-assisted system 108. In some embodiments, in order to generate the list of tasks, the machine learning model 112 is configured to detect anomalous behavior from the expected baselines defined by the rubric 116 and generate tasks to correct the anomalous behavior.

[0040] Where the output 114 is a post-interaction evaluation, the computing system 102 may be configured to associate the output 114 with the human-machine interaction event and the rubric 116 may comprise an evaluation rubric. In these embodiments, the post-interaction evaluation may include one or more scoring metrics that indicate a degree to which the human-machine interaction event conformed to expected baselines directed by the rubric 116. The scoring metrics may include a total overall score for the interaction event and / or sub-scores for different elements of the interaction in relation to different category or assessment areas as described herein. The scoring metrics may include a percentage or other similar numerical value documenting the degree of conformity to the expected baselines. The expected baselines may include a set of facets to be presented during the human-machine interaction event as described in more detail herein and / or other expectations for the interaction event. In some embodiments, the post-interaction evaluation may include recommendations for improving future human-machine interactions to conform to expected baselines directed by the rubric 116. The recommendations may include citations to learning content and exercises included in the one or more reference documents 118.

[0041] As shown in FIG. IB, the human-machine interaction data 110 may comprise a plurality of different data streams 110A-1 ION from one or more data sources (including data sources having different modalities). A first data stream 110A may comprise time-series data relating to operation of the computer-assisted system 108 during the human-machine interaction event.

[0042] A second data stream HOB may be indicative of language presented as part of the human-machine interaction event. For example, the second data stream HOB may comprise audio data, audio to text data, etc. for language presented by the operator O or another presenter of the human-machine interaction event. In some embodiments, the display device 119 may be configured to capture the second data stream HOB.

[0043] A third data stream 110C may comprises image data. In particular, the image data can include environmental image data such as an external video data stream, such as a video data stream from a video camera located within the environment (e.g., affixed to a wall or worn or otherwise carried by a person), showing an environment in which the humanmachine interaction is taking place. The environment may include a medical environment (e.g., operating room), a training location for the computer-assisted system 108, a sales location for the computer-assisted system 108, etc. The image data may also include a procedure video data stream, such as an endoscope image data stream. In some embodiments, the procedure video data stream may be a distinct data stream from the environmental image data. The image data can also include image data of the operator O or other presenter of the human-machine interaction event.

[0044] The plurality of data streams may include a plurality of additional data streams 110N that may have data modalities the same as or different from the first data stream 110A, the second data stream 110B, and the third data stream 110C. For example, in some embodiments, the plurality of additional data streams 110N may include text data, audio data,2D data from sensors, time-of-flight data, etc. Furthermore, in some embodiments, the plurality of additional data streams 110N may include data streams from sensors of the computer-assisted system 108 that are configured to generate additional environmental data for the environment in which the operator O and / or the computer-assisted system 108 are located. This additional environmental data may include video / image data and / or depth / point cloud data generated from DeepSight or other type of ultrasound sensors. The depth / point cloud data may include three-dimensional point cloud data.

[0045] In some embodiments, data elements in each of the plurality of data streams are aligned to associate data elements related to a common action represented in the humanmachine interaction data 110 (e.g., an operation made by the computer-assisted system 108 and associated commentary about the operation made by the operator O or another presenter).

[0046] As shown in FIG. IB, the machine learning model 112 includes a plurality of projection layers 120. The plurality of projection layers 120 are configured to receive the human-machine interaction data 110 and convert the underlying data into machine readable formats suitable for processing by the machine learning model 112. The machine-readable formats can include token elements representative of the human-machine interaction data 110. In some embodiments, the token elements are natural language descriptions of the corresponding data stream. In some embodiments, inputting the plurality of data streams into the machine learning model 112 includes inputting the aligned data elements of the data streams.

[0047] The plurality of projection layers 120 may include a distinct layer for each of the plurality of data streams or, alternatively, some of the plurality of projection layers 120 can be configured to receive multiple data streams of the plurality of data streams such as ones having the same modality. For example, a first projection layer 120A can be configured to receive the first data stream 110A, a second projection layer 120B can be configured to receive the second data stream 110B, a third projection layer 120C can be configured to receive the third data stream 110C, and corresponding projection layers 120N can be configured to receive the additional data streams 110N. However, in cases where some of the plurality of projection layers 120 are configured to receive multiple data streams, the first projection layer 120A may be configured to receive multiple time-series type data streams and the third projection layer 120C may be configured to receive multiple image type data streams.

[0048] In either case, each of the plurality of data streams may be input into one of the plurality of projection layers 120 and passed onto additional layers 122 before being receivedby an output layer. The output layer aggregates outputs from each of the plurality of projection layers 120 and the additional layers 122 to generate the output 114. As such, the output 114 is based on outputs from the projection layers 120 for at least a portion of the plurality of data streams.

[0049] The plurality of projection layers 120 may comprise various ML type layers such as embedding layers, transformers, CNNs, etc. that process the input data streams of the humanmachine interaction data 110 into values useable by the additional layers 122 of the machine learning model 112. The plurality of projection layers 120 and / or the additional layers 122 may include self-attention and / or cross attention layers. These attention layers are trained to modify received inputs into refined embeddings that represent contextual relationships between the data points in the input data stream (self-attention) or between different data streams (cross attention).

[0050] In general, the third projection layer 120C for the image data stream 202C is configured to transform the image data into embedded data useable by the remaining portions of the machine learning model 112. For example, the third projection layer 120C can include a convolutional neural network (CNN) or a large multimodal model (LMM) trained to process and classify image data.

[0051] The additional layers 122 may comprise various Al layers known in the art. Such layers include additional multiheaded self-attention or cross-attention layers, feed forward layers, etc. In some embodiments, the additional layers 122 may be omitted and the output 114 may be directly connected to the plurality of projection layers 120. However, in other embodiments, the additional layers 122 and may comprise layers of a pretrained Al model such as a large language model (LLM) or similar. This pretrained Al model can include third party provided models that are either fine-tuned to process the human-machine interaction data 110 or that are used as is without any additional tunning or training.

[0052] When the pretrained Al model utilized is an LLM, the plurality of projection layers 120 are configured to transform and / or project the input plurality of data streams into machine readable tokens representing textual descriptions of the underlying data elements that can be fed into the general word embedding layers of the pretrained LLM. The plurality of projection layers 120 may be configured to generate these textual descriptions via dedicated training processes. For example, the first projection layer 120A for the first data stream 110A may be configured to modify the time- series data into machine readable tokens that represent text descriptions of system events corresponding to the time-series data. Furthermore, in some embodiments, the text descriptions output by one or more of theplurality of projection layers 120 may be cumulative descriptions and classifications of the input data stream or streams. For example, in some embodiments, one of the plurality of projection layers 120 may include a behavior analyzer configured to output machine readable tokens that represent text descriptions of a behavior of the operator O or other presenter for the human-machine interaction event. In these embodiments, the input to the behavior analyzer may includes the second data stream 110B and the third data stream 110C. It should be appreciated that in some cases different data streams of the human-machine interaction data 110 may be used multiple times as the inputs into different ones of the plurality of projection layers 120.

[0053] In some embodiments, such as those where the additional layers 122 comprise a pretrained LLM, the rubric 116 and the one or more reference documents 118 may be directly input into the additional layers 122 as shown in FIG. IB. However, it should be appreciated that in other cases, one or both of the rubric 116 and the one or more reference documents 118 may be input into associated ones of the plurality of projection layers 120 before proceeding to the additional layers 122. However, in either case, the computing system 102 may be configured to obtain the one or more reference documents 118 (e.g., to select from a database or other data storage system) based upon at least a portion of the plurality of data streams of the human-machine interaction data 110 and the rubric 116. For example, the particular one or more reference documents 118 that are obtained may be different depending on a condition indicated by the plurality of data streams, a type of the human-machine interaction, a model type of the computer-assisted system 108, a procedure being demonstrated with the computer-assisted system 108, etc. Accordingly, the computing system 102 may be configured to implement retrieval-augmented generation (RAG) techniques to obtain the rubric 116 and / or the one or more reference documents 118. Furthermore, in these embodiments, the rubric 116 may define and / or direct how the machine learning model 112 is to utilize the one or more reference documents 118 to generate the output 114.

[0054] FIG. 1C is a block diagram showing a training process of the machine learning model 112. Training data 124 is generated and fed into an initialized version of the machine learning model 112 (e.g., a version where the parameter values are randomized) and processed through the plurality of projection layers 120 and the additional layers 122 until a training version of the output 114 is generated. The training version of the output 114 may then be compared to ground truth data 126 to identify an error 128 associated with the training version of the output 114. The error 128 may then be used to update each parameter value in the machine learning model 112 using backpropagation techniques as known in theart. This process of generating training output values from the training data 124, and backpropagating the resulting error 128 to update the parameter values of the machine learning model 112 may be repeated until a threshold condition indicating validity of the machine learning model 112 is achieved. This threshold can include a predetermined or minimum number of iterations through the process and / or value indicating negligible improvement to the error 128 value from further updates to the parameter values of the machine learning model 112.

[0055] As shown in FIG. 1C, in some embodiments, the training data 124 may be generated from at least a portion of the one or more reference documents 118 and a portion of historical human-machine interaction data 130 and another portion of the one or more reference documents 118 and the historical human-machine interaction data 130 may be used to generate the ground truth data 126. In these embodiments, the training of the machine learning model 112 may be considered to be “self- supervised” in that the ground truth data 126 are next token values expected from the input training data 124 used to calculate error for backpropagation. In this manner the one or more reference documents 118 and / or the historical human-machine interaction data 130 may be consumed into a latent space of the machine learning model 112. It should be appreciated that, in some embodiments, the training data 124 and ground truth data 126 generated for the self- supervised learning process may include only data from the historical human-machine interaction data 130.

[0056] However, additional supervised training may also be performed for the machine learning model 112. The supervised training may be accomplished with labeled versions of the historical human-machine interaction data 130. For example, the historical humanmachine interaction data 130 may be associated (e.g., “labeled”) with ground truth data 126 that comprises known evaluations and / or known guidance. Furthermore, it should be appreciated that the other tuning and training techniques may also be utilized such as unsupervised learning, semi-supervised learning, reinforcement learning etc.

[0057] Furthermore, the machine learning model 112 may be further tuned and updated based on data compiled during operation of the system 100 to provide evaluations and / or guidance to a user. In particular, this data can include additional human-machine interaction data, the evaluations and / or guidance that was provided to the user, and feedback on the effectiveness of the guidance that are used to improve the accuracy and effectiveness of the evaluations and / or guidance that are output from the machine learning model 112.

[0058] FIG. 2 illustrates an embodiment of the computer-assisted system 108. In particular, FIG. 2 shows a manipulator system embodiment of the computer-assisted system 108. Thesystem can be used, for example, in surgical, diagnostic, therapeutic, biopsy, or non-medical procedures, and is generally indicated by the reference numeral 108. As shown in FIG. 2, the computer-assisted system 108 can include one or more manipulator assemblies 202 for operating one or more medical instrument systems 204 in performing various procedures on a patient P positioned on a table T in a medical environment 201. For example, the manipulator assembly 202 can drive catheter or end effector motion, can apply treatment to target tissue, and / or can manipulate control members. The manipulator assembly 202 can be teleoperated, non-teleoperated, or a hybrid teleoperated and non-teleoperated assembly with select degrees of freedom of motion that can be motorized and / or teleoperated and select degrees of freedom of motion that can be non-motorized and / or non-teleoperated. An operator input system 206, which can be inside or outside of the medical environment 201, generally includes one or more input devices for controlling manipulator assembly 202. Manipulator assembly 202 supports medical instrument system 204 and can optionally include a plurality of actuators or motors that drive inputs on medical instrument system 204 in response to commands from a control system 212. The actuators can optionally include drive systems that when coupled to medical instrument system 204 can advance medical instrument system 204 into a naturally or surgically created anatomic orifice. Other drive systems can position, pose, or otherwise move the distal end of medical instrument in multiple orientations or degrees of freedom, which can include three degrees of linear motion (e.g., linear motion along the X, Y, Z Cartesian axes) and in three degrees of rotational motion (e.g., rotation about the X, Y, Z Cartesian axes). The manipulator assembly 202 can support various other systems for irrigation, treatment, or other purposes. Such systems can include fluid systems (including, for example, reservoirs, heating / cooling elements, pumps, and valves), generators, lasers, interrogators, and ablation components.

[0059] Computer-assisted system 108 also includes a display system 210 for displaying an image or representation of the surgical site and medical instrument system 204 generated by an imaging system 209 which can include an imaging system, such as an endoscopic imaging system, and / or an imaging system configured to capture video data of an environment in which the operator O and / or the computer-assisted system 108 is disposed. In some embodiments, an image sensor of the imaging system 209 is included in a wearable device associated with the operator O. The display system 210 may include the display device 119 of FIG. 1A. The outputs of the imaging system 209 can comprise a portion of the humanmachine interaction data 110 as described in more detail herein (e.g., the third data stream 110C). Display system 210 and operator input system 206 can be oriented so the operator Ocan control medical instrument system 204 and operator input system 206 with the perception of telepresence. A graphical user interface can be displayable on the display system 210 and / or a display system of an independent planning workstation.

[0060] In some examples, the endoscopic imaging system components of the imaging system 209 can be integrally or removably coupled to medical instrument system 204. However, in some examples, a separate imaging device, such as an endoscope, attached to a separate manipulator assembly can be used with medical instrument system 204 to image the surgical site. The endoscopic imaging system 209 can be implemented as hardware, firmware, software, or a combination thereof which interact with or are otherwise executed by one or more computer processors, which can include a processor system 214 of the control system 212.

[0061] Computer-assisted system 108 can also include a sensor system 208. The sensor system 208 can include a position / location sensor system (e.g., an actuator encoder or an electromagnetic (EM) sensor system) and / or a shape sensor system (e.g., an optical fiber shape sensor) for determining the position, orientation, speed, velocity, pose, and / or shape of the medical instrument system 204. The sensor system 208 can also include temperature, pressure, force, or contact sensors or the like. The outputs of the sensor system 208 can comprise a portion of the human-machine interaction data 110 as described in more detail herein. In particular, the outputs of the sensor system 208 can comprise the time-series data of the first data stream 110A, which corresponds to one or more of event data, kinematics data, or force data for the manipulator system. For example, the event data may include data indicating a number of instrument installs, insertions, removals, etc. and / or other events associated with computer-assisted system 108. The kinematics data may include data regarding movement of manipulator assembly 202 and instruments attached thereto by the operator O and / or the computer-assisted system 108. The force data may include indications of one or more forces exerted on the manipulator assembly 202 and instruments attached thereto.

[0062] Computer-assisted system 108 can also include the control system 212. Control system 212 includes memory 216 and the processor system 214 for effecting control between medical instrument system 204, operator input system 206, sensor system 208, and display system 210. Control system 212 also includes programmed instructions (e.g., a non-transitory machine-readable or computer-readable mediums storing the instructions) to implement a procedure using the manipulator assembly 202, including for navigation, steering, imaging, engagement feature deployment or retraction, applying treatment to target tissue (e.g., via theapplication of energy), or the like. The control system 212 may include the computing system 102 or be a distinct system coupled to the computing system 102 via wireless or wired methods known in the art.

[0063] Control system 212 can optionally further include a virtual visualization system to provide guidance to operator O when controlling medical instrument system 204 during a human-machine interaction. In some embodiments, the guidance includes virtual navigation guidance performing and / or demonstrating a capability of the computer-assisted system 108 based upon reference to an acquired pre-human-machine interaction or intra- human-machine interaction dataset of anatomic passageways. The virtual visualization system processes images of the surgical site imaged using imaging technology such as computerized tomography (CT), magnetic resonance imaging (MRI), fluoroscopy, thermography, ultrasound, optical coherence tomography (OCT), thermal imaging, impedance imaging, laser imaging, nanotube X-ray imaging, and / or the like.

[0064] In some embodiments, the human-machine interaction event may include a training exercise related to the manipulator system embodiment of the computer-assisted system 108 as shown in FIG. 2. In these embodiments, the evaluation version of the output 114 may include entries that relate to the manipulator system and the list of tasks in the guidance version of the output 114 may include tasks to perform using the manipulator system.

[0065] In embodiments where the guidance includes the list of tasks, the machine learning model 112 may be configured to determine a current facet being presented based on facet definitions contained in the rubric 116 and output a task to control the manipulator system embodiment of the computing system 102 based on the determination of the current facet being presented. Similarly, the machine learning model 112 may be configured to determine a future facet for presentation based on the facet definitions contained in the rubric 116 and output tasks to control the manipulator system embodiment based on the determination of the future facet. Any of these tasks may be presented on a portion of the display system 210 as part of an interactable prompt that when activated performs a corresponding task of the manipulator system embodiment. Specifically, the control system 212 may be configured to perform one or more of the tasks in response to detecting an interaction with the prompt. Additionally or alternatively, the control system 212 may be configured to perform one or more of the tasks in response to detecting the output 114 of the machine learning model 112.

[0066] FIG. 3 shows a computer-implemented method 300 for generating a postinteraction evaluation for a human-machine interaction using an artificial intelligence model(such as machine learning model 112). The method 300 may be executed by the computing system 102, the control system 212, or a similar computing system having a processing unit and a memory, such as a remote server communicatively coupled to the computing system 102 or control system 212.

[0067] At block 310, the computer-implemented method 300 includes obtaining post- human-machine interaction data (e.g., human-machine interaction data 110) representative of a human-machine interaction event with respect to a computer-assisted system (e.g., computer-assisted system 108). The post-human-machine interaction data comprises a plurality of data streams from one or more data sources, a first data stream of the plurality of data streams (e.g., first data stream 110A) comprises time-series data relating to operation of the computer-assisted system during the human-machine interaction event, and a second data stream of the plurality of data streams (e.g., second data stream 110B) is indicative of language presented as part of the human-machine interaction event. A third data stream of the plurality of data streams (e.g., third data stream HOC) may comprise image data. The image data may be environmental image data and / or procedure image data. The computer-assisted system may include a manipulator system. In these embodiments, the time-series data may include one or more of event data, kinematics data, or force data for the manipulator system and the human-machine interaction event may include a training exercise related to the manipulator system.

[0068] At block 320, the computer-implemented method 300 includes inputting each of the plurality of data streams into a projection layer of a machine learning model (e.g., plurality of projection layers 120 of machine learning model 112). The machine learning model is configured to output an evaluation based on an evaluation rubric (e.g., rubric 116), one or more reference documents (e.g., one or more reference documents 118), and outputs from the projection layers for at least a portion of the plurality of data streams. In some embodiments, the one or more reference documents may consumed into the latent space of the machine learning model using a self- supervised learning process where the model is tuned to predict one portion of a reference document from another part of the document that is input into the model. The training inputs may further include historical human-machine interaction event data that is associated with known evaluations. The one or more reference documents may include one or more of product user manuals, demonstrator training documents, instrument documentation, clinical documents, operational documents, procedural guidelines, troubleshooting guidelines, or human-machine interaction instructions. The machine learning model may comprise a large language model. The respective projection layer for the first datastream may be configured to modify the time- series data into machine readable tokens that represent text descriptions of system events corresponding to the time-series data.

[0069] At block 330, the computer-implemented method 300 includes receiving an evaluation for the human-machine interaction event as an output of the machine learning model. The evaluation may include one or more scoring metrics that indicate a degree to which the human-machine interaction event conformed to expected baselines directed by the evaluation rubric. The expected baseline may include a set of facets to be presented during the human-machine interaction event. The evaluation may also include recommendations for improving future human-machine interactions to conform to expected baselines directed by the evaluation rubric. The recommendations may include citations to learning content and exercises included in the one or more reference documents.

[0070] At block 340, the computer-implemented method 300 includes associating the output evaluation with the human-machine interaction event. For example, a database may be configured to store a plurality of evaluations associated with operators, targets (e.g., a person to which the operator is demonstrating the computer assisted system) and / or corporate entities. Accordingly, associating the output evaluation with the human-machine interaction event may include updating a record in the database associated with one or more entities involved in the human-machine interaction event with an indication of the event, event metadata (e.g., date, entities involved, a result), and the output evaluation. As a result, one can analyze a plurality of evaluations to assess, for example, operator development across human-machine interactions to identify training opportunities and / or modify the reference documents and / or rubrics.

[0071] The computer-implemented method 300 may also include obtaining the one or more reference documents based upon at least a portion of the plurality of data streams and the evaluation rubric and inputting the one or more reference documents into the machine learning model alongside the evaluation rubric and each of the plurality of data streams. In these embodiments, the evaluation rubric may define how the machine learning model is to utilize the one or more reference documents to generate the evaluation.

[0072] One of the projection layers may include a behavior analyzer configured to output machine readable tokens that represent text descriptions of behavior of a presenter for the human-machine interaction event. In these embodiments, the computer-implemented method 300 may include inputting the second data stream and a third data stream of the plurality of data streams into the behavior analyzer, the third data stream comprising image data of the presenter.

[0073] The computer-implemented method 300 may also include aligning data elements in each of the plurality of data streams to associate data elements related to a common action. In these embodiments, inputting the plurality of data streams into the machine learning model comprises inputting the aligned data elements of the data streams.

[0074] FIG. 4 shows a computer-implemented method 400 for generating intra-interaction guidance for a human-machine interaction using an artificial intelligence model (such as machine learning model 112). The method 400 may be executed by the computing system 102, the control system 212, or a similar computing system having a processing unit and a memory, such as a remote server communicatively coupled to the computing system 102 or control system 212.

[0075] At block 410, the computer-implemented method 400 includes obtaining a plurality of data streams from one or more data sources (e.g., human-machine interaction data 110). A first data stream of the plurality of data streams (e.g., first data stream 110A) comprises timeseries data relating to operation of the manipulator system by the control system as part of a human-machine interaction event for the computer-assisted system, and a second data stream of the plurality of data streams (e.g., second data stream 110B) is indicative of language presented as part of the human-machine interaction event. A third data stream of the plurality of data streams (e.g., third data stream 110C) may comprise image data. The image data may be environmental image data and / or procedure image data. The time- series data may include one or more of event data, kinematics data, or force data for the manipulator system. The human-machine interaction event may include a training exercise related to the manipulator system.

[0076] At block 420, the computer-implemented method 400 includes inputting each of the plurality of data streams into a respective projection layer of a machine learning model (e.g., plurality of projection layers 120 of machine learning model 112). The machine learning model is configured to output guidance based on a guidance rubric (e.g., rubric 116), one or more reference documents (e.g., one or more reference documents 118), and outputs from the projection layers for at least a portion of the plurality of data streams. In some embodiments the one or more reference documents may be consumed into a latent space of the machine learning model using a self-supervised process. The training inputs may further include historical human-machine interaction event data that is associated with known guidance. The one or more reference documents may include one or more of product user manuals, demonstrator training documents, instrument documentation, clinical documents, operational documents, procedural guidelines, troubleshooting guidelines, or human-machine interactioninstructions. The machine learning model may comprise a large language model. The respective projection layer for the first data stream may be configured to modify the timeseries data into machine readable tokens that represent text descriptions of system events corresponding to the time-series data.

[0077] At block 430, the computer-implemented method 400 includes receiving guidance as an output of the machine learning model. The guidance includes recommendations to conform the human-machine interaction event to expected baselines. The guidance may include a list of tasks to perform. The list of tasks may include tasks to perform using the manipulator system. In some embodiments, the computer-implemented method 400 includes determining a current facet being presented based on facet definitions contained in the guidance rubric and outputting a task to control the manipulator system based on the determination of the current facet being presented. The machine learning model may also be configured to determine a future facet for presentation based on facet definitions contained in the guidance rubric and output tasks to control the manipulator system based on the determination of the future facet. The computer-implemented method 400 may include presenting, via the display device, an interactable prompt to perform the tasks to control the manipulator system and controlling the manipulator system to perform the tasks in response to detecting an interaction with the prompt. Controlling the manipulator system to perform the tasks may also be done in response to detecting the output of the machine learning model. The list of tasks may include tasks to present a facet of the manipulator system. The machine learning model may be configured to detect anomalous behavior from the expected baselines defined by the guidance rubric, and generate the list of tasks to correct the anomalous behavior. The list of tasks may include a diagnostic task to perform, the diagnostic task configured to assist in identifying possible malfunctions in the computer-assisted system.

[0078] At block 440, the computer-implemented method 400 includes presenting the guidance on a display device operably coupled to the control system (e.g., display device 119). The display device may be included in a virtuality reality (VR) headset, an augmented reality (AR) headset, a mixed reality (MR) headset, or smart glasses. In some embodiments the computer-implemented method 400 includes recording the second data stream with the display device.

[0079] The computer-implemented method 400 may also include obtaining the one or more reference documents based upon at least a portion of the plurality of data streams and the guidance rubric and inputting the one or more reference documents into the machine learning model alongside the guidance rubric and each of the plurality of data streams. Inthese embodiments, the guidance rubric may define how the machine learning model is to utilize the one or more reference documents to generate the guidance.

[0080] One of the projection layers may include a behavior analyzer configured to output machine readable tokens that represent text descriptions of behavior of a presenter for the human-machine interaction event. In these embodiments, the computer- implemented method 400 may include inputting the second data stream and a third data stream of the plurality of data streams into the behavior analyzer, the third data stream comprising image data of the presenter.

[0081] The computer-implemented method 400 may also include aligning data elements in each of the plurality of data streams to associate data elements related to a common action. In these embodiments, inputting the plurality of data streams into the machine learning model comprises inputting the aligned data elements of the data streams.

[0082] FIG. 5 shows a computer-implemented method 500 for using the machine learning model 112 to generate the output 114, where the output 114 includes an evaluation of identified examples demonstrated or discussed as part of the human-machine interaction event. The method 500 may be executed by the computing system 102, the control system 212, or a similar computing system having a processing unit and a memory, such as a remote server communicatively coupled to the computing system 102 or control system 212. The method 500 may be perform during a human-machine interaction to provide intra-interaction guidance to the operator O.

[0083] At block 502, the computer-implemented method 500 includes dividing the humanmachine interaction data 110 into different windows. In some embodiments, the windows may be overlapping rolling windows of a fixed duration (e.g., 10 seconds, 30 seconds, 1 minute, etc.).

[0084] At blocks 504 and 506, the computer-implemented method 500 includes updating or creating a summary for each of the windows using the machine learning model 112. The summary may describe the various actions performed by the operator O during the window. For example, if the operator O provides verbal descriptions of facets of the computer-assisted system 108, the summary may describe those descriptions.

[0085] At block 508, the computer-implemented method 500 includes using the machine learning model 112 to identify examples (e.g., various facets and / or capabilities supported by the computer-assisted system 108 that are being demonstrated and / or performed by discussed by the operator O or other presenter) within the summaries.

[0086] At blocks 510 and 512, the computer-implemented method 500 includes evaluating each of the identified examples using the machine learning model 112, the rubric 116, and the one or more reference documents 118. To this end, each example may be associated with a separate rubric 116 and / or reference documents 118 to be used during the evaluation process. It should be appreciated that because the same example may appear in multiple windows, in some embodiments, the machine learning model 112 may utilize one or more prior evaluations for the same example generated during a prior execution of block 512 as an input.

[0087] FIG. 6 shows a computer-implemented method 600 for using the machine learning model 112 to generate the output 114, where the output 114 includes an evaluation of identified examples demonstrated or discussed as part of the human-machine interaction event. The method 600 may be executed by the computing system 102, the control system 212, or a similar computing system having a processing unit and a memory, such as a remote server communicatively coupled to the computing system 102 or control system 212.

[0088] The computer-implemented method 600 is similar to the computer-implemented method 500 except that the method 600 may be performed post-human-machine interaction. Accordingly, unlike for the method 500, at block 602, the human-machine interaction data 110 can be segmented such that the complete set of data associated with each example can be grouped together prior to the evaluation thereof. It should be appreciated that because multiple examples may be discussed contemporaneously, the same portion of the humanmachine interaction data 110 may appear in multiple sets of the segmented data.

[0089] After the human-machine interaction data 110 has been segmented, e computer- implemented method 600 include performing blocks 510 and 512 of the computer- implemented method 500 on the identified examples.

[0090] FIG. 7 shows a computer-implemented method 700 for using the machine learning model 112 to generate the output 114, where the output 114 includes a determination of whether specific required facets have been presented as part of the human-machine interaction event. The method 700 may be executed by the computing system 102, the control system 212, or a similar computing system having a processing unit and a memory, such as a remote server communicatively coupled to the computing system 102 or control system 212.

[0091] At block 702, the computer-implemented method 700 includes using the machine learning model 112 to identify required facets from the rubric 116 and the one or more reference documents 118.

[0092] At block 704 and 706, the computer-implemented method 700 includes, for each required facet identified, using the machine learning model 112 to check if the required facets identified were performed as indicated by the human-machine interaction data 110.

[0093] It is understood that the blocks of the methods 300, 400, 500, 600, and 700 need not occur strictly in the order shown.

[0094] Although the systems, methods, devices, and components thereof, have been described in terms of exemplary embodiments, they are not limited thereto. The detailed description is to be construed as exemplary only and does not describe every possible embodiment of the invention because describing every possible embodiment would be impractical, if not impossible. Numerous alternative embodiments could be implemented, using either current technology or technology developed after the filing date of this patent that would still fall within the scope of the claims defining the invention.

[0095] Those skilled in the art will recognize that a wide variety of modifications, alterations, and combinations can be made with respect to the above-described embodiments without departing from the scope of the invention, and that such modifications, alterations, and combinations are to be viewed as being within the ambit of the inventive concept.

Claims

What is claimed is:

1. A computer system comprising: one or more processors; and one or more non-transitory, computer-readable media storing instructions that, when executed by the one or more processors, cause the computer system to: obtain human-machine interaction data representative of a human-machine interaction event with respect to a computer-assisted system, wherein: the human-machine interaction data comprises a plurality of data streams from one or more data sources, a first data stream of the plurality of data streams comprises time-series data relating to operation of the computer-assisted system during the human-machine interaction event, and a second data stream of the plurality of data streams is indicative of language presented as part of the human-machine interaction event; input each of the plurality of data streams into a projection layer of a machine learning model, wherein the machine learning model is configured to output an evaluation based on an evaluation rubric, one or more reference documents, and outputs from the projection layers for at least a portion of the plurality of data streams; receive an evaluation for the human-machine interaction event as an output of the machine learning model; and associate the output evaluation with the human-machine interaction event.

2. The computer system of claim 1 wherein the one or more reference documents are consumed into a latent space of the machine learning model using a self- supervised process.

3. The computer system of claim 2 wherein training inputs for the machine learning model include historical human-machine interaction event data that is associated with known evaluations.

4. The computer system of claim 1 wherein the instructions further cause the one or more processors to:obtain the one or more reference documents based upon at least a portion of the plurality of data streams and the evaluation rubric; and input the one or more reference documents into the machine learning model alongside the evaluation rubric and each of the plurality of data streams.

5. The computer system of claim 4 wherein the evaluation rubric defines how the machine learning model is to utilize the one or more reference documents to generate the evaluation.

6. The computer system of claim 1, wherein the one or more reference documents include one or more of product user manuals, demonstrator training documents, instrument documentation, clinical documents, operational documents, procedural guidelines, troubleshooting guidelines, or human-machine interaction instructions.

7. The computer system of claim 1, wherein the computer-assisted system includes a manipulator system.

8. The computer system of claim 7, wherein the time-series data includes one or more of event data, kinematics data, or force data for the manipulator system.

9. The computer system of claim 7, wherein the human-machine interaction event includes a training exercise related to the manipulator system.

10. The computer system of claim 1, wherein a third data stream of the plurality of data streams comprises image data.

11. The computer system of claim 10, wherein the image data is environmental image data or procedure image data.

12. The computer system of claim 1, wherein the machine learning model comprises a large language model.

13. The computer system of claim 12, wherein the respective projection layer for the first data stream is configured to modify the time-series data into machine readable tokens that represent text descriptions of system events corresponding to the time-series data.

14. The computer system of claim 12, wherein one of the projection layers includes a behavior analyzer configured to output machine readable tokens that represent text descriptions of behavior of a presenter for the human-machine interaction event.

15. The computer system of claim 14, wherein the input to the behavior analyzer includes the second data stream and a third data stream of the plurality of data streams, the third data stream comprising image data of the presenter.

16. The computer system of any one of claims 1 to 15, wherein the evaluation includes one or more scoring metrics that indicate a degree to which the human-machine interaction event conformed to expected baselines directed by the evaluation rubric.

17. The computer system of claim 16, wherein the expected baseline includes a set of facets to be presented during the human-machine interaction event.

18. The computer system of any one of claims 1 to 15, wherein the evaluation includes recommendations for improving future human-machine interactions to conform to expected baselines directed by the evaluation rubric.

19. The computer system of claim 18, wherein the recommendations include citations to learning content and exercises included in the one or more reference documents.

20. The computer system of any one of claims 1 to 15, wherein data elements in each of the plurality of data streams are aligned to associate data elements related to a common action.

21. The computer system of claim 20, wherein inputting the plurality of data streams into the machine learning model comprises inputting the aligned data elements of the data streams.

22. A computer-implemented method comprising: obtaining human-machine interaction data representative of a human-machine interaction event with respect to a computer-assisted system, wherein: the human-machine interaction data comprises a plurality of data streams from one or more data sources, a first data stream of the plurality of data streams comprises time-series data relating to operation of the computer-assisted system during the human-machine interaction event, and a second data stream of the plurality of data streams is indicative of language presented as part of the human-machine interaction event; inputting each of the plurality of data streams into a projection layer of a machine learning model, wherein the machine learning model is configured to output an evaluation based on an evaluation rubric, one or more reference documents, and outputs from the projection layers for at least a portion of the plurality of data streams; receiving an evaluation for the human-machine interaction event as an output of the machine learning model; and associating the output evaluation with the human-machine interaction event.

23. The computer- implemented method of claim 22 wherein the one or more reference documents are consumed into a latent space of the machine learning model using a selfsupervised process.

24. The computer-implemented method of claim 23 wherein training inputs for the machine learning model include historical human-machine interaction event data that is associated with known evaluations.

25. The computer- implemented method of claim 22 further comprising: obtaining the one or more reference documents based upon at least a portion of the plurality of data streams and the evaluation rubric; and inputting the one or more reference documents into the machine learning model alongside the evaluation rubric and each of the plurality of data streams.

26. The computer-implemented method of claim 25 wherein the evaluation rubric defines how the machine learning model is to utilize the one or more reference documents to generate the evaluation.

27. The computer-implemented method of claims 22, wherein the one or more reference documents include one or more of product user manuals, demonstrator training documents, instrument documentation, clinical documents, operational documents, procedural guidelines, troubleshooting guidelines, or human-machine interaction instructions.

28. The computer- implemented method of claim 22, wherein the computer-assisted system includes a manipulator system.

29. The computer-implemented method of claim 28, wherein the time-series data includes one or more of event data, kinematics data, or force data for the manipulator system.

30. The computer-implemented method of claim 28, wherein the human-machine interaction event includes a training exercise related to the manipulator system.

31. The computer- implemented method of claim 22, wherein a third data stream of the plurality of data streams comprises image data.

32. The computer- implemented method of claim 31, wherein the image data is environmental image data or procedure image data.

33. The computer- implemented method of claim 22, wherein the machine learning model comprises a large language model.

34. The computer-implemented method of claim 33, wherein the respective projection layer for the first data stream is configured to modify the time-series data into machine readable tokens that represent text descriptions of system events corresponding to the timeseries data.

35. The computer-implemented method of claim 33, wherein one of the projection layers includes a behavior analyzer configured to output machine readable tokens that represent text descriptions of behavior of a presenter for the human-machine interaction event.

36. The computer-implemented method of claim 35, further comprising: inputting the second data stream and a third data stream of the plurality of data streams into the behavior analyzer, the third data stream comprising image data of the presenter.

37. The computer- implemented method of any one of claims 22 to 36, wherein the evaluation includes one or more scoring metrics that indicate a degree to which the humanmachine interaction event conformed to expected baselines directed by the evaluation rubric.

38. The computer- implemented method of claim 37, wherein the expected baseline includes a set of facets to be presented during the human-machine interaction event.

39. The computer-implemented method of any one of claims 22 to 36, wherein the evaluation includes recommendations for improving future human-machine interactions to conform to expected baselines directed by the evaluation rubric.

40. The computer-implemented method of claim 39, wherein the recommendations include citations to learning content and exercises included in the one or more reference documents.

41. The computer- implemented method of any one of claims 22 to 36, further comprising aligning data elements in each of the plurality of data streams to associate data elements related to a common action.

42. The computer- implemented method of claim 41, wherein inputting the plurality of data streams into the machine learning model comprises inputting the aligned data elements of the data streams.

43. A non-transitory machine-readable medium comprising a plurality of machine- readable instructions that when executed by one or more processors are adapted to cause the one or more processors to perform the method of any one of claims 22-42.

44. A computer-assisted system comprising: a manipulator system; and a control system operably coupled to the manipulator system, wherein the control system is configured to: obtain a plurality of data streams from one or more data sources, wherein: a first data stream of the plurality of data streams comprises time-series data relating to operation of the manipulator system or the control system as part of a human-machine interaction event for the computer-assisted system, and a second data stream of the plurality of data streams is indicative of language presented as part of the human-machine interaction event; input each of the plurality of data streams into a respective projection layer of a machine learning model, wherein the machine learning model is configured to output guidance based on a guidance rubric, one or more reference documents, and outputs from the projection layers for at least a portion of the plurality of data streams; receive guidance as an output of the machine learning model, wherein the guidance includes recommendations to conform the human-machine interaction event to expected baselines; and present the guidance on a display device operably coupled to the control system.

45. The computer-assisted system of claim 44, wherein the display device is included in a virtuality reality (VR) headset, an augmented reality (AR) headset, a mixed reality (MR) headset, or smart glasses.

46. The computer-assisted system of claim 44, wherein the second data stream is recorded by the display device.

47. The computer-assisted system of claim 44, wherein the one or more reference documents are consumed into a latent space of the machine learning model using a selfsupervised process.

48. The computer-assisted system of claim 47, wherein training inputs for the machine learning model include historical human-machine interaction event data that is associated with known guidance.

49. The computer-assisted system of claim 44, wherein the control system is further configured to: obtain the one or more reference documents based upon at least a portion of the plurality of data streams and the guidance rubric; and input the one or more reference documents into the machine learning model alongside the guidance rubric and each of the plurality of data streams.

50. The computer-assisted system of claim 49, wherein the guidance rubric defines how the machine learning model is to utilize the one or more reference documents to generate the output guidance.

51. The computer-assisted system of claim 44, wherein the guidance includes a list of tasks to perform.

52. The computer-assisted system of claim 51, wherein the list of tasks includes tasks to perform using the manipulator system.

53. The computer-assisted system of claim 52 wherein the machine learning model is configured to: determine a current facet being presented based on facet definitions contained in the guidance rubric; and output a task to control the manipulator system based on the determination of the current facet being presented.

54. The computer-assisted system of claim 44 wherein the machine learning model is configured to:determine a future facet for presentation based on facet definitions contained in the guidance rubric; and output tasks to control the manipulator system based on the determination of the future facet.

55. The computer-assisted system of claim 53 or 54, wherein the control system is further configured to: present, via the display device, an interactable prompt to perform the tasks to control the manipulator system; and control the manipulator system to perform the tasks in response to detecting an interaction with the prompt.

56. The computer-assisted system of claim 53 or 54, wherein the control system is further configured to: control the manipulator system to perform the tasks in response to detecting the output of the machine learning model.

57. The computer-assisted system of claim 51, wherein the list of tasks includes tasks to present a facet of the manipulator system.

58. The computer-assisted system of claim 51, wherein to generate the list of tasks, the machine learning model is configured to (i) detect anomalous behavior from the expected baselines defined by the guidance rubric; and (ii) generate the list of tasks to correct the anomalous behavior.

59. The computer-assisted system of claim 51, wherein the list of tasks includes a diagnostic task to perform, the diagnostic task configured to assist in identifying possible malfunctions in the computer-assisted system.

60. The computer-assisted system of claim 44, wherein the one or more reference documents include one or more of product user manuals, demonstrator training documents, instrument documentation, clinical documents, operational documents, procedural guidelines, troubleshooting guidelines, or human-machine interaction instructions.

61. The computer-assisted system of claim 44, wherein the time-series data includes one or more of event data, kinematics data, or force data for the manipulator system.

62. The computer-assisted system of claim 44, wherein the human-machine interaction event includes a training exercise related to the manipulator system.

63. The computer-assisted system of claim 44, wherein a third data stream of the plurality of data streams comprises image data.

64. The computer-assisted system of claim 63, wherein the image data is environmental image data or procedure image data.

65. The computer-assisted system of claim 44, wherein the machine learning model comprises a large language model.

66. The computer-assisted system of claim 65, wherein the respective projection layer for the first data stream is configured to modify the time-series data into machine readable tokens that represent text descriptions of system events corresponding to the time-series data.

67. The computer-assisted system of claim 65, wherein one of the projection layers includes a behavior analyzer configured to output machine readable tokens that represent text descriptions of behavior of a presenter for the human-machine interaction event.

68. The computer-assisted system of claim 67, wherein the input to the behavior analyzer includes the second data stream and a third data stream of the plurality of data streams, the third data stream comprising image data of the presenter.

69. The computer-assisted system of claim 44, wherein data elements in each of the plurality of data streams are aligned to associate data elements related to a common action.

70. The computer-assisted system of claim 69, wherein inputting the plurality of data streams into the machine learning model comprises inputting the aligned data elements of the data streams.

71. A computer- implemented method comprising: obtaining a plurality of data streams from one or more data sources, wherein: a first data stream of the plurality of data streams comprises time-series data relating to operation of a manipulator system by a control system as part of a human-machine interaction event for a computer-assisted system, and a second data stream of the plurality of data streams is indicative of language presented as part of the human-machine interaction event; inputting each of the plurality of data streams into a respective projection layer of a machine learning model, wherein the machine learning model is configured to output guidance based on a guidance rubric, one or more reference documents, and outputs from the projection layers for at least a portion of the plurality of data streams; receiving guidance as an output of the machine learning model, wherein the guidance includes recommendations to conform the human-machine interaction event to expected baselines; and presenting the guidance on a display device operably coupled to the control system.

72. The computer- implemented method of claim 71, wherein the display device is included in a virtuality reality (VR) headset, an augmented reality (AR) headset, a mixed reality (MR) headset, or smart glasses.

73. The computer- implemented method of claim 71, further comprising recording the second data stream with the display device.

74. The computer- implemented method of claim 71, wherein the one or more reference documents are consumed into a latent space of the machine learning model using a selfsupervised process.

75. The computer- implemented method of claim 74, wherein training inputs for the machine learning model include historical human-machine interaction event data that is associated with known guidance.

76. The computer-implemented method of claim 71, further comprising:obtaining the one or more reference documents based upon at least a portion of the plurality of data streams and the guidance rubric; and inputting the one or more reference documents into the machine learning model alongside the guidance rubric and each of the plurality of data streams.

77. The computer- implemented method of claim 76, wherein the guidance rubric defines how the machine learning model is to utilize the one or more reference documents to generate the output guidance.

78. The computer- implemented method of claim 71, wherein the guidance includes a list of tasks to perform.

79. The computer- implemented method of claim 78, wherein the list of tasks includes tasks to perform using the manipulator system.

80. The computer-implemented method of claim 79 further comprising: determining a current facet being presented based on facet definitions contained in the guidance rubric; and outputting an task to control the manipulator system based on the determination of the current facet being presented.

81. The computer- implemented method of claim 71 wherein the machine learning model is configured to: determine a future facet for presentation based on facet definitions contained in the guidance rubric; and output tasks to control the manipulator system based on the determination of the future facet.

82. The computer- implemented method of claim 80 or 81, further comprising: presenting, via the display device, an interactable prompt to perform the tasks to control the manipulator system; and controlling the manipulator system to perform the tasks in response to detecting an interaction with the prompt.

83. The computer-implemented method of claim 80 or 81, further comprising: controlling the manipulator system to perform the tasks in response to detecting the output of the machine learning model.

84. The computer-implemented method of claim 78, wherein the list of tasks includes tasks to present a facet of the manipulator system.

85. The computer-implemented method of claim 78, wherein the machine learning model is configured to (i) detect anomalous behavior from the expected baselines defined by the guidance rubric, and (ii) generate the list of tasks to correct the anomalous behavior.

86. The computer-implemented method of claim 78, wherein the list of tasks includes a diagnostic task to perform, the diagnostic task configured to assist in identifying possible malfunctions in the computer-assisted system.

87. The computer- implemented method of claim 71, wherein the one or more reference documents include one or more of product user manuals, demonstrator training documents, instrument documentation, clinical documents, operational documents, procedural guidelines, troubleshooting guidelines, or human-machine interaction instructions.

88. The computer-implemented method of claim 71, wherein the time-series data includes one or more of event data, kinematics data, or force data for the manipulator system.

89. The computer- implemented method of claim 71, wherein the human-machine interaction event includes a training exercise related to the manipulator system.

90. The computer- implemented method of claim 71, wherein a third data stream of the plurality of data streams comprises image data.

91. The computer- implemented method of claim 90, wherein the image data is environmental image data or procedure image data.

92. The computer- implemented method of claim 71, wherein the machine learning model comprises a large language model.

93. The computer-implemented method of claim 92, wherein the respective projection layer for the first data stream is configured to modify the time-series data into machine readable tokens that represent text descriptions of system events corresponding to the timeseries data.

94. The computer-implemented method of claim 92, wherein one of the projection layers includes a behavior analyzer configured to output machine readable tokens that represent text descriptions of behavior of a presenter for the human-machine interaction event.

95. The computer- implemented method of claim 94, wherein the input to the behavior analyzer includes the second data stream and a third data stream of the plurality of data streams, the third data stream comprising image data of the presenter.

96. The computer- implemented method of claim 71, further comprising aligning data elements in each of the plurality of data streams to associate data elements related to a common action.

97. The computer- implemented method of claim 96, wherein inputting the plurality of data streams into the machine learning model comprises inputting the aligned data elements of the data streams.

98. A non-transitory machine-readable medium comprising a plurality of machine- readable instructions that when executed by one or more processors are adapted to cause the one or more processors to perform the method of any one of claims 71-97.

Citation Information

Patent Citations

  • AI robot system applied to collaborative system

    CN115185367A

  • Training and / or utilizing machine learning models for use in natural language based robotic control

    JP2023525676A