Systems and methods for generating trainee guidance using adaptive artificial intelligence

Adaptive AI systems analyze multi-modal data to generate personalized training guidance, addressing the limitations of traditional medical training by adapting to trainee progress and outcomes, enhancing training efficiency and effectiveness.

WO2026155783A2PCT designated stage Publication Date: 2026-07-23INTUITIVE SURGICAL OPERATIONS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INTUITIVE SURGICAL OPERATIONS INC
Filing Date
2025-10-10
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing medical training procedures rely heavily on expert instructors, lack real-world application training, and fail to dynamically adapt to trainee progress and training outcomes, leading to inefficient and ineffective training.

Method used

A computer system utilizing adaptive artificial intelligence to generate trainee guidance by analyzing multi-modal data, including performance assessment and user inputs, to provide personalized and dynamic training content and methodologies.

Benefits of technology

Enables quicker and more effective training by providing personalized guidance that adapts to the trainee's progress and training outcomes, improving surgical skills and outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025050414_23072026_PF_FP_ABST
    Figure US2025050414_23072026_PF_FP_ABST
Patent Text Reader

Abstract

A computer system may obtain multi-modal data relating to observed actions of a trainee, generate a trainee performance assessment by analyzing the multi-modal data. A trainee performance assessment constitution including rules for assessing performance of the observed actions is input into a trainee performance assessment machine learning model to control how the trainee performance assessment machine learning model generates the trainee performance assessment. The computer system may generate a trainee guidance by analyzing the trainee performance assessment and the multi-modal data. A trainee guidance constitution is input into a trainee guidance machine learning model. The trainee guidance constitution includes rules that control how the trainee guidance machine learning model determines the guidance content and the guidance methodology of the trainee guidance. The computer system may provide, to the trainee, the guidance content via the guidance methodology.
Need to check novelty before this filing date? Find Prior Art

Description

Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC SYSTEMS AND METHODS FOR GENERATING TRAINEE GUIDANCE USING ADAPTIVE ARTIFICIAL INTELLIGENCE CROSS-REFERENCE TO REEATED APPEICATIONS

[0001] This application claims priority to and the benefit of the filing date of provisional U.S. Patent Application No. 63 / 706,355 entitled “SYSTEMS AND METHODS FOR GENERATING TRAINEE GUIDANCE USING ADAPTIVE ARTIFICIAL INTELLIGENCE,” filed on October 11, 2024. The entire contents of the provisional application are hereby expressly incorporated herein by reference.TECHNICAL FIELD

[0002] The present disclosure relates generally to providing trainee guidance for computer-assisted systems and more particularly to generating trainee guidance for a medical procedure using adaptive artificial intelligence.BACKGROUND

[0003] Training people to become proficient in new and evolving medical procedures such as those that utilize computer-assisted systems has traditionally been a complicated individualized process that relies on direct instruction from a limited group of expert instructors. In response, training processes that deploy automated systems or simulators have been developed to enable training without reliance on direct expert instruction. These training processes include virtual reality simulators used during the first step of training before letting the trainee use the actual computer assisted system in dry and wet lab environments. The training processes also include training curriculums that use physical, virtual reality, or augmented reality simulators such as the Fundamentals of Laparoscopic Surgery (FLS).Additionally, expert-in-the-loop training system for dual console surgical systems have also been proposed.

[0004] However, these revised training procedures still suffer from several drawbacks. In particular, these procedures may (1) still rely on the presence of an expert during training, (2) not provide the training on the actual equipment used in real world applications, and / or (3) be static training regiments that do not dynamically adapt the training content and methodologies to account for (i) the progress of a trainee over time and (ii) data regarding training outcomes for other, similar, trainees.

[0005] Therefore, there is a need for a system and training process that provides improved automated guidance to a trainee during a training procedure. Such techniques may allow for quicker more effective training and result in improved surgical outcomes.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC SUMMARY

[0006] In some aspects, the techniques described herein relate to a computer system including: one or more processors; and one or more non-transitory, computer-readable media storing instructions that, when executed by the one or more processors, cause the computer system to: obtain multi-modal data relating to observed actions of a trainee during a training procedure; generate a trainee performance assessment by analyzing, via a trainee performance assessment machine learning model, the multi-modal data, wherein a trainee performance assessment constitution including rules for assessing performance of the observed actions is input into the trainee performance assessment machine learning model to control how the trainee performance assessment machine learning model generates the trainee performance assessment; generate a trainee guidance by analyzing, via a trainee guidance machine learning model, the trainee performance assessment and the multi-modal data, wherein: a trainee guidance constitution is input into the trainee guidance machine learning model, and the trainee guidance constitution includes rules that control how the trainee guidance machine learning model determines guidance content and a guidance methodology for the trainee guidance; and provide, to the trainee, the guidance content via the guidance methodology.

[0007] In some additional aspects, the techniques described herein relate to a computer-implemented method including: obtaining multi-modal data relating to observed actions of a trainee during a training procedure; generating a trainee performance assessment by analyzing, via a trainee performance assessment machine learning model, the multi-modal data, wherein a trainee performance assessment constitution including rules for assessing performance of the observed actions is input into the trainee performance assessment machine learning model to control how the trainee performance assessment machine learning model generates the trainee performance assessment; generating a trainee guidance by analyzing, via a trainee guidance machine learning model, the trainee performance assessment and the multi-modal data, wherein: a trainee guidance constitution is input into the trainee guidance machine learning model, and the trainee guidance constitution includes rules that control how the trainee guidance machine learning model determines guidance content and a guidance methodology of the trainee guidance; and providing, to the trainee, the guidance content via the guidance methodology.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC

[0008] In some further aspects, a non-transitory machine-readable medium comprising a plurality of machine-readable instructions that when executed by one or more processors are adapted to cause the one or more processors to perform any of the methods described herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] FIG. 1 is a diagram of a computer-assisted system in accordance with one or more embodiments.

[0010] FIG. 2A is a schematic diagram of a system for generating trainee guidance.

[0011] FIG. 2B is a schematic diagram of another system for trainee guidance.

[0012] FIG. 2C is a schematic diagram of a system for providing multi-modal data for use in the systems shown in FIG. 2A and FIG. 2B.

[0013] FIG. 2D is a schematic diagram showing different methodologies for providing the trainee guidance generated by the systems of FIG. 2A or FIG. 2B.

[0014] FIG. 3 is a schematic diagram of the structure of the various machine learning models described herein.

[0015] FIG. 4 is a flow diagram of a method for generating instrument guidance utilizing adaptive artificial intelligence.

[0016] Examples of the present disclosure and their advantages are best understood by referring to the detailed description that follows. It should be appreciated that like reference numerals are used to identify like elements illustrated in one or more of the figures, wherein showings therein are for purposes of illustrating examples of the present disclosure and not for purposes of limiting the same.DETAILED DESCRIPTION

[0017] In this description, specific details are set forth describing some embodiments consistent with the present disclosure. Numerous specific details are set forth in order to provide a thorough understanding of the embodiments. It will be apparent, however, to one skilled in the art that some embodiments may be practiced without some or all of these specific details. The specific embodiments disclosed herein are meant to be illustrative but not limiting. One skilled in the art may realize other elements that, although not specifically described here, are within the scope and the spirit of this disclosure. In addition, to avoid unnecessary repetition, one or more features shown and described in association with one embodiment may be incorporated into other embodiments unless specifically described otherwise or if the one or more features would make an embodiment non-functional.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC

[0018] Further, the terminology in this description is not intended to limit the invention. For example, spatially relative terms-such as “beneath”, “below”, “lower”, “above”, “upper”, “proximal”, “distal”, and the like-may be used to describe the relation of one element or feature to another element or feature as illustrated in the figures. These spatially relative terms are intended to encompass different positions (i.e., locations) and orientations (i.e., rotational placements) of the elements or their operation in addition to the position and orientation shown in the figures. For example, if the content of one of the figures is turned over, elements described as “below” or “beneath” other elements or features would then be “above” or “over” the other elements or features. A device may be otherwise oriented and the spatially relative descriptors used herein interpreted accordingly. Likewise, descriptions of movement along and around various axes include various special element positions and orientations. In addition, the singular forms “a”, “an”, and “the” are intended to include the plural forms as well, unless the context indicates otherwise. Additionally, the terms “comprises”, “comprising”, “includes”, and the like specify the presence of stated features, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups.Components described as coupled may be electrically or mechanically directly coupled, or they may be indirectly coupled via one or more intermediate components.

[0019] Elements described in detail with reference to one embodiment, implementation, system, or module may, whenever practical, be included in other embodiments, implementations, systems, or modules in which they are not specifically shown or described. For example, if an element is described in detail with reference to one embodiment and is not described with reference to a second embodiment, the element may nevertheless be claimed as included in the second embodiment. Thus, to avoid unnecessary repetition in the following description, one or more elements shown and described in association with one embodiment, implementation, or application may be incorporated into other embodiments, implementations, or aspects unless specifically described otherwise, unless the one or more elements would make an embodiment or implementation non-functional, or unless two or more of the elements provide conflicting functions.

[0020] In some instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0021] This disclosure describes various devices, elements, and portions of computer-assisted systems and elements in terms of their state in three-dimensional space. As usedIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC herein, the term “position” refers to the location of an element or a portion of an element (e.g., three degrees of translational freedom in a three-dimensional space, such as along Cartesian x-, y-, and z-coordinates). As used herein, the term “orientation” refers to the rotational placement of an element or a portion of an element (e.g., three degrees of rotational freedom in three-dimensional space, such as about roll, pitch, and yaw axes, represented in angle-axis, rotation matrix, quaternion representation, and / or the like). As used herein, and for a device with a kinematic series, such as with a repositionable structure with a plurality of links coupled by one or more joints, the term “proximal” refers to a direction toward a base of the kinematic series, and “distal” refers to a direction away from the base along the kinematic series.

[0022] As used herein, the term “pose” refers to the multi-degree of freedom (DOF) spatial position and orientation of a coordinate system of interest attached to a rigid body. In general, a pose includes a pose variable for each of the DOFs in the pose. For example, a full 6-DOF pose for a rigid body in three-dimensional space would include 6 pose variables corresponding to the 3 positional DOFs (e.g., x, y, and z) and the 3 orientational DOFs (e.g., roll, pitch, and yaw). A 3-DOF position only pose would include only pose variables for the 3 positional DOFs. Similarly, a 3-DOF orientation only pose would include only pose variables for the 3 rotational DOFs. Further, a velocity of the pose captures the change in pose over time (e.g., a first derivative of the pose). For a full 6-DOF pose of a rigid body in three-dimensional space, the velocity would include 3 translational velocities and 3 rotational velocities. Poses with other numbers of DOFs would have a corresponding number of velocities translational and / or rotational velocities.

[0023] This disclosure occasionally refers to the disclosed techniques being applied to “patients” undergoing a “medical procedure” or “operation.” It should be appreciated that these references are not intended to limit the application of the disclosed techniques to applied medicine contexts. For example, the described techniques can be applied to facilitate physician training, equipment testing and / or calibration, and / or other contexts. Accordingly, any reference to the term “patient” is done for ease of explanation and also envisions the application of the described techniques to a generic “subject.”

[0024] The word “task” is used herein to refer to a discrete portion of procedure that may be autonomously, semi-autonomously, or manually implemented in furtherance of a procedure. For example, a task may be to move an endoscope to a particular portion, to advance an instrument to a particular depth, to replace an instrument coupled to aIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC manipulator, and so on. In some embodiments, a task is associated with component tasks to accomplish an overall goal. For example, a task to analyze a worksite may include component tasks related to moving an endoscope to view the worksite, advancing an instrument to predetermined depth, and enabling a functionality supported by the instrument.

[0025] Aspects of this disclosure are described in reference to computer-assisted systems, which can include devices that are teleoperated, externally manipulated, autonomous, semiautonomous, and / or the like. Further, aspects of this disclosure are described in terms of an implementation using a teleoperated surgical system, such as the da Vinci® Surgical System commercialized by Intuitive Surgical, Inc. of Sunnyvale, California. Knowledgeable persons will understand, however, that inventive aspects disclosed herein may be embodied and implemented in various ways, including teleoperated and non-teleoperated, and medical and non-medical embodiments and implementations. Implementations on da Vinci® Surgical Systems are merely exemplary and are not to be considered as limiting the scope of the inventive aspects disclosed herein. For example, techniques described with reference to surgical instruments and surgical methods may be used in other contexts. Thus, the instruments, systems, and methods described herein may be used for humans, animals, portions of human or animal anatomy, industrial systems, general robotic, or teleoperated systems. As further examples, the instruments, systems, and methods described herein may be used for non-medical purposes including industrial uses, general robotic uses, sensing or manipulating non-tissue work pieces, cosmetic improvements, imaging of human or animal anatomy, gathering data from human or animal anatomy, setting up or taking down systems, training medical or non-medical personnel, and / or the like. Additional example applications include use for procedures on tissue removed from human or animal anatomies (with or without return to a human or animal anatomy) and for procedures on human or animal cadavers. Further, these techniques can also be used for medical treatment or diagnosis procedures that include, or do not include, surgical aspects.

[0026] The disclosure generally relates to systems and methods for training and intelligently utilizing adaptive artificial intelligence (Al) or machine learning (ML) techniques and models to generate trainee guidance with respect to a training procedure. In general, the trainee guidance may provide guidance content to the trainee in one or multiple methodologies based on an assessment of multi-modal data relating to observed actions of the trainee during the training procedure.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC

[0027] In some embodiments, the trainee guidance may at least partially relate to use of a computer-assisted system such as a computer-assisted medical system to be used during a medical procedure. In these embodiments, the trainee guidance may include material specific to the computer-assisted system such as video, audio, and, or haptic feedback that instructs the trainee on how to perform or improve performance of tasks using the computer-assisted system. Furthermore, the trainee guidance may also, in whole or in part, include material that does not directly involve the computer-assisted system such as instructions for how to perform surgical procedures that do not utilize a computer-assisted system.

[0028] The adaptive Al and ML techniques and models described herein relate to dynamically adapting the inputs to machine learning models as a function of multi-modal data to be processed by the models and / or user inputs. In particular, the adapted inputs may include adaptive constitutions with modifiable instruction sets that direct how the machine learning model is to process the multi-modal data. As a result, the performance of the machine learning models can be adapted without the cost and time of tuning the underlying models.

[0029] To adapt a constitution, in some embodiments, the systems described herein may construct the instruction sets from different sets of predefined rules that are stored in a data store. For example, the sets of predefined rules may include rule sets relating to assessing trainee performance, selecting content of the trainee guidance, selecting a presentation methodology, identifying trends in the multi-modal data, identifying correlations, different trainee skill levels, detecting an activity of the training procedure, generating tasks for the training procedure, etc. For each type of rule set, the data store may include different versions of the rules to utilize in different scenarios. For example, the rule sets may correspond to different skill levels, different trainee response profiles, different training procedure difficulties, and so on.

[0030] The system may also include in the constitution a fixed set of rules that control how the target machine learning model should generally operate to produce the expected output. The fixed set of rules may also include instructions that control how the target machine learning model should consider the other predefined rules adaptively inserted into the constitution and define other overarching concerns to be considered by the model when generating the relevant outputs. Accordingly, the systems described herein may be configured to process multi-model input data and / or user inputs to identify an appropriate set of rules and update a constitution in accordance therewith. In some embodiments, if the systems do notIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC detect an appropriate set of rules, the systems may utilize a generative Al model to generate one or more rules intended to implement the desired function of the target ML model.

[0031] By modifying the constitutions based on the training scenario, the systems described herein are able to dynamically optimize the outputs of a given model for a specific task at hand, environment, and / or conditions without having to fine-tune the models and systems.

[0032] FIG. 1 is a simplified diagram of an example computer-assisted system 100, according to various embodiments. The computer-assisted system 100 may be a computer-assisted medical system for assisting with performing tasks for medical procedures. Further, the computer-assisted system 100 may utilize adaptable Al models for providing preoperative and / or intra-operative guidance in relation to the computer-assisted system 100 as described herein. In some examples, the computer-assisted system 100 is a teleoperated system. In medical examples, the computer-assisted system 100 can be a teleoperated medical system such as a surgical system. As shown, the computer-assisted system 100 includes a follower device 104 that can be teleoperated by being controlled by one or more leader devices (also called “leader input devices” when designed to accept external input), described in greater detail below. Systems that include a leader device and a follower device are referred to as leader-follower systems, and also sometimes referred to as master-slave systems. Also shown in FIG. 1 is an input system that includes a workstation 102 (e.g., a console), and in various embodiments the input system can be in any appropriate form and may or may not include the workstation 102.

[0033] In the example of FIG. 1, the workstation 102 includes one or more leader input devices 106 that are designed to be contacted and manipulated by an operator 108. For example, the workstation 102 may comprise one or more leader input devices 106 for use by the hands, the head, or some other body part(s) of operator 108. The leader input devices 106 in this example are supported by the workstation 102 and can be mechanically grounded. In some embodiments, an ergonomic support 110 (e.g., forearm rest) can be provided on which the operator 108 can rest his or her forearms. In some examples, the operator 108 can perform tasks at a worksite within a workspace near the follower device 104 during a procedure, by commanding the follower device 104 using the leader input devices 106. In a medical example, the worksite may be a surgical worksite associated with a patient.

[0034] A display device 112 is also included in the workstation 102. The display device 112 may be configured to display images for viewing by the operator 108. The display deviceIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC 112 can be moved in various DOFs to accommodate the viewing position of the operator 108 and / or to provide control functions. In embodiments where the display device 112 provides control functions, the leader input devices 106 may include the display device 112. In the example of the computer-assisted system 100, displayed images may depict a worksite at which the operator 108 is performing various tasks by manipulating the leader input devices 106 and / or the display device 112. In some examples, images displayed by display device 112 may be received by the workstation 102 from one or more imaging devices arranged at a worksite. In other examples, the images displayed by the display device 112 may be generated by the display device 112 (or by a different connected device or system), such as for virtual representations of tools, the worksite, or for user interface components. As will be explained below, in some embodiments the display device 112 may display pre-operative and / or intra-operative guidance such as instrument guidance in relation to the computer-assisted system 100 or the medical procedure being performed thereby.

[0035] In examples, the display device 112 may be a touch-screen device and may receive user-input via the touch screen. The touch screen may provide a user interface that a user may select from various options to provide one or more user preferences for performing a procedure or medical task. Additionally, the display device 112 may display one or more images such as an intraoperative or pre-operative image of a patient, organ, surgical site etc., and the user may provide user input via the touch screen to indicate one or more regions or elements displayed in the intraoperative or pre-operative images. The display device 112 may provide one or more images or user interfaces and a user may interact with and provide user indications and input via another device such as a keyboard, mouse, audio device, etc. For example, a user may provide text input such as natural language text via a keyboard to provide a user input to the system. A user may use a mouse or another similar device to click on or indicate selection of an option or user feedback based on one or more images presented by the display device. The workstation 102 may further include a microphone that may capture and record audio of a user providing indications of user preferences and user inputs to the system. One or more processors, or controllers as discussed further herein, may then derive natural language text from the recorded audio data to derive a user input to the system 100. In examples, the user input may further be determined from the audio data as a user responding to one or more prompts to confirm, reject, or otherwise indicate a user preference or input. Additionally, user input may be provided via one or more sources such as video, haptic input, data banks, instruments (e.g., by a user changing a setting on an instrument orIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC auxiliary device), sensors, etc. It should be appreciated that while the foregoing describes obtaining the user input from the display device 112 of the workstation 102, in other embodiments, the user input may be obtained by other display devices associated with the operating room that are interacted with by personnel other than the operator 108.

[0036] As illustrated, the computer-assisted system 100 also includes a follower device 104 that can be commanded by the workstation 102. In a medical example, the follower device 104 can be located near an operating table (e.g., a table, bed, or other support) on which a patient can be positioned. In some medical examples, the workspace is provided on an operating table, e.g., on or in a patient, simulated patient, or model, training dummy, etc. (not shown). As illustrated, the follower device 104 may include a plurality of repositionable structures 120 (sometimes referred to as “manipulator arms” in robotic embodiments). In some embodiments, the repositionable structures 120 may include a plurality of links that are rigid members and joints that can be individually actuated as part of a kinematic series. Additionally, each of the repositionable structures 120 is configured to couple to an instrument 122. While FIG. 1 illustrates a follower device 104 that has four repositionable structures 120a- 120d, in other embodiments, the follower device 104 may include one, two, three, four, five, six, or additional or fewer repositionable structures 120a-120d.

[0037] The instrument 122 can include, for example, a working portion 126 and one or more structures for supporting and / or driving the working portion 126. Example working portions 126 include end effectors that physically contact or manipulate material, energy application elements that apply electrical, RF, ultrasonic, or other types of energy, sensors that detect characteristics of the workspace environment (such as temperature sensors, imaging devices, etc.), and the like. In various embodiments, examples of instruments 122 include, without limitation, a sealing instrument, a cutting instrument, a sealing-and-cutting instrument, an energy instrument for applying energy, a gripping instrument e.g., clamps, jaws), a stapler, an imaging instrument such as one using optical, RF, or ultrasonic imaging modalities, a sensing instrument, an irrigation instrument, a suction instrument, and / or the like. In addition, the instrument 122 may include a transmission mechanism 128 that can be coupled to a drive assembly 130 of the respective repositionable structure 120a- 120d. The drive assembly 130 may include a drive and / or other mechanisms controllable from workstation 102 that transmit forces to the transmission mechanism 128 to articulate or otherwise actuate the instrument 122.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC

[0038] As illustrated, each instrument 122 may be mounted to a portion of a respective repositionable structure 120a-120d. In FIG. 1, this is shown with the drive assembly 130 physically coupled to the transmission mechanism 128. The distal portion of each repositionable structure 120a-120d further includes a cannula mount 124 to which a cannula (not shown) is mounted. When a cannula is mounted to the cannula mount 124, a shaft of the instrument 122 passes through the cannula and into a workspace.

[0039] In various embodiments, one or more of the working portions 126 of the instruments 122 may include an imaging device for capturing images. The imaging device may include any sensing technology capable of acquiring an image. Example imaging instruments include an optical endoscope, a hyperspectral camera, an ultrasonic sensor, etc. Imaging instruments may comprise monoscopic imagers, stereoscopic imagers, and / or the like. Imaging devices based on radiofrequency domains may capture images in any frequency spectrum, including visible light, infrared light, ultraviolet light, and / or the like. The imaging device may include an illumination source to light the region being imaged. In embodiments where the working portions 126 of one or more of the instruments 122 include an imaging device, the instrument 122 may be configured to capture images of a portion of the workspace for display via the display device 112.

[0040] In some embodiments, the repositionable structures 120a- 120d and / or instruments 122 can be controlled to move the working portion 126 in response to manipulation of the leader input devices 106 by the operator 108. Accordingly, the repositionable structures 120a-120d and / or instruments 122 may be said to “follow” the leader input devices 106 through teleoperation. This enables the operator 108 to perform tasks at the worksite using the repositionable structures 120a- 120d and / or instruments 122. For a surgical example, the operator 108 can direct the repositionable structures 120a- 120d of the follower device 104 to move the working portions 126 as part of a surgical procedure performed at an internal surgical site that is entered via one or more minimally invasive apertures or natural orifices. It should be appreciated that, in some embodiments, the follower device 104 may include nonteleoperated components that the operator 108 or other medical professional must manually manipulate to a desired pose.

[0041] In some embodiments, a repositionable structure 120a of the computer-assisted system 100 may be configured to support a working portion 126a that includes an imaging device (also referred to herein as an “imaging device 126a”). For convenience, an instrument 122 that includes an imaging device is also referred to as an “imaging instrument” herein.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC The control system 140 may be configured to command the repositionable structure 120a and / or the imaging instrument 122 comprising the imaging device 126a to automatically position and / or orient (“pose”) the field of view (FOV) of the imaging device 126a to provide images of the workspace and / or other instruments 122.

[0042] In the illustrated embodiment, a control system 140 is communicatively coupled to the workstation 102. In other embodiments, the control system 140 may be provided as a component of the workstation 102 and / or the follower device 104. During teleoperation, as the operator 108 moves the leader input device(s) 106, one or more sensors configured to detect the leader input device(s) 106 generate spatial and / or orientation movement data that is provided to control system 140. The control system 140 may interpret the spatial and / or orientation information to determine and / or provide control signals to the follower device 104 to control the movement of repositionable structures 120a- 120d, instruments 122, and / or working portions 126. In addition to the components of the follower device 104, in some embodiments, the control system 140 is configured to interpret inputs received from the workstation 102 to control operation of one or more auxiliary devices (not depicted) utilized in a procedure. For example, the workstation 102 may be used to control a pose of a surgical bed or operation of an insufflator.

[0043] In one embodiment, the control system 140 supports one or more wired communication protocols, (e.g., Ethernet, USB, and / or the like) and / or one or more wireless communication protocols (e.g., Bluetooth, IrDA, HomeRF, IEEE 1102.11, DECT, Wireless Telemetry, and / or the like) for communications between the control system 140 and the workstation 102 and / or the follower device 104.

[0044] In some embodiments, the control system 140 may be implemented at one or more computing systems. For example, one or more computing systems may be used to control the follower device 104. As another example, one or more computing systems may be used to control components of the workstation 102, such as movement of a display device 112.

[0045] As illustrated, the control system 140 includes a processor system 150, a memory 160, and an artificial intelligent (Al) assist module 180. The memory 160 may store a control module 170. The processor system 150 may include one or more processors having different processing architectures for processing instructions. For example, the one or more processors may be one or more cores or micro-cores of a multi-core processor, a central processing unit (CPU), a microprocessor, a field-programmable gate array (FPGA), an application- specificIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC integrated circuit (ASIC), a digital signal processor (DSP), a graphics processing unit (GPU), a tensor processing unit (TPU), and / or the like.

[0046] In some embodiments, the processor system 150 includes circuity to support one or more communication interfaces (e.g., Bluetooth interface, infrared interface, network interface, optical interface, etc.). Additionally, a communication interface of control system 140 may include an integrated circuit for connecting the control system 140 to a network (not shown) (e.g., a local area network (LAN), a wide area network (WAN) such as the Internet, mobile network, or any other type of network) and / or to another device, such as the workstation 102 and / or the follower device 104.

[0047] Additionally, the memory 160 may include non-persistent storage (e.g., volatile memory, such as random access memory (RAM), cache memory), persistent storage (e.g., a hard disk, an optical drive such as a compact disk (CD) drive or digital versatile disk (DVD) drive, a flash memory, a floppy disk, a flexible disk, a magnetic tape, any other magnetic medium, any other optical medium, programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a FLASH-EPROM, and / or any other memory chip or cartridge. The non-persistent storage and persistent storage are examples of non-transitory, tangible machine-readable media that can store executable code that, when run by one or more processors (e.g., processor system 150), can cause the one or more processors to perform one or more of the techniques and / or methods disclosed herein.

[0048] The Al assist module 180 may implement one or more machine learning models and / or training protocols. For example, the Al assist module 180 may implement one or more neural networks, deep learning models, decision trees, support vector machines, linear regression, generative Al models, reinforced learning models, random forests, Naive Bayes models, large language models (LLMs), generative adversarial networks, foundation models, image recognition models, linear discriminant analysis models, creative applications, autoregressive models, supervised or unsupervised learning models, multimodal models, vision language models (VLMs), vision foundation models (VFMs), large multi-modal models (LMMs), Transformer models (including Robotic Transformer models), or another machine learning or Al model for performing the methods described herein. The structure of the one or more machine learning is described in more detail with respect to FIGs. 2A-3. The Al assist module 180 may include dedicated processors and memory for storing and performing Al processes, or the Al assist module 180 may utilize resources of the processorIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC system 150 and the memory 160 to store and / or perform any processing or tasks required to perform the methods described herein.

[0049] Additionally, the control system 140 may also include one or more input devices (such as a touchscreen, keyboard, mouse, microphone, touchpad, electronic pen, or any other type of input device) and / or output devices (such as a display device, a speaker, external storage, a printer, or any other output device). In some embodiments, the control system 140 may be implemented on a particular node of a distributed computing system (e.g., a cloud computing system). As another example, different functionalities associated with the control system 140 may be implemented on different nodes of the distributed computing system. Further, one or more elements of the aforementioned control system 140 may be located at a remote location and connected to the other elements over a network.

[0050] In an endoscopic surgery example, the imaging instrument comprising the imaging device 126a may be inserted into the patient prior to the other instruments 122, including a second instrument 122b comprising a second working portion 126b. The second instrument 122b can include any appropriate working portion 126b, and can even include a second imaging device. Accordingly, the imaging device 126a may be maneuvered to be positioned to identify a target to which other instruments may interact with as part of another task. The control system 140 may, for example, automatically command the corresponding repositionable structures 120a and 120b to position respective instruments 122a and 122b to perform one or more tasks in tandem, or sequentially based on the specific task, instruments, and positions of the repositionable structures 120a and 120b. In examples, the control system 140 may perform Al processes and algorithms via the Al assist module 180 to provide the preoperative or intra-operative guidance as described herein. The control system 140, via the Al assist module 180, may also perform various task selection and identification processes such as those disclosed in U.S. Provisional application 63 / 667,234 titled “Multi-Task Al System For Dynamic and Intelligent Robotic Task Planning Based On User Input”, which is incorporated by reference herein in its entirety. Accordingly, the control system 140 is able to identify and control repositionable structures to perform tasks for a medical procedure in a variety of scenarios using adaptable Al.

[0051] In some embodiments, the computer-assisted system 100 may be used as part of a training procedure as described herein. The training procedure may include real practice procedures where the follower device 104 is utilized on cadavers or other suitable human analogs and simulated procedures where the follower device 104 is disconnected from theIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC workstation 102 and the control system 140. Instead, in simulated procedure embodiments, a simulation module as described herein may simulate responses from the follower device 104 to inputs from the operator 108 on the workstation 102. Furthermore, while examples described herein generally relate to training procedures implemented using the computer-assisted system 100, it is envisioned that the disclosed techniques can be implemented in conventional procedures that do not utilize computer-assisted robotic systems. As will be explained below, in these embodiments, the Al models described herein may be configured in a similar manner, but receive fewer data modalities as part of the multi-modal data streams described herein.

[0052] FIG. 2A is a schematic diagram of a system 200A for using adaptive Al for generating trainee guidance in relation to a training procedure. The system 200A includes a plurality of software or hardware modules that may be executed by a processing unit 201 to generate the trainee guidance. The processing unit 201 may include the control system 140 and the Al assist module 180 of FIG. 1 or other processing components known in the art. For example, the processing unit 201 may include one or more processors, each of which may be a programmable microprocessor or the like that executes software instructions stored in a memory unit 203 (e.g., memory 160 of FIG. 1 or another suitable memory unit not directly associated with the control system 140 of FIG. 1) to execute some or all of the functions of the system 200A as described herein. The processing unit 201 may include one or more graphics processing units (GPUs) and / or one or more central processing units (CPUs), for example. Alternatively, or in addition, one or more processors in the processing unit 201 may be other types of processors (e.g., application- specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc.), and some of the functionality of the system 200A as described herein may instead be implemented in hardware.

[0053] As shown in FIG. 2A, the system 200A includes a trainee performance assessment machine learning model 202 and a trainee guidance machine learning model 204, which are both executable by the processing unit 201.

[0054] Generally, the trainee performance assessment machine learning model 202 is configured to generate a trainee performance assessment 212 by analyzing the training procedure multi-modal data 206 as directed by a trainee performance assessment constitution 208.

[0055] The trainee performance assessment machine learning model 202 may include a set of interconnected nodes, layers, trained parameter values (e.g., multiplicative weights,Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC additive bias, etc.), etc. The trained parameters may be derived via backpropagation techniques in a training process that uses historical data inputs. An example architecture of the trainee performance assessment machine learning model 202 is described in more detail in connection with FIG. 3.

[0056] In some embodiments, the trainee performance assessment machine learning model 202 is configured using a multi-modal transformer type architecture that receives inputs of different modalities relating to the training procedure. However, it should be appreciated that this architecture may be substituted or supplemented with other Al architectures including, but not limited to, convolutional neural network (CNN) architectures, recurrent / recursive neural network (RNN) architectures, sorting / clustering architectures, etc.

[0057] The training procedure multi-modal data 206 input into the trainee performance assessment machine learning model 202 and the trainee guidance machine learning model 204 may generally include observation data relating to observed actions of one or more trainees during the training procedure. The observation data may include real time data streams or similar from the computer-assisted system 100 or a simulation module, image data of the training during the training procedure, audio data of the training during the training procedure, other user input data received by the user interface 219 during the training procedure, procedure image data, etc.

[0058] With simultaneous reference to FIG. 2C, the observation data portions of the trainee procedure multi-modal data 206 are described in more detail. As shown in FIG. 2C, the observation data portions of the training procedure multi-modal data 206 may be provided by the computer-assisted system 100, a simulation module 232, an image sensor 233, and the user interface 219.

[0059] As shown in FIG. 2C, the observation data portions of the multi-modal data 206 may include data received from the computer-assisted system 100, such as real force data 206A, real event data 206B, real kinematics data 206C, and real procedure image data 206D.. In particular, the real force data 206A may include data indicating forces exerted on the instruments 122 and / or the repositionable structures 120 during the training procedure. The real event data 206B may include data indicating events such as instrument installs, insertions, removals, etc. for the computer-assisted system 100 during the training procedure. The real kinematics data 206C may include kinematic data indicating the pose and / or movements of the instruments 122 and / or the repositionable structures 120 of the computer-assisted system 100 during the training procedure. The real procedure image data 206D mayIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC include image data of the procedure from one or more imaging devices of the computer-assisted system 100 (e.g., image data the is generated by an endoscopic instrument 122).

[0060] In embodiments where the training procedure is a simulated training procedure, the observation data portions of the multi-modal data 206 may include simulated data generated by the simulation module 232, such as simulated force data 206E, simulated event data 206F, simulated kinematics data 206G, and simulated procedure image data 206H. The simulated data may be simulated equivalents of the real data streams described above. In these embodiments, the simulated data may be obtained via an application programing interface of the simulation module 232.

[0061] As shown in FIG. 2C, the observation data portions of the multi-modal data 206 may also include external video data 2061 showing the location where the training procedure is occurring and the trainee while undergoing the training procedure. The external video data 2061 may be generated by an image sensor 233 located in the room in which the training procedure is occurring.

[0062] In some embodiments, the observation data portions of the multi-modal data 206 may also include the user input data 206 J from the user interface 219. The user input data 206 J may include audio text or other user inputs captured by the user interface 219. In some embodiments, the user interface 219 may be a intra-simulation user interface presented by the simulation module 232.

[0063] The training procedure multi-modal data 206 may also include background data. The background data may include data obtained from the data store 218, such as trainee data related to a particular trainee being trained as part of the training procedure, historical trainee data related to training of past trainees during past training procedure, procedure data that defines generally accepted practices for different types of procedures, historical event data from the computer-assisted system 100 that are associated with different performance assessments (e.g., good, bad, acceptable, etc.), historical video data from past procedures, etc.

[0064] In some embodiments, the trainee data may include a historical trainee performance assessment for one or more prior performances of a training exercise that is associated with a training skillset and is to be performed multiple times during the training procedure. For example, the historical trainee performance assessment may indicate that the trainee failed to successfully complete the training exercise. In some embodiments, the trainee data may include trainee profiles from a trainee database that indicate historical trainee assessments associated with historical training procedures for the particular individual trainee beingIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC trained as part of the training procedure or other trainees that have similar characteristics as the particular individual trainee. The trainee profiles may indicate one or more skill levels related to different aspects of various procedures (e.g., suturing, instrument navigation, tool utilization, etc.), which the processing unit 201 may use to select the training procedure to perform, the number and type of training exercises to perform as part of the training procedure, and / or the set of rules to include in the constitutions described herein.

[0065] Additionally, the background data may include intra-procedure background data generated by the trainee performance assessment machine learning model 202, trainee guidance machine learning model 204, or other model described herein (see e.g., activity detection machine learning model 220 and task generation machine learning model 222 shown in FIG. 2B). For example, the intra-procedure data may include past trainee performance assessments 212 and past trainee guidance 214 generated in relation to prior tasks of the training procedure.

[0066] The trainee performance assessment constitution 208 directs the trainee performance assessment machine learning model 202 on how to analyze the training procedure multi-modal data 206 to generate the trainee performance assessment 212. In particular, trainee performance assessment constitution 208 includes a plurality of rules that control or otherwise instruct the trainee performance assessment machine learning model 202 on how to generate the trainee performance assessment 212.

[0067] The plurality of rules may include different rules that relate to different portions of the training procedure multi-modal data 206. For example, the trainee performance assessment constitution 208 may include rules that direct how the trainee performance assessment machine learning model 202 evaluates a performance of the observed actions of the trainee (e.g., the observation data described herein) when generating the trainee performance assessment 212. In some embodiments, the rules in the trainee performance assessment constitution 208 may include trend identifying rules that direct the trainee performance assessment machine learning model 202 to identify and evaluate trends across one or more prior performances of a training exercise as part of the training procedure. In some embodiments, the trend identifying rules may include rules that direct the trainee performance assessment machine learning model 202 to determine a correlation between prior performance of the training exercise and past trainee guidance 214 to evaluate the effectiveness of the past trainee guidance 214. Similarly, the trend identifying rules may also include rules that direct the trainee performance assessment machine learning model 202 toIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC determine a correlation between training materials provided as part of past trainee guidance 214 and a subsequent performance of the training exercise to evaluate the effectiveness of the training materials.

[0068] The rules in the trainee performance assessment constitution 208 may also include a set of trainee skill level rules that control how the trainee performance assessment machine learning model 202 generates the trainee performance assessment 212. For example, the trainee performance assessment constitution 208 may direct the trainee performance assessment machine learning model 202 how to evaluate the one or more skill levels indicated by trainee profiles or other sections of the training procedure multi-modal data 206 when assessing training performance. As a result, the trainee performance assessment 212 may be tailored to the trainee skill level.

[0069] The trainee performance assessment constitution 208 may also include trainee observation rules that directs how the trainee performance assessment machine learning model 202 analyzes the body language, words, tone, etc. used by the trainee when performing the observed actions represented by the training procedure multi-modal data 206. These rules may inform the trainee performance assessment machine learning model 202 how comfortable the trainee is performing the training procedure.

[0070] The trainee performance assessment constitution 208 may also include sets of fixed rules that are present in each version thereof. For example, the trainee performance assessment constitution 208 may include rules to process the training procedure multi-modal data 206 as it is generated to monitor and assess the trainee’s performance of the training procedure in accordance with other rules included in the constitution 208.

[0071] The trainee performance assessment 212 output from the trainee performance assessment machine learning model 202 may include numbers, text, values, ranges, etc. that represent objective performance indicators of the observed actions of the trainee (e.g., the observation data of the training procedure multi-modal data 206). For example, the trainee performance assessment 212 may include natural language text describing whether the observed actions of the trainee demonstrate a good, bad, or neutral execution of the training procedure (and / or component tasks thereof). In some embodiments, the trainee performance assessment machine learning model 202 may generate the trainee performance assessment 212 by referencing the background data portions of the training procedure multi-modal data 206.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC

[0072] Furthermore, in some embodiments, the trainee performance assessment 212 may include a score that indicates where the trainee’s performance is located on a preconfigured spectrum of performances. In these embodiments, the trainee performance assessment constitution 208 may contain a rubric that defines the kinds of observation data associated with different score values or ranges of values that the trainee performance assessment machine learning model 202 references when generating the score included in the trainee performance assessment 212.

[0073] The trainee guidance machine learning model 204 is configured to generate the trainee guidance 214 by analyzing the training procedure multi-modal data 206 and the trainee performance assessment 212 as directed by trainee guidance constitution 210. The trainee guidance machine learning model 204 may include a set of interconnected nodes, layers, trained parameter values (e.g., multiplicative weights, additive bias, etc.), etc. The trained parameters may be set derived b ackpropagation techniques in a training process that uses historical data inputs. An example architecture of the trainee guidance machine learning model 204 is described in more detail in connection with FIG. 3.

[0074] In some embodiments, the trainee guidance machine learning model 204 is configured using a multi-modal transformer type architecture that receives inputs of different modalities relating to the training procedure. However, it should be appreciated that this architecture may be substituted or supplemented with other Al architectures including, but not limited to, convolutional neural network (CNN) architectures, recurrent / recursive neural network (RNN) architectures, sorting / clustering architectures, etc.

[0075] The trainee guidance constitution 210 directs the trainee guidance machine learning model 204 on how to analyze the training procedure multi-modal data 206 and the trainee performance assessment 212 to generate the trainee guidance 214. In particular, the trainee guidance constitution 210 includes a plurality of rules that control or otherwise instruct the trainee guidance machine learning model 204 on how to generate the trainee guidance 214.

[0076] The plurality of rules may include different rules that relate to different portions of the training procedure multi-modal data 206. For example, the trainee guidance constitution 210 may include rules that direct how the trainee guidance machine learning model 204 determines guidance content and guidance methodology (e.g., the channel via which the guidance content is provided of the trainee guidance 214.

[0077] Additionally, the rules in the trainee guidance constitution 210 may include a set of trainee skill level rules that control how the trainee guidance machine learning model 204Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC generates the trainee guidance 214 in view of one or more skill levels indicated by any trainee profiles included in the training procedure multi-modal data 206. In some embodiments, set of trainee skill level rules in the trainee guidance constitution 210 may control a frequency or a methodology of guidance provided by the trainee guidance machine learning model 204 during execution of the training procedure.

[0078] For example, the rules of the trainee guidance constitution 210 may include rules related to analyzing the training procedure multi-modal data 206 to detect that the trainee is likely to perform the training procedure (and / or a task thereof) incorrectly. In this example, the trainee guidance constitution 210 may adapt the intra-procedure training guidance 214 based on the trainee skill level. For example, the rules may indicate that a less skilled trainee should receive additional guidance to ensure proper performance of the training procedure. On the other hand, the rules may indicate that a more skilled trainee is permitted to incorrectly complete the training procedure and provide a prompt after the completion of the training procedure that causes the trainee to self-reflect on where the training procedure went wrong.

[0079] Additionally, the trainee guidance constitution 210 may also include tonal rules that direct the tone, emotion, sentiment, prosody, or diction of the guidance content output from the trainee guidance machine learning model 204. In some embodiments, as the trainee is provided with guidance in different tonalities, the systems update a trainee profile to reflect the positive and negative response to different tonal characteristics. Accordingly, the rules may cause the trainee guidance machine learning model 204 to output the guidance in a tonality that is most likely to have a positive impact on the trainee’s growth. Generally, the tonal rules may include rules to: (1) ask thought-provoking, open-ended questions that challenge a trainee’s preconceptions and encourage the trainee to engage in deeper reflection and critical thinking; (2) actively listen to the trainee’s responses, paying careful attention to their underlying thought processes and making a genuine effort to understand the trainee’s perspectives; (3) choose guidance methodology based on provided objective performance indicators of the trainee; (4) guide the trainee in exploration of topics by encouraging the trainee to discover answers independently, rather than providing direct answers, to enhance the trainee’s reasoning and analytical skill; (5) promote critical thinking by encouraging the trainee to question assumptions, evaluate evidence, and consider alternative viewpoints in order to arrive at well-reasoned conclusions; and / or (6) demonstrate humility byIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC acknowledging limitations and uncertainties, modeling a growth mindset and exemplifying the value of lifelong learning when instructing the trainee.

[0080] In some embodiments, the trainee guidance constitution 210 may also include rules that direct the trainee guidance machine learning model 204 to identify, based on the trainee guidance machine learning model 204 and the trainee performance assessment 212, training materials for the trainee to review before performing a subsequent training exercise or training procedure. In these examples, the trainee guidance 214 may include an indication of the identified training materials (e.g., an indication or link to a location of such materials in the data store 218) or the training materials (and / or excerpt therefrom) itself. In some further embodiments, the trainee guidance machine learning model 204 may be configured to implement retrieval augmented generation (RAG) techniques on a corpus of training documents to generate training materials adapted to the trainee and the trainee performance assessment 212.

[0081] The trainee guidance 214 output from the trainee guidance machine learning model 204 includes the guidance content and / or the guidance methodology. In scenarios where multiple trainees are performing the training procedure, the trainee guidance 214 may also indicate the trainee to which the guidance is to be provided. The guidance content includes elements generated and output from the trainee guidance machine learning model 204 that are configured to instruct or assist the trainee on improving execution of current or future aspects of the training procedure such as future training exercises. For example, the guidance content may include a prompt to the trainee and / or a text, audio, and / or video description of how to perform a specific action associated with the training procedure. The guidance methodology includes a particular modality selected by the trainee guidance machine learning model 204 in which to provide guidance content. For example, the guidance methodology may include one or more of visual guidance, audio guidance, robotic guidance, augmented reality guidance, virtual reality guidance, and a modification of a training exercise. It should also be appreciated that the form and details of the guidance content may be dependent on the selected guidance methodology. Additional details of the guidance content and methodology are described in connection with FIG. 2D.

[0082] It should be appreciated that the trainee guidance 214 may also include temporal indications as to when the guidance content is to be provided to the trainee. For example, the temporal indications may indicate whether to provide the guidance as an intervention to prevent failure of the training procedure and / or demonstrate proper performance of theIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC training procedure, or to provide the guidance as a retrospective guidance after the completion of the training procedure. As another example, the trainee guidance 214 may be provided in relation to a performance of a particular task or exercise of the training procedure before the trainee moves on to a next task or exercise.

[0083] As discussed above, in some embodiments the disclosed constitutions may be adapted to the trainee and / or training procedure. Accordingly, the system 200A includes a constitution generator 216 to generate, instantiate, and / or adapt the trainee performance assessment constitution 208 and the trainee guidance constitution 210 by dynamically selecting the plurality of rules to include therein. For example, the constitution generator is configured to select the plurality of rules based on user input provided via the user interface 219 (e.g., user input data 206 J shown in FIG. 2C), a trainee profile maintained in the data store 218, and / or data included in the training procedure multi-modal data 206.

[0084] To receive user-directed inputs, the user interface 219 may include a chatbot interface and / or an intraoperative imaging interface via which the user is able to provide inputs related to the training procedure. For example, the user interface 219 may guide the user through the process of selecting a training procedure in accordance with a training plan and configure the disclosed constitutions to perform the selected training procedure. In some scenarios, the user interface 219 may further enable the user to provide instructions to modify the training procedure and / or the trainee guidance 214. For example, the user may input instructions to not receive the trainee guidance 214 until after the procedure is complete, to receive robotic-assisted guidance during the training procedure, and so forth.

[0085] The constitution generator 216 may then be configured to process the user-provided (and / or otherwise obtained) inputs to select the appropriate sets of rules to include in the disclosed constitutions. Accordingly, the constitution generator 216 may include an algorithmic software or hardware module executable by the processing unit 201 according to instructions stored in the memory unit 203. Alternatively, the constitution generator 216 may include a machine learning model with trained or tuned parameter values for selecting the rules for the trainee performance assessment constitution 208, trainee guidance constitution 210, and other constitutions as described herein.

[0086] Once the trainee guidance machine learning model 204 generates the trainee guidance 214, the processing unit 201 provides the trainee guidance 214 to the user interface 219 (and / or other output devices associated with the trainee guidance methodology) to direct performance of the training procedure by the trainee. The user interface 219 may include aIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC display unit such as the display device 112 included in the workstation 102 of the computer-assisted system 100 of FIG. 1. Additionally, or alternatively, in some embodiments the display unit may include one or more display devices located at the location where the training procedure is taking place (see e.g., the display 238 shown in FIG. 2D).

[0087] As shown in FIG. 2A, the trainee guidance 214 may also be provided to the trainee performance assessment machine learning model 202 and the trainee guidance machine learning model 204 so that the models may consider the effectiveness of the trainee guidance 214 in modifying or correcting the observed actions of the trainee when generating future trainee performance assessments 212 and future trainee guidance 214.

[0088] FIG. 2B is a schematic diagram of another system 200B for trainee guidance. The system 200B includes a plurality of software or hardware modules that may be executed by the processing unit 201 to generate the trainee guidance 214. This plurality of software or hardware modules may include some or all of the software or hardware modules of the system 200A shown in FIG. 2A and the associated inputs and outputs thereof. For example, the system 200B may include the trainee performance assessment machine learning model 202, the trainee guidance machine learning model 204, the training procedure multi-modal data 206, the trainee performance assessment 212, the trainee guidance 214, the constitution generator 216, the data store 218, and the user interface 219 as described in connection with FIG. 2A. Additionally, it should also be appreciated that the system 200B also includes the trainee performance assessment constitution 208 and the trainee guidance constitution 210 of FIG. 2A as inputs into the trainee performance assessment machine learning model 202 and trainee guidance machine learning model 204, respectively. Furthermore, the training procedure multi-modal data 206 are also input into the trainee performance assessment machine learning model 202 and the trainee guidance machine learning model 204 in the system 200B in the manner described in connection with the system 200A of FIG. 2A.

[0089] As shown in FIG. 2B, the system 200B also includes additional software or hardware modules and associated inputs and outputs beyond those of the system 200A shown in FIG. 2A. The additional elements of the system 200b may include an activity detection machine learning model 220 and a task generation machine learning model 222, which are both executable by the processing unit 201.

[0090] The activity detection machine learning model 220 is configured to generate the activity indicator 228 by analyzing the training procedure multi-modal data 206 as directed by the activity detection constitution 224. The activity detection machine learning model 220Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC may comprise a set of interconnected nodes, layers, trained parameter values (e.g., multiplicative weights, additive bias, etc.), etc. The trained parameters may be set via backpropagation techniques in a training process that uses historical data inputs. An example architecture of the activity detection machine learning model 220 is described in more detail in connection with FIG. 3.

[0091] In some embodiments, the activity detection machine learning model 220 is configured using a multi-modal transformer type architecture that can receive inputs of different modalities relating to the training procedure. However, it should be appreciated that this architecture may be substituted or supplemented with other Al architectures including, but not limited to, convolutional neural network (CNN) architectures, recurrent / recursive neural network (RNN) architectures, sorting / clustering architectures, etc.

[0092] The activity detection constitution 224 directs the activity detection machine learning model 220 on how to analyze the training procedure multi-modal data 206 to generate the activity indicator 228. In particular, the activity detection constitution 224 includes a plurality of rules that control or otherwise instruct the activity detection machine learning model 220 on how to generate the activity indicator 228. The plurality of rules may include different rules that relate to different portions of the training procedure multi-modal data 206. For example, the activity detection constitution 224 may include rules that direct how the activity detection machine learning model 220 analyzes the training procedure multimodal data 206 to identify a particular activity of the training procedure that is being performed by the trainee. In some embodiments, the activity detection constitution 224 may direct the activity detection machine learning model 220 to analyze the training procedure multi-modal data 206 by referencing the background data portions of the training procedure multi-modal data 206.

[0093] As another example, the activity detection constitution 224 may include rules that define particular activities that can be performed by the trainee. Accordingly, the rules related to outputting the activity indicator 228 may include rules to output a particular activity defined in the activity detection constitution 224. The activity definitions may include associations with different values of the training procedure multi-modal data 206 that indicate performance of the defined activity. It should be appreciated that the particular activities defined in the activity detection constitution 224 may by dynamically included by the constitution generator 216 based on the particular training procedure being performed...Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC

[0094] The system 200B may analyze the activity indicator 228 to determine one or more tasks to perform in furtherance of the training procedure. Accordingly, the task generation machine learning model 222 is configured to generate the tasks 230 by analyzing the training procedure multi-modal data 206 and / or the activity indicator 228 as directed by the task generation constitution 226. The task generation machine learning model 222 may comprise a set of interconnected nodes, layers, trained parameter values (e.g., multiplicative weights, additive bias, etc.), etc. The trained parameters may be set via backpropagation techniques in a training process that uses historical data inputs. An example architecture of the task generation machine learning model 222 is described in more detail in connection with FIG. 3.

[0095] In some embodiments, the task generation machine learning model 222 is configured using a multi-modal transformer type architecture that can receive inputs of different modalities relating to the training procedure. However, it should be appreciated that this architecture may be substituted or supplemented with other Al architectures including, but not limited to, convolutional neural network (CNN) architectures, recurrent / recursive neural network (RNN) architectures, sorting / clustering architectures, etc.

[0096] The task generation constitution 226 directs the task generation machine learning model 222 on how to analyze the training procedure multi-modal data 206 and / or the activity indicator 228 to generate the tasks 230. In particular, the task generation constitution 226 includes a plurality of rules that control or otherwise instruct the task generation machine learning model 222 on how to generate the tasks 230. The plurality of rules may include different rules that relate to different portions of the training procedure multi-modal data 206. and / or the activity indicator 228.

[0097] For example, the task generation constitution 226 may include rules that direct the task generation machine learning model 222 to generate the tasks 230 that implement autonomous and / or semi-autonomous functionality associated with the computer-assisted system 100 via a Robotic Transformer model (such as the Robotics Transformer model 242 of Fig. 2D). For example, the system 200B may implement the automated task generation and performance techniques described in U.S. Application No. 63 / 667,234 filed on July 3, 2024, there entire disclosure of which is hereby incorporated by reference. In addition to operational data described in the ’234 application, the task generation machine learning model 222 may also analyze historical training data to identify appropriate tasks. For example, the task generation machine learning model 222 may analyze training materials that indicate task sequencing, historical training assessments associated with a trainee, historicalIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC training assessments associated with other trainees with similar skill levels as the trainee, and so on.

[0098] In some embodiments, the automated tasks are implemented via a simulation supported by the simulation module 232. In these embodiments, the output tasks may be communicated to an API of the simulation module 232 such that a simulation engine implements the indicated tasks within a simulated environment.

[0099] In addition to generating autonomously-implemented tasks, the task generation machine learning model 222 may also output tasks to be performed by the trainee. For example, the task generation machine learning model 222 may generate specific tasks for the trainee to perform in a manual and / or semi-autonomous manner. In some embodiments the rules in the task generation constitution 226 may include rules that direct how the task generation machine learning model 222 determines a granularity, duration, or sequence of the tasks 230 based on the training procedure multi-modal data 206. For example, the task generation machine learning model 222 may generate a task to ablate a target anatomy to a highly skilled trainee. On the other hand, the task generation machine learning model 222 may generate multiple tasks to perform the same ablation activity to a less skilled trainee (e.g., advance an instrument to particular region, re-orient the instrument to a particular angle, activate a particular function associated with the instrument, etc.).

[0100] The constitution generator 216 may generate or instantiate the activity detection constitution 224 and the task generation constitution 226 by dynamically selecting the plurality of rules to include therein based on user input received by the user interface 219 (e.g., user input data 206J shown in FIG. 2C) and / or data present in the training procedure multi-modal data 206. In particular, the constitution generator 216 may be configured to identify and select the plurality of rules to include in the activity detection constitution 224 and the task generation constitution 226 from the data store 218 or other similar data storage devices or systems based on the training procedure multi-modal data 206 and / or the received user input on the user interface 219 or other input module of the system 200 in any manner as described herein.

[0101] Further to the constitution generator 216 of the system 200A, the constitution generator 216 of the system 200B may include additional rules in the trainee performance assessment constitution 208 and the trainee guidance constitution 210 that relate to analysis and utilization of the activity indicator 228 and the tasks 230. For example, the constitution generator 216 may include rules in the trainee performance assessment constitution 208 thatIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC direct the trainee performance assessment machine learning model 202 to generate the trainee performance assessment 212 with reference to the activity indicator 228. Similarly, the constitution generator 216 may include rules in the trainee guidance constitution 210 that direct the trainee guidance machine learning model 204 to generate the trainee guidance 214 with reference to the activity indicator 228 and the tasks 230.

[0102] With reference now to FIG. 2D, the different guidance methodologies and content of the trainee guidance 214 output from the trainee guidance machine learning model 204 will be described in more detail. As shown in FIG. 2D, the guidance methodology included in the trainee guidance 214 may include an audio methodology channel 234, a visual methodology channel 236, and / or a robotic methodology channel 237.

[0103] Though not shown in FIG. 2D, in some embodiments, the guidance methodology may also include changing the training exercise or otherwise modifying future aspects of the training procedure. For example, this type of guidance may change a simulated scenario in which the training procedure is performed in a simulation environment supported by the simulation module 232. In particular, the guidance may cause additional or fewer obstacles to be present within the simulation environment, introduce unexpected anatomical conditions, change a difficulty of navigating one or more lumens, and so on. In some embodiments, the changes to the training exercise may be based on other portions of the trainee guidance provide to the trainee (e.g., content provided via the audio methodology channel 234, a visual methodology channel 236, and / or a robotic methodology channel 237).

[0104] The audio methodology channel 234 may include outputting audio guidance to the trainee via the user interface 219 (see FIGS. 2A and 2B). The audio guidance may include the guidance content generated by the trainee guidance machine learning model 204. For example, the audio may direct the trainee to adjust a real or simulated endoscope camera, use a particular instrument or aspect of the follower device 104 on a simulated or real test subject, etc.

[0105] The visual methodology channel 236 may include generating an overlay on the real procedure image data 206D, simulated procedure image data 206H, or the external video data 2061. The overlay may include text providing specific instructions of the guidance content (e.g., the same or similar text as provided in the audio methodology channel 234) as well as indicators 240 of important areas of the real procedure image data 206D, simulated procedure image data 206H, or the external video data 2061, etc. for the trainee to pay attention to. In some embodiments, the visual methodology channel 236 may also include presenting aIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC training video that includes more detailed step by step instruction to the trainee on how to perform a task of the training procedure, such as a task that the trainee failed to accomplish satisfactorily (e.g., a task where the trainee performance assessment 212 indicates a poor performance by the trainee).

[0106] The robotic methodology channel 237 may include the processing unit 201 and / or the control system 140 operating one or more motor controllers of user input components (e.g., the leader input devices 106) in a manner that demonstrates how the workstation 102 is manipulated to perform a particular task or exercise of the training procedure. In some embodiments, the robotic methodology channel 237 may include the processing unit 201 and / or the control system 140 changing a level of autonomy associated with the autonomous aspects of the training procedure implemented via a Robotic Transformer model 242. For example, the robotic methodology channel may include causing the control system to autonomously advance an instrument 122 towards a target region of the subject.

[0107] As described above, in some embodiments, the training guidance 214 may implement a change to the training exercise (such as by changing a simulation supported by the simulation module 232). In these embodiments, the change in the training exercise may include one or more of increasing a difficulty of the training exercise, decreasing the difficulty of the training exercise, advancing the trainee to a next training procedure in a sequence of training procedures, modifying a simulation environment, changing a task granularity, etc.

[0108] It should be appreciated that the trainee guidance 214 may employ various combinations of the different guidance methodologies described in connection with FIG. 2D together. For example, the audio methodology channel 234 and visual methodology channel 236 may be employed together to present audio describing indications in the visual guidance. Furthermore, the audio methodology channel 234 and visual methodology channel 236 may be combined with the robotic methodology channel 237 to provide audio and visual explication for the controlled movements of the user input system 106.

[0109] FIG. 3 shows an example machine learning model architecture 300 that can be implemented for one or more of the trainee performance assessment machine learning model 202, the trainee guidance machine learning model 204, the activity detection machine learning model 220, and the task generation machine learning model 222. The example model architecture 300 is generally configured to accept one or more inputs (such as the training procedure multi-modal data 206) and output a model output 302.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC

[0110] In some embodiments, the example model architecture 300 utilizes the same underlying machine learning model (e.g., a LLM or LMM having a particular set of trained parameter weights) when utilized to implement different machine learning modules described herein. In these embodiments, the particular rules included in the constitution 304 causes the different machine learning modules to produce different outputs. That said, in some embodiments, different machine learning modules may implement different underlying machine learning models having different parameter weights.

[0111] As shown in FIG. 3, the machine learning model architecture 300 generally includes embedding or projection layers 306, additional layers 308, and an output layer 310.

[0112] The embedding or projection layers 306 may be configured to receive and process different portions of the training procedure multi-modal data 206 and / or the previous model outputs 302 for further processing and analysis by the additional layers 308 and the output layer 310. In some embodiments, the embedding or projection layers 306 may be configured to process different modalities of data (e.g., the real or simulated operational training procedure data) into a single input format (e.g., natural language text) for processing by the additional layers 308.

[0113] For example, image data (e.g., real procedure image data 206D, simulated procedure image data 206H, and external video data 2061 shown in FIG. 2C) in the training procedure multi-modal data 206 may be input into a visual language model (VLM) of the embedding or projection layers 306 to convert the image data to natural language descriptions of what is depicted in the image data. In this example, the VLM may be configured to identify objects, such as surgical instruments and devices (e.g., surgical beds or tables, medical devices, display devices, etc.), individuals and personnel (e.g., clinicians, doctors, medical technicians, etc.) depicted by the external video data 2061.

[0114] As another example, particular data streams of the training procedure multi-modal data 206 may be input into a transformer model to identify time-dependent trends for that data stream. As one example, a series of trainee performance assessments 212 of the trainee performing the same training exercise may be input into a transformer model to detect trends in skill development. In this example, the Transformer model may be able to detect patterns that indicate that a trainee is improving over time. Accordingly, the embedding layer 306 may generate a natural language description of this progress such that the trainee guidance machine learning model 204 may advance the trainee along a training plan (e.g., present a new training procedure or make the current training procedure more difficult).Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC

[0115] As another example, the Transformer model may be able to detect that a trainee responded positively to certain types of trainee guidance 214 and generate natural language embeddings indicative of the detected correlations. In this example, the embedding layer 306 may generate a natural language description of this correlation such that the trainee guidance machine learning model 204 generates future trainee guidance 214 in a similar manner.

[0116] It should be appreciated that for some data streams of the training procedure multimodal data 206, the data are already in a natural language text from that can be processed by the additional layers 308 and / or the output layer 310. For some of these data streams, the data may be input, without modification, to the additional layers 308 and the output layer 310. For others of these data streams, the data may still be processed by a respective embedding or projection layer 306. For example, the embedding layer 306 may be configured to summarize or extract key details of the natural language text before further processing by the additional layers 308 and the output layer 310. This may confine the set of data received by the additional layers 308 to a predefined context window for the additional layers 308 to improve processing efficiency.

[0117] The parameters of the embedding or projection layers 306 associated with each stream of the training procedure multi-modal data 206 may be independently trained or tuned from the other parameters of embedding or projection layers 306, the additional layers 308, and / or the output layers 310.

[0118] Additionally, it should be appreciated that the different machine learning modules described herein may include different embedding layers 306 to process the same stream of training procedure multi-modal data 206. For example, a scene recognition embedding model of the trainee performance assessment machine learning model 202 may generate embeddings of the real procedure image data 206D indicating a position of an instrument relative to an expected position for performing the activity indicated by the activity indicator 228 to assess instrument positioning. On the other hand, a scene recognition model of the task generation machine learning model 222 may identify one or more objects depicted in the real procedure image data 206D (e.g., a lesion) to identify tasks that should be performed to address the depicted object.

[0119] The additional layers 308 may comprise various Al or ML type layers known in the art. Such layers include multiheaded self-attention layers, cross-attention layers, multi-layer perceptrons, feed forward layers, softmax layers, etc. In some embodiments, the additional layers 308 may comprise layers of a pretrained Al model such as an LLM or similar. ThisIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC pretrained Al model can include third party provided models that are either fine-tuned to process the training procedure multi-modal data 206 or that are used as is without any additional tuning or training (e.g., via a call to a third party provided application programming interface or via a call to a private self-hosted instance of the model). In particular, the pretrained Al model may be configured to receive inputs in the natural language text format output from the embedding or projection layers 306. However, it should be appreciated that other variations of the additional layers 308 with different input formats are also possible. As described above, in some embodiments, the different machine learning modules described herein may implement different pretrained Al models as the additional layers 308.

[0120] As shown in FIG. 3, the constitution 304 may be input into the additional layers 308 to direct how the additional layers 308 process the natural language text inputs generated by the embedding or projection layers 306. However, it should be appreciated that in other cases, the constitution 304 may be input an associated one of the embedding or projection layers 306 before proceeding to the additional layers 308.

[0121] The output layer 310 aggregates outputs from each proceeding layer of the machine learning model architecture 300 (e.g., the embedding or projection layers 306 and the additional layers 308) to generate the model output 302 (such as the trainee performance assessment 212, the trainee guidance 214, the activity indicator 228, or the tasks 230).

[0122] With collective reference to FIGS. 1-2D, an example training procedure for a trainee using the computer-assisted system 100 using the system 200A or system 200B will be described. It should be appreciated that the following description represent one example training scenario, and the details may vary across training scenarios depending on trainee performance and / or the particular training procedure.

[0123] Initially, the processing unit 201 and / or the control system 140 determines instructions and details for the training procedure. As described herein, the training procedure may include a plurality of training exercises that are configured to instruct the trainee on how to perform various tasks using the computer-assisted system 100. The training procedure may be configured as a test where the trainee initially undergoes guided instruction for one or more iterations of each exercise in the training procedure and then is prompted to perform the training exercises without guidance until the trainee achieves a performance score or assessments sufficient to pass the test (e.g., trainee performance assessment 212 meets a preconfigured passing score).Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC

[0124] Once the training procedure details are configured, the processing unit 201 and / or the control system 140 may recall relevant background information from the data store 218 to be included in the training procedure multi-modal data 206. The relevant background information may include details specific to the trainee such as past assessment from prior training procedures, general details on the tasks or exercises in the training procedure details, a training plan for the trainee, and other materials as described herein.

[0125] Next, the processing unit 201 and / or the control system 140 directs the constitution generator 216 to generate the various constitutions described herein. For example, the constitution generator 216 may generate the trainee performance assessment constitution 208 for the particular training procedure. In particular, the constitution generator 216 may include rules for assessing the expected tasks to be performed as part of the training procedures. In some embodiments, the constitution generator 216 selects these rules based on skill levels indicated in a trainee profile, a trainee training plan, and / or one or more historical trainee assessments.

[0126] As another example, the constitution generator 216 may generate the trainee guidance constitution 210 to configure how the trainee guidance machine learning model 204 generates guidance for the trainee. The constitution generator 216 may generate rules based on trainee skill levels and / or indications in the trainee profile indicating correlations between guidance types (e.g., tonality, methodology, etc.) and the corresponding impact on trainee performance.

[0127] As another example, the constitution generator 216 may generate the task generation constitution 226 and / or the activity detection constitution 224 to include rules related to the specific tasks and / or activities associated with the training procedure.

[0128] After the processing unit 201 and / or the control system 140 finishes initializing the constitutions, the processing unit 201 and / or the control system 140 may indicate to the trainee that it is time to begin the training procedure. In response, the trainee begins to operate the workstation 102 to perform a first task or exercise of the training procedure. As the trainee performs the first task, the computer-assisted system 100, the simulation module 232, the image sensor 233, and / or the user interface 219 generate observational data portions of the training procedure multi-modal data 206.

[0129] The processing unit 201 and / or the control system 140 routes the training procedure multi-modal data 206, including both the observational data portions for the first task and the relevant background data, to the various machine learning modules described herein. ForIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC example, the processing unit 201 and / or the control system 140 may route the training procedure multi-modal data 206 to the activity detection machine learning model 220 to generate, based on the activity detection constitution 224, an activity indicator 228 associated with the first task and the task generation machine learning model 222 to generate, based on the task generation constitution 226, one or more tasks 230 that are to be implemented autonomously or semi-autonomously by the control system 140 and / or the simulation module 232.

[0130] The processing unit 201 and / or the control system 140 then routes the training procedure multi-modal data 206 for the first task, and the activity indicator 228 to the trainee performance assessment machine learning model 202 to generate, based on the trainee performance assessment constitution 208, a trainee performance assessment 212 for the first task.

[0131] The processing unit 201 and / or the control system 140 then routes the training procedure multi-modal data 206 for the first task and the trainee performance assessment 212 to the trainee guidance machine learning model 204 to generate, based on the trainee guidance constitution 210, the trainee guidance 214. As described herein, the trainee guidance 214 may indicate one or more methodologies for providing the guidance (e.g., audio, visual, robotic, simulation configuration, etc.). Accordingly, the processing unit 201 and / or the control system 140 may implement or provide the trainee guidance via the appropriate methodology (e.g., outputting the guidance via the user interface 219, changing the simulation scenario, etc.).

[0132] The processing unit 201 and / or the control system 140 may then continue to monitor trainee performance of the training procedure to generate subsequent activity indicators 228, tasks 230, trainee performance assessments 212, and / or trainee guidance 214. This iterative process may repeat until the trainee has passed all assessment portions of the training procedure or until the trainee otherwise stops the training procedure. In the event that the trainee stops the training procedure before completion, the processing unit 201 and / or the control system 140 may save details on the trainee’s progress (including the trainee’s responsiveness to different guidance content and methodologies) in the trainee profile maintained at the data store 218 for later usage.

[0133] Similarly, after the trainee has fully completed the training procedure, the processing unit 201 and / or the control system 140 may compile the training procedure multimodal data 206 to generate a set of post-training procedure data for storage in the data storeIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC 218 and / or another training database. The post-training procedure data may include all the trainee performance assessments 212 generated during the training procedure and indications of the trainee guidance 214 provided therewith. The post-training procedure data in the training database may be used to update the skill levels and / or guidance preferences of the trainee such that the constitution generator 216 generates constitutions with rule sets adapted to the trainee during future training procedures. In some embodiments, a machine learning model training system may also utilize the port-training procedure data as training data to tune one or more of the trainee performance assessment machine learning model 202 and the trainee guidance machine learning model 204.

[0134] FIG. 4 is a flow diagram of a computer-implemented method 400 for generating trainee guidance using adaptive Al. The method 400 may be performed by a processor system or a control system (such as the processor system 150 and control system 140 of FIG.1 and the processing unit 201 of FIGS. 2A and 2B). In some embodiments, the control system may implement an Al-assist module (such as the Al-assist module 180) to perform the functionality described with respect to the machine learning models.

[0135] At block 410, the method 400 includes obtaining multi-modal data (e.g., training procedure multi-modal data 206) relating to observed actions of a trainee during a training procedure. The multi-modal data may include one or more of image data, audio data, kinematics data, simulation output data, and trainee data.

[0136] At block 420, the method 400 includes generating a trainee performance assessment (e.g., trainee performance assessment 212) by analyzing, via a trainee performance assessment machine learning model (e.g., trainee performance assessment machine learning model 202), the multi-modal data. A trainee performance assessment constitution (e.g., trainee performance assessment constitution 208) including rules for assessing performance of the observed actions is input into the trainee performance assessment machine learning model to control how the trainee performance assessment machine learning model generates the trainee performance assessment. The trainee performance assessment machine learning model may include a large language model (LLM) or a large multi-modal model (LMM).

[0137] At block 430, the method 400 includes generating a trainee guidance (e.g., trainee guidance 214) by analyzing, via a trainee guidance machine learning model (e.g., trainee guidance machine learning model 204), the trainee performance assessment and the multimodal data. A trainee guidance constitution (e.g., trainee guidance constitution 210) is input into the trainee guidance machine learning model. The trainee guidance constitution includesIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC rules that control how the trainee guidance machine learning model determines guidance content and a guidance methodology of the trainee guidance. The guidance content includes a description of how to perform a specific action associated with the training procedure. The trainee guidance constitution may include tonal rules that direct a tone, emotion, sentiment, prosody, or diction of the guidance content output from the trainee guidance machine learning model. The guidance content may include a prompt to the trainee and the observed actions include a response to the prompt. The guidance methodology may include one or more of a visual guidance, audio guidance, robotic guidance, augmented reality guidance, virtual reality guidance, and a modification of a training exercise. The guidance content may include an indication of training materials for the trainee to review before performing a subsequent training exercise or training procedure.

[0138] At block 440, the method 400 includes providing, to the trainee, the guidance content via the guidance methodology (e.g., audio methodology channel 234, the visual methodology channel 236, the robotic methodology channel 237, etc.).

[0139] The method 400 may include receiving, via a user input system, one or more user inputs indicative of operation of simulated or physical robotic equipment. The user input system includes one or more motor controllers for controlling a pose of the user input system. The method 400 may also include controlling the one or more motor controllers to physically move the user input system in a manner that demonstrates how the user input system is manipulated during the training procedure to implement the robotic guidance. In some embodiments, the training procedure is a semiautonomous procedure. In these embodiments, the method 400 may include causing a controller operatively coupled to one or more repositionable structures to implement autonomous aspects of the training procedure in coordination with the observed actions of the trainee to implement the robotic guidance.

[0140] In some embodiments, the training procedure is a simulated training procedure. In these embodiments, the method 400 may include simulating, via a simulation module, responses to interactions of the trainee with a user input system to respond to simulated conditions presented within a simulation environment. A portion of the multi-modal data includes outputs of the simulation module. The outputs of the simulation module may include one or more of simulated event data, simulated image data, simulated kinematics data, or simulated force data, the guidance methodology is a visual methodology and providing the guidance content comprises overlaying the guidance content onto the simulated image data.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC The guidance methodology is a visual methodology and providing the guidance content comprises overlaying the guidance content onto the simulated image data.

[0141] In some embodiments, the observed actions are performed as part of a training exercise associated with a training skillset, wherein the training exercise is to be performed multiple times during the training procedure. In these embodiments, the method 400 may include inputting a prior trainee performance assessment for one or more prior performances of the training exercise into at least one of the trainee performance assessment machine learning model and the trainee guidance machine learning model. The prior trainee performance assessment may indicate that the trainee failed to successfully complete the training exercise. The trainee performance assessment constitution may include rules that direct the trainee performance assessment machine learning model to identify trends across the one or more prior performances of the training exercise and generate the training performance assessment based on the identified trends. In these embodiments, to identify the trends, the trainee performance assessment constitution may include rules that direct the trainee performance assessment machine learning model to determine a correlation between prior performance of the training exercise and the trainee guidance.

[0142] To determine the correlation, the trainee performance assessment constitution may include rules that direct the trainee performance assessment machine learning model to determine a correlation between guidance methodologies of the prior training guidance and performance of the training exercise subsequent to the respective prior training guidance. To determine the correlation, the trainee performance assessment constitution may include rules that direct the trainee performance assessment machine learning model to determine a correlation between guidance contents of the prior training guidance and performance of the training exercise subsequent to the respective prior training guidance. The trainee guidance constitution may include rules that control how the trainee guidance machine learning model generates the training guidance based on the determined correlation. The method 400 may include updating a trainee profile for the trainee to indicate the determined correlation. The method 400 may also include changing the training exercise based on the training guidance provided to the trainee. The change in the training exercise may include one or more of increasing a difficulty of the training exercise, decreasing the difficulty of the training exercise, advancing the trainee to a next training procedure in a sequence of training procedures, modifying a simulation environment, changing a task granularity, changing a level of autonomy in performing the training exercise.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC

[0143] In some embodiments, the method 400 includes obtaining a trainee profile from a trainee database. The trainee profile may indicate historical trainee assessments associated with historical training procedures. The trainee profile may also indicate one or more skill levels and the method 400 may include selecting the training procedure to perform based on the one or more skill levels. In these embodiments, the trainee performance assessment constitution may include a first set of trainee skill level rules that control how the trainee performance assessment machine learning model generates the trainee performance assessment in view of the one or more skill levels indicated by the trainee profile.Additionally, the trainee guidance constitution may include a second set of trainee skill level rules that control how the trainee guidance machine learning model generates the trainee guidance in view of the one or more skill levels indicated by the trainee profile. The second set of trainee skill level rules included in the trainee guidance constitution may control a frequency or a modality of intervention guidance provided by the trainee guidance machine learning model during execution of the training procedure. The intervention guidance may provide guidance for correcting one or more actions that will likely lead to incorrect performance of the training procedure. The method 400 may also include identifying one or more additional trainee profiles in the trainee database associated with similar skill levels as the one or more skill levels indicated by the trainee profile and including or modifying a set of trainee skill rules in one or more of the trainee performance assessment constitution or the trainee guidance constitution based on the trainee profile and the one or more additional trainee profiles.

[0144] The method 400 may include generating an activity indicator by analyzing, via an activity detection machine learning model, the multi-modal data. An activity detection constitution including rules for identifying the observed actions is input into the activity detection machine learning model to control how the activity detection machine learning model analyzes the multi-modal data to detect the observed actions. The method 400 may also include inputting the activity indicator into one or more of the trainee performance assessment machine learning model and the trainee guidance machine learning model.

[0145] The method 400 may include generating one or more tasks in furtherance of the training procedure by analyzing, via a task generation machine learning model, the multimodal data. A task generation constitution including rules for generating tasks is input into the task generation machine learning model to control how the task generation machine learning model analyzes the multi-modal data to generate the one or more tasks. The methodIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC 400 may also include inputting an indication of the one or more tasks into the trainee guidance machine learning model to generate portions of the guidance content relating to performing the one or more tasks.

[0146] The method 400 may include obtaining post-training procedure data on execution of the training procedure. The post-training procedure data may include (i) a trainee assessment for the training procedure for one or more training exercises associated therewith and (ii) an indication of the trainee guidance provided therewith. In these embodiments the method 400 may include storing the post-training procedure data in a training database. The post-training procedure data in the training database may be used to tune one or more of the trainee performance assessment machine learning model, the trainee guidance machine learning model, the trainee performance assessment constitution, or the trainee guidance constitution to optimize trainee performances during execution of subsequent training procedures.

[0147] One or more components of the examples discussed in this disclosure, may be implemented in software for execution on one or more processors of a computer system. The software may include code that when executed by the one or more processors, configures the one or more processors to perform various functionalities as discussed herein. The code may be stored in a non-transitory computer readable storage medium (e.g., a memory, magnetic storage, optical storage, solid-state storage, etc.). The computer readable storage medium may be part of a computer readable storage device, such as an electronic circuit, a semiconductor device, a semiconductor memory device, a read only memory (ROM), a flash memory, an erasable programmable read only memory (EPROM); a floppy diskette, a CD-ROM, an optical disk, a hard disk, or other storage device. The code may be downloaded via computer networks such as the Internet, Intranet, etc. for storage on the computer readable storage medium. The code may be executed by any of a wide variety of centralized or distributed data processing architectures. The programmed instructions of the code may be implemented as a number of separate programs or subroutines, or they may be integrated into a number of other aspects of the systems described herein. The components of the computing systems discussed herein may be connected using wired and / or wireless connections. In some examples, the wireless connections may use wireless communication protocols such as Bluetooth, near-field communication (NFC), Infrared Data Association (IrDA), home radio frequency (HomeRF), IEEE 502.11, Digital Enhanced Cordless Telecommunications (DECT), and wireless medical telemetry service (WMTS).Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC

[0148] Various general-purpose computer systems may be used to perform one or more processes, methods, or functionalities described herein. Additionally or alternatively, various specialized computer systems may be used to perform one or more processes, methods, or functionalities described herein. In addition, a variety of programming languages may be used to implement one or more of the processes, methods, or functionalities described herein.

[0149] While certain examples and examples have been described above and shown in the accompanying drawings, it is to be understood that such examples and examples are merely illustrative and are not limited to the specific constructions and arrangements shown and described, since various other alternatives, modifications, and equivalents will be appreciated by those with ordinary skill in the art.

Claims

Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC What is claimed is:

1. A computer system comprising:one or more processors; andone or more non-transitory, computer-readable media storing instructions that, when executed by the one or more processors, cause the computer system to:obtain multi-modal data relating to observed actions of a trainee during a training procedure;generate a trainee performance assessment by analyzing, via a trainee performance assessment machine learning model, the multi-modal data, wherein a trainee performance assessment constitution including rules for assessing performance of the observed actions is input into the trainee performance assessment machine learning model to control how the trainee performance assessment machine learning model generates the trainee performance assessment;generate a trainee guidance by analyzing, via a trainee guidance machine learning model, the trainee performance assessment and the multi-modal data, wherein:a trainee guidance constitution is input into the trainee guidance machine learning model, andthe trainee guidance constitution includes rules that control how the trainee guidance machine learning model determines guidance content and a guidance methodology for the trainee guidance; andprovide, to the trainee, the guidance content via the guidance methodology.

2. The computer system of claim 1, wherein the training procedure is a simulated training procedure.

3. The computer system of claim 2, further comprising:a simulation module configured to simulate responses to interactions of the trainee with a user input system to respond to simulated conditions presented within a simulation environment, wherein a portion of the multi-modal data comprises outputs of the simulation module.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC 4. The computer system of claim 3, wherein the outputs of the simulation module include one or more of simulated event data, simulated image data, simulated kinematics data, or simulated force data.

5. The computer system of claim 4, wherein the guidance methodology is a visual methodology and to provide the guidance content, the instructions, when executed by the one or more processors, cause the computer system to:overlay the guidance content onto the simulated image data.

6. The computer system of claim 1, wherein the guidance content includes an indication of training materials for the trainee to review before performing a subsequent training exercise or training procedure.

7. The computer system of claim 1, wherein the observed actions are performed as part of a training exercise associated with a training skillset, wherein the training exercise is to be performed multiple times during the training procedure.

8. The computer system of claim 7, wherein a prior trainee performance assessment for one or more prior performances of the training exercise are input into at least one of the trainee performance assessment machine learning model and the trainee guidance machine learning model.

9. The computer system of claim 8, wherein the prior trainee performance assessment indicates that the trainee failed to successfully complete the training exercise.

10. The computer system of claim 8, wherein the trainee performance assessment constitution includes rules that direct the trainee performance assessment machine learning model to identify trends across the one or more prior performances of the training exercise and generate the training performance assessment based on the identified trends.

11. The computer system of claim 10, wherein to identify the trends the trainee performance assessment constitution includes rules that direct the trainee performance assessment machine learning model to determine a correlation between prior performance ofIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC the training exercise and prior trainee guidance output by the trainee guidance machine learning model.

12. The computer system of claim 11, wherein the instructions, when executed by the one or more processors, cause the computer system to:update a trainee profile for the trainee to indicate the determined correlation.

13. The computer system of claim 11, wherein to determine the correlation, the trainee performance assessment constitution includes rules that direct the trainee performance assessment machine learning model to determine a correlation between guidance methodologies of the prior training guidance and performance of the training exercise subsequent to the respective prior training guidance.

14. The computer system of claim 11, wherein to determine the correlation, the trainee performance assessment constitution includes rules that direct the trainee performance assessment machine learning model to determine a correlation between guidance contents of the prior training guidance and performance of the training exercise subsequent to the respective prior training guidance.

15. The computer system of claim 11, wherein the trainee guidance constitution includes rules that control how the trainee guidance machine learning model generates the training guidance based on the determined correlation.

16. The computer system of claim 7, wherein, the instructions, when executed by the one or more processors, cause the computer system to:modify the training exercise based on the training guidance provided to the trainee.

17. The computer system of claim 16, wherein to modify the training exercise the one or more processors, cause the computer system to do one or more of increasing a difficulty of the training exercise, decreasing the difficulty of the training exercise, advancing the trainee to a next training procedure in a sequence of training procedures, modifying a simulation environment, changing a task granularity, changing a level of autonomy in performing the training exercise.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC18. The computer system of claim 1, wherein the instructions, when executed by the one or more processors, cause the one or more processors to:obtain a trainee profile from a trainee database, wherein the trainee profile indicates historical trainee assessments associated with historical training procedures.

19. The computer system of claim 18, wherein the trainee profile indicates one or more skill levels, and the computer system is configured to select the training procedure to perform based on the one or more skill levels.

20. The computer system of claim 19, wherein one or more of:(i) the trainee performance assessment constitution includes a first set of trainee skill level rules that control how the trainee performance assessment machine learning model generates the trainee performance assessment in view of the one or more skill levels indicated by the trainee profile, and(ii) the trainee guidance constitution includes a second set of trainee skill level rules that control how the trainee guidance machine learning model generates the trainee guidance in view of the one or more skill levels indicated by the trainee profile.

21. The computer system of claim 20, wherein the second set of trainee skill level rules included in the trainee guidance constitution control a frequency or a modality of intervention guidance provided by the trainee guidance machine learning model during execution of the training procedure.

22. The computer system of claim 21, wherein the intervention guidance provides guidance for correcting one or more actions that will likely lead to incorrect performance of the training procedure.

23. The computer system of claim 19, wherein the instructions, when executed by the one or more processors, cause the one or more processors to:identify one or more additional trainee profiles in the trainee database associated with similar skill levels as the one or more skill levels indicated by the trainee profile; andIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC include or modify a set of trainee skill rules in one or more of the trainee performance assessment constitution or the trainee guidance constitution based on the trainee profile and the one or more additional trainee profiles.

24. The computer system of claim 1, wherein the trainee performance assessment machine learning model comprises a large language model (LLM) or a large multi-modal model (LMM).

25. The computer system of claim 1, wherein the instructions, when executed by the one or more processors, cause the one or more processors to:generate an activity indicator by analyzing, via an activity detection machine learning model, the multi-modal data, wherein an activity detection constitution including rules for identifying the observed actions is input into the activity detection machine learning model to control how the activity detection machine learning model analyzes the multi-modal data to detect the observed actions; andinput the activity indicator into one or more of the trainee performance assessment machine learning model and the trainee guidance machine learning model.

26. The computer system of claim 1, wherein the instructions, when executed by the one or more processors, cause the one or more processors to:generate one or more tasks in furtherance of the training procedure by analyzing, via a task generation machine learning model, the multi-modal data, wherein a task generation constitution including rules for generating tasks is input into the task generation machine learning model to control how the task generation machine learning model analyzes the multi-modal data to generate the one or more tasks; andinput an indication of the one or more tasks into the trainee guidance machine learning model to generate portions of the guidance content relating to performing the one or more tasks.

27. The computer system of claim 1, wherein the guidance content includes a description of how to perform a specific action associated with the training procedure.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC 28. The computer system of claim 1, wherein the trainee guidance constitution includes tonal rules that direct a tone, emotion, sentiment, prosody, or diction of the guidance content output from the trainee guidance machine learning model.

29. The computer system of claim 1, wherein the guidance content includes a prompt to the trainee and the observed actions include a response to the prompt.

30. The computer system of claim 1, wherein the guidance methodology includes one or more of visual guidance, audio guidance, robotic guidance, augmented reality guidance, virtual reality guidance, and a modification of a training exercise.

31. The computer system of claim 30, further comprising:a user input system configured to receive one or more user inputs indicative of operation of simulated or physical robotic equipment, wherein the user input system includes one or more motor controllers for controlling a pose of the user input system;wherein to implement the robotic guidance, the computer system is configured to control the one or more motor controllers to physically move the user input system in a manner that demonstrates how the user input system is manipulated during the training procedure.

32. The computer system of claim 30, further comprising:one or more repositionable structures; anda controller operatively coupled to the one or more repositionable structures; wherein:the training procedure is a semiautonomous procedure; andto implement the robotic guidance, the computer system is configured to cause the controller to implement autonomous aspects of the training procedure in coordination with the observed actions of the trainee.

33. The computer system of claim 1, wherein the multi-modal data includes one or more of image data, audio data, kinematics data, simulation output data, and trainee data.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC 34. The computer system of any of claims 1-33, wherein the instructions, when executed by the one or more processors, cause the computer system to:obtain post-training procedure data on execution of the training procedure, the posttraining procedure data including (i) a trainee assessment for the training procedure for one or more training exercises associated therewith and (ii) an indication of the trainee guidance provided therewith; andstore the post-training procedure data in a training database.

35. The computer system of claim 34, wherein the post-training procedure data in the training database is used to tune one or more of the trainee performance assessment machine learning model, the trainee guidance machine learning model, the trainee performance assessment constitution, or the trainee guidance constitution to optimize trainee performances during execution of subsequent training procedures.

36. A computer-implemented method comprising:obtaining multi-modal data relating to observed actions of a trainee during a training procedure;generating a trainee performance assessment by analyzing, via a trainee performance assessment machine learning model, the multi-modal data, wherein a trainee performance assessment constitution including rules for assessing performance of the observed actions is input into the trainee performance assessment machine learning model to control how the trainee performance assessment machine learning model generates the trainee performance assessment;generating a trainee guidance by analyzing, via a trainee guidance machine learning model, the trainee performance assessment and the multi-modal data, wherein:a trainee guidance constitution is input into the trainee guidance machine learning model, andthe trainee guidance constitution includes rules that control how the trainee guidance machine learning model determines guidance content and a guidance methodology of the trainee guidance; andproviding, to the trainee, the guidance content via the guidance methodology.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC 37. The computer-implemented method of claim 36, wherein the training procedure is a simulated training procedure.

38. The computer-implemented method of claim 37, further comprising:simulating, via a simulation module, responses to interactions of the trainee with a user input system to respond to simulated conditions presented within a simulation environment, wherein a portion of the multi-modal data comprises outputs of the simulation module.

39. The computer-implemented method of claim 38, wherein the outputs of the simulation module include one or more of simulated event data, simulated image data, simulated kinematics data, or simulated force data.

40. The computer-implemented method of claim 39, wherein the guidance methodology is a visual methodology and providing the guidance content comprises overlaying the guidance content onto the simulated image data.

41. The computer-implemented method of claim 36, wherein the guidance content includes an indication of training materials for the trainee to review before performing a subsequent training exercise or training procedure.

42. The computer-implemented method of claim 36, wherein the observed actions are performed as part of a training exercise associated with a training skillset, wherein the training exercise is to be performed multiple times during the training procedure.

43. The computer-implemented method of claim 42, further comprising:inputting a prior trainee performance assessment for one or more prior performances of the training exercise into at least one of the trainee performance assessment machine learning model and the trainee guidance machine learning model.

44. The computer-implemented method of claim 43, wherein the prior trainee performance assessment indicates that the trainee failed to successfully complete the training exercise.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC45. The computer-implemented method of claim 43, wherein the trainee performance assessment constitution includes rules that direct the trainee performance assessment machine learning model to identify trends across the one or more prior performances of the training exercise and generate the training performance assessment based on the identified trends.

46. The computer- implemented method of claim 45, wherein to identify the trends, the trainee performance assessment constitution includes rules that direct the trainee performance assessment machine learning model to determine a correlation between prior performance of the training exercise and the trainee guidance.

47. The computer-implemented method of claim 46 further comprising:updating a trainee profile for the trainee to indicate the determined correlation.

48. The computer-implemented method of claim 46, wherein to determine the correlation, the trainee performance assessment constitution includes rules that direct the trainee performance assessment machine learning model to determine a correlation between guidance methodologies of the prior training guidance and performance of the training exercise subsequent to the respective prior training guidance.

49. The computer-implemented method of claim 46, wherein to determine the correlation, the trainee performance assessment constitution includes rules that direct the trainee performance assessment machine learning model to determine a correlation between guidance contents of the prior training guidance and performance of the training exercise subsequent to the respective prior training guidance.

50. The computer-implemented method of claim 46, wherein the trainee guidance constitution includes rules that control how the trainee guidance machine learning model generates the training guidance based on the determined correlation.

51. The computer-implemented method of claim 42, further comprising:modifying the training exercise based on the training guidance provided to the trainee.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC 52. The computer-implemented method of claim 51, wherein modifying the training exercise includes one or more of increasing a difficulty of the training exercise, decreasing the difficulty of the training exercise, advancing the trainee to a next training procedure in a sequence of training procedures, modifying a simulation environment, changing a task granularity, changing a level of autonomy in performing the training exercise.

53. The computer-implemented method of claim 36, further comprising:obtaining a trainee profile from a trainee database, wherein the trainee profile indicates historical trainee assessments associated with historical training procedures.

54. The computer-implemented method of claim 53, wherein the trainee profile indicates one or more skill levels and further comprising:selecting the training procedure to perform based on the one or more skill levels.

55. The computer-implemented method of claim 54, wherein one or more of:(i) the trainee performance assessment constitution includes a first set of trainee skill level rules that control how the trainee performance assessment machine learning model generates the trainee performance assessment in view of the one or more skill levels indicated by the trainee profile, and(ii) the trainee guidance constitution includes a second set of trainee skill level rules that control how the trainee guidance machine learning model generates the trainee guidance in view of the one or more skill levels indicated by the trainee profile.

56. The computer- implemented method of claim 55, wherein the second set of trainee skill level rules included in the trainee guidance constitution control a frequency or a modality of intervention guidance provided by the trainee guidance machine learning model during execution of the training procedure.

57. The computer-implemented method of claim 56, wherein the intervention guidance provides guidance for correcting one or more actions that will likely lead to incorrect performance of the training procedure.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC 58. The computer-implemented method of claim 54, further comprising:identifying one or more additional trainee profiles in the trainee database associated with similar skill levels as the one or more skill levels indicated by the trainee profile; and including or modifying a set of trainee skill rules in one or more of the trainee performance assessment constitution or the trainee guidance constitution based on the trainee profile and the one or more additional trainee profiles.

59. The computer-implemented method of claim 36, wherein the trainee performance assessment machine learning model comprises a large language model (LLM) or a large multi-modal model (LMM).

60. The computer-implemented method of claim 36, further comprising:generating an activity indicator by analyzing, via an activity detection machine learning model, the multi-modal data, wherein an activity detection constitution including rules for identifying the observed actions is input into the activity detection machine learning model to control how the activity detection machine learning model analyzes the multi-modal data to detect the observed actions; andinputting the activity indicator into one or more of the trainee performance assessment machine learning model and the trainee guidance machine learning model.

61. The computer- implemented method of claim 36, further comprising:generating one or more tasks in furtherance of the training procedure by analyzing, via a task generation machine learning model, the multi-modal data, wherein a task generation constitution including rules for generating tasks is input into the task generation machine learning model to control how the task generation machine learning model analyzes the multi-modal data to generate the one or more tasks; andinputting an indication of the one or more tasks into the trainee guidance machine learning model to generate portions of the guidance content relating to performing the one or more tasks.

62. The computer-implemented method of claim 36, wherein the guidance content includes a description of how to perform a specific action associated with the training procedure.Intuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC63. The computer-implemented method of claim 36, wherein the trainee guidance constitution includes tonal rules that direct a tone, emotion, sentiment, prosody, or diction of the guidance content output from the trainee guidance machine learning model.

64. The computer-implemented method of claim 36, wherein the guidance content includes a prompt to the trainee and the observed actions include a response to the prompt.

65. The computer- implemented method of claim 36, wherein the guidance methodology includes one or more of visual guidance, audio guidance, robotic guidance, augmented reality guidance, virtual reality guidance, and a modification of a training exercise.

66. The computer-implemented method of claim 65, further comprising:receiving, via a user input system, one or more user inputs indicative of operation of simulated or physical robotic equipment, wherein the user input system includes one or more motor controllers for controlling a pose of the user input system; andcontrolling the one or more motor controllers to physically move the user input system in a manner that demonstrates how the user input system is manipulated during the training procedure to implement the robotic guidance.

67. The computer- implemented method of claim 65, wherein the training procedure is a semiautonomous procedure and further comprising:causing a controller operatively coupled to one or more repositionable structures to implement autonomous aspects of the training procedure in coordination with the observed actions of the trainee to implement the robotic guidance.

68. The computer- implemented method of claim 36, wherein the multi-modal data includes one or more of image data, audio data, kinematics data, simulation output data, and trainee data.

69. The computer-implemented method of any of claims 36-68, further comprising: obtaining post-training procedure data on execution of the training procedure, the post-training procedure data including (i) a trainee assessment for the training procedure forIntuitive Docket No.: P06953-WO Attorney Docket No.: 33685 / 70520 / PC one or more training exercises associated therewith and (ii) an indication of the trainee guidance provided therewith; andstoring the post-training procedure data in a training database.

70. The computer-implemented method of claim 69, wherein the post-training procedure data in the training database is used to tune one or more of the trainee performance assessment machine learning model, the trainee guidance machine learning model, the trainee performance assessment constitution, or the trainee guidance constitution to optimize trainee performances during execution of subsequent training procedures.

71. A non-transitory machine-readable medium comprising a plurality of machine-readable instructions that when executed by one or more processors are adapted to cause the one or more processors to perform the computer- implemented method of any one of claims 36-70.