Intelligent and real-time task guidance system for surgical operating rooms

A computer-assisted system with machine learning models analyzes diverse data streams to generate and assign tasks, improving surgical efficiency and accessibility.

WO2026015586A1PCT designated stage Publication Date: 2026-01-15INTUITIVE SURGICAL OPERATIONS INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/US2025/036892
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2025-07-09
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Conventional medical systems fail to efficiently process large sets of different data modalities to determine task assignments for medical procedures, often relying on limited factors and reducing the effectiveness of personnel and equipment utilization.

Method used

A computer-assisted system utilizing machine learning models to analyze multiple data streams, including system and environmental data, to generate and assign tasks to personnel, enhancing task generation and assignment processes.

Benefits of technology

Improves surgical outcomes and task workflows by reducing reliance on skilled practitioners, enabling wider access to surgical treatment and diagnosis across various medical domains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025036892_15012026_PF_FP_ABST
    Figure US2025036892_15012026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods are described for determining and assigning tasks for performing medical procedures. The system may be configured to receive a plurality of data streams related to a medical procedure, wherein the plurality of data streams includes one or more of system data, medical environment data, and indications of personnel performing the medical procedure; analyze, using a task generation machine learning model, the plurality of data streams to generate natural language output relating to one or more tasks to be performed in furtherance of the medical procedure, wherein one or more inputs into the task generation machine learning model includes inputting embeddings of the plurality of data streams; analyze, via a task assignment machine learning model, the one or more tasks to assign the tasks to respective personnel; and provide indications to the respective personnel for performing the respective tasks assigned to the respective personnel.
Need to check novelty before this filing date? Find Prior Art

Description

INTELLIGENT AND REAL-TIME TASK GUIDANCE SYSTEM FOR SURGICAL OPERATING ROOMSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of the filing date of provisional U.S. Patent Application No. 63 / 669,553 entitled “INTELLIGENT AND REAL-TIME TASK GUIDANCE SYSTEM FOR SURGICAL OPERATING ROOMS,” filed on Inly 10, 2024. The entire contents of the provisional application are hereby expressly incorporated herein by reference.TECHNICAL FIELD

[0002] The present disclosure relates generally to computer-assisted systems and more particularly to training and utilizing artificial intelligence to determine and assign tasks associated with medical procedures to individuals for performing the procedures.BACKGROUND

[0003] Medical procedures often require various tasks to performed by individuals or devices. Conventional systems for monitoring medical environments do not provide any functionality relating to determining what tasks need to be performed for a medical procedure and generating outputs relating to specific tasks to be performed. For example, current medical environment task assistance methods do not take in, and cannot efficiently process, large sets of different data modalities, to determine task assignments for optimizing a procedure. Task assignment is typically performed by personnel or systems that consider very few factors which may lead to inefficient execution of various tasks.

[0004] Typically, an electronic controller for determining task assignment does not consider many relevant different types of data, and only considers available personnel and equipment. Even so, the systems may only consider a subset of available equipment as input by a user, which reduces the ability of systems and practitioners to effectively and efficiently perform procedures. Accordingly, it may be beneficial to implement machine learning models to assist in the analysis of the data streams to improve the task generation and assignment processes.

[0005] Accordingly, there is a need for improved techniques that enable automatic determination and assignment of tasks for performing medical procedures. Such techniques can allow improved surgical outcomes, improved task workflows, and reduced reliance on highly skilled practitioners with niche skillsets who may otherwise not be available for an operation or procedure. These improvements further allow for wider access to surgical treatment and diagnosis across a broad range of medical and clinical domains.SUMMARY

[0006] The following presents a simplified summary of various examples described herein and is not intended to identify key or critical elements or to delineate the scope of the claims.

[0007] In some aspects, the techniques described herein relate to a system to assist personnel in performing a medical procedure using a computer-assisted medical system, wherein the system is configured to receive a plurality of data streams related to a medical procedure, wherein the plurality of data streams includes one or more of system data, medical environment data, and indications of personnel performing the medical procedure; analyze, using a task generation machine learning model, the plurality of data streams to generate natural language output relating to one or more tasks to be performed in furtherance of the medical procedure, wherein one or more inputs into the task generation machine learning model includes inputting embeddings of the plurality of data streams; analyze, via a task assignment machine learning model, the one or more tasks to assign the tasks to respective personnel; and provide indications to the respective personnel for performing the respective tasks assigned to the respective personnel.

[0008] In some aspects, the techniques described herein relate to a method for assisting personnel in performing a medical procedure, the method comprising: receiving a plurality of data streams, wherein the plurality of data streams includes one or more of system data, medical environment data, and indications of personnel performing the medical procedure; analyzing, using a task generation machine learning model, the plurality of data streams to generate natural language output relating to one or more tasks to be performed in furtherance of the medical procedure, wherein one or more inputs into the task generation machine learning model includes inputting embeddings of the plurality of data streams; analyzing, via a task assignment machine learning model, the one or more tasks to assign the tasks to respective personnel; and providingindications to the respective personnel for performing the respective tasks assigned to the respective personnel.

[0009] In some aspects, the techniques described herein relate to a computer-readable media storing instructions that, when executed by one or more processors, causes the one or more processors to perform any of the methods described herein.

[0010] It is to be understood that both the foregoing general description and the following detailed description are illustrative and explanatory in nature and are intended to provide an understanding of the present disclosure without limiting the scope of the present disclosure. In that regard, additional aspects, features, and advantages of the present disclosure will be apparent to one skilled in the art from the following detailed description.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] FIG. 1 is a diagram of a computer-assisted system in accordance with one or more embodiments.

[0012] FIG. 2 is a schematic diagram of a system for determining and assigning tasks for performing a medical procedure.

[0013] FIG. 3 is a schematic diagram of a robotic action module for generating actions tokens to control a repositionable structure.

[0014] FIG. 4 is a schematic diagram of an example task generation module for determining task specific data streams for controlling a repositionable structure.

[0015] FIG. 5 is a schematic diagram of an example task selection module for determining task specific data streams for controlling a repositionable structure.

[0016] FIG. 6 is a schematic diagram of an example task assignment module for assigning tasks to perform a procedure.

[0017] FIG. 7A is a schematic diagram illustrating a process for detokenizing action tokens.

[0018] FIGs. 7B and 7C depict de-tokenized commands for controlling a repositionable structure.

[0019] FIG. 8 is a schematic diagram of an example training module for extending the techniques to a training scenario.

[0020] FIG. 9 is a flow diagram of a method for performing task assignment of tasks for a medical procedure.

[0021] Examples of the present disclosure and their advantages are best understood by referring to the detailed description that follows. It should be appreciated that like reference numerals are used to identify like elements illustrated in one or more of the figures, wherein showings therein are for purposes of illustrating examples of the present disclosure and not for purposes of limiting the same.DETAILED DESCRIPTION

[0022] In this description, specific details are set forth describing some embodiments consistent with the present disclosure. Numerous specific details are set forth in order to provide a thorough understanding of the embodiments. It will be apparent, however, to one skilled in the ail that some embodiments may be practiced without some or all of these specific details. The specific embodiments disclosed herein arc meant to be illustrative but not limiting. One skilled in the art may realize other elements that, although not specifically described here, are within the scope and the spirit of this disclosure. In addition, to avoid unnecessary repetition, one or more features shown and described in association with one embodiment may be incorporated into other embodiments unless specifically described otherwise or if the one or more features would make an embodiment non-functional.

[0023] Further, the terminology in this description is not intended to limit the invention. For example, spatially relative terms-such as “beneath”, “below”, “lower”, “above”, “upper”, “proximal”, “distal”, and the like-may be used to describe the relation of one element or feature to another element or feature as illustrated in the figures. These spatially relative terms are intended to encompass different positions (i.e., locations) and orientations ( / .<?.. rotational placements) of the elements or their operation in addition to the position and orientation shown in the figures. For example, if the content of one of the figures is turned over, elements described as “below” or “beneath” other elements or features would then be “above” or “over” the other elements or features. A device may be otherwise oriented and the spatially relative descriptors used herein interpreted accordingly. Likewise, descriptions of movement along and around various axes include various special element positions and orientations. In addition, the singular forms “a”, “an”, and “the” are intended to include the plural forms as well, unless the context indicatesotherwise. Additionally, the terms “comprises”, “comprising”, “includes”, and the like specify the presence of stated features, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups. Components described as coupled may be electrically or mechanically directly coupled, or they may be indirectly coupled via one or more intermediate components.

[0024] Elements described in detail with reference to one embodiment, implementation, system, or module may, whenever practical, be included in other embodiments, implementations, systems, or modules in which they are not specifically shown or described. For example, if an element is described in detail with reference to one embodiment and is not described with reference to a second embodiment, the element may nevertheless be claimed as included in the second embodiment. Thus, to avoid unnecessary repetition in the following description, one or more elements shown and described in association with one embodiment, implementation, or application may be incorporated into other embodiments, implementations, or aspects unless specifically described otherwise, unless the one or more elements would make an embodiment or implementation non-functional, or unless two or more of the elements provide conflicting functions.

[0025] In some instances, well known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0026] This disclosure describes various devices, elements, and portions of computer-assisted systems and elements in terms of their state in three-dimensional space. As used herein, the term “position” refers to the location of an element or a portion of an element (e.g., three degrees of translational freedom in a three-dimensional space, such as along Cartesian x-, y-, and z- coordinates). As used herein, the term “orientation” refers to the rotational placement of an element or a portion of an element (e.g., three degrees of rotational freedom in three-dimensional space, such as about roll, pitch, and yaw axes, represented in angle-axis, rotation matrix, quaternion representation, and / or the like). As used herein, and for a device with a kinematic series, such as with a repositionable structure with a plurality of links coupled by one or more joints, the term “proximal” refers to a direction toward a base of the kinematic series, and “distal” refers to a direction away from the base along the kinematic series.

[0027] As used herein, the term “pose” refers to the multi-degree of freedom (DOF) spatial position and orientation of a coordinate system of interest attached to a rigid body. In general, a pose includes a pose variable for each of the DOFs in the pose. For example, a full 6-DOF pose for a rigid body in three-dimensional space would include 6 pose variables corresponding to the 3 positional DOFs (e.g., x, y, and z) and the 3 orientational DOFs (e.g., roll, pitch, and yaw). A 3- DOF position only pose would include only pose variables for the 3 positional DOFs. Similarly, a 3-DOF orientation only pose would include only pose variables for the 3 rotational DOFs. Further, a velocity of the pose captures the change in pose over time (e.g., a first derivative of the pose). For a full 6-DOF pose of a rigid body in three-dimensional space, the velocity would include 3 translational velocities and 3 rotational velocities. Poses with other numbers of DOFs would have a corresponding number of velocities translational and / or rotational velocities.

[0028] This disclosure occasionally refers to the disclosed techniques being applied to “patients” undergoing a “medical procedure.” It should be appreciated that these references are not intended to limit the application of the disclosed techniques to applied medicine contexts. For example, the described techniques can be applied to facilitate physician training, equipment testing and / or calibration, and / or other contexts. Accordingly, any reference to the term “patient” is done for ease of explanation and also envisions the application of the described techniques to a generic “subject.”

[0029] The word “task” is used herein to refer to a discrete portion of procedure that may be autonomously, semi-autonomously, or manually implemented in furtherance of a procedure. For example, a task may be to move an endoscope to a particular portion, to advance an instrument to a particular depth, to replace an instrument coupled to a manipulator, and so on. In some embodiments, a task is associated with component tasks to accomplish an overall goal. For example, a task to analyze a worksite may include component tasks related to moving an endoscope to view the worksite, advancing an instrument to predetermined depth, and enabling a functionality supported by the instrument.

[0030] Aspects of this disclosure are described in reference to computer-assisted systems, which can include devices that are teleoperated, externally manipulated, autonomous, semiautonomous, and / or the like. Further, aspects of this disclosure are described in terms of an implementation using a teleoperated surgical system, such as the da Vinci® Surgical System commercialized byIntuitive Surgical, Inc. of Sunnyvale, California. Knowledgeable persons will understand, however, that inventive aspects disclosed herein may be embodied and implemented in various ways, including teleoperated and non-teleoperated, and medical and non-medical embodiments and implementations. Implementations on da Vinci® Surgical Systems are merely exemplary and are not to be considered as limiting the scope of the inventive aspects disclosed herein. For example, techniques described with reference to surgical instruments and surgical methods may be used in other contexts. Thus, the instruments, systems, and methods described herein may be used for humans, animals, portions of human or animal anatomy, industrial systems, general robotic, or teleoperated systems. As further examples, the instruments, systems, and methods described herein may be used for non-medical purposes including industrial uses, general robotic uses, sensing or manipulating non-tissue work pieces, cosmetic improvements, imaging of human or animal anatomy, gathering data from human or animal anatomy, setting up or taking down systems, training medical or non-medical personnel, and / or the like. Additional example applications include use for procedures on tissue removed from human or animal anatomies (with or without return to a human or animal anatomy) and for procedures on human or animal cadavers. Further, these techniques can also be used for medical treatment or diagnosis procedures that include, or do not include, surgical aspects.

[0031] FIG. 1 is a simplified diagram of an example computer-assisted system 100, according to various embodiments. The computer-assisted system 100 may be a computer-assisted medical system for assisting with performing tasks for medical procedures. Further, the computer-assisted system 100 may determine and assign tasks for performing medical tasks as described herein. In some examples, the computer-assisted system 100 is a teleoperated system. In medical examples, the computer-assisted system 100 can be a teleoperated medical system such as a surgical system. As shown, the computer-assisted system 100 includes a follower device 104 that can be teleoperated by being controlled by one or more leader devices (also called “leader input devices” when designed to accept external input), described in greater detail below. Systems that include a leader device and a follower device are referred to as leader-follower systems, and also sometimes referred to as master-slave systems. Also shown in FIG. 1 is an input system that includes a workstation 102 (e.g., a console), and in various embodiments the input system can be in any appropriate form and may or may not include the workstation 102.

[0032] In the example of FIG. 1 , the workstation 102 includes one or more leader input devices 106 that arc designed to be contacted and manipulated by an operator 108. For example, the workstation 102 may comprise one or more leader input devices 106 for use by the hands, the head, or some other body part(s) of operator 108. The leader input devices 106 in this example are supported by the workstation 102 and can be mechanically grounded. In some embodiments, an ergonomic support 110 (e.g., forearm rest) can be provided on which the operator 108 can rest his or her forearms. In some examples, the operator 108 can perform tasks at a worksite within a workspace near the follower device 104 during a procedure, by commanding the follower device 104 using the leader input devices 106. In a medical example, the worksite may be a surgical worksite associated with a patient.

[0033] A display device 112 is also included in the workstation 102. The display device 112 may be configured to display images for viewing by the operator 108. The display device 112 can be moved in various DOFs to accommodate the viewing position of the operator 108 and / or to provide control functions. In embodiments where the display device 112 provides control functions, the leader input devices 106 may include the display device 112. In the example of the computer-assisted system 100, displayed images may depict a worksite at which the operator 108 is performing various tasks by manipulating the leader input devices 106 and / or the display device 112. In some examples, images displayed by display device 112 may be received by the workstation 102 from one or more imaging devices arranged at a worksite. In other examples, the images displayed by the display device 112 may be generated by the display device 112 (or by a different connected device or system), such as for virtual representations of tools, the worksite, or for user interface components. As will be explained below, in some embodiments the display device 112 may display one or more tasks for the operator 108 to perform with respect to any component of the computer-assisted system 100.

[0034] As illustrated, the computer-assisted system 100 also includes a follower device 104 that can be commanded by the workstation 102. In a medical example, the follower device 104 can be located near an operating table (e.g., a table, bed, or other support) on which a patient can be positioned. In some medical examples, the workspace is provided on an operating table, e.g., on or in a patient, simulated patient, or model, training dummy, etc. (not shown). As illustrated, the follower device 104 may include a plurality of repositionable structures 120 (sometimes referred to as “manipulator arms” in robotic embodiments). In some embodiments, the repositionablestructures 120 may include a plurality of links that are rigid members and joints that can be individually actuated as part of a kinematic series. Additionally, each of the rcpositionablc structures 120 is configured to couple to an instrument 122. While FIG. 1 illustrates a follower device 104 that has four repositionable structures 120a- 120d, in other embodiments, the follower device 104 may include one, two, three, four, five, six, or additional or fewer repositionable structures 120a-120d.

[0035] The instrument 122 can include, for example, a working portion 126 and one or more structures for supporting and / or driving the working portion 126. Example working portions 126 include end effectors that physically contact or manipulate material, energy application elements that apply electrical, RF, ultrasonic, or other types of energy, sensors that detect characteristics of the workspace environment (such as temperature sensors, imaging devices, etc.), and the like. In various embodiments, examples of instruments 122 include, without limitation, a sealing instrument, a cutting instrument, a sealing-and-cutting instrument, an energy instrument for applying energy, a gripping instrument (e.g., clamps, jaws), a stapler, an imaging instrument such as one using optical, RF, or ultrasonic imaging modalities, a sensing instrument, an irrigation instrument, a suction instrument, and / or the like. In addition, the instrument 122 may include a transmission mechanism 128 that can be coupled to a drive assembly 130 of the respective repositionable structure 120a-120d. The drive assembly 130 may include a drive and / or other mechanisms controllable from workstation 102 that transmit forces to the transmission mechanism 128 to articular or otherwise actuate the instrument 122.

[0036] As illustrated, each instrument 122 may be mounted to a portion of a respective repositionable structure 120a- 120d. In FIG. 1, this is shown with the drive assembly 130 physically coupled to the transmission mechanism 128. The distal portion of each repositionable structure 120a- 120d further includes a cannula mount 124 to which a cannula (not shown) is mounted. When a cannula is mounted to the cannula mount 124, a shaft of the instrument 122 passes through the cannula and into a workspace.

[0037] In various embodiments, one or more of the working portions 126 of the instruments 122 may include an imaging device for capturing images. The imaging device may include any sensing technology capable of acquiring an image. Example imaging instruments include an optical endoscope, a hyperspectral camera, an ultrasonic sensor, etc. Imaging instruments may comprisemonoscopic imagers, stereoscopic imagers, and / or the like. Imaging devices based on radiofrequency domains may capture images in any frequency spectrum, including visible light, infrared light, ultraviolet light, and / or the like. The imaging device may include an illumination source to light the region being imaged. In embodiments where the working portions 126 of one or more of the instruments 122 include an imaging device, the instrument 122 may be configured to capture images of a portion of the workspace for display via the display device 112.

[0038] In some embodiments, the repositionable structures 120a- 120d and / or instruments 122 can be controlled to move the working portion 126 in response to manipulation of the leader input devices 106 by the operator 108. Accordingly, the repositionable structures 120a- 120d and / or instruments 122 may be said to “follow” the leader input devices 106 through teleoperation. This enables the operator 108 to perform tasks at the worksite using the repositionable structures 120a- 120d and / or instruments 122. For a surgical example, the operator 108 can direct the repositionable structures 120a-120d of the follower device 104 to move the working portions 126 as part of a surgical procedure performed at an internal surgical site that is entered via one or more minimally invasive apertures or natural orifices. It should be appreciated that, in some embodiments, the follower device 104 may include non-teleoperated components that the operator 108 or other medical professional must manually manipulate to a desired pose.

[0039] In some embodiments, a repositionable structure 120a of the computer-assisted system 100 may be configured to support a working portion 126a that includes an imaging device (also referred to herein as an “imaging device 126a”). For convenience, an instrument 122 that includes an imaging device is also referred to as an “imaging instrument” herein. The control system 140 may be configured to command the repositionable structure 120a and / or the imaging instrument 122 comprising the imaging device 126a to automatically position and / or orient (“pose”) the field of view (FOV) of the imaging device 126a to provide images of the workspace and / or other instruments 122.

[0040] In the illustrated embodiment, a control system 140 is communicatively coupled to the workstation 102. In other embodiments, the control system 140 may be provided as a component of the workstation 102 and / or the follower device 104. During teleoperation, as the operator 108 moves the leader input device(s) 106, one or more sensors configured to detect the leader input device(s) 106 generate spatial and / or orientation movement data that is provided to control system140. The control system 140 may interpret the spatial and / or orientation information to determine and / or provide control signals to the follower device 104 to control the movement of rcpositionablc structures 120a-120d, instruments 122, and / or working portions 126. In addition to the components of the follower device 104, in some embodiments, the control system 140 is configured to interpret inputs received from the workstation 102 to control operation of one or more auxiliary devices (not depicted) utilized in a procedure. For example, the workstation 102 may be used to control a pose of a surgical bed or operation of an insufflator.

[0041] In one embodiment, the control system 140 supports one or more wired communication protocols, (e.g., Ethernet, USB, and / or the like) and / or one or more wireless communication protocols (e.g., Bluetooth, IrDA, HomeRF, IEEE 1 102.11 , DECT, Wireless Telemetry, and / or the like) for communications between the control system 140 and the workstation 102 and / or the follower device 104.

[0042] In some embodiments, the control system 140 may be implemented at one or more computing systems. For example, one or more computing systems may be used to control the follower device 104. As another example, one or more computing systems may be used to control components of the workstation 102, such as movement of a display device 112.

[0043] As illustrated, the control system 140 includes a processor system 150, a memory 160, and an artificial intelligent (Al) assist module 180. The memory 160 may store a control module 170. The processor system 150 may include one or more processors having different processing architectures for processing instructions. For example, the one or more processors may be one or more cores or micro-cores of a multi-core processor, a central processing unit (CPU), a microprocessor, a field-programmable gate array (FPGA), an application- specific integrated circuit (ASIC), a digital signal processor (DSP), a graphics processing unit (GPU), a tensor processing unit (TPU), and / or the like.

[0044] In some embodiments, the processor system 150 includes circuity to support one or more communication interfaces (e.g., Bluetooth interface, infrared interface, network interface, optical interface, etc.). Additionally, a communication interface of control system 140 may include an integrated circuit for connecting the control system 140 to a network (not shown) (e.g., a local area network (LAN), a wide area network (WAN) such as the Internet, mobile network, or any othertype of network) and / or to another device, such as the workstation 102 and / or the follower device 104.

[0045] Additionally, the memory 160 may include non-persistent storage (e.g., volatile memory, such as random access memory (RAM), cache memory), persistent storage (e.g., a hard disk, an optical drive such as a compact disk (CD) drive or digital versatile disk (DVD) drive, a flash memory, a floppy disk, a flexible disk, a magnetic tape, any other magnetic medium, any other optical medium, programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a FLASH-EPROM, and / or any other memory chip or cartridge. The non-persistent storage and persistent storage are examples of non-transitory, tangible machine- readable media that can store executable code that, when ran by one or more processors (e.g., processor system 150), can cause the one or more processors to perform one or more of the techniques and / or methods disclosed herein.

[0046] The Al assist module 180 may implement one or more machine learning models and / or training protocols therefor. For example, the Al assist module 180 may implement one or more neural networks, deep learning models, decision trees, support vector machines, linear’ regression, generative Al models, reinforced learning models, random forests, Naive Bayes models, large language models (LLMs), generative adversarial networks, foundation models, image recognition models, linear discriminant analysis models, creative applications, autoregressive models, supervised or unsupervised learning models, multimodal models, vision language models (VLMs), vision foundation models (VFMs), large multi-modal models (LMMs), Transformer models (including Robotic Transformer models), or another machine learning or Al model for performing the methods described herein. The structure of the one or more machine learning is described in more detail with respect to FIGs. 2-7C. The Al assist module 180 may include dedicated processors and memory for storing and performing Al processes, or the Al assist module 180 may utilize resources of the processor system 150 and the memory 160 to store and / or perform any processing or tasks required to perform the methods described herein.

[0047] It should be appreciated that while FIG. 1 depicts the Al assist module 180 as being a component of the control system 140 of the computer-assisted system 100, in other embodiments, some or all of the structures and / or functionality described with respect to the Al assist module 180 may be implemented at another computing system (not depicted) associated with operatingroom. This computing system may include a processing unit and a memory unit similar for implementing the various Al modules disclosed herein in a similar’ manner as described with respect to the processing system 150 and the memory 160. In some embodiments, one or more of the modules described herein related to assigning tasks to OR personnel may be implemented a computing system external to and communicatively coupled with the control system 140 may be configured to provide guidance to OR personnel. In these embodiments, the multimodal data obtained by the control system 140 and / or the outputs of any module implemented at the control system 140 may be communicated to the external Al assist module 180 for processing thereat.

[0048] Further, in some embodiments, the Al assist module 180 may be implemented completely independent of the computer-assisted system 100. For example, some procedures may not be performed using a computer-assisted system (e.g., via manually-operated equipment or tools). In these embodiments, it should be appreciated that while there may be fewer data modalities available as inputs to the various Al modules disclosed herein, the general functionality of these modules may be unchanged. For example, the Al modules described herein may be able to perform the described functionality based on operating room video data, laparoscopic imaging data, and / or other sensor data associated with the procedure and / or operating room otherwise available for processing by the Al assist module 180. Accordingly, while the instant disclosure includes references to the control system 140 performing functionality related to the Al assist module 180, it should be appreciated that these references envision a similar implementation at an alternative computing system associated with the operating room.

[0049] Additionally, the control system 140 and / or the external computing system may also include one or more input devices (such as a touchscreen, keyboard, mouse, microphone, touchpad, electronic pen, or any other type of input device) and / or output devices (such as a display device, a speaker, external storage, a printer, or any other output device). In some embodiments, the Al assist module 180 may be implemented on a particular node of a distributed computing system (e.g., a cloud computing system). As another example, different modules associated with the Al assist module 180 may be implemented on different nodes of the distributed computing system. Further, one or more elements of the aforementioned the Al assist module 180 may be located at a remote location and connected to the other elements over a network.

[0050] In an endoscopic surgery example, the imaging instrument comprising the imaging device 126a may be inserted into the patient prior to the other instruments 122, including a second instrument 122b comprising a second working portion 126b. The second instrument 122b can include any appropriate working portion 126b, and can even include a second imaging device. Accordingly, the imaging device 126a may be maneuvered to positioned to identify a target to which other instruments may interact with as part of another task. The control system 140 may, for example, automatically command the corresponding repositionable structures 120a and 120b to position respective instruments 122a and 122b to perform one or more tasks in tandem, or sequentially based on the specific task, instruments, and positions of the repositionable structures 120a and 120b. In examples, the Al assist module 180 may perform Al processes and algorithms to identify tasks for performing a medical procedure, and to further assign the tasks to respective personnel for performing the tasks. Additionally, the Al assist module 180 may identify certain tasks to be performed automatically by one or more robotic systems and / or devices, semi- automatically with user assistance or input to the system 100, and tasks that are to be performed entirely manually by personnel. It should be appreciated that some manual tasks may be performed independent of the computer-assisted system 100. For example, some tasks generated by the Al assist module 180 may include sterilizing instruments or a worksite, obtaining an instrument for use later in a procedure, adjusting a pose of the subject, and so on.

[0051] For tasks that are to be performed semi-automatically or manually, the Al assist module 180 may further analyze the available personnel to assign the task to the most appropriate individual for performing the task. For example, the Al assist module 180 may assign tasks to personnel that have suitable skills and / or training to perform the task. As another example, the Al assist module 180 may assign tasks from among suitable personnel based on their availability (e.g., whether or not the person is performing another task).

[0052] The disclosed techniques enable the Al assist module 180 to perform real-time operations (e.g., within 5ms, within 10ms, within 20ms, etc.). Accordingly, the Al assist module 180 is able to identify and assign tasks to appropriate personnel in response to dynamic conditions during the course of a medical procedure. Such techniques improve the efficiency of operating the computer-assisted system or instrument, simplify user control of the computer-assisted system 100 (or other surgical equipment), improve the efficiency of the medical procedure by streamlining task workflow, and may reduce the required personnel in a medical environment for performing aprocedure. Further, although a surgical example is shown, the disclosed techniques provide an improvement to the computer-assisted system 100 or alternative computing systems in the non- surgical aspects of the procedure, and can be used to improve computer-assisted systems applied in non-medical contexts.

[0053] While embodiments described herein relate to implementing the disclosed techniques intraoperatively, the disclosed techniques may be extended to other contexts. As one example, the techniques disclosed herein may be extended to schedule personnel for the various procedures that are to be performed by a facility. In this example, the techniques for assigning personnel to specific tasks are adapted to assign personnel to specific procedures based on similar considerations. As another example, the techniques disclosed herein may be applied in a training scenario where the medical procedure is performed using a training object instead of a human subject. In these examples, the data stream input into the disclosed models may be based, in part, upon historical or simulated data. The Al assist module 180 may then generate and assign tasks to personnel to develop the skills of individual persons and / or the coordination between teams of personnel at performing a procedure. As a result, the techniques can generate training scenarios and assign tasks in a matter to optimize personnel training.

[0054] FIG. 2 is a schematic diagram of a system 200 for determining and assigning tasks for performing a medical procedure. The tasks may be assigned to be entirely or partially performed manually or by a robotic structure or system (such as the repositionable structure 120). The system 20 includes a series of modules that may be executed by the Al assist module 180 of FIG. 1 to determine and assign one or more tasks associated with a medical procedure. The system 200 includes sources of multi-modal data including one or more multi-modal data streams 202, a task generation module 210, a task selection module 220, a task assignment module 230, and a robotic action module 250. It should be appreciated that in other embodiments, additional or fewer modules may be implemented. Further, in other embodiments, one or more of the described modules may be combined into a single module.

[0055] The system 200 includes one or more sources of multi-modal data 202. For example, the multimodal data 202 may include a data stream 202a indicative of force exerted upon an instrument, a data stream 202b indicative of system events generated by the repositionable structure and / or a control system thereof, a data stream 202c indicative of kinematic dataassociated with the repositionable structure and / or instruments or auxiliary devices associated therewith, an external video data stream 202d indicative of one or more objects and / or personnel within a medical environment, a procedure video data stream 202e (such as image data generated by an endoscope or a laparoscopic image sensor), a data stream 202f indicative of personnel within the medical environment and data associated therewith (e.g., a name, role, qualifications, procedure history, a network address of an associated device, etc.), and / or other sources of data that indicate a state of a procedure facilitated by the repositionable structures (such as 3D point cloud or depth data generated by a time-of-flight sensor, stereoscopic image sensors, and / or structured light sensors). In examples, the data streams 202a- 202f may include system data (e.g., data associated with system events, system capabilities, etc.), and data associated with a medical environment (e.g., via video data, image data, available instruments (either within the medical environment or in a storeroom of the medical facility, etc.). In examples, the multimodal data 202 may include one or more data streams that provide a 3D point cloud of medical environments or rooms at which the procedure is performed.

[0056] The system 200 may be configured to route the multimodal data 202 to a task generation module 210 configured to output one or more tasks that may be performed in furtherance of the procedure based on a procedure state represented by the input multimodal data 202. The task generation module 210 may identify tasks that are to be performed in the near term (e.g., tasks that respond to conditions detected in an input set of multi-modal data) and / or tasks that are to be performed in the long term (e.g., tasks that will need to be performed later in the procedure after one or more near-term tasks are performed). The task generation module 210 may further generate tasks to be performed in the future for a medical procedure scheduled to be performed at a later time or date. By generating long-term tasks, the system 200 is able to track preparedness for performing the long-term task and generate additional tasks to prepare the control system for performing the long-term task as the procedure advances closer to the appropriate time for performing the long-term tasks. In some embodiments, the task generation module 210 may include component models to facilitate the analysis. For example, the task generation module 210 may include a scene recognition model 212 to parse image data to generate natural language descriptions of the scene represented by the image data (and the corresponding locations in the image data) and a task generation model 215 configured to actually generate the one or more tasks.As one example, the scene recognition model 212 may detect and identify objects in a medical environment such as an operating room (OR) to generate an inventory of objects in the OR.

[0057] In some embodiments, a data stream input into the task generation module 210 may include an embedding indicative of the medical procedure, or of one or more phases of the medical procedure. Technique for detecting the type and / or phase of a medical procedure are described in U.S. Provisional Application Nos. 63 / 663,539 filed June 24, 2024, and 63 / 663,996 filed June 25, 2024, the entire disclosures of which are hereby incorporated by reference. Accordingly, in these embodiments, the phase information may be input into the task generation model 215 to generate tasks based on the particular phase for the medical procedure. Additionally, as described elsewhere herein, the phase of the medical procedure may be used when updating a procedure schedule for a facility in view of a high-priority procedure. Accordingly, in these embodiments, the phase information may be output to a scheduling system for the facility.

[0058] The system 200 may then provide the output tasks (e.g., natural language descriptions of an action to be performed) to a task selection module 220 to select which tasks should actually be implemented by the control system and / or an operator thereof. As illustrated, the task selection module may include a task filtering model 222 configured to filter the tasks generated by the task generation model 215 and a task selection model 225 configured to select a one or more valid tasks to be performed by the control system and / or an operator thereof. The system 200 may then provide the selected tasks to the task assignment module 230, and / or the robotic action module 250.

[0059] The task assignment module 230 may be configured to analyze the one or more selected tasks to assign the tasks to respective personnel. Additionally, one or more data streams of the plurality of data streams 202 may be input into the task assignment module 230 to assist in the task assignment analysis. The task assignment module 230 includes a task analysis model 232 configured to analyze to identify the which personnel (including combinations of various personnel) are capable of implementing the selected tasks and a task assignment model 245 configured to assign the task to respective personnel. The task analysis model 232 may identify different requirements for each task (e.g., personnel skillsets and qualifications, instruments, instrument and system capabilities, etc.). In some embodiments, a particular task may require multiple personnel to perform different aspects of a given task. Accordingly, the task analysismodel 232 may be configured to divide the task into subtasks that are assigned to different personnel.

[0060] The task assignment module 230 further includes a task assignment model 235 that analyzes the personnel requirements for the input tasks (and / or subtasks thereof) and assigns the various tasks to appropriate personnel. It should be appreciated that if the appropriate personnel are not present within the medical environment, or not available to perform a task (e.g., the personnel are assigned to perform another task when the unassigned task is to be initiated and / or completed), the task assignment model 235 may assign the task to a robotic structure or system (if possible), or may attempt to assign the task to other personnel. Additionally, if required the personnel is not present or available, the task assignment module 230 may send an indication to the task generation module 210 and / or the task selection module 220 to generate and / or select alternative tasks that can be performed based on the current task assignments.

[0061] In examples, the outputs of task selection module 220 and / or the task assignment module 230 may be provided to the operator, such as via the workstation 102 of FIG. 1, an operating room display (such as a display coupled to an external computing system that implements at least a portion of the Al assist module 180), or other display device associated with the procedure. Accordingly, the operator of the system 200 may be different persons than the operator 108 of the computer-assisted system 100 (e.g., a head nurse, a robotic coordinator, etc.). The operator may then approve of the tasks to be performed and / or the task assignments prior to implementing the selected tasks. If the operator disapproves the selected tasks, or specific task assignments, the task selection module 220 may provide alternative tasks that can be performed, and / or the task assignment module 230 may provide different assignments of the tasks. In some embodiments, if the operator disapproves the assignment of the tasks, the operator may be able to manually assign the task to personnel or a robotic system via the display device 112.

[0062] After the tasks assignment is generated and / or confirmed, the system 200 may be configured to provide indications of the assigned tasks to the respective personnel implementing the task. In some embodiments, the system presents the task assignment on a display screen present in the OR. Additionally or alternatively, the system 200 may transmit the task assignment to a wearable or other personal electronic device associated with the assigned personnel. In these embodiments, the system 200 may include a personnel record that indicates a network address fora wearable or other personal electronic device carried by the various personnel at the facility. Accordingly, the system may query the corresponding records to route the task assignments to the appropriate devices.

[0063] It should be appreciated that some tasks may be implemented semi-autonomously. Accordingly, to perform these semi- autonomous tasks, the system 200 includes the robotic action module 250. As illustrated, the robotic action module 250 may be configured to receive the selected semi- autonomous tasks from the task selection module 220, task assignments from the task assignment module 230, and / or the multimodal data 202. The robotic action module 250 may then process the multimodal data 202 and the selected tasks to generate action tokens to control the manipulator system to perform the selected tasks. For example, the action tokens may be natural language descriptions of a robotic action to be performed by the manipulator system in furtherance of a selected task. The robotic action module 250 may embed the input data streams 202 via an embedding stage 310 for input into an RT model 315. The robotic action module 250 may then convert the action tokens to command one or more components of a computer-assisted robotic device (such as the follower device 104 of FIG. 1) to perform one or more of the selected tasks.

[0064] It should be appreciated that in some embodiments, one or more of the modules 210, 220, 230, may be combined into a single module, for example, by implementing chain of thought (CoT) or other recurrent prompting techniques such that a single prompt performs the described functionality corresponding to the individual modules. Additionally, in other embodiments, the system 200 may include additional modules that generate inputs to the robotic action module 250 to implement autonomous tasks. For example, the system 200 may include an instrument or arm selection module that assigns tasks to a particular’ instrument or arm, or a data modality selection module that selects which data streams are embedded and input into the RT model 315. One such embodiment of a system that includes these additional modules is described in U.S. Provisional Application No. 63 / 663,539 filed June 24, 2024.

[0065] FIG. 3 is a schematic diagram of a robotic action module 250 (such as the robotic action module 255 of FIG. 2), for generating actions tokens to control a repositionable structure (such as the repositionable structures 120 of FIG. 1). The robotic action module 250 may be implemented as part of the Al assist module 180 of FIG. 1. As described with respect to FIG. 2, the roboticaction module 250 may be configured to receive as inputs one or more data streams forming the multi-modal data 202, an indication of a selected task 254, and an indication of a one or more specific assigned tasks 256.

[0066] As illustrated, the robotic action module 250 includes an embedding stage 310 that embeds the various data modalities (e.g., text, image, audio, video, sensor data, etc.) into a common data space for input into an RT model 315. The embedding stage 310 may be configured to project the input data streams into a common vector space. In examples, the embedding stage 310 may implement one or more embedding methods including, without limitation, a word2vec, GloVe, ELMo, BERT, principal component analysis, singular value decomposition, a transformer model, Doc2Vec, paragraph vectors, convolutional neural networks, pre-trained models for image embedding, node embeddings, iknowledge graph embeddings (e.g., TransE, TrensR, DistMult, etc.), word embeddings, graph embeddings, entity embeddings, or another type of embedding supported by the RT model 315. It should be appreciated that while the embedding stage 310 is depicted as a single block, each data stream 202 may be associated with a respective embedding model for generating embeddings of the respective data type included within the embedding stage 310. Further, while embedding stage 310 is depicted as a component of the robotic action module 250, in some embodiments, the embeddings generated by the embedding stage 310 are input into the various models of the task generation module 210, the task selection module 220, and the task assignment module 230.

[0067] The various embedding models implemented by the embedding stage 310 may be trained and / or fine-tuned using historical data of prior procedures. As one example, an embedding model to analyze the procedure video data stream 202e may be trained to identify objects (e.g., instruments, ports, anatomical features, etc.) that are expected to be seen during the procedure. In this example, the embedding model for the procedure video data stream 202e may be trained and / or fine-tuned using historical image data in which the pixels representative of the corresponding objects are labels. The embedding model may then be trained or tuned in any suitable manner that minimizes loss with respect to the set of ground truth labels.

[0068] As another example, an embedding model for the kinematic data stream 202c, the force data stream 202a, and / or the events data stream 202b may be trained using labeled segments from historical procedures performed using the same type of manipulator system associated with thesystem 200. For example, after a historical procedure has been performed, the set of kinematic, force, and / or event data captured during the historical procedure may be correlated, compiled, and aligned in time. For some types of conditions, the set of data may then be labeled by identifying time segments at which a particular condition occurred (e.g., an instrument or end effector being operated in a particular manner, performance of a particular’ phase of a procedure, transition to a new phase of a procedure, etc.). In some embodiments, the condition may relate to identifying a particular phase of a medical procedure. By building a library of labeled time segments, an embedding model can be trained to identify the conditions in the corresponding data streams when implemented in the robotic action module 250. In these embodiments, the embedding model may include a time-series analysis model (such as a Transformer model or a recurrent neural network (RNN)) configured to capture the temporal relationship between data within the labeled time segments.

[0069] For other conditions (such as alerts, errors, anomalies, abrupt and / or significant changes in operation, or other instantaneous conditions), the representation of that condition in the historical data may be labeled such that the embedding model can be trained to detect the labeled conditions. In either case, the training system may implement self- supervised learning to identifying the conditions and / or fine-tuning based on the procedure type and / or manipulator system such that the identification is tailored to the procedure / manipulator system associated with the system 200.

[0070] The robotic action module 250 then provides the embedded data output from the embedding stage 310 as an input to the RT model 315. The RT model 315 may implement any type of robotic transformer architecture configured to convert embeddings into action tokens for control of a robotic system. One such suitable architecture is the RT-2 architecture by DeepMind, but other RT model architectures are envisioned. The RT model 315 is configured to process the embedded data from the embedding stage 310 to generate action tokens 320 for controlling the instruments, auxiliary devices, and / or repositionable structures to perform the selected tasks. The action tokens 320 may be textual descriptions of robotic actions that the selected arm, instrument, and / or auxiliary device are to perform to implement the selected task. It should be appreciated that the particular tokens supported by the manipulator system may vary between models. Accordingly, the system 200 may be coupled to a library of RT models 315 each fine-tuned to generate action tokens supported by a different manipulator system type. As a result, the particularoutputs of the tuned RT model 315 are adapted to the actual manipulator system associated with the system 200.

[0071] The robotic action module 250 then inputs the action tokens 320 to a robotic controller 325 that de-tokenizes the action tokens 320 into actual control commands the control operation of instruments, arms, and / or auxiliary devices according to the assigned tasks 256. In examples, to perform an assigned task 256, a de-tokenized action token may include a plurality of commands including movements of one or more repositionable structures, control of an action of an instrument (e.g., power up, power down, inject, extract, incise, etc.), control of an auxiliary structure (e.g., surgical bed, display device, etc.), or another command for performing a given task. In some embodiments, the de-tokenized control commands are signals suitable for various types of robotic control architectures that may be implemented at the manipulator system. For example, the de-tokenized control commands may be a signal input into a proportional-integral-derivative (PID) controller that controls a particular joint, or a model predictive control (MPC) signal, and / or other types of control signals. As described further herein, the de-tokenized commands may be specific to a given robotic system or machine. For example, a given system may include manipulators with more degrees of freedom of motion than another system, or an instrument, such as a camera, may have a wider field of view than another camera, and the de-tokenized commands may differ depending on the various capabilities and parameters for a given robotic system.

[0072] FIG. 4 is a schematic diagram of the task generation module 210 for determining tasks that can be performed based on a current state of the manipulator system and / or procedure as reflected by the multimodal data streams 202. As shown in FIG. 2, the task generation module 210 is configured to receive the multi-modal data streams 202 as an input. Similar to the embedding stage 310 of the robotic action module 250, the task generation module 210 may embed the data streams 202 for input into a task generation model 215. In some embodiments, the embedding models used to embed the data streams 202 are the same as the embedding models included in the embedding stage 310 of the robotic action module 250.

[0073] For example, the embedding model for an endoscopic data stream may be routed to a vision-language model (VLM) such as scene recognition model 212 configured to output natural language descriptions of a scene depicted by the image data output by an endoscopic instrument. In surgical applications, for example, the VLM may be configured to identify objects, such assurgical instruments and devices (e.g., surgical beds or tables, medical devices, display devices, etc.), individuals and personnel (e.g., clinicians, doctors, medical technicians, etc.) depicted by an image data stream.

[0074] As another example, the embedding model of the event data stream may be configured to analyze the time series event data using a transformer model 214 to generate an output token identifying a particular phase of the procedure. It should be appreciated that while the term “transformer model” is used herein, in other embodiments, other model architectures suitable for analyzing temporal dependencies may be implemented. These output tokens may be input into the task generation model 215 to generate tasks that are in furtherance of operation typically associated with the identified phase. Similarly, the transformer model 214 may be configured to detect the completion of a phase or sub-phase of the procedure and generate an output token identifying the transition. These output tokens may be input into the task generation model 215 to begin generating tasks to implement actions associated with a subsequent phase or subphase. Additionally, the transformer model 214 may be configured to detect anomalous operation associated with phase of the procedure and output token indicative of the anomaly (e.g., a token indicative that anomalous operation is occurring, a token indicating a particular type of anomaly that is occurring, and so on). These output tokens may be input into the task generation model 215 to generate tasks to correct the anomaly.

[0075] In addition to generating tokens based on time-dependencies in the event data stream, the transformer model 214 may also be configured to output tokens based on an instantaneous representation in the event data stream. To this end, the event data in the event data stream may indicate the complete state of the robotic components of the manipulator system at any given time. Accordingly, the transformer model 214 may be configured to identify equipment operatively coupled to the control system (e.g., manipulator arms, instruments, auxiliary devices) to generate an inventory of devices. In some embodiments, the inventory may also be input to the task generation model 215 such that the task generation model 215 understands that capabilities of the control system and generates tasks that can be implemented thereby.

[0076] In addition to identifying equipment currently coupled to the control system, an embedding model for another data stream may be configured to detect equipment that is otherwise available for use in the procedure (e.g., sterilized equipment on-hand in an operating room,equipment available elsewhere at the site). For example, the on-hand equipment may be detected via an operating room video data stream or via an equipment scheduling data stream. Additionally, the task generation model 215 may receive data indicative of personnel that are currently available, will be available for a scheduled procedure, or are required to perform one or more tasks via the personnel data stream 202f.

[0077] The tokens generated by the embedding models may then be input into a task selection model 215. The task selection model 215 may be a LMM configured to directly accept image data streams or an LLM that accepts the natural language description(s) of the image data stream(s) generated by the VLM 212. The task selection model may then generate tasks in furtherance of the detected (or subsequent) phase of the procedure based on the available equipment, personnel, and any other conditions indicated by the input data streams 202.

[0078] It should be appreciated that the task generation model 215 may be configured to output multiple tasks to implement the actions associated with a particular phase of the procedure. For example, to implement a procedural phase associated with advancing the instruments to a worksite, the task generation model 215 may generate tasks associated with aligning the instruments with respective ports and advancing the instruments towards the worksite. It should be appreciated that the task generation model may be configured to generate multiple different sets of tasks for implementing the procedural phase. To this end, the different sets of tasks may be performed to implement a procedural phase. Similarly, the tasks within a set of tasks may be performed in different sequences. This enables the various different approaches to performing a procedural phase to be evaluated such that an optimal set of tasks is ultimately selected for implementation.

[0079] FIG. 5 is a schematic diagram of the task selection module 220 configured to select one or more tasks and / or sets of tasks generated by the task generation module 210 for implementation by the control system. As illustrated, the task selection module 220 may select the tasks in two stages.

[0080] In the first stage, a task filtering model 222 is implemented to categorize the generated tasks for filtering. For example, the task filtering model 222 may categorize each task is capable of being performed autonomously, semi-autonomously, manually, or not at all. Accordingly, tasks that are not capable of being performed may be filtered out. It should be appreciated that while the task generation model 215 is generally configured to generate tasks in view of the configurationof the manipulator system, the generative nature of LLMs and LMMs may still result in the generation of tasks that cannot be performed. Accordingly, the task filtering model 222 may function as a sanity check mechanism on the task generation model 215 to ensure only feasible tasks are implemented.

[0081] In a second stage, the remaining tasks are input into a task selection model 225. In some embodiments, the task selection model 225 is an LLM. The task selection model 225 is configured to analyze the potential tasks or sets of tasks to be performed and select a preferred task or set of tasks to implement. For example, the task selection model 225 may be configured to evaluate the tasks or set of tasks using one or more evaluation metrics (e.g., time to perform the task, confidence in ability to autonomously perform the task, an ease of manual implementation, a confidence in safe performance of the task, abilities or personnel, available personnel, or relevant tasks).

[0082] The task selection module 220 then outputs the selected task(s) and provides the selected task(s) to the task assignment module 230, and the robotic action module 250.

[0083] FIG. 6 is a schematic diagram of an example task assignment module 230 of FIG. 2 for assigning manual and / or semiautonomous tasks (such as the tasks selected by the task selection module 220) associated with a medical procedure. Accordingly, the task assignment module 230 may be communicative coupled to the task selection module 220.

[0084] As illustrated, the task assignment module 230 may include a task analysis model 232 configured to analyze the selected tasks and determine personnel to perform the selected tasks. In some embodiments, the task analysis model 232 is an LLM. In these embodiments, the task analysis model 232 may be configured to generate a prompt to the LLM that asks the LLM to identify the various personnel, equipment, robotic systems, etc. required to perform the input task. For example, the task analysis model 232 may identify a type of task associated with the input task (e.g., a repositioning task, an end effector activation task, a configuration task, etc.) and identify which personnel and / or skillsets or qualifications thereof are required to perform the input task. As one example, an anesthetist may be required for performing certain tasks (administering the anesthesia), while the anesthetist is not required for other tasks (manually repositioning a manipulator).

[0085] In some embodiments, the input also includes a “constitution” or a set of rules defining how the LLM is to analyze the tasks to identify the required personnel for each task. For example,the constitution may include qualifications rules that indicate skills required to perform a task. As one specific example, the rules for a task related to creating a port on a subject may include a first subtask to be performed by surgeon to create an incision that will be used for the port and a second subtask to be performed by a nurse to dress the incision. Accordingly, when the input task and the constitution are input to the LLM, the LLM may output the skills required to perform the task (or subtasks) to assign appropriate personnel that satisfy the skill requirements.

[0086] After identifying the requirements to perform the task and / or subtasks thereof, the task assignment module 230 may input the tasks and the data streams 202 into the task assignment model 235. It should be appreciated that the task assignment model 235 may be the same or different LLM as the task analysis model 232. In embodiments where the same LLM is used for both models 232, 235, the functionality of the models 232, 235 may be achieved via a single prompt to the LLM. In these embodiments, the constitution input to the LLM with the prompt may include additional rules that govern the task assignment process.

[0087] For example, one set of rules related to the task assignment process may relate to personnel rules that indicate skills associated with personnel in the medical environment or otherwise associated with the facility. To this end, the personnel rules may provide context to the LLM such that the LLM is able to ensure that only qualified personnel are assigned to perform tasks (or subtasks thereof). In some embodiments, the system 200 is configured to dynamically update the personnel rules as personnel enter / exit the OR. For example, the system 200 may employ techniques for identifying personnel, querying a personnel database to identify skills associated with the identified personnel, and updating the personnel rules to include indications the personnel skills based on the data maintained in the personnel database. One such system for tracking personnel as they enter / exit an OR is described in U.S. Provisional Patent Application 63 / 526,971.

[0088] Additionally, the constitution may include personnel rules for personnel that are not currently present in the OR (e.g., personnel located elsewhere at the facility and / or on call). This may enable the LLM to identify that additional personnel are needed to perform a task (e.g., a phase of the procedure could be more efficiently executed with the inclusion of an additional person, a specialist is needed to address an unexpected condition) and attempt to assign the task to the additional personnel. In some embodiments, personnel presence is determined based upon thedata steams 202 and / or based on one or more prior tasks assigned to the personnel by the system200.

[0089] Based on the inputs, the task analysis model 232 may be configured to output personnel assigned to each task (or subtask thereof). If the required personnel are available, the task assignment module 230 may then output the tasks assignments for communication to the personnel assigned thereto. For, example, the system 200 may output the task assignments via one or more speakers or displays in the OR. Additionally, the system 200 may provide indications of the assigned tasks to personnel via personal electronic devices associated with the assigned personnel, such as a tablet, smartphone, laptop, a wearable device etc. As another example, if the required personnel are not located in the OR, but are otherwise available and located at the facility, the system 200 may automatically generate a paging message informing the assigned personnel of the need associated with the current medical procedure. It should be appreciated that the task assignment module 230 may queue the originally input task until detecting that the assigned personnel is ready to implement the task.

[0090] In scenarios where the assigned personnel are unavailable due to being assigned to another task, the task assignment module 230 may generate an output to the task generation module 210 and / or task selection module 220 requesting that new tasks are generated and / or alternative tasks are selected to realize a current phase of the medical procedure.

[0091] In some embodiments, if the task involves semi-autonomous implementation (or an autonomous implementation with an assigned instrument), the task assignment module 230 may provide an indication of the task to the robotic action module 250 for implementation at the appropriate time. Additionally, the task assignment module 230 may be in communication with one or more devices configured to provide indications of an assigned task to one or more personnel. For example, the task assignment module 230 may output the task assignments to a display device for operator approval. If the operator disapproves the task assignment, the task assignment module 230 may generate one or more alternative task assignments. Alternatively, the operator may provide a user indication of a preferred task assignment and the task analysis model 232 may output a ranking of potential assignments for a given task based on the user indicated preference. The task assignment model 235 may then further assign the tasks to personnel based on the user preference.

[0092] It should be appreciated that while the above description of the task assignment module 230 relates to assigning personnel to perform tasks, in other embodiments, similar techniques can be applied to assign particular- instruments to tasks based on analysis of instrument availability as indicated, for example, by an instrument availability data stream or prior instrument assignments. In these embodiments, the task analysis model 232 may additionally output an instrument type required to perform the task for input into the task assignment model 235. Accordingly, in these embodiments, the task assignment module 230 may output an indication of the assigned instrument to the robotic action module 250 for implementation thereby.

[0093] FIG. 7A is a schematic diagram illustrating detokenizing action tokens to control a repositionable structure (such as the repositionable structures 120 of FIG. 1 ), an instrument, or an auxiliary device to implement a semi-autonomous task. Detokenization is a process under which a system converts an action token 320 (e.g., a text string or commands) output by the robotics transformer model 315 into actual robotic commands 330 (an “action sequence”) to control the indicated equipment, for example, by changing the pose or by activating or otherwise controlling a functionality supported by the indicated equipment.

[0094] As described above, the RT model 315 may be fine-tuned and / or selected based on the particular manipulator system model. As such, knowledge of the capabilities of the manipulator system are incorporated into the RT model 315 itself. Similarly, the RT model 315 may be configured to maintain a list of current equipment coupled to the control system. Accordingly, the action tokens 320 output by the RT model 315 may include a textual indication of the equipment that is to perform the action. For example, an output action token 320 may state “Move Instrument X coupled to Arm A to target work area” or “Use Instrument Y to cauterize incision.” Accordingly, the action tokens 320 and detokenization are specific to a robotic system or model based on the capabilities of a given robotic system.

[0095] In some embodiments, the action tokens 320 may be further tailored to the current state of the manipulator system. For example, the RT model 315 may be configured to accept a kinematic data stream as an input such that the output action tokens indicate a pose for the instrument and / or arm in the robotic coordinate system. Accordingly, as one example, rather than outputting an action token 320 indicating that an instrument should move to the target work area, the output action token 320 may instead indicate “Move Instrument X to position (x, y, z) andorientation a, P, y).” As a result, the action tokens 320 output by the RT model 315 are tailored to the specific state of the manipulator system improving the ability of the control system to accurately interpret and implement autonomously-generated control commands.

[0096] Similarly, because the RT model 315 is aware of the functionality supported by the instruments coupled to the control system, the RT model 315 is able to output commands that are specifically tailored to the implementing instrument. For example, rather than outputting an action token that states “Clamp blood vessel,” the RT model 315 may instead output an action token 320 that states “Place gripper around blood vessel at position (x, y, z) and engage grip to exert 50 pascals of force.” As a result, the RT model 315 is able to generate action tokens 320 that have improved alignment with the actual functionality supported by the control system, thereby enabling more precision in the generation of the action tokens 320 and more robust usage of instrument functionality.

[0097] After the RT model 315 generates the action tokens 320, the action tokens 320 are input into the robotic controller 325 to convert the natural language action tokens into de-tokenized commands 330 (e.g., parameterized commands, such as PID parameter values or MPC parameter values, used by the control system actually realize the instructed actions). With simultaneous reference to FIGS. 7B and 7C, depicted are example parameterized command structures utilized by the control system.

[0098] The de-tokenized command 730 of FIG. 7B is configured to control operation of a gripper instrument that has 6 motive degrees of freedom (DOFs). Accordingly, the de-tokenized command 730 may include deltas by which the control system is to change each DOF. Additionally, the de-tokenized command 830 includes additional DOFs related to the functionality supported by the instrument. For example, in the illustrated example related to a gripper, the “gripper” DOF may indicate an amount of force the gripper should exert. Similarly, the de- tokenized command 731 of FIG. 7C is configured to control operation of a fluid extraction instrument that has 5 motive degrees of freedom (DOFs) and one functional DOF. It should be appreciated that the length of the de-tokenized control commands 330 may vary based on the number of motive and functional DOFs supported by the controlled equipment. Accordingly, the length of the de-tokenized commands 330 have fewer unused parameters, thereby enabling more efficient usage of control buses.

[0099] In some embodiments, the de-tokenized commands 330 may also control one or more user feedback devices, such as a display device. For example, a dc-tokcnizcd command 330 may be configured to cause a display unit to provide an instruction for executing a task that is to be performed manually or semi-autonomously by a user.

[0100] While FIGs. 2-7C describe a process for performing task assignment to personnel, instruments, and robotic systems to implement a one or more tasks, it should be appreciated that the disclosed process may be repeated throughout the procedure until its completion. As a result, each stage of the procedure may be discretized into tasks, that are converted into de-tokenized commands such that any portion of the procedure can be implemented via closed-loop autonomous control, semi-autonomous control, or manually by personnel. It should be appreciated that in some embodiments, the modules 210, 220, and / or 230 may analyze the multimodal data stream 202 to anticipate future tasks, and assign the future tasks, that are predicted to be performed based on a current state. Thus, the tasks analyzed by the modules 210, 220, and / or 230may differ from the task being implemented by the robotic action module 250 or by personnel. As a result, the control system is able to anticipate future actions and future task assignments in order to ensure closed- loop control of the overall procedure occurs more efficiently.

[0101] In some embodiments, the control system may further include a scheduler module configured to generate a sequence of expected tasks (and their corresponding data modalities and / or assigned personnel, instruments, robotic systems, or auxiliary devices) and their corresponding triggers for execution. Accordingly, when the scheduler detects, based on the multimodal data streams 202, that a trigger event has occurred, the scheduler can route one or more subsequent tasks to robotic action module 250 or provide indications of a specific task to respectively assigned personnel. It should be appreciated that if the multimodal data streams 202 indicate that conditions have changed (e.g., an anomalous condition is detected), the scheduler can adjust the priority and / or sequencing to prioritize addressing the current conditions.

[0102] As described herein, in some embodiments, the disclosed techniques may be applied in a training environment to simulate one or more scenarios that personnel may encounter during a live medical procedure. FIG. 8 depicts an alternate version of the system 200 that may be implemented to facilitate personnel training with one or more training objects (e.g., a trainingdummy). In particular, in the embodiment of FIG. 8, the system 200 includes a training module 205 between the multimodal data streams 202 and the task generation module 210.

[0103] The training module 205 may be configured to present training data 203 to the task generation module as a data stream 202 that is analyzed to generate one or more tasks. In some embodiments, the training data 203 is a new standalone data stream generated during a historical performance of the medical procedure. In other embodiments, the training data 203 includes features by which one or more of the multi-modal data streams 202 should be modified. For example, the training data 203 may include image data of an unexpected feature the operator may encounter during performance of the medical procedure. The training data 203 may also include synthetic image data generated from one or more OR simulation engines and different simulation parameters associated therewith (e.g., a number of people in the room and their roles, a room layout, equipment availability, procedure type, procedure phase, etc.). In some embodiments, the training module 205 may perform a feature transfer algorithm (e.g., a neural style transfer algorithm) to transpose the unexpected feature into image data included in a data stream 202. As a result, the task generation module 210 may generate tasks as if that feature was actually present in the training object. Similarly, the modified data stream may be presented in a display viewed by the operator such that the operator is interface with the manipulator system in a manner as if the feature was actually present.

[0104] As another example, a training scenario may be an unexpected equipment failure or other unexpected limit on operation of the manipulator system. In these examples, the training module 205 may generate a natural language description of the failure such that the task generation module 210 does not generate tasks that involve the impacted equipment. Additionally, the training module 205 may modify operation of the control system 140 such that the control commands simulate the response of the manipulator (or lack thereof) in the training scenario.

[0105] Additionally, the training module 205 may include an evaluation model 207 configured to determine the training scenario and / or evaluate operation of the manipulator system during the training scenario. In some embodiments, the evaluation model 207 is an LLM. Accordingly, the evaluation model 207 may be configured to analyze the skills, qualification, procedure history, etc. associated with the personnel performing the training procedure to select a particular training scenario. To this end, the training module 205 may be communicatively coupled to a database(not depicted) at which historical and / or simulated data for a plurality of different scenarios is maintained. Accordingly, upon determining a particular training scenario to implement, the evaluation model 207 may cause the model to obtain the appropriate training data to utilize as the training data 203.

[0106] In addition to selecting the training scenarios, the evaluation model 207 may be configured to evaluate performance of the training procedure during the training scenario. To this end, the database may additionally maintain a rubric associated with each training scenario. Accordingly, the training module 205 may be configured to input the rubric and the data streams 202 into the evaluation model 207 to generate an evaluation of the personnel in accordance with the evaluation criteria set forth in the corresponding rubric.

[0107] FIG. 9 is a flow diagram of a method 900 for performing task assignment of tasks for a medical procedure. The method 900 may be performed by a processor system that implements an Al assist module (such as the Al assist module 180). In implementations, the method 900 may further be performed with a system including one or more repositionable structures that may be operably coupled to one or more instruments (such as the instruments 122) and controlled via a control system (such as the control system 140).

[0108] The method 900 may begin at block 902 when the Al assist module receives a plurality of data streams (such as the multimodal data streams 202) from one or more data sources. The plurality of data streams may include multi-modal data streams of different types of data, as provided by different devices or systems. For example, the plurality of multi-modal data streams may include data from a video camera, audio data, force sensor data, system events (such as the events data stream 202b), endoscopic image data (such as the procedure video data stream 202e), operating room image data (such as the external video data stream 202d), kinematics data (such as the kinematics data stream 202c), haptics data, force data (such as the force data stream 202a), personnel data (such as the personnel data stream 202f), shape sensing data, 3D point cloud data, tissue impedance data, environmental data, personnel procedure history, phase information, task information, personnel training data, instrument availability data, etc. Accordingly, the plurality of data streams may include environmental data including one or more of data indicative of an interaction between the system and its environment, force data associated with instrument contact with patient tissue, force data associated with feedback to one or more manipulators, personnelpresent in a medical environment, and data indicative of system collisions with objects in the environment. The plurality of data streams may include data indicative of schedules of personnel to determine at what times personnel may be available to perform tasks associated with a procedure. In some embodiments, the plurality of data streams may also include user preference data.

[0109] At block 904, the Al assist module may analyze, via a task generation machine learning model, the plurality of data streams to generate natural language output relating to one or more tasks to be performed in furtherance of the medical procedure. Additionally, the Al assist module may identify one or more tasks to be performed by an instrument operatively coupled to one or more repositionable structures or by a user in a manual or semi-autonomous manner.

[0110] To analyze the plurality of data streams, the Al assist module may be configured to obtain and input a task generation constitution into the task generation machine learning model. The task generation constitution may be configured to define a set of rules for analyzing the data streams. For example, the task generation constitution may include foundational rules that define allowable robotic actions (e.g., limitations on instrument operation for patient and user safety, such as Asimov’s three rules), safety rules that define safe operation of the computer-assisted medical system (e.g., spatial limitations, sanitization standards, etc.), embodiment rules that define capabilities and limitations of the computer-assisted medical system, associated instruments and devices (e.g., degrees of freedom of robotic arm movement, imaging capabilities of endoscopes and imaging devices, volume capabilities of extraction and injection devices, etc.), and / or user preference rules (e.g., rules defined and input by a user or personnel for performing a procedure or task). Accordingly, the task generation machine learning model may generate the natural language outputs according to the set of rules of the task generation constitution.

[0111] As described herein, inputting the plurality of data streams into the task generation machine learning model may include generating embedding of the plurality of data streams (such as via the embedding stage 310). For example, the Al assist module may provide an event data stream into a transformer model (such as the transformer model 214) trained to identify event or procedural milestones. Accordingly, an embedding of a data stream may be a natural language description of the medical procedure and / or phase thereof. Additionally, in some examples, the Al assist module may analyze the plurality of data streams to determine one or more current tasksbeing performed and, based on the current task being performed, generate one or more additional tasks to be performed.

[0112] At block 906, the Al assist module may analyze, via a task assignment machine learning model, the one or more tasks to assign the tasks to respective personnel. To assign the one or more tasks, the Al assist module may input, into the task assignment machine learning model, a task assignment constitution defining a set of rules for task assignment. For example, the task assignment constitution may include qualifications rules that indicate skills required to perform a task and / or personnel rules that indicate skills associated with personnel in the medical environment. In some embodiments, the Al assist module adaptively updates the personnel rules of the task assignment constitution as personnel enter and exit the medical environment.

[0113] The Al assist module may further analyze the plurality of data streams to identify a completion of a task assigned to a particular person to generate natural language output relating to one or more additional tasks to be performed in furtherance of the medical procedure. Additionally, the Al assist module may analyze, via the task assignment machine learning model, the one or more additional tasks to be performed to assign the one or more additional tasks to respective personnel. In some embodiments, the task assignment machine learning model may classify the one or more tasks as being able to be performed autonomously, semi-autonomously, manually, or not at all, and assign respective personnel to the semi-autonomous and / or manual tasks based on the task assignment constitution.

[0114] In some embodiments, the Al assist module may determine, via the task assignment machine learning model, that additional personnel are needed to perform a task or that personnel assigned a task are not within the medical environment. In these embodiments, the Al assist module may be configured to provide an indication indicative of the determination to a user, the additional personnel, or the personnel not within the medical environment.

[0115] At block 908, the Al assist module may provide indications to the respective personnel for performing the respective tasks assigned to the respective personnel. For example, the indications may comprise one or more of a natural language output, a visual indication, and an audio indication. The Al assist module may provide the indication of an assigned task to the respective personnel via a device associated with respective assigned personnel (e.g., a tablet,smartphone, computer workstation, smart watch, pager, wearable device, etc.), via the computer- assisted medical system, and / or via a display device within the medical environment.

[0116] In some embodiments, the techniques described with respect to the method 900 are extended to assign personnel to procedures as part of a scheduling system. In these embodiments, the plurality of data streams may include a data stream indicative of procedures to be performed at a facility and a data stream indicative of personnel data for the facility. For example, the Al assist module may detect, via a data stream, a new procedure to be performed and, responsive to the detection, assign personnel to the new procedure. In some embodiments, the Al assist module may analyze, via a procedure assignment machine learning model, the data stream indicative of personnel data to assign personnel to respective procedures to be performed. In some embodiments, the Al assist module inputs, into the procedure assignment machine learning model, a procedure assignment constitution defining a set of rules for procedure assignment. These rules may be similar to the those included in the task assignment constitution used to assign personnel to specific tasks in furtherance of a procedure. Accordingly, the Al assist module may determine personnel qualifications for procedure assignments in a similar manner as determining personnel qualifications for task assignments.

[0117] In addition to qualification matching, the procedure assignment constitution may include rules about procedure prioritization such that personnel can be re-assigned based on priority. For example, an emergency procedure to be performed in an emergency room may be a higher priority for some personnel than other scheduled procedures. In some embodiments, the procedure rules in the procedure assignment constitution also include information on personnel requirements for various phases of the procedure. Thus, the procedure assignment machine learning model may estimate a narrower time window when personnel is to be assigned to a procedure such that they can be reassigned when not needed. Additionally, Al assist modules that implement the method 900 may be configured to report a current status of the procedure to a central system. As a result, a data stream may include phase information associated with current medical procedures. Accordingly, the control system may be configured to reassign, via the procedure assignment machine learning model, personnel based on procedure priority and / or a current phase of the current medical procedures.

[0118] Regardless, after assigning the personnel to the procedure, the Al assist module may be configured to generate, based on the personnel assigned to the respective procedures, a schedule for performing the respective procedures. The Al assist module may then transmit data indicative of the schedule to assigned personnel. The Al assist module may also store the schedule in a memory or provide the schedule to a centralized device for the facility.

[0119] In some embodiments, the techniques described with respect to the method 900 are extended to facilitate personnel training. Accordingly, in these examples, the Al assist module may also include a training module (such as the training module 205) to integrate the plurality of data streams with simulated data or historical data representative of one or more scenarios that may arise during a procedure (such as the training data 203). Accordingly, during these embodiments, the method 900 may be performed using a training object instead of a human subject.

[0120] In these embodiments, the Al assist module may analyze, via the task generation machine learning model, the plurality of data streams to identify one or more training tasks to be performed in response to the one or more scenarios. For example, the one or more scenarios may reflect an unexpected event for the procedure, such as one or more of a detection of an unexpected feature associated with a subject or an equipment malfunction. Accordingly, the training module may be configured select scenarios on which to train the personnel (such as by using the evaluation model 207). More particularly, the training module may select the scenarios based on personnel data indicative of personnel skills, procedure history, or other characteristics or a training plan associated with personnel or a facility. In some embodiments the personnel data and / or the training plan are input along with a training constitution to a LLM to select a scenario. Regardless, the control system may be configured to obtain the simulated data or historical data for the selected scenario from a stored scenario library to utilize as the training data.

[0121] During performance of the training procedure, a data stream indicative of operational data that indicates of operation of the one or more repositionable structures may be modified by the training module to reflect a state of the training scenario. For example, the training module may map the operational data to the historical data or simulated data. To this end, the training module may modify the operational data to include a portion of the historical or simulated data(such as by performing feature transfer techniques or tracking a simulated position of an instrument).

[0122] After the training module finishes modifying the one or more data streams to implement the selected training scenario, the Al assist module may analyze, via the task generation machine learning model, the data streams to identify one or more training tasks to be performed in response to the one or more scenarios. This may occur in substantially the same manner as described with block 904. The Al assist module may then analyze, via the task assignment machine learning model, the identified training tasks to assign the identified training tasks to respective personnel. This may occur in substantially the same manner as described with respect to block 906. Subsequently, the AT assist module may provide indications to the respective personnel for performing the respective training tasks assigned to the respective personnel. This may occur in substantially the same manner as described with respect to block 908.

[0123] During the training procedure, the training module may be configured to analyze the one or more data streams to generate a metric to evaluate performance during the selected scenarios. For example, the Al assist module may input a projection of the data streams and a rubric into a LLM to generate the metric. Based on the metric, the training module may generate feedback data. The feedback data may be provided in real-time to via the means by which task assignments are provided and / or after the completion of the training procedure.

[0124] One or more components of the examples discussed in this disclosure, such as control system 140, may be implemented in software for execution on one or more processors of a computer system. The software may include code that when executed by the one or more processors, configures the one or more processors to perform various functionalities as discussed herein. The code may be stored in a non-transitory computer readable storage medium (e.g., a memory, magnetic storage, optical storage, solid-state storage, etc.). The computer readable storage medium may be part of a computer readable storage device, such as an electronic circuit, a semiconductor device, a semiconductor memory device, a read only memory (ROM), a flash memory, an erasable programmable read only memory (EPROM); a floppy diskette, a CD-ROM, an optical disk, a hard disk, or other storage device. The code may be downloaded via computer networks such as the Internet, Intranet, etc. for storage on the computer readable storage medium. The code may be executed by any of a wide variety of centralized or distributed data processingarchitectures. The programmed instructions of the code may be implemented as a number of separate programs or subroutines, or they may be integrated into a number of other aspects of the systems described herein. The components of the computing systems discussed herein may be connected using wired and / or wireless connections. In some examples, the wireless connections may use wireless communication protocols such as Bluetooth, near-field communication (NFC), Infrared Data Association (IrDA), home radio frequency (HomeRF), IEEE 502.11, Digital Enhanced Cordless Telecommunications (DECT), and wireless medical telemetry service (WMTS).

[0125] Various general-purpose computer systems may be used to perform one or more processes, methods, or functionalities described herein. Additionally or alternatively, various specialized computer systems may be used to perform one or more processes, methods, or functionalities described herein. In addition, a variety of programming languages may be used to implement one or more of the processes, methods, or functionalities described herein.

[0126] While certain examples and examples have been described above and shown in the accompanying drawings, it is to be understood that such examples and examples are merely illustrative and are not limited to the specific constructions and arrangements shown and described, since various other alternatives, modifications, and equivalents will be appreciated by those with ordinary skill in the art.

Claims

WHAT IS CLAIMED IS:

1. A system to assist personnel in performing a medical procedure using a computer- assisted medical system, wherein the system is configured to: receive a plurality of data streams related to a medical procedure, wherein the plurality of data streams includes one or more of system data, medical environment data, and indications of personnel performing the medical procedure; analyze, using a task generation machine learning model, the plurality of data streams to generate natural language output relating to one or more tasks to be performed in furtherance of the medical procedure, wherein one or more inputs into the task generation machine learning model includes inputting embeddings of the plurality of data streams; analyze, via a task assignment machine learning model, the one or more tasks to assign the tasks to respective personnel; and provide indications to the respective personnel for performing the respective tasks assigned to the respective personnel.

2. The system of claim 1, wherein the plurality of data streams includes one or more of endoscopic image data, operating room image data, laparoscopic image data, personnel procedure history data, phase information data, and personnel training data.

3. The system of claim 1, wherein the system is further configured to obtain a task generation constitution for input into the task generation machine learning model, the task generation constitution defining a set of rules for analyzing the one or more data streams.

4. The system of claim 3, wherein to generate the natural language output relating to the one or more tasks, the system is configured to: generate the natural language outputs according to the set of rules of the task generation constitution.

5. The system of claim 3, wherein the task generation constitution further includes user preference rules.

6. The system of claim 1 , wherein to analyze the plurality of data streams using the task generation machine learning model, the system is further configured to: analyze the plurality of data streams to determine one or more current tasks being performed; and analyze, using the task generation machine learning model, the one or more current tasks being performed to generate the natural language output.

7. The system of claim 1, wherein to assign the one or more tasks to respective personnel, the system is configured to: input, into the task assignment machine learning model, a task assignment constitution defining a set of rules for task assignment.

8. The system of claim 7, wherein the task assignment constitution includes procedural rules that define rules related to specific actions that are performed in furtherance of medical procedures.

9. The system of claim 7, wherein the task assignment constitution includes qualifications rules that indicate skills required to perform a task.

10. The system of claim 7, wherein the task assignment constitution includes personnel rules that indicate skills associated with personnel in the medical environment.

11. The system of claim 10, wherein the personnel rules are adaptively updated as personnel enter and exit the medical environment.

12. The system of claim 1, wherein to provide the indications, the system is configured to: transmit an indication of the assigned task to a device associated with the assigned personnel.

13. The system of claim 1, wherein the system is further configured to:determine, via the task assignment machine learning model, that additional personnel are needed to perform a task or that personnel assigned a task arc not within the medical environment.

14. The system of claim 13, wherein the system is further configured to: provide an indication indicative of the determination to a user, the additional personnel, or the personnel not within the medical environment.

15. The system of claim 1, wherein to analyze the one or more tasks the system is configured to: classify the one or more tasks as being able to be performed autonomously, semi- autonomously, manually, or not at all.

16. The system of claim 1, wherein the indications comprise one or more of a natural language output, a visual indication, and an audio indication.

17. The system of claim 1, wherein the indications are provided via a computer-assisted medical system.

18. The system of claim 1, wherein the indications are provided via a device associated with the respective personnel.

19. The system of claim 1, wherein the system is further configured to: analyze, via the task generation machine learning model, the data streams to identify a completion of a task assigned to a particular person to generate natural language output relating to one or more additional tasks to be performed in furtherance of the medical procedure; and analyze, via the task assignment machine learning model, the one or more additional tasks to be performed to assign the one or more additional tasks to respective personnel.

20. The system of any one of claims 1 to 19, wherein:the plurality of data streams includes a data stream indicative of procedures to be performed at a facility and a data stream indicative of personnel data for the facility; and the system is further configured to: analyze, via a procedure assignment machine learning model, the data stream indicative of personnel data to assign personnel to respective procedures to be performed; and generate, based on the personnel assigned to the respective procedures, a schedule for performing the respective procedures.

21. The system of claim 20, the plurality of data streams includes a data stream indicative of procedures to be performed at the facility, and the system is configured to: analyze, via the procedure assignment machine learning model, the data stream indicative of instrument data to assign instruments to respective procedures to be performed.

22. The system of claim 20, wherein to analyze the plurality of data streams to assign personnel to the one or more tasks the system is further configured to: input, into the procedure assignment machine learning model, a procedure assignment constitution defining a set of rules for procedure assignment.

23. The system of claim 20, wherein the system is further configured to: detect, via a data stream, a new procedure to be performed; and responsive to the detection, assign, via the procedure assignment machine learning model, personnel to the new procedure.

24. The system of claim 23, wherein the new procedure is a high-priority procedure, and to assign personnel to the new procedure, the system is further configured to: reassign, via the procedure assignment machine learning model, personnel assigned to another scheduled procedure to the high-priority procedure.

25. The system of claim 24, wherein the plurality of data streams includes phase information associated with current medical procedures and wherein the system is further configured to:reassign, via the procedure assignment machine learning model, personnel based on a current phase of the current medical procedures.

26. The system of claim 20, wherein the system is further configured to: transmit data indicative of the schedule to assigned personnel.

27. The system of any one of claims 1 to 19, wherein the plurality of data streams includes simulated data or historical data representative of one or more scenarios that may arise during a procedure, and wherein the system is further configured to: analyze, via the task generation machine learning model, the data streams to identify one or more training tasks to be performed in response to the one or more scenarios; analyze, via the task assignment machine learning model, the identified one or more training tasks to assign the identified training tasks to respective personnel; and provide indications to the respective personnel for performing the respective training tasks assigned to the respective personnel.

28. The system of claim 27, wherein the one or more scenarios reflect an unexpected event for the procedure.

29. The system of claim 28, wherein the unexpected event is one or more of a detection of an unexpected feature associated with a subject or an equipment malfunction.

30. The system of claim 27, wherein the system is further configured to obtain the simulated data or historical data from a stored scenario library.

31. The system of claim 30, wherein the simulated data or historical data includes image data for output via a display device of the system.

32. The system of claim 27, further comprising one or more repositionable structures configured to support respective instruments, wherein the medical procedure is performed using a training object.

33. The system of claim 32, wherein a data stream is operational data indicative of operation of the one or more repositionable structures and the system is configured to: map the operational data to the historical data or simulated data.

34. The system of claim 27, wherein the system includes a training module configured to select scenarios on which to train the personnel.

35. The system of claim 34, wherein the training module selects scenarios based on personnel data indicative of personnel skills, procedure history, or other characteristics.

36. The system of claim 34, wherein the training module selects scenarios based on a training plan associated with personnel or a facility.

37. The system of claim 34, wherein the training module is configured to generate a metric to evaluate performance during the selected scenarios.

38. The system of claim 37, wherein the training module generates feedback data based on the metric.

39. A method to assist personnel in performing a medical procedure, the method comprising: receiving a plurality of data streams, wherein the plurality of data streams includes one or more of system data, medical environment data, and indications of personnel performing the medical procedure; analyzing, using a task generation machine learning model, the plurality of data streams to generate natural language output relating to one or more tasks to be performed in furtherance of the medical procedure, wherein one or more inputs into the task generation machine learning model includes inputting embeddings of the plurality of data streams; analyzing, via a task assignment machine learning model, the one or more tasks to assign the tasks to respective personnel; andproviding indications to the respective personnel for performing the respective tasks assigned to the respective personnel.

40. The method of claim 39, wherein the plurality of data streams includes one or more of endoscopic image data, operating room image data, laparoscopic image data, personnel procedure history data, phase information data, and personnel training data.

41. The method of claim 39, further comprising inputting a task generation constitution into the task generation machine learning model, the constitution defining a set of rules for analyzing the plurality of data streams.

42. The method of claim 41, wherein generating the natural language output relating to the one or more tasks is performed according to the set of rules of the task generation constitution.

43. The method of claim 41, wherein the task generation constitution further includes user preference rules.

44. The method of claim 39, wherein analyzing the plurality of data streams using the task generation machine learning model comprises: analyzing the plurality of data streams to determine one or more current tasks being performed; and analyze, using the task generation machine learning model, the one or more current tasks being performed to generate the natural language output.

45. The method of claim 39, wherein assigning the tasks to respective personnel comprises: inputting, into the task assignment machine learning model, a task assignment constitution defining a set of rules for task assignment.

46. The method of claim 45, wherein the task assignment constitution includes procedural rules that define rules related to specific actions that are performed in furtherance of medical procedures.

47. The method of claim 45, wherein the task assignment constitution includes qualifications rules that indicate skills required to perform a task.

48. The method of claim 45, wherein the task assignment constitution includes personnel rules that indicate skills associated with personnel in a medical environment.

49. The method of claim 48, wherein the personnel rules are adaptively updated as personnel enter and exit the medical environment.

50. The method of claim 39, where providing the indications comprises: transmitting an indication of the assigned task to a device associated with the assigned personnel.

51. The method of claim 39, further comprising: determining, via the task assignment machine learning model, that additional personnel are needed to perform a task or that personnel assigned a task is not within the medical environment.

52. The method of claim 51, further comprising: providing an indication indicative of the determination to a user, the additional personnel, or the personnel not within the medical environment.

53. The method of claim 39, wherein analyzing the one or more tasks comprises: classifying the one or more tasks as being able to be performed autonomously, semi- autonomously, manually, or not at all.

54. The method of claim 39, wherein the indications comprise one or more of a natural language output, a visual indication, an audio indication.

55. The method of claim 39, wherein the indications are provided via a computer-assisted medical system.

56. The method of claim 39, wherein the indications are provided via a device associated with the respective personnel.

57. The method of claim 39, further comprising: analyzing, via the task generation machine learning model, the data streams to identify a completion of a task assigned to a particular person to generate natural language output relating to one or more additional tasks to be performed in furtherance of the medical procedure; and analyzing, via the task assignment machine learning model, the one or more additional tasks to be performed to assign the additional tasks to respective personnel.

58. The method of any one of claims 39 to 57, wherein the plurality of data streams includes a data stream indicative of procedures to be performed at a facility, and a data stream indicative of personnel data for the facility; and the method further comprises: analyzing, via a procedure assignment machine learning model, the data stream indicative of personnel data to assign personnel to respective procedures to be performed; and generating, based on the personnel assigned to the respective procedures, a schedule for performing the respective procedures.

59. The method of claim 58, wherein the plurality of data streams includes an instrument availability data stream indicative of availability of instruments used in the procedure, and the method further comprises: analyzing, via the procedure assignment machine learning model, the instrument availability data stream to assign instruments to respective procedures to be performed.

60. The method of claim 58, wherein analyzing the plurality of data streams to assign personnel to the tasks comprises: inputting, into the procedure assignment machine learning model, a procedure assignment constitution defining a set of rules for procedure assignment.

61. The method of claim 60, wherein the method further comprises: detecting, via a data stream, a new procedure to be performed; and responsive to the detection, assigning, via the procedure assignment machine learning model, personnel to the new procedure.

62. The method of claim 61, wherein the new procedure is a high-priority procedure, and assigning personnel to the new procedure comprises: reassigning, via the procedure assignment machine learning model, personnel assigned to another scheduled procedure to the high-priority procedure.

63. The method of claim 62, wherein the plurality of data streams includes phase information associated with current medical procedures, and the method further comprises: reassigning, via the procedure assignment machine learning model, personnel based on a current phase of the current medical procedures.

64. The method of claim 58, further comprising transmitting data indicative of the schedule to assigned personnel.

65. The method of any one of claims 39 to 57, wherein the plurality of data streams includes simulated data or historical data representative of one or more scenarios that may arise during a procedure, and wherein the method further comprises: analyzing, via the task generation machine learning model, the data streams to identify one or more training tasks to be performed in response to the one or more scenarios; analyzing, via the task assignment machine learning model, the identified one or more training tasks to assign the identified training tasks to respective personnel; and providing indications to the respective personnel for performing the of respective training tasks assigned to the respective personnel.

66. The method of claim 65, wherein the one or more scenarios reflect an unexpected event for the procedure.

67. The method of claim 66, wherein the unexpected event is one or more of a detection of an unexpected feature associated with a subject or an equipment malfunction.

68. The method of claim 65, further comprising obtaining the simulated data or historical data from a stored scenario library.

69. The method of claim 68, wherein the simulated data or historical data includes image data for output via a display device of the system.

70. The method of claim 65, wherein the medical procedure is performed using a training object.

71. The method of claim 70, wherein a data stream is operational data indicative of operation of one or more repositionable structures of a computer-assisted system and the method comprises; mapping the operational data to the historical data or simulated data.

72. The method of claim 65, further comprising: selecting, via a training module, scenarios on which to train the personnel.

73. The method of claim 72, wherein selecting the scenarios comprises: selecting, via the training module, scenarios based on personnel data indicative of personnel skills, procedure history, or other characteristics.

74. The method of claim 72, wherein selecting the scenarios comprises: selecting, via the training module, scenarios based on based on a training plan associated with personnel or a facility.

75. The method of claim 73, further comprising:generating, via a training module, a metric to evaluate performance during the selected scenarios.

76. The method of claim 75, further comprising: generating, via the training module, feedback data based on the metric.

77. One or more non-transitory, computer-readable media storing instructions that, when executed by one or more processors, causes the one or more processors to perform the method of any one of claims 39-76.

Citation Information

Cited By

  • Multi-task ai system for dynamic and intelligent robotic task planning based on a user specified persona

    WO2026102066A1