Virtual medical session interactions with non-powered objects

By using non-powered objects to map user interactions to virtual components in extended reality, the system addresses the inefficiencies of powered robotic systems, achieving energy-efficient and compute-efficient virtual interactions in robotic medical procedures.

WO2026006568A1PCT designated stage Publication Date: 2026-01-02INTUITIVE SURGICAL OPERATIONS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/035450
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-06-26
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing robotic medical systems face challenges in maintaining efficient and reliable operation due to the complexity and energy inefficiency of implementing robotic medical procedures, particularly when using powered objects that require extensive computational resources.

Method used

A system utilizing non-powered physical objects to map user interactions to virtualized components in an extended reality environment, using sensor data to track interactions and eliminate the need for electronics on these objects, thereby reducing computational intensity and improving energy efficiency.

Benefits of technology

This approach allows for energy-efficient and compute-efficient virtual interactions by transforming user interactions with non-powered objects into virtual poses or configurations, enhancing the operational efficiency of robotic medical systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025035450_02012026_PF_FP_ABST
    Figure US2025035450_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Virtual medical session interactions with non-powered objects is provided. A processor can identify a setting for an application to virtualize a medical session with a system and an indication of an object representing a component of the system. The processor can select, based on the object and the setting, a model to process data from a sensor capturing user interactions with the object. The processor can determine, using the model and the data, a user interaction with the object. The processor can transform the determined user interaction with the physical object to at least one of a virtual pose or a virtual configuration for a virtual component that represents the component of the computer-assisted robotic system. The processor can display the virtual component according to the virtual pose or the virtual configuration.
Need to check novelty before this filing date? Find Prior Art

Description

VIRTUAL MEDICAL SESSION INTERACTIONS WITH NON-POWERED OBJECTSCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of priority under 35 U.S.C. § 119 to U.S. Provisional Patent Application No. 63 / 665,878, filed June 28, 2024, which is hereby incorporated by reference herein in its entirety.BACKGROUND

[0002] Medical procedures can be performed in a medical environment, such as an operating room. As the amount and variety of equipment in the operating room increases and medical procedures increase in complexity, it can be challenging to maintain the equipment operating efficiently and reliably.SUMMARY

[0003] The technical solutions of this disclosure addresses challenges in virtualizing robotic medical procedures via non-powered or communicatively uncoupled objects used as proxies to detect user interactions on virtual components of a robotic medical system. A nonpowered object can refer to or include an object or component that may not have the capability to be powered (e.g., via electricity). Implementing robotic medical surgery virtualization can be energy inefficient and compute intensive, adding to the complexity of the implementation. The technical solutions reduce the computational intensity and improve energy efficiency by using non-powered physical objects to map user interactions to virtualized components in an extended reality (XR) environment. The system can use sensor data to track user’s interactions with the object, allowing for the physical object to be free from such electronics while providing virtual interactions in an energy and compute efficient manner.

[0004] At least one aspect of the technical solutions relates to a system. The system can include one or more processors coupled with memory. The one or more processors can identify a an application to virtualize at least a portion of a medical session where a computer-assisted robotic system is used. The one or more processors can identify an indication of a physical object that represents a component associated with the computer-assisted robotic system. The one or more processors can select, based on the physical object, one or more models trained with machine learning to process data from one or more optical sensors that capture userinteractions with the physical object. The one or more processors can determine, using the one or more models and based on the data collected from the one or more optical sensors, a user interaction with the physical object. The one or more processors can transform the determined user interaction with the physical object to at least one of a virtual pose or a virtual configuration for a virtual component that represents the component of the computer-assisted robotic system. The one or more processors can display, on a display unit attached to computer- assisted robotic system, the virtual component according to the at least one of the virtual pose or the virtual configuration.

[0005] At least one aspect of the technical solutions relates to a system. The system can include one or more processors coupled with memory. The one or more processors can identify a setting for an application to virtualize at least a portion of a medical session where a computer-assisted robotic system is used. The one or more processors can identify an indication of a physical object that represents a component associated with the computer- assisted robotic system. The one or more processors can select, based on the physical object and the setting, one or more models trained with machine learning to process data from one or more optical sensors that capture user interactions with the physical object. The one or more processors can determine, using the one or more models and based on the data collected from the one or more optical sensors, a user interaction with the physical object. The one or more processors can transform the determined user interaction with the physical object to at least one of a virtual pose or a virtual configuration for a virtual component that represents the component of the computer-assisted robotic system. The one or more processors can display, on a display unit attached to computer-assisted robotic system, the virtual component according to the at least one of the virtual pose or the virtual configuration.

[0006] At least one aspect of the technical solutions relates to a method. The method can include one or more processors identifying an application to virtualize at least a portion of a medical session where a computer-assisted robotic system is used. The method can include the one or more processors identifying an indication of a physical object that represents a component associated with the computer-assisted robotic system. The method can include the one or more processors selecting, based on the physical object, one or more models trained with machine learning to process data from one or more optical sensors that capture user interactions with the physical object. The method can include the one or more processors determining, using the one or more models and based on the data collected from the one or more optical sensors, a user interaction with the physical object. The method can includetransforming, by the one or more processors, the determined user interaction with the physical object to at least one of a virtual pose or a virtual configuration for a virtual component that represents the component of the computer-assisted robotic system. The method can include the one or more processors displaying, on a display unit attached to computer-assisted robotic system, the virtual component according to the at least one of the virtual pose or the virtual configuration.

[0007] At least one aspect of the technical solutions relates to a non-transitory computer- readable medium storing processor executable instructions. The instructions, when executed by one or more processors, can cause the one or more processors to identify an application to virtualize at least a portion of a medical session where a computer-assisted robotic system is used. The instructions, when executed by one or more processors, can cause the one or more processors to identify an indication of a physical object that represents a component associated with the computer-assisted robotic system. The instructions, when executed by one or more processors, can cause the one or more processors to select, based on the physical object, one or more models trained with machine learning to process data from one or more optical sensors that capture user interactions with the physical object. The instructions, when executed by one or more processors, can cause the one or more processors to determine, using the one or more models and based on the data collected from the one or more optical sensors, a user interaction with the physical object. The instructions, when executed by one or more processors, can cause the one or more processors to transform the determined user interaction with the physical object to at least one of a virtual pose or a virtual configuration for a virtual component that represents the component of the computer-assisted robotic system. The instructions, when executed by one or more processors, can cause the one or more processors to display, on a display unit attached to computer-assisted robotic system, the virtual component according to the at least one of the virtual pose or the virtual configuration.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The accompanying drawings are not intended to be drawn to scale. Like reference numbers and designations in the various drawings indicate like elements. For purposes of clarity, not every component can be labeled in every drawing. In the drawings:

[0009] FIG. 1 depicts an example system for virtual medical session interactions with an non-powered object.

[0010] FIG. 2 illustrates an example of a physical object that can be used for user interaction and virtualization or simulation of a component of a robotic medical system.

[0011] FIG. 3 illustrates an example of a surgical system, in accordance with some aspects of the technical solutions.

[0012] FIG. 4 depicts an example block diagram of an example computer system is shown, in accordance with some embodiments.

[0013] FIG. 5 depicts an example method for virtual medical session interactions with an non-powered object.DETAILED DESCRIPTION

[0014] Following below are more detailed descriptions of various concepts related to, and implementations of, systems, methods, apparatuses for virtual medical session interactions with non-powered objects. The various concepts introduced above and discussed in greater detail below can be implemented in any of numerous ways.

[0015] Although the present disclosure is discussed in the context of a surgical procedure, in various aspects, the technical solutions of this disclosure can be applicable to other medical treatments, sessions, environments or activities, as well as non-medical activities where extended reality functionality is utilized in robotics. For instance, technical solutions can be applied in any environment, application or industry in which activities, operations, processes or acts are performed with tools or instruments that can be captured on video and for which ML modeling can be used deliver XR reality content for a virtual component of a robotic medical system based on user interactions with an object representing the component.

[0016] This technology is generally directed to using non-powered objects and mechanisms to perform virtual interactions using a virtual portal (e.g., a grounded virtual portal, or a freeform virtual reality head mounted display). The technology allows for quickly configuring a virtual portal to adapt to different use-cases in software with minimal hardware modifications.

[0017] For example, the controllers used on a surgeon console can be heavily instrumented with various electronic sensors that are attached to a physical frame. These sensors can provide haptic / force feedback. However, in a system based on a free-form VR HMD, the users can move their hands in free space, which means that any feedback is either visual or light vibrations on the controllers, which can lead to ergonomic discomfort.

[0018] This technology can use sensor data, such data of 2D / 3D sensors to track users’ hand poses and positions within space, as well as the users’ interactions with a physical linkage. Since the sensor data can visually track the pose and interactions with the physical linkage, these technical solutions can eliminate the need for sensors on the physical linkages themselves. By eliminating the need for sensors on the physical linkages, the technical solutions can use software to allow for various types of interactions with physical linkages.

[0019] FIG. 1 depicts an example system for virtual medical session interactions with an non-powered object. For example, system 100 can deliver extended reality content for a virtual component of a robotic medical system, based on user interactions with an object representing the component. Example system 100 can include a combination of hardware and software for providing non-powered objects and mechanisms for larger virtual interactions with physical constraints. The system 100 can provide any XR interactions, including virtual reality (VR) or augmented reality (AR) interactions using non-powered objects to represent components associated with a robotic medical system.

[0020] Example system 100 can include a medical environment 102, including one or more RMSs 120 communicatively coupled with one or more head mounted devices (HMDs) 122 and one or more data processing systems (DPSs) 130 via one or more networks 101. Medical environment 102 can include one or more sensors 104, data capture devices 110, medical instruments 112, visualization tools 114, displays 116 and robotic medical systems (RMS) 120, which can also be referred to collectively as components 190. Components 190 can also include an object 180, which can also be deployed in a medical environment 102 or outside of it.

[0021] Data processing system (DPS) 130 can be communicatively coupled with an RMS 120, via a network 101. DPS 130 can generate, provide and display virtualized interactions (e.g., interactions provided in the XR environment or domain) involving the non-powered object 180 used to represent any of the components 190 or their operations in the medical environment 102. DPS 130 can include one or more virtualizer applications 132 display virtualized version of a component 190 (e.g., as a virtual component 138) based on user interactions with an object 180. The DPS 130 can include one or more interaction determiners 160 to monitor, detect or determine user’s interactions with the object 180 representing the component 190. The DPS 130 can include one or more virtuality transformers 150 to transform user interactions with the object 180 into object representations 152 (e.g., poses or configurations of the object). The DPS 130 can include one or more machine learning (ML)environments 140 for providing machine learning aspects of the solution including one or more model selectors 142 for selecting one or more ML models 144. ML environment can include one or more weighting functions 146 to provide weighting to different ML models 144 or data 172. Virtualizer application 132 can include one or more of settings 134, object information 136 or virtual components 138. Virtuality transformer 150 can include one or more object representations 152 and interactive elements 154. The DPS 130 can include one or more data repositories 170 that can include, store and provide data 172, such as the data collected from the sensors 104 from the medical environment 102. The DPS 130 can include one or more interfaces 174.

[0022] System 100 can include a head-mounted device (HMD) 122 that can be communicatively coupled with a DPS 130 or an RMS 120 and allow a user to utilize the virtualization or XR related features of the system 100. Head-mounted device (HMD) 122 can include one or more displays 116 and sensors 104. HMD 122 can include one or more eye trackers 124, hand trackers 126 and voice controllers 128.

[0023] Robotic medical system 120, also referred to as an RMS 120 or a computer-assisted robotic system 120, can be deployed in any medical environment 102. Medical environment 102 can include any space or facility for performing medical procedures, maneuvers or tasks in medical sessions, including for example a surgical facility, or an operating room. Medical environment 102 can include medical instruments 112 (e.g., surgical tools used for various surgical maneuvers or tasks) which the RMS 120 can facilitate or utilize for performing surgical patient procedures, whether invasive, non-invasive, or any in-patient or out-patient procedures. Robotic medical system 120 can be centralized or distributed across a plurality of components, computing devices or systems, such as computing system 400 (e.g., used on servers, network devices or cloud computing products) to implement various functionalities of the RMS 120, including network communication or processing of data streams 162 across various devices over the network 101.

[0024] The medical environment 102 can include one or more data capture devices 110 (e.g., optical devices, such as cameras) as well as sensors 104 (e.g., detectors or sensing devices) for making measurements and capturing data 172. Data 172 can include any sensor data, such as images or videos of a surgery, measurements of distances (e.g., depth), temperature, stress (e.g., pressure of vibration), light, motion, humidity, motion, velocity, acceleration, force or material (e.g., gas) concentration. Data 172 can include kinematics data on any movement of medical instruments 112, users (e.g., medical staff), or any devices in themedical environment 102. Data 172 can include any events data, such as information on instances of a failure of a medical instrument 112, a collision involving a medical instrument or a collision involving an anatomy of a patient, as well as occurrences of installation, configuration or selection of any devices in the medical environment 102.

[0025] The medical environment 102 can include one or more visualization tools 114 to gather the captured data 172 and process the data for display to the user (e.g., a surgeon, a medical professional or a technician utilizing the system 100) via one or more (e.g., touchscreen) displays 116 or displays of an HMD 122. A display 116 can present or display data 172 or information generated based on data 172 (e.g., outputs of virtualizer application 132) during the course of a medical procedure (e.g., a surgery), or AR, VR (e.g., XR) session in which a user utilizes an object 180 to simulate operation of one or more components 190 of the medical environment.

[0026] System 100 can include any number of sensors 104 dispersed throughout the medical environment 102. Sensors 104 can include electronic or electrical components, a combination of electronic and software components, mechanical or electromechanical components, or any combination thereof. Various sensors 104 can include devices, systems or components detecting, measuring and / or monitoring a variety of signals in a variety of applications, such as temperature, pressure, location, proximity, light, humidity, motion, acceleration, velocity, distance, depth, infrared (IR), magnetic field, electric field, pressure or movement of gas, images or video stream (e.g., from a camera acting as a sensor 104), sounds (e.g., microphone), force, touch, moisture, radiation, pH, humidity or vital signs of a person. For instance, a sensor 104 can include a global positioning system sensor or transceiver for location, one or more image sensors for capturing images, one or more accelerometers, one or more gyroscopes, one or more magnetometers, or another suitable form of sensor that detects motion and / or location. Sensor 104 can include a camera capturing images or video of components 190, or a device measuring a distance between the sensor 104 and object 180 or providing data from which such distance can be determined by the DPS 130 (e.g., virtuality transformer 150). Sensors 104 can include one or more gyroscopes that can detect rotational movement (e.g., pitch, yaw, roll) and one or more accelerometers can measure translational movement (e.g., forward / back, up / down, left / right). Sensor 104 can detect, determine, measure or quantify a motion, such of a physical object, such as a hand gesture a user gesture, the contour of the hand, a movement of an eye (e.g., location of an iris), a user interaction, a voice command or any other user action.

[0027] Data processing system 130 can include any combination of hardware and software for providing XR (e.g., AR or VR) representations of virtualized interactions in a medical environment 102. DPS 130 can include the functionality for providing virtualized representations of interactions with any components 190 using one or more objects 180. DPS 130 can include any computing device (e.g., computing system 400) and can include one or more servers, virtual machines or can be part of, or include a cloud computing environment. The data processing system 130 can be provided via a centralized computing device (e.g., computer system 400), or can be provided via distributed computing components, such as including multiple, logically grouped servers and facilitating distributed computing techniques. DPS 130 can be provided or included within or using, any portion (e.g., computer system 400) of an RMS 120 or an HMD 122 or any other device. DPS 130 can be provided on a logical group of servers, such as a data center, a server farm or a machine farm, which can include virtual machines, which can be geographically distributed or dispersed. DPS 130 can be provided as a single entity or a single platform.

[0028] DPS 130 can be operatively coupled, or associated with, the medical environment 102, directly or via a network 101. The network 101 can be any type or form of network, such as a body area network (BAN), a personal area network (PAN), a local-area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), or the Internet. Network 101 can include wireless communication links between an HMD 122 and any combination of devices or components of a medical environment 102 or any portion of a DPS 130. The network 101 can be a type of a broadcast network, a telecommunications network, a data communication network, a computer network, Wireless Fidelity (Wi-Fi) network, a Bluetooth network, a cellular network (e.g., 4G, 5G or 6G network), or other types of wired or wireless networks.

[0029] The DPS 130, or components thereof, can be located at least partially at the location of the surgical facility associated with the medical environment 102 or remotely therefrom. Elements of the DPS 130, can be accessible via portable devices such as laptops, mobile devices, wearable smart devices or HMD 122. The DPS 130, or components thereof, can include other or additional elements that can be considered desirable to have in performing the functions described herein. The DPS 130, or components thereof, can include, or be associated with, one or more components or functionality of a computing including, for example, one or more processors coupled with memory that can store instructions, data or commands for implementing the functionalities of the DPS 130 discussed herein.

[0030] Data repository 170 of the DPS 130 can include any memory or a storage device for storing data 172, or any other information (e.g., instructions, computer code, parameters or functionalities for implementing DPS 130 functionality). Data repository 170 can include or be implemented in a main memory 415, ROM 420 or a storage device 425. Data repository 170 can be communicatively coupled with a processor 410 and can include or store data files, data structures, arrays, values, or other information that facilitates operation of the DPS 130. The data repository 170 can include one or more local or distributed databases and can include a database management system. The data repository 170 can include, maintain, or manage one or more data streams of data 172, including data collected by one or more data capture devices 110, such as a set of 3D sensors from a variety of angles or vantage points with respect to an object 180.

[0031] Data 172 can include any data stream (e.g., a series of data packets of a particular type or a form) that can be generated by a particular device (e.g., a sensor 104 or a data capture device 110). Data 172 can include any plurality of data packets (e.g., one or more streams) of data from any number sensors 104. Data 172 can include information or data of still-image or video frames, image or video frames of a camera device on a medical instrument 112, such as an endoscopic device. Data 172 can include sensors 104, including timestamped data corresponding to force, torque or biometric data, haptic feedback data, endoscopic images or data, ultrasound images or videos and any other sensor data. Data 172 can include a stream of kinematics data, including any data indicative of temporal positional coordinates of a device, or indicative of movement of a medical instrument 112 on an RMS 120. Data 172 can include a stream of data packets corresponding to events, such as events indicative of, or corresponding to, errors in operation or faults, collisions of a medical instrument 112, occurrences of installation, uninstallation, engagement or disengagement, setting or unsetting of any medical instrument 112 on an RMS 120.

[0032] The system 100 can include one or more data capture devices 110 capturing data 172. Data capture devices 110 can include medical imaging devices (e.g., magnetic resonance imaging, computed tomography, ultrasound machines or fluoroscopy devices). Data capture devices 110 can include endoscopic cameras from laparoscopes or arthroscopes. Data capture devices 110 can include optical tracking systems, such as stereoscopic cameras or infrared cameras, electromagnetic tracking systems, microscope systems, force or tactile feedback devices, biometric data capture devices (e.g., electrocardiogram or electroencephalogram). Data capture devices 110 can include any sensors 104, including video cameras for collectingany data 172 that can be used for machine learning, including detection of objects 180 and determining or virtualizing interactions involving the objects 180. Data capture devices 110 can include cameras or other image capture devices for capturing videos or images from a particular viewpoint within the medical environment 102. The data capture devices 110 can be positioned, mounted, or otherwise located to capture content from any viewpoint that facilitates the data processing system capturing various surgical tasks or actions. Data capture devices 110 can be used to detect or recognize user gestures or actions, and determine distances to, from and between one or more objects 180 of a medical environment 102.

[0033] Data capture devices 110 can include any of a variety of sensors 104, such as detectors, cameras, video imaging devices, infrared imaging devices, visible light imaging devices, intensity imaging devices, including black, color and grayscale imaging devices, depth imaging devices, (e.g., stereoscopic imaging devices, time-of-flight imaging devices, etc.), medical imaging devices such as endoscopic imaging devices, ultrasound imaging devices, etc., non-visible light imaging devices, any combination or sub-combination of the above mentioned imaging devices, or any other type of imaging devices that can be suitable for the purposes described herein. Data capture devices 110 can include cameras that a surgeon can use to perform a surgery and observe manipulation components within a purview of field of view suitable or selected for the given task performance. Data capture devices 110 can include depth cameras, color image cameras, stereo cameras (e.g., using two or more lenses for capturing 3D information), infrared (IR) cameras, structured light sensor, light detection and ranging (LIDAR) sensors or ultrasonic sensors.

[0034] For instance, data capture devices 110 can capture, detect, or acquire sensor data 172, such as videos or images, including for example, still images, video images, vector images, bitmap images, other types of images, or combinations thereof. The data capture devices 110 can capture the images at any suitable predetermined capture rate or frequency. Settings, such as zoom settings or resolution, of each of the data capture devices 110 can vary as desired to capture suitable images from any viewpoint. For instance, data capture devices 110 can have fixed viewpoints, locations, positions, or orientations. The data capture devices 110 can be portable, or otherwise configured to change orientation or telescope in various directions. The data capture devices 110 can be part of a multi-sensor architecture including multiple sensors, with each sensor being configured to detect, measure, or otherwise capture a particular parameter. Data capture devices 110 can generate sensor data 172 from any type and form of a sensor, such as a positioning sensor, a biometric sensor, a velocity sensor, anacceleration sensor, a vibration sensor, a motion sensor, a pressure sensor, a light sensor, a distance sensor, a current sensor, a focus sensor, time-of-flight sensor, optical flow sensor, proximity sensor, a temperature or pressure sensor or any other type and form of sensor used for providing data on medical instruments or tools 112, or data capture devices (e.g., optical devices).

[0035] Display 116 can show, illustrate or play data 172, such as a video stream, in which medical tools 112 at or near surgical sites are shown. For example, display 116 can display a rectangular image of a surgical site along with at least a portion of medical tools 112 (e.g., instruments) being used to perform surgical tasks. Display 116 can provide compiled or composite images generated by the visualization tool 114 from a plurality of data capture devices 110 to provide visual feedback from one or more points of view. Display 116 can include a display used an HMD 122, for delivering visual content in live view mode (e.g., a video camera view of a medical environment 102). Display 116 can include be used for an AR mode in which augmented reality objects are inserted and displayed alongside video camera captured components 190 of a portion of the medical environment 102. Display 116 can be used for a VR mode in which virtualized components 138 or objects are displayed in a virtual representation of a portion of the medical environment 102.

[0036] Display 116 can be configured to be attached to computer-assisted RMS 120 and to display or present any virtual components 138 according to their object representations 152 (e.g., at least one of the virtual pose or the virtual configuration). Display 116 can display or present an interface 174 which can be or include a user interface of a virtualizer application 132 presenting AR or VR simulation of a medical environment along with the virtual component 138 (e.g., component 190 represented by the object 180). Display 116 can display or present the virtual component according to the at least one of the second virtual pose or the second virtual configuration. The RMS 120 and the display 116 can be physically grounded to a reference platform, such as a reference floor, or each other. For instance, the display 116 can be physically attached to the RMS 120 or a component associated with the RMS 120. The display 116 can update or present updated virtual components 138, interactive elements 154 or object representations 152 in the interface 174. The display 116 can present or display user interactions determined or implemented using an interactive element 154 that comprises at least one of a mechanical button, a foot pedal, or a hot spot on the physical object. The display 116 can be included within and coupled to the HMD 122 and provide interface 174 including an AR or VR representation of the illustrated or simulated environment along with any objectinformation 136, virtual components 138, interactive elements 154 or object representations 152 and illustrate changes of interactions based on user interactions with the object.

[0037] The visualization tool 114 that can be configured or designed to receive any number of different data streams from any number of data capture devices 110 and combine them into frames displayed on a display 116. The visualization tool 114 can be configured to receive data stream components and combine the plurality of data stream components into a single data stream. The data stream can be a data stream for displaying any XR content, such as augmented reality features (e.g., virtual components 138) in a camera view of a medical environment. The visualization tool 114 can receive a visual sensor data from one or more medical tools 112, sensors or cameras with respect to a surgical site or an area in which a surgery is performed. The visualization tool 114 can incorporate, combine or utilize multiple types of data (e.g., positioning data of a medical tool 112 along sensor readings of pressure, temperature, images vibration or any other data) to generate an output to present on a display 116. Visualization tool 114 can present locations of medical tools 112 along with locations of any reference points or surgical sites, including locations of anatomical parts of the patient (e.g., organs, glands or bones).

[0038] Medical tools 112 can be any type and form of tool or instrument used for surgery, medical procedures or a tool in an operating room or environment. Medical tool 112 can be imaged by, associated with or include an image capture device. For instance, a medical tool 112 can be a tool for making incisions, a tool for suturing a wound, an endoscope for visualizing organs or tissues, an imaging device, a needle and a thread for stitching a wound, a surgical scalpel, forceps, scissors, retractors, graspers, or any other tool or instrument to be used during a surgery. Medical tools 112 can include hemostats, trocars, surgical drills, suction devices or any instruments for use during a surgery. The medical tool 112 can include other or additional types of therapeutic or diagnostic medical imaging implements. The medical tool 112 can be configured to be installed in, coupled with, or manipulated by an RMS 120, such as by manipulator arms or other components for holding, using and manipulating the medical instruments or tools 112.

[0039] RMS 120 can be a computer-assisted system configured to perform a surgical or medical procedure or activity on a patient via or using or with the assistance of one or more robotic components or medical tools 112. RMS 120 can include and utilize computer- controlled robotic arms, high-resolution imaging devices, and real-time feedback sensors 104, all of which can be utilized to execute precise surgical movements using medical instruments112. RMS 120 can be controlled or used by a user (e.g., a surgeon) to implement tasks, such as using medical instruments 112 to make incisions, suture wounds, remove tissues and insert medical components or systems. RMS 120 can include any number of manipulator arms for grasping, holding or manipulating various medical tools 112 and performing computer-assisted medical tasks using medical tools 112 controlled by the manipulator arms.

[0040] HMD 122 can include any number of sensors 104 to track head movements or eye trackers 124, to facilitate the virtual or augmented content to respond dynamically to the user's perspective as well as receive user inputs or selections. HMD 122 can be utilized by users, such as medical professionals, to participate in medical procedures remotely. HMD 122 can be worn on a user’s head or can be fixed to a specific location, providing stability, and display extended reality (XR) that includes augmented, virtual, live-view or mixed realities, depicting various features of the medical environment 102. HMD 122 can include sensors 104 providing information about the user's location, orientation, and gaze direction, allowing the HMD 122 to generate corresponding views and images.

[0041] HMD 122 can include sensors 104, a communication interface, and a visualization tool 114 to detect its location, orientation, and the user's gaze direction. Using this information, the HMD 122 can render images representing the current state of the medical environment 102 and integrate real-time data from different capture devices 110. HMD 122 can include the functionality to deliver extended or mixed reality content, including by rendering images, videos, or audios based on user interactions. HMD 122 can be included in a medical environment 102 and can be used to simulate or virtualize an object 180 or interactions with the objects that can be rendered or displayed.

[0042] Eye tracker 124 can include any combination of hardware and software for measuring, detecting, monitors and analyzing a user's gaze, including eye movements and locations. Eye tracker 124 can include the functionality to use the user’s gaze (e.g., movements and locations of the eyes) to determine various user commands or actions, such as selections of features displayed. Eye tracker 124 can utilize sensor data (e.g., camera images) or infrared sensors to track the user's eye movements in real-time, or gestures indicated by eye movements or movement of any other part of the user’s body, allowing for accurate detection of where the user is looking with respect to the features displayed in the extended reality on the display 116. Eye tracker 124 can interpret the user’s visual focus and trigger actions such as selecting options, navigating through menus, activating controls, or even initiating specific commands without the need for physical input devices.

[0043] Hand tracker 126 can include any combination of hardware and software that uses data from various sensors 104 to monitor and analyze the movements of any part of a user’s body, such as users’ hands, arms, head movement or any other portion of the user’s body. Hand tracker 126 can include or utilize computer vision techniques (e.g., using a hand tracking ML model 144 trained to track user’s wrist and fingers) to use images or videos of user’s hands to detect or discern various gestures that the user makes. Hand tracker 126 can track movements of user’s hands, arms, head, torso, legs or any other portion of body of the user. Body movements can be tracked individually (e.g., movements of hands, arms, eyes) or collectively (e.g., body of a plurality of body parts moving together). Hand tracker 126 can operate in real time and allow the user to seamlessly interact with DPS 130. For instance, the user can use the hand tracker 126 to (e.g., via sensors 104) to control medical tools 112 in the medical environment 102. For instance, the user can user the hand tracker 126 to move, control or manipulate the object 180 (e.g., within medical environment 102 or outside of the medical environment 102). Hand tracker 126 can detect gestures, such as finger movements for scrolling or selecting items, hand waving for navigation or moving of objects, pinching or expanding fingers for zooming in or out, and making specific hand movements or shapes to trigger specific commands or actions. Hand tracker 126 can detect motions or movements of the object 180, selection or interaction with interactive elements 154 or markers 182.

[0044] Voice controller 128 can include any combination of hardware and software that uses voice for controlling devices. Voice controller 128 can include a microphone to detect a user’s verbal command or an instruction. Voice controller 128 can detect statements, such as code words, to trigger a command for the user to control RMS 120 via HMD 122. For example, a voice controller 128 can facilitate the user to provide a verbal instruction to an HMD 122 to implement an action, such as a switch between different modes of operation (e.g., live-view video or an XR mode of operation).

[0045] Object 180 can include any physical object which can be used by the DPS 130 to generate virtual interactions involving virtual components 138 representing components 190 in the medical environment 102. Object 180 can include a non-powered physical object that can be physically handled, moved or manipulated physically by a user, or can be gestured towards, touched or otherwise interacted with in physical or a virtual sense. Object 180 can include an object that is sized and shaped or otherwise configured to be handled by the hand of a user, such as an elongate piece of material, a component or a structure. Object 180 can include, for instance, a rod (e.g., a cylindrical rod), a stick, a pencil or a pen, a handle, or any other objectthat can be handled or moved by a user’s hand. Object 180 can include any non-powered object a user can physically touch, press or move.

[0046] A non-powered object can refer to or include an object or component that may not have the capability to be powered (e.g., via electricity). The non-powered object may lack or not include wireless communication interfaces or electronic sensors. For example, the nonpowered object may lack WIFI, Bluetooth, or nearfield communication interfaces. The nonpowered object may lack or not include light sources or otherwise generate or emit electromagnetic waves. The non-powered object may lack transducers.

[0047] Object 180 can also include objects in the medical environment 102 in which an RMS 120 is deployed. For instance, object 180 can include any component or device coupled with or included in an RMS 120, including any sensors 104, data capture devices 110, medical instruments 112, visualization tool 114 or display 116. Object 180 can include components or portions of a patient table, a medical instrument 112 or a medical device or a system. Object 180 can include a portion of a wall of an operating room which the user can utilize for interaction or gesturing, a picture or a window, a frame or an artifact of a wall or an area.

[0048] Object 180 can include markers 182. A marker 182 can include any one or more features, components or shapes acting as a mark or identifier of the object 180. Marker 182 can include one or more reflective components, such as retroreflective stickers or features for providing or reflecting visible, UV or IR light to a data capture device 110. Marker 182 can include a mark of any shape, such as a triangle, a rectangle, a square, a star, a circle or any other shape. Marker 182 can include an arrangement of a plurality of marks, such as a line, a pattern or a particular order of marks (e.g., retroreflective components) that can be used to identify the object 180. For instance, the physical object 180 can include a marker 182 with a pattern, which can be used by interaction determiner 160, virtuality transformer 150 or model selector 142 for making its determinations (e.g., using ML models 144). The marker 182 can include a machine-readable code, such as a QR code or barcode.

[0049] Object 180 can include interactive elements 154. An interactive element 154 can include any feature, shape, area or an element of an object 180 that can be used to indicate, trigger, use or utilize virtualized interactions. For example, interactive element 154 can include dummy buttons, hot spots (e.g., areas on an object 180), levers or handles or foot pedals that can be indicative of particular actions or activities of a virtualized component 190 (e.g., virtual component 138 represented in a simulated AR or VR environment in a virtualizer application132). Interactive elements 154 can include for example a hot spot on the object 180 that can be defined based on a relative location of the area identified as a hotspot with respect to markers 182 on the object 180.

[0050] Interface 174 can include any interface for interacting with a user of a DPS 130 or any of its components. Interface 174 can include a user interface that can receive indications or inputs from users, such as user selections, input values or parameters. Interface 174 can receive indications via interaction determiner 160 which can infer, determine, detect or identify user selections based on user gestures (e.g., from a hand tracker 126) or eye selections (e.g., based on eye tracker 124). Interface 174 can include a graphical user interface (GUI) that can be displayed on a display 116 (e.g., such as on display of a medical environment 102 or a display of an HMD 122).

[0051] Virtualizer application 132 can include any application for virtualizing medical procedures. Virtualizer application 132 can include or display in a VR or AR mode of operation any feature or component 190 (e.g., entire medical environment with the RMS 120 and the user interacting with the RMS 120, shapes and locations of medical instruments 112, sensors 104, data capture devices 110, visualization tools 114 and displays 116). Virtualizer application 132 can include any application executed on a DPS 130 for virtualizing medical procedures or sessions implemented using RMS 120. Virtualizer application 132 can be fully or partially executed or have its output provided on an HMD 122 or a display 116. Virtualizer application 132 can receive indications or inputs via an interface 174, such as user selections that can be identified, detected or determined based on hand gestures (e.g., detected by a hand tracker 126) or eye gestures (e.g., eye movements, focus or gaze direction) detected by an eye tracker 124. Virtualizer application 132 can be configured or set based on one or more settings 134 for configuring or setting the application, its functionality or its operation. Virtualizer application 132 can be configured or setup to provide training for a user that is tailored to different phases of medical sessions, such as for example an operating room setup, a preoperative setup, an intraoperative medical procedure or a post-operative clean-up.

[0052] The data processing system 130 can include a virtualizer application 132 designed, constructed and operational to identify a setting 134 for an application (e.g., virtualizer application 132) to virtualize at least a portion of a medical session (e.g., pre-operation, intraoperative, or post-operative). For example, the portion of the medical session can correspond to a portion in which a computer-assisted robotic system 120 is used to perform an aspect of the medical session or medical procedure. To identify the setting 134, the virtualizer application132 can receive an input from an interface indicating a type of mode of operation for the virtualizer application 132. The virtualizer application 132 can identify the setting 134 from a profile of a user that can identify the type of functionality or mode of operation of the virtualizer application 132. The virtualizer application 132 can identify the setting 134 automatically based on a condition (e.g., type of object 180 identified using one or more ML models 144, such as for example, in response to detecting that the object 180 is a pen, the virtualizer application 132 can automatically select a training mode.

[0053] Virtualizer application 132 can include the functionality to identify a setting 134 for the virtualizer application 132 so as to configure the application to virtualize at least a portion of a medical session in which a computer-assisted RMS 120 is utilized. Settings 134 can include any setting for configuring or adjusting the operation of the virtualizer application 132. Setting 134 can identify, configure or define a type of the virtualizer application 132 being deployed or utilized. Settings 134 can identify, configure or define the virtualizer application 132 as a training simulation application (e.g., an application for implementing a simulation of a remote robotic surgery). In some cases, the training application or simulation can be provided as a grounded mechanism where any one or more of: an object 180, display 116 or RMS 120 can be grounded or attached to a grounding reference point (e.g., each other or a platform). In some cases, the training application or simulation can be provided as an ungrounded mechanism where the object 180 or display 116 may not be attached or otherwise fixedly coupled to a grounding reference point (e.g., each other, a platform, or the RMS 120).

[0054] Grounded can refer to the object being mechanically coupled to a physical reference point such that an amount of movement of the object 180 relative to the physical reference point is constrained, fixed, or limited. For example, the object 180 can be mechanically or fixedly attached to the grounded reference point such that the position of the object 180 cannot change, but the orientation (e.g., yaw, pitch and roll) of the object can change. In some cases, the object 180 can be grounded such that the position can change a predetermined amount (e.g., plus or minus 1 inch, 2 inches, 3 inches, 4 inches, or some other amount) relative to the grounded reference point.

[0055] Ungrounded can refer to the object 180 lacking any physical or mechanical coupling to a reference point. In some cases, ungrounded can refer to the object 180 lacking a rigid, mechanical coupling to the reference point, while being flexibly or elastically coupled to the reference point (e.g., tethered via a flexible rope, cable, or wire).

[0056] In some cases, the same object 180 can be used in a grounded or ungrounded manner depending on the type of virtual application 132 or experience provided by the virtual application 132. For example, a user with a first skill level can use the object 180 in a grounded virtual experience provided by a training virtual application 132, while a second user with a second skill level can use the object 180 in an ungrounded virtual experience provided by the training virtual application 132. The second skill level may be greater than the first skill level, thereby allowing the second user to challenge their skills or hone new skills. In some cases, the virtual application 132 can provided a grounded or ungrounded virtual experience with object 180 in accordance with a hospital site, characteristic of a medical environment or operating room, type of medical procedure, user preference, or type of object 180. For example, if the virtual application 132 is providing a virtual experience that includes simulating a surgical session for a surgeon, then the virtual application 132 can provide (or determine to provide or be requested to provide) a grounded virtual experience with a grounded object 180. If, for example, the virtual application 132 is providing a virtual experience that includes simulating preparation of the operating room (e.g., preparing the robotic medical system) by a staff member for the operating room, then the virtual application 132 can provide (or determine to provide or be requested to provide) an ungrounded virtual experience with an ungrounded object 180.

[0057] Settings 134 can identify, configure or define the virtualizer application 132 as a telepresence and telerobotic systems for facilitating remote robotic procedures (e.g., application for implementing robotic surgery). Settings 134 can identify, configure or define a virtual or augmented reality interactive or educational application for learning about particular surgical aspects or features. Settings 134 can include a setting file that can define various settings or parameters, including settings for a robotic medical system, use cases, phases in a medical session, or types of medical procedures. Setting file can be customized based on user preferences.

[0058] Virtualizer application 132 can receive, detect, identify or utilize any object information 136. Object information 136 can include any information or data on an object 180. Object information can include information indicative of an object 180, such as an indication. For example, virtualizer application 132 can receive a user selection, input or a setting (e.g., via a user interface 174) as an indication of the object 180. For instance, virtualizer application 132 can receiver or identify an indication of a physical object 180. The indication can identify the physical object 180 that represents a component associated with the computer-assisted RMS120. Object information 136 (e.g., indication of the object 180) can be determined or detected using any combination of virtuality transformer 150 determining virtual components 138 or object representations 152 using one or more ML models 144.

[0059] Virtualizer application 132 can receive object information 136 about multiple objects 180 based on the data 172 or can determine or identify objects 180 based on object information 136 about multiple objects 180. For instance, the sensor data 172 can be input into one or more ML models 144 (e.g., via interaction determiner 160 or virtuality transformer 150). The sensor data 172 can include information (e.g., sensor measurements or camera images from multiple angles) about one or more objects 180, based on markers 182 (e.g., individual markers or marker arrangements) or interactive elements 154. The interaction determiner or the virtuality transformer 150 can determine, detect or identify the one or more objects 180 from the data and provide indication of the identified one or more objects 180 to the virtualizer application.

[0060] For instance, virtualizer application 132 can detect that a second physical object 180 is swapped in for an earlier used physical object 180. The second physical object 180 can include a second one or more markers 182 that are different than a first one or more markers 182 of the physical object. For example, virtualizer application 132 can utilize one or more ML models 144 to detect the type of the object 180 utilized based on the unique markers 182 of each of the objects 180, allowing the DPS 130 to distinguish between the objects 180. Virtualizer application 132 can determine a type of the physical object based on the pattern of the one or more markers 182 detected via the data 172 from the one or more optical sensors 104 or data capture devices 110.

[0061] Virtualizer application 132 can display or present one or more virtual components 138. Virtual component 138 can include any virtual representation of a component 190 that an object 180 represents. For instance, virtual component 138 can include a virtual version of a medical instrument 112 that can be represented by a physical object 180 being manipulated, handled or moved by a user and illustrated in the virtualizer application 132. For instance, virtual component 138 can include a virtual version of an arm manipulator of an RMS 120 that can be represented by a physical object 180 being manipulated, handled or moved by a user and illustrated in the virtualizer application 132. For instance, virtual component 138 can include a virtual version of a sensor 104, a data capture device 110 or a medical system or a component used to assist with the surgery, which can be represented by a physical object 180 being manipulated, handled or moved by a user and illustrated in the virtualizer application 132.

[0062] Interaction determiner 160 can include any combination of hardware and software for detecting, recognizing or determining objects 180 and user interactions with the objects 180. Interaction determiner 160 can include functionality for utilizing ML models 144 trained on data 172 to detect, identify, determine or predict objects 180 and interactions of the user with the one or more objects 180. Interaction determiner 160 can identify or detect objects based on data 172 on markers 182 or interactive elements 154 of objects 180. For instance, interaction determiner 160 can use ML models 144 to detect the number of markers 182, the type of markers 182 (e.g., shapes and sizes) or the arrangement of markers 182 (e.g., relative orientation of markers with respect to a reference point or each other) and based on this detection, identify or determine the object 180.

[0063] Interaction determiner 160 can utilize data 172 to detect user interactions (e.g., user handling) of the object 180. For example, interaction determiner 160 can determine user interactions based on the determined positioning of the user’s body parts (e.g., fingers, hands, joints, arms, legs, torso, head or eyers) with respect to the object 180. Interaction determiner 160 can determine the user interactions based on the determined motion or movement of the user from a timestamped series of data 172 (e.g., one or more data streams) where user’s body or body parts positioning is determined in relation to the object 180 over time.

[0064] Interaction determiner 160 can be configured such that the interaction determiner 160 detects or determines the user interaction with the physical object as the interactions changes the object representation 152. The object representations 152 can include any virtualized or simulated representations of the object 180, as well as any representation of the positioning, pose, location or action applied to the object 180. For example, object representations 152 can include at least one of a pose of the physical object or a configuration of the physical object. The pose of the physical object 180 (e.g., object representation 152) can include at least one of a position of the physical object relative to a reference frame, or an orientation of the physical object relative to the reference frame.

[0065] Interaction determiner 160 can determine or detect object representations 152, such as a configuration of the physical object 180 which can correspond to a state of activation of an interactive element 154. The interactive element 154 can include, for example, one or more of a mechanical button associated with the physical object 180 (e.g., a spring-loaded button on the object 180), a foot pedal associated with the physical object 180 (e.g., a foot pedal for the object 180 or a RMS 120), or a hot spot associated with the physical object 180 (e.g., an area or a region on the object 180 for user interaction or signaling of an action). For instance, a userhandling the object 180 can activate any of the interactive elements (e.g., press on a dummy foot pedal, hold a dummy button on the object or touch a hot spot on the object) and the interaction determiner 160 can determine a particular action (e.g., mouse click or a user selection of a feature or component 190 displayed in the virtualizer application 132).

[0066] Interaction determiner 160 can determine a series of user interactions with the object 180 and provide virtualized object representations 152 and any interactions of the user with interactive elements 154 on the object 180 through time. For instance, interaction determiner 160 can use one or more ML models 144 to determine, based on second data 172 (e.g., updated sensor data) collected from the one or more optical or other sensors, a second user interaction with the physical object 180. The second user interaction can be subsequent to the first user interaction. The second user interaction can be followed by a third user interaction following the second user interaction which the interaction determiner 160 can determine using the one or more ML models 144. The series of user interactions can indicate a particular action or task being performed, such as a task of a medical procedure or a phase of the medical procedure. The interaction determiner 160 can utilize the one or more ML models to detect or determine the action, the task or the phase or a multi-phase medical procedure.

[0067] The interaction determiner 160 can update the object representation 152 based on subsequent detected user interactions. For instance, the interaction determiner 160 can change, based on the determined second or third user interaction with the physical object 180, any object representation 152, such as least one of the virtual pose or the virtual configuration for the virtual component to at least one of a second virtual pose or a second virtual configuration. The interaction determiner 160 can use one or more ML models 144 to determine a gesture of a hand of a user interacting with the physical object 180. The interaction determiner 160 can determine a type of an interactive element 154 of the physical object 180 (e.g., a button-type interactive element, a pedal-type interactive element, a hot spot type interactive element or any other type of interactive element 154). The interaction determiner 160 can determine or detect, using the one or more models 144 and based on the data from the one or more optical sensors (e.g., data 172), the user interaction with the interactive element 154 of the physical object, such as pressing or selection of a dummy button (e.g., an non-powered button on an nonpowered object 180), a foot pedal or a selection of a hot spot on the object.

[0068] The interaction determiner 160 can determine that a portion of the physical object 180 is occluded. The interaction determiner 160 can identify the portion of the physical object 180 that is occluded from the data 172 from the one or more optical sensors, such as to takeaction and use data 172 from other sensors 104 from which the object is not occluded. For instance, in response to determining that a particular camera (e.g., 110 or 104) has an obstructed or incomplete view of a portion of an object 180, the interaction determiner 160 can use data 172 from another source (e.g., sensor 104 or data capture device 110), or can more heavily rely on data 172 from the other source from which the object is not occluded (e.g., using weighting function 146).

[0069] Weighting function 146 can include any combination of hardware and software, such as computer code function, for determining weights of particular ML models 144 or data 172 used for making determinations. Weighting function 146 can include functionality for making determination which data 172 stream or ML model 144 is to be weighted more or less for making a determination, detection or providing output. Weighting function 146 can, for example, weigh data stream of data 172 in which an object 180 is not occluded (e.g., from the field of view of a camera) more than data 172 in which the object is occluded. Weighting function 146 can determine one or more ratios for allocating the weight to two or more sources of data 172 or ML models 144 used for making a determination, such as a determination of an object 180 (e.g., virtual component 138, object information 136, object representations 152, interactive elements 154, user interactions by interaction determiner 160, or any other determination that can be made using ML models 144).

[0070] Object representations 152 can include any representations of a configuration or a pose of an object 180 to be virtualized. Object representations 152 can be positional, movement or actional configurations or poses of the object 180 determined based on user interactions with the object 180 determined by the interaction determiner 160. Object representation 152 can include a pose or a configuration of an object 180, such a virtualization of the object (e.g., virtual component 138) that is tilted, shifted, moved, positioned towards a particular direction or in a particular position. Object representation 152 can include the virtualized object (e.g., virtual component 138) that is inverted, bent or slid side to side or moved up or down. Object representation 152 can include the virtualized object (e.g., virtual component 138) that is activated, selected, turned on, deactivated, unselected or turned off, such as by selecting a virtual c spot o. Object representation 152 can include the virtualized object (e.g., virtual component 138) that is configured, such as by selection, clicking or touching of a mechanical button on the object 180, a hot spot (e.g., a location or an area on the object 180), or a foot pedal (e.g., on a RMS 120).

[0071] Virtuality transformer 150 can include any combination of hardware and software for detecting or identifying objects 180 from data 172 and representing, implementing, virtualizing or transforming user interactions determined by interaction determiner 160. Virtuality transformer 150 can represent or transform user interactions with respect to virtual components 138 (e.g., virtualized object 180 or based on object information 136) into object representations 152, such as virtual poses or virtual configurations. Virtuality transformer 150 can generate object representations 152 of the object 180.

[0072] Object representations 152 can include virtual representations corresponding to the object 180 or its transformations, interactions or configurations. Object representations 152 can include a rendering, a drawing or an illustration of a component 190 represented by the object 180 or its handling or movements. The component 190 representing the object can include any medical instrument 112, a medical device, a portion of a robotic medical system 120, such as a manipulator arm of the robot, any of which can be represented by the object 180. Interactive elements 154 of the object 180 can be utilized to simulate buttons or features of the component 190 as the object representation 152 provides a virtual version of the component 190 handled, managed, moved or manipulated based on user interactions with the object 180 as determined by the interaction determiner.

[0073] Virtuality transformer 150 can update the display unit (e.g., display 116) to provide updated representation of the virtualized object representations 152 of the object 180. The updates to the virtualized object representations 152 of the object 180 can be done responsive to the user interaction detected by the interaction determiner 160. For instance, the user interaction can be implemented using an interactive element 154. The interactive element 154 can include at least one of a mechanical button, such as a spring-loaded button or a dummy button on the object 180. The interactive element 154 can include a non-powered element, such as the interactive element 154 can include a foot pedal, such as a dummy pedal shaped or positioned as a pedal of an RMS 120 that the user (e.g., surgeon) can press to activate a virtualizer application 132 or its feature or a function. The interactive element 154 can include a hot spot on the physical object 180. The hot spot interactive element 154 can include an area or a region of the object 180 or a portion of a medical environment with respect to which user interaction carries a meaning or a significance of a particular action, such as a selection or activation of a feature or a button or activation of a functionality. The interaction determiner 160 can detect such interactions with respect to the interactive element 154 and update display (e.g., interface 174) to present the updated user interaction (e.g., object representation 152) on adisplay 116. In addition to optical sensors, the DPS 130 can use audio sensors or time of flight sensors 104 for determining or using interactive element 154 operation. Such sensors 104 can be used to validate the pose or interaction detected with the optical sensor. For example, if the data from the optical sensor indicates a button press, the data from the audio sensor can validate the button press by detecting audio signature or sound the corresponds to the button press.

[0074] Virtuality transformer 150 can generate object representations 152 based on the user interactions. Virtuality transformer 150 can transform the determined user interaction with the physical object 180 to at least one of a virtual pose or a virtual configuration for a virtual component 138 that represents the component 190 of the computer-assisted robotic system. For example, virtual component 138 can include virtual illustration or object representation 152 of the object 180. Virtual component 138 can include an illustration or a simulation of a component 190 represented by the object 180. Virtual component 138 can include a shape and a size of the component (e.g., a medical instrument 112 or a part of an RMS 120) illustrated within an illustration or simulation of a medical environment.

[0075] Machine learning (ML) environment 140 can include any combination of hardware and software for providing ML functionality. ML environment 140 can include ML training functions for training ML models 144. For example, ML environment 140 can include supervised or unsupervised Al or ML training functionalities for training ML models 144 to track and model user’s hand position, orientation and gestures, foot position and orientation and collisions or interactions of the user with the workspace and input devices. ML environment 140 can include any number of ML models trained to recognize and detect objects 180 and object interactions from a user. ML environment 140 can include ML functionalities, such as similarity functions, encoders and decoders for processing input data and providing outputs. ML environment 140 can include model selector 142 and weighting functions for selecting ML models 144 for particular objects 180.

[0076] ML environment 140 can include the functionality to train the ML models 144 for various specific functionalities. For instance, ML environment 140 can include model trainers to train models for specific use cases, such as specific types of objects, specific medical procedures, specific type of determinations. For instance, one ML model can make determinations using the entire body or hand of a user, while another, more specifically trained ML model 144 can focus on the movement of the user’s fingers, thumb or index and middle finger locations to make determinations. The trainers can train ML models for soft body object detection, deformation in object, button presses, rigid body interaction, hover over button or hotspot or to validate visual detection based on audio signals (e.g., detect mechanical click sound of button to validate button press)

[0077] When generating virtual representation of a medical environment 102, the ML models can be trained to insert into the virtual representation of the medical environment additional features or devices than they are present in the physical medical environment. For example, an ML model 144 can be trained to show the placement of the foot pedal at different distances, and show virtual hands at different locations to trick the user into using their left or right foot to hit the foot pedal. The ML model 144 can be trained to not show or illustrate in the VR where the user’s hand is actually located in the physical reality. Instead, ML models 144 can be trained to generate VR representation of the user’s hands at a different location in the VR than their location in the physical space. Virtualizer application 132 can implement offsets in three dimensions (e.g., x-axis, y-axis and z-axis) between user’s body placement in the physical space and the virtual space to provide a view for the user that the user’s hands or feet are interacting with a different part of the system then they are in the reality.

[0078] For example, the virtuality transformer or the virtualizer application can utilize ML models 144 to provide a type of simulation that corresponds to an online simulation platform in which the ML model 144 can track the thumb, index, and middle finger with the wrist in order to determine user interactions. The online simulation platform type of simulation can facilitate enhancing skills relevant to the user of a computer-assisted robotic system, as well as provide a virtual environment where users can practice and refine their surgical techniques through realistic simulations. The online simulation platform type of simulation can provide an interactive, guided scenario that can be accessible to users through a computing device or the computer-assisted robotic system. The ML model 144 can be trained to place more weight on the tracking of the fingers and the wrist than other parts of the hand. The ML model 144 can be trained to track the user’s feet and the foot pedal and to control the camera. For example, the ML model 144 can be trained to track the whole hand and the direction of the palm / wrist as a priority over individual fingers. For example, the ML model 144 can be trained to track foot rollers, which can be used to navigate within the room.

[0079] The setting of the application can be used by the application to select or determine a type of simulation to provide, or select one or more models to provide the type of simulations. Types of simulations can include, for example, an online simulation platform, a simulation for a type of computer-assisted robotic system, a simulation for a type of medical procedure, etc. For example, the type of simulation can correspond to an interaction with a type of computer-assisted robotic system, or computer-assisted robotic system, used to perform a type of medical procedure. The computer-assisted robotic system can be a type that performs a biopsy, such as a minimally invasive biopsy in a peripheral lung using an articulating catheter that can navigate complex airway paths. In yet another example, the type of simulation can be for a procedural planning and navigation system, such as a system used to generate detailed 3D maps of patients’ airways using imaging data in order to identify and reach target biopsy sites accurately.

[0080] The setting of the application can indicate different phases of the application, which the application can use to tailer the simulation to the indicated phase. For example, the application can, using the setting, tailor the simulation or training to one or more phases of the medical procedures. Phases of the medical procedure can include, for example, operating room setup, pre-operative setup, intraoperative medical procedure, post-operative clean-up, etc.

[0081] ML model 144 can include any model for determining or recognizing user interactions with an object 180. ML model 144 can include any machine learning or artificial intelligence model trained on a data set of data 172 to identify objects 180, render, generate or provide virtual components 138 or object representations 152 and detect or determine user interactions with the object 180. ML model 144 can be trained to determine one or more user interactions with an object 180 based on the data 172 (e.g., data collected from the one or more sensors 104, such as optical sensors). ML models 144 can be trained with machine learning to process data 172 using one or more optical sensors that capture user interactions with the physical object.

[0082] ML model 144 can be configured to be triggered, controlled, managed, called by or run by, or on behalf of, any one or more of a virtualizer application 132, virtuality transformer 150, model selector 142 and interaction determiner 160. For instance, any determinations, detections, identifications or outputs provided by the virtualizer application 132, virtuality transformer 150, model selector 142 or interaction determiner 160 can be implemented using one or more ML models 144 which these components can call, such as by using application programming interface (API) calls.

[0083] ML model 144 can be configured (e.g., trained or set up) to determine user interactions for an interaction determiner 160. For instance, ML model 144 can determine one or more first or second user interactions with the physical object 180 based data 172. For instance, ML model 144 can determine a first user interaction based on or using first data 172 (e.g., corresponding to a first time period) and also determine a second user interaction basedon or using second data 172 (e.g., corresponding to a second time period following the first time period). The data 172 can be collected from the one or more data capture devices 110, such as optical sensors 104. The second user interaction can be subsequent to the first user interaction. ML model 144 can be used (e.g., by interaction determiner 160) to determine object information 136, virtual components 138 or object representations 152 (e.g., with respect to any interactive elements 154) based on the first and second interactions.

[0084] ML models 144 can be trained or configured to implement or provide any range of outputs or determinations with respect to any component of the DPS 130. For instance, one or more ML models 144 can determine a gesture of a hand of a user interacting with the physical object. For example, ML models 144 can detect or identify objects 180 based on markers or interactive elements 154. For example, ML models 144 can determine interactions of the user with respect to the objects, or engagement with any interactive element 154. For example, ML models 144 can generate or draw an illustration or a simulation of a 2D or a 3D representation of a medical environment including the object 180 being drawing as a virtual component 138 (e.g., representing a particular component 190). ML model 144 can present object representations 152 including poses, movements, positioning or configuration of the virtual component 138 (e.g., based on the movements of the object 180 in the physical space). ML models 144 can provide or generate AR or VR simulations or virtualizations of object 180 (e.g., virtual components 138 and object representations 152) along with any details of the medical environment 102, including patient location or anatomy. ML models 144 can be trained to determine confidence scores of their determinations, such as a confidence score of a detected or identified object 180, virtual component 138, object representation 152, user interaction or any other determination. A confidence score can include a value, parameter or a determination of a level of confidence that the determination or output is accurate or correct.

[0085] ML models 144 can include supervised or unsupervised Al or ML algorithms to track and model hand position, hand orientation, hand gestures, foot position, foot orientation, and collisions / interactions of the user with the workspace and input devices. For instance, the ML environment can utilize multiple ML models 144, such as a first ML model (e.g., a larger model trained on a more general subject matter or field) and then one or more smaller ML models 144 that can be customized or trained on smaller fields and that can build off the larger ML model. For example, the larger ML model 144 can track hands and feet, while the customized, smaller ML models 144 can build off the output from the larger models to track specific, more granular features, based on a use case, such as thumb, index, and middle fingerwith the wrist for simulation such as the online simulation platform type of simulation. The technical solutions can combine multiple ML models 144, such as a hand model and an object IR model and apply different weights to different ML models 144, based on the confidence score of outputs of the ML models, which can be dependent upon visibility, data reliability or occlusion. By combining such categorized ML model structure, the technical solutions can reduce ML model drifting or hallucinations and improve the ML determination accuracy.

[0086] ML models 144 can be trained to adapt to different types of use cases, involving different applications, phases of medical session, types of medical procedures, or types of objects. Based on the type of use case, certain features or functions (e.g., ML models 144) can be weighted more by the weighting function than other functions or models, thereby reducing noise and improving computing efficiency, accuracy and the reliability. This can be accomplished, for example, by model selector selecting certain ML models 144 that are trained for the desired features, using weighting functions or otherwise applying weights, filtering out certain types of data (e.g., in response to detecting data errors or object occlusion), or activating a select set of sensors based on the use case.

[0087] Model selector 142 can include any combination of hardware and software for selecting ML models 144 for particular functionalities. Model selector 142 can select ML models 144 for particular operations. Model selector 142 can select ML models 144 for particular objects 180. Model selector 142 can utilize one or more ML models 144 to select a first ML model 144 configured to process visible light captured by the one or more optical sensors and select a second ML model 144 configured to process infrared light captured by the one or more optical sensors. Model selector 142 can select a third ML model 144 configured to receive first output from the first ML model 144 and second output from the second ML model 144. The third ML model 144 can determine an object representation 152, such as a pose of the physical object based on the first output and the second output by combining the first output with the second output to generate a third output. The third ML model 144 can determine the pose of the physical object based on the third output. The pose can correspond to an orientation or a position of the physical object 180. The third ML model 144 can determine the pose of the physical object based on the weighted output from a weighting function 146. The pose can correspond to a position of the physical object 180 within a time interval corresponding to a plurality of poses of the physical object 180.

[0088] Model selector 142 can utilize a weighting function 146 to select ML models 144. Weighting function 146 can be used to determine confidence scores. For instance, ML model144 can identify a first confidence score for the first output from the first model and identify a second confidence score for the second output from the second model. The model selector 142 or a third ML model 144 can apply the weighting function to the first output and the second output based on the first confidence score and the second confidence score. By applying a weighting function 146 to the first output and the second output to generate a weighted output.

[0089] The weight can increase or decrease based on ML model determinations. For example, the weight can be adjusted or determined, responsive to the determination that the portion is occluded. For example, in response to determining that one data 172 used by one ML model 144 is occluded or obstructed from view, the weight can be weight applied to a second output from a second ML model 144 configured to process a different data 172. For instance, data 172 from a visible light camera can be occluded or obstructed, and ML model 144 can determine to apply a weight more heavily (e.g., above a threshold) towards a second ML model 144 that utilizes infrared light relative, to the first output from the first model.

[0090] Weighting function 146 can determine confidence scores for various model outputs. Weighting function 146 can identify a first confidence score for the first output from the first ML model 144 and identify a second confidence score for the second output from the second ML model 144. An ML model 144 can apply the weighting function 146 to the first output and the second output based on the first confidence score for the output from the first ML model 144 and the second confidence score for the output from the second ML model 144. The first and the second ML models 144 can be trained to track at least one of hand movement, eye movement, or foot movement based on one or more different data 172 (e.g., data streams from different data capture devices 110).

[0091] The weighting function 146 can be used to determine, using the second one or more models and based on the data from the one or more optical sensors, a user interaction with the second physical object based at least on the second one or more markers 182. The marker 182 shapes, sizes or patterns of markers 182 can be used to identify objects 180. The weighting function can determine (e.g., by weighing one or more outputs or selecting one or more ML models 144), using the one or more ML models 144 and based on the user interaction with the physical object 180, a first task of a plurality of tasks of the medical session. The weighting function 146 can be used to determine, using the one or more models and based on the data, a pose of the physical object (e.g., by weighing one or more outputs or selecting one or more ML models 144). The weighting function 146 can determine, using the one or more models and the pose of the physical object, a second task of the plurality of tasks of the medical session.

[0092] Model selector 142 can select one or more ML models 144 to use based on the physical object 180 or based on the setting 134. For instance, some ML models 144 can be used for a first object 180 type or a first setting 134 and some ML models 144 can be used for a second object 180 type or a second setting 134. Model selector 142 can select the one or more models based on the component 190 used in the medical session performed using the computer-assisted robotic system. The component 190 can be represented by the physical object 180. For instance, as the user uses or interacts with the physical object 180, in the XR (e.g., VR or AR) representation of the medical environment in the virtualizer application 132, the selected component 190 can be represented in its virtual form as a virtual component 138 as the user utilizes the physical object 180 to interact and simulate with the virtual component 138 displayed in the virtualizer application 132 in the display 116 (e.g., of a HMD 122).

[0093] Model selector 142 can select the one or more ML models 144 based on the type of the physical object 180. Model selector 142 can select the one or more ML models 144 based on the type of interactive element 154. Model selector 142 can select a first ML model 144 configured to process visible light captured by the one or more optical sensors. Model selector 142 can select a second ML model 144 configured to process infrared light captured by the one or more optical sensors 104. Model selector 142 can select a third ML model 144 configured to receive first output from the first ML model 144 and second output from the second ML model 144 and use the third ML model 144 to determine the object representation 152 (e.g., a pose of the physical object), based on the first output and the second output. Model selector 142 can select, based on the second physical object, a second one or more ML models 144 trained with machine learning to process data from the one or more optical sensors 104 that capture user interactions with the second physical object, when second object 180 is swapped for the prior used, first object 180.

[0094] FIG. 2 depicts an example 200 of a physical object 180 that can be used for user interaction and virtualization or simulation of a component 190 of a robotic medical system 120. Example 200 shows a physical object 180 shaped as an elongate component, piece of material or a member having interactive elements 154 and markers 182. The surface of the object 180 or any interactive elements 154 or markers 182 can include physical features, such as ribbing, ridges or roughened surface or other physical structures to provide a point of reference to the user while touching the object 180 and looking at a display 116 (e.g., looking away from the object 180). Surface structures on the object can mark the locations of the markers 182 or interactive elements 154 (e.g., locations of buttons, hot spots or areas of usertouch or interactions). Surface structures can allow the user to more easily interact with the object 180.

[0095] Markers 182 can be arranged in a pattern or a configuration, such as a line, triangle, rectangle. Markers 182 can include any size or shape or material (e.g., fluorescent material or a retroreflective material) allowing for improved detection by a data capture device 110. Markers 182 can be on any part of the object 180 (e.g., any or all of the ends, or the middle). Interactive elements 154 can also serve or include markers 182 or vice versa.

[0096] FIG. 3 depicts a surgical system 300, in accordance with some aspects of the technical solutions. The surgical system 300 can be an example of a medical environment 102. The surgical system 300 may include a robotic medical system 305 (e.g., the robotic medical system 120), a user control system 310, and an auxiliary system 315 communicatively coupled one to another. A visualization tool 320 (e.g., the visualization tool 114) may be connected to the auxiliary system 315, which in turn may be connected to the robotic medical system 305. Thus, when the visualization tool 320 is connected to the auxiliary system 315 and this auxiliary system is connected to the robotic medical system 305, the visualization tool may be considered connected to the robotic medical system. In some embodiments, the visualization tool 320 may additionally or alternatively be directly connected to the robotic medical system 305.

[0097] The surgical system 300 may be used to perform a computer-assisted medical procedure on a patient 325. In some embodiments, surgical team may include a surgeon 330A and additional medical personnel 330B-330D such as a medical assistant, nurse, and anesthesiologist, and other suitable team members who may assist with the surgical procedure or medical session. The medical session may include the surgical procedure being performed on the patient 325, as well as any pre-operative (e.g., which may include setup of the surgical system 300, including preparation of the patient 325 for the procedure), and post-operative (e.g., which may include clean up or post care of the patient), or other processes during the medical session. Although described in the context of a surgical procedure, the surgical system 300 may be implemented in a non-surgical procedure, or other types of medical procedures or diagnostics that may benefit from the accuracy and convenience of the surgical system.

[0098] The robotic medical system 305 can include a plurality of manipulator arms 335 A- 335D to which a plurality of medical tools (e.g., the medical tool 112) can be coupled or installed. Each medical tool can be any suitable surgical tool (e.g., a tool having tissueinteraction functions), imaging device (e.g., an endoscope, an ultrasound tool, etc.), sensing instrument (e.g., a force-sensing surgical instrument), diagnostic instrument, or other suitable instrument that can be used for a computer-assisted surgical procedure on the patient 325 (e.g., by being at least partially inserted into the patient and manipulated to perform a computer- assisted surgical procedure on the patient). Although the robotic medical system 305 is shown as including four manipulator arms (e.g., the manipulator arms 335A-335D), in other embodiments, the robotic medical system can include greater than or fewer than four manipulator arms. Further, not all manipulator arms can have a medical tool installed thereto at all times of the medical session. Moreover, in some embodiments, a medical tool installed on a manipulator arm can be replaced with another medical tool as suitable.

[0099] One or more of the manipulator arms 335A-335D and / or the medical tools attached to manipulator arms can include one or more displacement transducers, orientational sensors, positional sensors, and / or other types of sensors and devices to measure parameters and / or generate kinematics information. One or more components of the surgical system 300 can be configured to use the measured parameters and / or the kinematics information to track (e.g., determine poses of) and / or control the medical tools, as well as anything connected to the medical tools and / or the manipulator arms 335A-335D.

[0100] The user control system 310 can be used by the surgeon 330A to control (e.g., move) one or more of the manipulator arms 335A-335D and / or the medical tools connected to the manipulator arms. To facilitate control of the manipulator arms 335A-335D and track progression of the medical session, the user control system 310 can include a display (e.g., the display 116 or 340) that can provide the surgeon 330A with imagery (e.g., high-definition 3D imagery) of a surgical site associated with the patient 325 as captured by a medical tool (e.g., the medical tool 112, which can be an endoscope) installed to one of the manipulator arms 335A-335D. The user control system 310 can include a stereo viewer having two or more displays where stereoscopic images of a surgical site associated with the patient 325 and generated by a stereoscopic imaging system can be viewed by the surgeon 330A. In some embodiments, the user control system 310 can also receive images from the auxiliary system 315 and the visualization tool 320.

[0101] The surgeon 330A can use the imagery displayed by the user control system 310 to perform one or more procedures with one or more medical tools attached to the manipulator arms 335A-335D. To facilitate control of the manipulator arms 335A-335D and / or the medical tools installed thereto, the user control system 310 can include a set of controls. These controlscan be manipulated by the surgeon 330A to control movement of the manipulator arms 335 A- 335D and / or the medical tools installed thereto. The controls can be configured to detect a wide variety of hand, wrist, and finger movements by the surgeon 330A to allow the surgeon to intuitively perform a procedure on the patient 325 using one or more medical tools installed to the manipulator arms 335A-335D.

[0102] The auxiliary system 315 can include one or more computing devices configured to perform processing operations within the surgical system 300. For example, the one or more computing devices can control and / or coordinate operations performed by various other components (e.g., the robotic medical system 305, the user control system 310) of the surgical system 300. A computing device included in the user control system 310 can transmit instructions to the robotic medical system 305 by way of the one or more computing devices of the auxiliary system 315. The auxiliary system 315 can receive and process image data representative of imagery captured by one or more imaging devices (e.g., medical tools) attached to the robotic medical system 305, as well as other data stream sources received from the visualization tool. For example, one or more image capture devices (e.g., the data capture devices 110) can be located within the surgical system 300. These image capture devices can capture images from various viewpoints within the surgical system 300. These images (e.g., video streams) can be transmitted to the visualization tool 320, which can then passthrough, forward, relay, or transmit those images to the auxiliary system 315 as a single combined data stream. The auxiliary system 315 can then transmit the single video stream (including any data stream received from the medical tool(s) of the robotic medical system 305) to present on a display (e.g., the display 116 or 340) of the user control system 310.

[0103] In some embodiments, the auxiliary system 315 can be configured to present visual content (e.g., the single combined data stream) to other team members (e.g., the medical personnel 330B-330D) who might not have access to the user control system 310. Thus, the auxiliary system 315 can include a display 340 configured to display one or more user interfaces, such as images of the surgical site, information associated with the patient 325 and / or the surgical procedure, and / or any other visual content (e.g., the single combined data stream). In some embodiments, display 340 can be a touchscreen display and / or include other features to allow the medical personnel 330A-330D to interact with the auxiliary system 315.

[0104] The robotic medical system 305, the user control system 310, and the auxiliary system 315 can be communicatively coupled one to another in any suitable manner. For example, in some embodiments, the robotic medical system 305, the user control system 310,and the auxiliary system 315 can be communicatively coupled by way of control lines 345, which can represent any wired or wireless communication link that can serve a particular implementation. Thus, the robotic medical system 305, the user control system 310, and the auxiliary system 315 can each include one or more wired or wireless communication interfaces, such as one or more local area network interfaces, Wi-Fi network interfaces, cellular interfaces, etc. It is to be understood that the surgical system 300 can include other or additional components or elements that can be needed or considered desirable to have for the medical session for which the surgical system is being used.

[0105] FIG. 4 depicts an example block diagram of an example computer system 400, also referred to as a computing system 400 or a computing environment 400, to be used in accordance with some embodiments. The computer system 400 can be any computing device used herein and can include or be used to implement a data processing system 130 or its components. The computer system 400 includes at least one bus 405 or other communication component or interface for communicating information between various elements of the computer system. The computer system further includes at least one processor 410 or processing circuit coupled to the bus 405 for processing information. The computer system 400 also includes at least one main memory 415, such as a random-access memory (RAM) or other dynamic storage device, coupled to the bus 405 for storing information, and instructions to be executed by the processor 410. The main memory 415 can be used for storing information during execution of instructions by the processor 410. The computer system 400 can further include at least one read only memory (ROM) 420 or other static storage device coupled to the bus 405 for storing static information and instructions for the processor 410. A storage device 425, such as a solid-state device, magnetic disk or optical disk, can be coupled to the bus 405 to persistently store information and instructions.

[0106] The computer system 400 can be coupled via the bus 405 to a display 430, such as a liquid crystal display, or active-matrix display, for displaying information. An input device 435, such as a keyboard or voice interface can be coupled to the bus 405 for communicating information and commands to the processor 410. The input device 435 can include a touch screen display (e.g., the display 430). The input device 435 can also include a cursor control, such as a mouse, a trackball, or cursor direction keys, for communicating direction information and command selections to the processor 410 and for controlling cursor movement on the display 430.

[0107] The processes, systems and methods described herein can be implemented by the computer system 400 in response to the processor 410 executing an arrangement of instructions contained in the main memory 415. Such instructions can be read into the main memory 415 from another computer-readable medium, such as the storage device 425. Execution of the arrangement of instructions contained in the main memory 415 causes the computer system 400 to perform the illustrative processes described herein. One or more processors in a multiprocessing arrangement can also be employed to execute the instructions contained in the main memory 415. Hard-wired circuitry can be used in place of or in combination with software instructions together with the systems and methods described herein. Systems and methods described herein are not limited to any specific combination of hardware circuitry and software.

[0108] Although an example computing system has been described in FIG. 4, the subject matter including the operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.

[0109] FIG. 5 depicts an example method for virtual medical session interactions with a non-powered object. For example, the method 500 can provide extended reality (XR) content for user interactions with a physical object representing a component of a robotic medical system. The method 500 can be performed by a system having one or more processors executing computer-readable instructions stored on a memory. The method 500 can be performed, for example, by any combination of features or techniques discussed in connection with example systems 100-400 and FIGs. 1-4. For instance, the method 500 can be implemented one or more processors 410 of a computing system 400 executing non-transitory computer-readable instructions stored on a memory (e.g., the memory 415, 420 or 425) and using data from a data repository 170 to implement the functionalities of a DPS 130.

[0110] The method 500 can be used to provide virtual representations of components of a robotic medical system represented by a physical object interacted or handled by a user using operations 505-540. At operation 505, the method can identify an application setting. At operation 510, the method can receive data about user interactions. At method 515, the method can determine if the selected ML model is suitable for the determinations. At 520, if the determination at 515 is negative, the method can select one or more ML models. At 525, the method can use the one or more ML models to identify the physical object for the component. At 530, the method can determine user interactions with the object. At 535, the method cantransform the user interactions into object representations. At 540, the method can display virtual component with object presentations and revert back to operation 510 to receive additional input data for a next determination.

[0111] At operation 505, the method can identify an application setting. The method can include the one or more processors coupled with a memory storing instructions, identifying a setting for a virtualizer application. The one or more processors can identify the setting based on a user input into a user interface, such as a user selection of a setting for the virtualizer application via an interactive input device (e.g., a mouse click or a selection). The setting can include a setting for a type of an application or a mode of the application. The type or mode of the application can include a VR simulation application for simulating a medical procedure, an AR application for improving visibility and efficiency of a surgical procedure on a patient or a training application for providing education or practice to a surgeon with respect to a simulated application.

[0112] User selection can pick an application from a plurality of applications. The selected application can provide virtual representations of components represented by a physical object handled by a user. The application can include a user interface for displaying XR contents in the display, such as a display of a head mounted device (HMD) used by a user or a display coupled with a robotic medical system.

[0113] At operation 510, the method can receive data about user interactions. The method can include the one or more processors receiving data collected from one or more sensors, such as optical sensors. The data can include readings or measurements from device capture devices, including cameras for capturing images or videos of the physical object or the user interaction with the object from a plurality of directions or angles. The data can include readings or data indicative of object movement or motion. The data can include readings or data indicative of positioning of the user’s hands or fingers with respect to a portion of the physical object. The data can include information on locations of markers or interactive elements on the object. The data an include size and shape of the object.

[0114] The data can include information about user interaction with any portion or a feature of an object. For example, the data can include information about user’s selection, touching or pointing towards an interactive element, such as a hot spot on the object or a button. The data can include information about user’s selection, touching or pointing of a foot pedal of the object or of an RMS. The data can include information about the motion of the object (e.g.,multiple video frames of the object and the user’s handling of the object). The data can include sound information that can be generated based on the user’s pressing of a button on an object.

[0115] At method 515, the method can determine if the selected ML model is suitable for the determinations. The method can the one or more processors utilizing an ML environment to provide any determinations regarding selections of the ML models. For example, a model selector of the ML environment can include the functionality to analyze data received at operation 510 and make a determination on selection of the one or more ML models for processing the data. The model selector can identify the type of the data to be used for the ML models and can selector identify the ML models to be used for determinations based on the data. For example, the model selector can analyze the data and determine one or more ML models for the particular object type. For instance, an object type can include a size and a shape of the object which can be selected based on the type of the component the object is meant to represent.

[0116] At an initial run, the ML model selector can select a suitable ML model for the selected application setting and the type of object. For example, the ML model can utilize an interaction determiner or a virtuality transformer to determine or identify the object being captured or described in the data. For instance, the model selector can utilize an ML model trained on identifying the objects to detect the object from the data, based for example, on the markers of the object or interactive elements of the object (e.g., their sizes, shapes or arrangements).

[0117] The model selector can determine if the current ML model is suitable for selection based on any one or more of: the physical object (e.g., type of the object determined or identified), the setting of the virtualizer application, or the type of data being available (e.g., the data determined not to be corrupted or occluded). The ML model can apply a weighting function for particular types of data (e.g., when data is occluded or corrupted) or for particular ML models to be used. If the model selector determines that the current selection of the ML model is not suitable for the next operation, the model selector model selector can move to operation 520 to select a new model. If the model selector determines that the current selection of the ML model is suitable, the ML model can move to operation 525.

[0118] At 520, if the determination at 515 is negative, the method can select one or more ML models suitable for the operations at 520-535. The one or more processors can implement a model selector to select a single ML model or a plurality of ML models for a plurality ofdeterminations of any of the operations or determinations made at 520-535. For instance, the model selector can select one or more ML models for detecting the physical object (e.g., based on the data, and the object type). For example, the model selector can select one or more ML models for determining user interactions with the object. For example, the model selector can select one or more ML models for transforming user interactions into object representations (e.g., poses or configurations of the object in response to the user interactions).

[0119] The model selector can select one or more ML models for a particular setting of an application. For example, one or more ML models can be selected for a first setting (e.g., a type of an application or a mode of operation of the virtualizer application), and one or more different ML models can be selected for a second or a different setting of the application. ML models can be selected by the model selector based on the weighting function applying different weights to different ML models, based on the data.

[0120] For example, a first data stream for a data from a first data capture device used by a first ML model is obstructed. A weighting function can utilize an ML model to determine that a view of the object is occluded or obstructed by an intervening structure. The weighting function can determine, in response to the occlusion or obstruction for a duration of the data, to utilize more heavily (e.g., apply more weight) to determinations of a second ML model whose data stream is not obstructed or occluded. For example, the weighting function can utilize both the first (e.g., occluded) and a second (e.g., not occluded) data for a determination by one or more ML models and instruct the one or more ML models to apply a first (e.g., reduced or lower than a threshold) weight to a first data stream and a second (e.g., increased or a higher than a threshold) weight for the second data stream.

[0121] At 525, one or more ML models can be utilized to identify the physical object of the component. For example, responsive to the determination at 515 in the affirmative or the new one or more ML models being selected at 520, the method can use the one or more ML models as currently selected to identify the physical object for the component. For example, the interaction determiner can detect, recognize or identify the physical object based at least one the data input into the one or more ML models. For example, the interaction determiner can identify the object as a portable and handheld object for simulating a first type of a components, such as handheld medical instruments. For example, the interaction determiner can identify the object as a portable or handheld object for simulating a first type of a components, such as handheld medical instruments. For example, the interaction determiner can identify the object as a grounded type of object for simulating a robotic medical system or auser station at the robotic medical system. For example, the interaction determiner can identify the object as a work-station for remotely operating the robotic medical system.

[0122] The interaction determiner can determine the object and determine the component that is represented by the object. The interaction determiner can use the one or more ML models trained to detect or identify the object and to generate or provide one or more virtual components corresponding to the component represented by the object. For example, the interaction determiner can identify the object 180 as a first type of an object (e.g., an object for simulation operation of a particular medical instrument). The interaction determiner can provide a virtual component corresponding to the given particular medical instrument to generate and present an XR version of the particular medical instrument, such as a virtual representation object or an augmented reality object, for presentation or display in a user interface of the virtualizer application.

[0123] At 530, the method can determine user interactions with the object. The method can include the one or more processors implementing an interaction determiner to identify, detect or determine one or more user interactions with the physical object. The interactions determiner can utilize or trigger one or more ML models to detect, determine or identify the user interactions based at least one or more sensor data input into the one or more ML models. For example, the user interactions can include data corresponding to a hand tracker that can track the user’s hands, arms, legs or any other portion of the user’s body to discern interaction between the user and the physical object. For example, the user interactions can include multiple camera streams of images from multiple angles showing at least a portion of the user and the object.

[0124] The method can include the interactions determiner determining any interaction between the user and the object. For example, the interactions determiner can determine that the user is holding the physical object. For example, the interactions determiner can determine that the user is interacting with one or more interactive elements. For example, the interactions determiner can determine that the user has made a selection on a particular interactive element of the object, such as (e.g., a dummy button, a hot spot or a foot pedal). For example, the interactions determiner can determine that the user has turned on or turned off the virtual component, has activated or disactivated a mode or an operation, has applied a particular input or provided a particular output or taken any other action using the interaction with the object (e.g., via an interactive element selection or user’s gestures or movements with respect to the object).

[0125] At 535, the method can transform the user interactions into object representations. The method can include the one or more processors determining object representations. The one or more processers can implement the virtuality transformer to determine the object representations based on the detected user interactions and transforming the virtual components illustrated in the virtualizer application into the corresponding object representation.

[0126] For example, the virtuality transformer can determine that the user interaction changed the pose, position or location of the object. Responsive to this determination, the virtuality transformer can update the virtual component corresponding to the object (e.g., represented by the object) in the virtualizer application to reflect the change in the pose, position or location. For example, the virtuality transformer can determine that the user interaction changed the configuration, setting or operation of the object or the represented virtual component 138 (e.g., based on a user interaction with an interactive element).Responsive to this determination, the virtuality transformer can update the virtual component in the virtualizer application to reflect the change in the configuration, setting or operation. For example, the virtuality transformer can determine that the user interaction changed the tilt, angle, shape (e.g., via bending) or arrangement of features on the object. Responsive to this determination, the virtuality transformer can update the virtual component corresponding to the object in the virtualizer application to reflect such changes.

[0127] At 540, the method can display virtual component with object presentations. The method can include the one or more processors displaying the virtual content corresponding to the object (e.g., virtual component) along with any object representations (e.g., changes to the pose or configuration) on a display. For example, the interaction determiner can utilize one or more ML models to redraw or modify the visual representation of the virtual component to reflect the changes to the object representation and update the visual image of the virtual component.

[0128] For example, the user interface can generate an updated image or rendering of the current state of the virtual component and revert back to operation 510 to receive additional input data for a next determination. The method 500 can continue therefore cycling through acts or operations 510-540 at a periodic rate, such as 30 frames per second, 45 frames per second or 60 frames per second, thereby providing a continuously updated visual representation of the user’s interaction with the virtual component, based on the user’s handling of the physical object.

[0129] The herein described subject matter sometimes illustrates different components contained within, or connected with, different other components. It is to be understood that such depicted architectures are illustrative, and that in fact many other architectures can be implemented which achieve the same functionality. In a conceptual sense, any arrangement of components to achieve the same functionality is effectively “associated” such that the desired functionality is achieved. Hence, any two components herein combined to achieve a particular functionality can be seen as “associated with” each other such that the desired functionality is achieved, irrespective of architectures or intermedial components. Likewise, any two components so associated can also be viewed as being “operably connected,” or “operably coupled,” to each other to achieve the desired functionality, and any two components capable of being so associated can also be viewed as being “operably couplable,” to each other to achieve the desired functionality. Specific examples of operably couplable include but are not limited to physically mateable or physically interacting components or wirelessly interactable or wirelessly interacting components or logically interacting or logically interactable components.

[0130] With respect to the use of plural or singular terms herein, those having skill in the art can translate from the plural to the singular or from the singular to the plural as is appropriate to the context or application. The various singular / plural permutations can be expressly set forth herein for sake of clarity.

[0131] It will be understood by those within the art that, in general, terms used herein, and especially in the appended claims (e.g., bodies of the appended claims) are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” etc.).

[0132] Although the figures and description can illustrate a specific order of method steps, the order of such steps can differ from what is depicted and described, unless specified differently above. Also, two or more steps can be performed concurrently or with partial concurrence, unless specified differently above. Such variation can depend, for example, on the software and hardware systems chosen and on designer choice. All such variations are within the scope of the disclosure. Likewise, software implementations of the described methods can be accomplished with standard programming techniques with rule-based logic and other logic to accomplish the various connection steps, processing steps, comparison steps, and decision steps.

[0133] It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation, no such intent is present. For example, as an aid to understanding, the following appended claims can contain usage of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles “a” or “an” limits any particular claim containing such introduced claim recitation to inventions containing only one such recitation, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” or “an” should typically be interpreted to mean “at least one” or “one or more”); the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should typically be interpreted to mean at least the recited number (e.g., the bare recitation of “two recitations,” without other modifiers, typically means at least two recitations, or two or more recitations).

[0134] Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together, etc.). In those instances where a convention analogous to “at least one of A, B, or C, etc.” is used, in general, such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B.”

[0135] Further, unless otherwise noted, the use of the words “approximate,” “about,” “around,” “substantially,” etc., mean plus or minus ten percent.

[0136] The foregoing description of illustrative implementations has been presented for purposes of illustration and of description. It is not intended to be exhaustive or limiting withrespect to the precise form disclosed, and modifications and variations are possible in light of the above teachings or can be acquired from practice of the disclosed implementations. It is intended that the scope of the invention be defined by the claims appended hereto and their equivalents.

Claims

CLAIMSWhat is claimed is:

1. A system, comprising: one or more processors, coupled with memory, to: identify an application to virtualize at least a portion of a medical session where a computer-assisted robotic system is used; identify an indication of a physical object that represents a component associated with the computer-assisted robotic system; select, based on the physical object, one or more models trained with machine learning to process data from one or more optical sensors that capture user interactions with the physical object; determine, using the one or more models and based on the data collected from the one or more optical sensors, a user interaction with the physical object; transform the determined user interaction with the physical object to at least one of a virtual pose or a virtual configuration for a virtual component that represents the component of the computer-assisted robotic system; and display, on a display unit attached to computer-assisted robotic system, the virtual component according to the at least one of the virtual pose or the virtual configuration.

2. The system of claim 1, wherein the user interaction with the physical object changes at least one of a pose of the physical object or a configuration of the physical object.

3. The system of claim 2, wherein the pose of the physical object comprises at least one of a position of the physical object relative to a reference frame, or an orientation of the physical object relative to the reference frame.

4. The system of claim 2, wherein the configuration of the physical object corresponds to a state of activation of at least one of a mechanical button associated with the physical object, a foot pedal associated with the physical object, or a hot spot associated with the physical object.

5. The system of claim 1, wherein the one or more processors are further configured to: determine, using the one or more models and based on second data collected from the one or more optical sensors, a second user interaction with the physical object, the second userinteraction subsequent to the user interaction; change, based on the determined second user interaction with the physical object, at least one of the virtual pose or the virtual configuration for the virtual component to at least one of a second virtual pose or a second virtual configuration; and display, on the display unit, the virtual component according to the at least one of the second virtual pose or the second virtual configuration.

6. The system of claim 1, wherein the computer-assisted robotic system and the display unit are physically grounded to a reference platform.

7. The system of claim 1, wherein the one or more processors are further configured to: select the one or more models based on the component used in the medical session performed using the computer-assisted robotic system, the component to be represented by the physical object.

8. The system of any one of claims 1, 6 or 7, wherein the physical object comprises a marker with a pattern, and the one or more processors are further configured to: receive the indication of the physical object; determine a type of the physical object based on the pattern of the marker detected via the data from the one or more optical sensors; and select the one or more models based on the type of the physical object.

9. The system of claim 1, wherein the one or more processors are further configured to: identify a setting for the application to virtualize at least the portion of the medical session; and select the one or more models based on the setting.

10. The system of claim 9, wherein the one or more processors are further configured to: determine the setting for the application from a profile of a user of the computer- assisted robotic system.

11. The system of claim 9, wherein the setting corresponds to a type of simulation to be provided by the application.

12. The system of claim 9, wherein the setting relates to a phase of the medical session.

13. The system of claim 1, wherein the one or more processors are further configured to: update the display unit responsive to the user interaction, wherein the user interaction is implemented using an interactive element that comprises at least one of a mechanical button, a foot pedal, or a hot spot on the physical object.

14. The system of claim 13, wherein the interactive element lacks any coupling with the one or more processors and comprising the one or more processors to determine, using the one or more models, a gesture of a hand of a user interacting with the physical object.

15. The system of claims 13 or 14, wherein the one or more processors are further configured to: determine a type of interactive element of the physical object; select the one or more models based on the type of interactive element; and detect, using the one or more models and based on the data from the one or more optical sensors, the user interaction with the interactive element of the physical object.

16. The system of claim 1, wherein the one or more processors are further configured to: select a first model configured to process visible light captured by the one or more optical sensors; select a second model configured to process infrared light captured by the one or more optical sensors; select a third model configured to receive first output from the first model and second output from the second model; and use the third model to determine a pose of the physical object based on the first output and the second output.

17. The system of claim 16, wherein the one or more processors are further configured to: combine the first output with the second output to generate a third output; and use the third model to determine the pose of the physical object based on the third output, wherein the pose corresponds to at least one of an orientation or a position of the physical object.

18. The system of claim 16, wherein the one or more processors are further configured to: apply a weighting function to the first output and the second output to generate a weighted output; and use the third model to determine the pose of the physical object based on the weighted output, wherein the pose corresponds to a position of the physical object within a time interval corresponding to a plurality of poses of the physical object.

19. The system of claim 18, wherein the one or more processors are further configured to: determine a portion of the physical object is occluded from the one or more optical sensors; and increase, responsive to the determination that the portion is occluded, a weight applied to the second output from the second model configured to process the infrared light relative to the first output from the first model.

20. The system of claim 18, wherein the one or more processors are further configured to: identify a first confidence score for the first output from the first model; identify a second confidence score for the second output from the second model; and apply the weighting function to the first output and the second output based on the first confidence score and the second confidence score.

21. The system of claim 1, wherein the one or more models are trained to track at least one of hand movement, eye movement, or foot movement.

22. The system of claim 1, wherein the display unit is a head mounted display.

23. The system of claim 1, wherein the one or more processors are further configured to: detect that a second physical object is swapped in for the physical object, the second physical object comprising a second one or more markers that are different than a first one or more markers of the physical object; select, based on the second physical object, a second one or more models trained with machine learning to process data from the one or more optical sensors that capture user interactions with the second physical object; and determine, using the second one or more models and based on the data from the one or more optical sensors, a user interaction with the second physical object based at least on thesecond one or more markers.

24. The system of claim 1, wherein the one or more processors are further configured to: determine, using the one or more models and the user interaction with the physical object, a first task of a plurality of tasks of the medical session; determine, using the one or more models and based on the data, a pose of the physical object; and determine, using the one or more models and the pose of the physical object, a second task of the plurality of tasks of the medical session.

25. A method, comprising: identifying, by one or more processors coupled with memory, an application to virtualize at least a portion of a medical session where a computer-assisted robotic system is used; identifying, by the one or more processors, an indication of a physical object that represents a component associated with the computer-assisted robotic system; selecting, by the one or more processors, based on the physical object, one or more models trained with machine learning to process data from one or more optical sensors that capture user interactions with the physical object; determining, by the one or more processors, using the one or more models and based on the data collected from the one or more optical sensors, a user interaction with the physical object; transforming, by the one or more processors, the determined user interaction with the physical object to at least one of a virtual pose or a virtual configuration for a virtual component that represents the component of the computer-assisted robotic system; and displaying, by the one or more processors, on a display unit attached to computer- assisted robotic system, the virtual component according to the at least one of the virtual pose or the virtual configuration.

26. The method of claim 25, wherein the user interaction with the physical object changes at least one of a pose of the physical object or a configuration of the physical object.

27. The method of claim 26, wherein the pose of the physical object comprises at least one of a position of the physical object relative to a reference frame, or an orientation of the physicalobject relative to the reference frame.

28. The method of claim 26, wherein the configuration of the physical object corresponds to a state of activation of at least one of a mechanical button associated with the physical object, a foot pedal associated with the physical object, or a hot spot associated with the physical object.

29. The method of claim 25, comprising: determining, by the one or more processors, using the one or more models and based on second data collected from the one or more optical sensors, a second user interaction with the physical object, the second user interaction subsequent to the user interaction; changing, by the one or more processors, based on the determined second user interaction with the physical object, at least one of the virtual pose or the virtual configuration for the virtual component to at least one of a second virtual pose or a second virtual configuration; and displaying, by the one or more processors, on the display unit, the virtual component according to the at least one of the second virtual pose or the second virtual configuration.

30. The method of claim 25, wherein the computer-assisted robotic system and the display unit are physically grounded to a reference platform.

31. The method of claim 25, comprising: selecting, by the one or more processors, the one or more models based on the component used in the medical session performed using the computer-assisted robotic system, the component to be represented by the physical object.

32. The method of any one of claims 25, 30 or 31, wherein the physical object comprises a marker with a pattern, and the method further comprises: receiving, by the one or more processors, the indication of the physical object; determining, by the one or more processors, a type of the physical object based on the pattern of the marker detected via the data from the one or more optical sensors; and selecting, by the one or more processors, the one or more models based on the type of the physical object.

33. The method of claim 25, comprising:updating, by the one or more processors, the display unit responsive to the user interaction, wherein the user interaction is implemented using an interactive element that comprises at least one of a mechanical button, a foot pedal, or a hot spot on the physical object.

34. The method of claim 33, wherein the interactive element lacks any coupling with the one or more processors and comprising the one or more processors to determine, using the one or more models, a gesture of a hand of a user interacting with the physical object.

35. The method of claims 33 or 34, comprising: determining, by the one or more processors, a type of interactive element of the physical object; selecting, by the one or more processors, the one or more models based on the type of interactive element; and detecting, by the one or more processors, using the one or more models and based on the data from the one or more optical sensors, the user interaction with the interactive element of the physical object.

36. The method of claim 35, comprising: selecting, by the one or more processors, a first model configured to process visible light captured by the one or more optical sensors; selecting, by the one or more processors, a second model configured to process infrared light captured by the one or more optical sensors; selecting, by the one or more processors, a third model configured to receive first output from the first model and second output from the second model; and using, by the one or more processors, the third model to determine a pose of the physical object based on the first output and the second output.

37. The method of claim 36, comprising: combining, by the one or more processors, the first output with the second output to generate a third output; and using, by the one or more processors, the third model to determine the pose of the physical object based on the third output, wherein the pose corresponds to at least one of an orientation or a position of the physical object.

38. The method of claim 36, comprising: applying, by the one or more processors, a weighting function to the first output and the second output to generate a weighted output; and using, by the one or more processors, the third model to determine the pose of the physical object based on the weighted output, wherein the pose corresponds to a position of the physical object within a time interval corresponding to a plurality of poses of the physical object.

39. The method of claim 38, comprising: determining, by the one or more processors, a portion of the physical object is occluded from the one or more optical sensors; and increasing, by the one or more processors, responsive to the determination that the portion is occluded, a weight applied to the second output from the second model configured to process the infrared light relative to the first output from the first model.

40. The method of claim 38, comprising: identifying, by the one or more processors, a first confidence score for the first output from the first model; identifying, by the one or more processors, a second confidence score for the second output from the second model; and applying, by the one or more processors, the weighting function to the first output and the second output based on the first confidence score and the second confidence score.

41. The method of claim 25, wherein the one or more models are trained to track at least one of hand movement, eye movement, or foot movement.

42. The method of claim 25, wherein the display unit is a head mounted display.

43. The method of claim 25, comprising: detecting, by the one or more processors, that a second physical object is swapped in for the physical object, the second physical object comprising a second one or more markers that are different than a first one or more markers of the physical object; selecting, by the one or more processors, based on the second physical object, a second one or more models trained with machine learning to process data from the one or more opticalsensors that capture user interactions with the second physical object; and determining, by the one or more processors, using the second one or more models and based on the data from the one or more optical sensors, a user interaction with the second physical object based at least on the second one or more markers.

44. The method of claim 25, comprising: identifying, by the one or more processors, a setting for the application to virtualize at least the portion of the medical session; and selecting, by the one or more processors, the one or more models based on the setting.

45. The method of claim 44, comprising: determining, by the one or more processors, the setting for the application from a profile of a user of the computer-assisted robotic system.

46. The method of claim 44, wherein the setting corresponds to a type of simulation to be provided by the application.

47. The method of claim 44, wherein the setting relates to a phase of the medical session.

48. The method of claim 25, comprising: determining, by the one or more processors, using the one or more models and the user interaction with the physical object, a first task of a plurality of tasks of the medical session; determining, by the one or more processors, using the one or more models and based on the data, a pose of the physical object; and determining, by the one or more processors, using the one or more models and the pose of the physical object, a second task of the plurality of tasks of the medical session.

49. A non-transitory computer-readable medium storing processor executable instructions, that when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 25-48.

Citation Information

Patent Citations

  • Augmented reality triggering of devices

    US20200302694A1

  • Mobile virtual reality system for surgical robotic systems

    US20210307831A1