Three dimensional virtual scene initialization of a medical procedure environment
A computing system generates realistic 3D virtual surgical scenes from monocular images using machine learning, addressing the lack of effective training methods for medical robotic systems by providing interactive simulations for surgeons to practice and explore new techniques.
Patent Information
- Application Number
- PCT/US2025/042392
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-19
- Filing Date
- 2025-08-18
- Publication Date
- 2026-02-26
AI Technical Summary
Existing medical robotic systems lack effective methods for generating realistic 3D virtual simulations of surgical scenes from monocular images, limiting training opportunities for surgeons and failing to replicate the exact clinical scenarios they will encounter.
A computing system uses machine learning models to generate 3D meshes and texture files from monocular surgical images, creating highly detailed and realistic 3D virtual scenes that can be interacted with in tutorial, competition, or generative modes, allowing surgeons to practice and explore new techniques in a safe environment.
Enables surgeons to train and practice surgical skills in a clinically accurate virtual environment, facilitating improved training and exploration of new techniques without the need for resource-intensive 3D scanning equipment.
Smart Images

Figure US2025042392_26022026_PF_FP_ABST
Abstract
Description
Atty. Dkt: 135039-0493 (P06965-WO)THREE DIMENSIONAL VIRTUAL SCENE INITIALIZATION OF A MEDICAL PROCEDURE ENVIRONMENTCROSS-REFERENCES TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority under 35 U.S.C. § 119 to U.S. Provisional Patent Application No. 63 / 684,585, filed August 19, 2024, which is hereby incorporated by reference herein in its entirety.BACKGROUND
[0002] A medical robotic system can include an instrument for performing a medical session or procedure. For example, the instrument can be used to perform surgery, therapy, or a medical evaluation. The medical robotic system can include an endoscope that captures a video of the medical procedure.SUMMARY
[0003] Technical solutions disclosed herein can include a three dimensional (3D) virtual scene reconstruction of a medical procedure environment. This technology can facilitate or automatically initiate a 3D virtual simulation that corresponds to or matches a scene in a medical procedure using an image of the scene. The computing system can generate at least one 3D mesh (e.g., a collection of vertices, edges, and faces) and at least one texture file (e.g., appearance of an anatomical structure, color of the anatomical structure, roughness or smoothness of the anatomical structure, state of anatomical structures) from an image of a medical procedure, such as a monocular surgical image, and generate or initialize the 3D virtual scene using the 3D mesh and texture file. The computing system can use machine learning models to generate a segmentation mask, an anatomy state mask, and a metric depth mask based on the image of the scene. With these masks, the system can construct the 3D mesh and texture file that can represent a 3D version of the scene. By using the masks, the computing system can produce a highly detailed and realistic 3D virtual scene. The system can allow a user to interact with the 3D virtual simulation in various modes. For example, the system can provide a tutorial mode, a competition mode, or a generative mode. The tutorial mode can limit, constrain, or prohibit certain types or amounts of input (e.g., the user may be required to follow the same trajectory made by an instrument during the actual medical procedure). The competition mode may loosen certain constraints to allow the user,14855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) in the simulation, to try to match or outperform performance metrics relative to the medical procedure. In the generative mode, the system can use generative artificial intelligence techniques to initiate a virtual simulation that is based on the frame of the scene in the medical procedure as well as other parameters. For example, in the generative mode, the system can change patient parameters, anatomical structure parameters (e.g., amount of bleeding, textures, or state), etc.
[0004] At least one aspect of the present disclosure is a system. The system can include one or more processors, coupled with memory, to identify a frame that captures a scene in a medical procedure performed with a robotic medical system. The one or more processors can generate, using one or more models trained with machine learning, one or more masks that segment an anatomical structure in the frame, label the anatomical structure in the frame, and indicate a depth value for each pixel in the frame. The one or more processors can construct 3D mesh and texture data based on the one or more masks. The one or more processors can provide, in an interactive virtual environment, a virtual scene rendered using the 3D mesh and texture data generated from the frame that captures the scene in the medical procedure.
[0005] The one or more processors can generate, using the one or more models, the one or more masks including segments and labels for anatomical structures in the frame. The one or more processors can construct the 3D mesh and texture data for the anatomical structures in the frame. The one or more processors can provide the virtual scene rendered using the 3D mesh and texture data for the anatomical structures in the frame.
[0006] The one or more masks can include at least one of a segmentation mask and a state mask.
[0007] The frame can correspond to a single view of the scene in the medical procedure.
[0008] The one or more processors can determine, based on the frame and using the one or more models, a state of the anatomical structure in the frame. The one or more processors can construct, based on the state, the 3D mesh and texture data. The one or more processors can provide, in the interactive virtual environment, the virtual scene rendered to represent the state of the anatomical structure.
[0009] The state can include at least one of burned, cut, or bleeding.24855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0010] The one or more processors can display the interactive virtual environment via a display device.
[0011] The display device can be coupled with a head-mounted display.
[0012] The one or more processors can receive a data stream including information about a state of a subject that undergoes the medical procedure. The one or more processors can construct the 3D mesh and texture data based on the state of the subject.
[0013] The one or more processors can access a data repository storing parameters for the subject. The one or more processors can construct the 3D mesh and texture data using the parameters for the subject.
[0014] The parameters for the subject can provide an indication of a likelihood of a state of the anatomical structure during the medical procedure.
[0015] The one or more processors can receive a kinematics data stream captured via one or more sensors associated with the scene in the medical procedure. The one or more processors can construct the 3D mesh and texture data using the kinematics data stream.
[0016] The kinematics data stream can indicate an amount force detected between an instrument of the robotic medical system and at least a portion of the anatomical structure in the scene.
[0017] The one or more processors can display a video of the medical procedure, the video including frames that includes the frame. The one or more processors can receive, via an interface, a request to establish the virtual scene in the interactive virtual environment for the frame of the frames.
[0018] The one or more processors can receive, via an interface, a request to initiate a simulation in the interactive virtual environment for the frame of the scene in the medical procedure.
[0019] The one or more processors can receive a request to establish the virtual scene with a view angle that is different from the view angle with which the frame is captured for the scene in the medical procedure.
[0020] The one or more processors can detect a trigger condition in the video. The one or34855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) more processors can determine, responsive to the trigger condition, to generate a prompt including a recommendation to establish the virtual scene for the medical procedure. The one or more processors can receive the request to establish the virtual scene responsive to the prompt.
[0021] The trigger condition can be based on at least one of a type of task performed with the robotic medical system in the video stream, an occurrence of a milestone event in the frame, or a performance metric associated with the medical procedure.
[0022] The one or more processors can establish a mode of interaction for the interactive virtual environment, wherein the mode of interaction includes at least one of i) a tutorial mode in which parameters of a subject associated with the medical procedure are fixed and a trajectory of movement of a virtual robotic medical system in the virtual scene is constrained based on the trajectory of movement of the robotic medical system during the medical procedure, ii) a competition mode in which the parameters of the subject associated with the medical procedure are fixed and the trajectory of movement of the virtual robotic system is unconstrained relative to the trajectory of movement of the robotic medical system during the medical procedure, or iii) a generative mode in which parameters for a virtual subject in the virtual scene are adjustable and the trajectory of movement of the virtual robotic system is unconstrained relative to the trajectory of movement of the robotic medical system during the medical procedure.
[0023] The one or more processors can receive, via an interface, a request to initiate a generative simulation based on the frame of the scene. The one or more processors can identify parameters for the generative simulation, wherein at least one of the parameters is different than a parameter of the scene in the medical procedure. The one or more processors can construct, using the one or more models, the virtual scene using the 3D mesh and texture data based on the one or more masks and the parameters for the generative simulation.
[0024] The one or more processors can detect, via the interface, an interaction in the virtual scene rendered in the interactive virtual environment. The one or more processors can generate, using the one or more models, one or more subsequent virtual scenes responsive to the interaction. The one or more processors can provide, in the interactive virtual environment, the one or more subsequent virtual scenes.
[0025] The one or more processors can identify a first value of a performance metric44855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) associated with the medical procedure performed with the robotic medical system. The one or more processors can identify a second value of the performance metric associated with an interaction in the virtual scene. The one or more processors can display the first value of the performance metric and the second value of the performance metric.
[0026] The one or more processors can determine, based on a comparison of the second value with a threshold, to reset the virtual scene in the interactive virtual environment.
[0027] The one or more processors can display the frame in a first graphical user interface element in the interactive virtual environment. The one or more processors can display the virtual scene in a second graphical user interface element in the interactive virtual environment.
[0028] At least one aspect of the present disclosure is directed to a method. The method can include identifying, by one or more processors coupled with memory, a frame that captures a scene in a medical procedure performed with a robotic medical system. The method can include generating, by the one or more processors, using one or more models trained with machine learning, one or more masks that segment an anatomical structure in the frame, label the anatomical structure in the frame, and indicate a depth value for each pixel in the frame. The method can include constructing, by the one or more processors, 3D mesh and texture data based on the one or more masks. The method can include providing, by the one or more processors, in an interactive virtual environment, a virtual scene rendered using the 3D mesh and texture data generated from the frame that captures the scene in the medical procedure.
[0029] At least one aspect of the present disclosure is directed to a non-transitory computer-readable medium storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to identify a frame that captures a scene in a medical procedure performed with a robotic medical system. The instructions can cause the one or more processors to generate, using one or more models trained with machine learning, one or more masks that segment an anatomical structure in the frame, label the anatomical structure in the frame, and indicate a depth value for each pixel in the frame. The instructions can cause the one or more processors to construct 3D mesh and texture data based on the one or more masks. The instructions can cause the one or more processors to provide, in an interactive virtual environment, a virtual scene rendered using the 3D mesh and54855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) texture data generated from the frame that captures the scene in the medical procedure.
[0030] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of various aspects and implementations, and provide an overview or framework for understanding the nature and character of the claimed aspects and implementations. The drawings provide illustration and a further understanding of the various aspects and implementations, and are incorporated in and constitute a part of this specification. The foregoing information and the following detailed description and drawings include illustrative examples and should not be considered as limiting.BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings are not intended to be drawn to scale. Like reference numbers and designations in the various drawings indicate like elements. For purposes of clarity, not every component may be labeled in every drawing. In the drawings:
[0032] FIG. 1 A depicts an example computing system to generate a 3D mesh from a frame of a medical procedure video.
[0033] FIG. IB depicts an example computing system including a segmentation function, depth estimator, and mesh generator to produce data for a simulation module.
[0034] FIG. 2 depicts an example computing system to reconstruct a medical procedure environment in a 3D virtual scene.
[0035] FIG. 3 depicts an example three dimensional virtual scene initialized using a frame of a medical procedure video.
[0036] FIG. 4 depicts example outcomes produced according to different actions performed in a virtual scene initialized using a frame of a medical procedure video.
[0037] FIG. 5 depicts an example method of reconstructing a medical procedure environment in a three dimensional virtual scene.
[0038] FIG. 6 depicts an example architecture of a computing system.64855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)DETAILED DESCRIPTION
[0039] Following below are more detailed descriptions of various concepts related to, and implementations of, methods, apparatuses, and systems for 3D virtual scene reconstruction of a medical procedure environment. The various concepts introduced above and discussed in greater detail below may be implemented in any of numerous ways.
[0040] This disclosure is generally directed to generating a 3D virtual simulation of a surgical scene using an image, such as 2-dimensional image or a monocular image. A medical robotic system can be operated by a surgeon to perform a medical procedure. However, training or practicing techniques to use the medical robotic system can vary and may not achieve a realistic clinical setting outside of a live operating room. This creates a challenging demand to enable surgeons to learn skills for operating the medical robotic system and explore new surgical techniques. Furthermore, an expert may supervise the training, but the expert may have a limited amount of time to facilitate training. In some cases, in order to train or learn new skills, a less experienced surgeon (e.g., a resident) can preform sections or steps of a medical procedure under the supervision of a primary surgeon.##
[0041] In some instances, a predefined training course or predetermined or generalized computer simulation can be used by a surgeon to train. However, these generalized computer simulations may only teach basic techniques, and may be confined to generalized training exercises, and not specific or real medical scenarios. Since they are generalized, these generalized computer simulations may not focus on the exact clinical scenarios surgeons will encounter or the surgeon may need to practice.
[0042] It can be challenging or difficult for a computing system to generate or initiate a 3D virtual scene using 2D images that accurately represents a medical procedure, e.g., a scene that includes an accurate and changing anatomical state (e.g., burns, cuts, bleeding, etc.) and the physical status or state of the patient that the medical procedure is performed on. For example, without 3D scanning equipment that captures 3D images or videos of a medical procedure, it can be difficult or resource intensive for the computing system to generate or initialize the 3D virtual scene from a 2D image. Furthermore, the computing system may need a significant or large amount of parameters or simulator configuration data to produce the 3D virtual environment. The computing system may need to utilize manually produced 3D objects or 3D modeling data to produce the 3D virtual simulation.74855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0043] To solve these, and other technical problems, technical solutions of this disclosure can include 3D virtual scene reconstruction of a medical procedure environment. This technology can facilitate or automatically initiate a 3D virtual simulation that corresponds to or matches a scene in a medical procedure using an image of the scene. The computing system can generate a 3D mesh (e.g., a collection of vertices, edges, and faces) and texture file (e.g., appearance of an anatomical structure, color of the anatomical structure, roughness or smoothness of the anatomical structure, state of anatomical structures) from an image of a medical procedure, such as a monocular surgical image, and generate or initialize the 3D virtual scene using the 3D mesh and texture file. Instead of generating the 3D virtual scene from resource intensive 3D scanning, the computing system can produce the 3D information from monocular images. In this regard, the computing system can integrate with the medical robotic equipment, without needing any extra devices. Furthermore, the computing system can produce the 3D information from tissue or anatomical state data, which results in a 3D virtual scene that is highly realistic.
[0044] To achieve this 3D virtual scene generation, the computing system can use machine learning models to generate a segmentation mask, an anatomy state mask, and a metric depth mask based on the image of the scene. With these masks, the system can construct the 3D mesh and texture file that can represent a 3D version of the scene - or virtual scene. By using the masks, the computing system can produce a highly detailed and realistic 3D virtual scene. The system can use the virtual scene to establish a 3D simulation in an interactive virtual environment. Thus, this technology can initiate a 3D virtual simulation that represents the scene in the medical procedure that is captured in a particular frame (e.g. a 2D monocular image).
[0045] Furthermore, the 3D virtual scene can be constructed from at least one image of the virtual scene. For example, a surgeon can select which image to reconstruct the 3D virtual scene from, and the computing system can generate the 3D virtual scene using the selected image. In this regard, the computing system can reproduce any clinical scenes according to the preferences of the surgeon in an agile manner. This virtual scene initialization at any moment or step of the medical procedure, which results in re-created the real-life surgical scene, provides valuable training advantages for surgeons, allowing the surgeons to retry or reperform the medical procedure in a virtual environment. These recreated scenes can allow surgeons to have access to clinically accurate training scenarios,84855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) and explore of new techniques in safe environment, such as pre-operative practice cases. The computing system can render the surgical scene in a user-specified perspective which can allow a surgeon to train, improve, explore techniques, and plan their next operation within this clinically derived perspective.
[0046] The system can allow a user to interact with the 3D virtual simulation in various modes. For example, the system can provide a tutorial mode, a competition mode, or a generative mode. The tutorial mode can limit, constrain, or prohibit certain types or amounts of input (e.g., the user may be required to follow the same trajectory made by an instrument during the actual medical procedure). The competition mode may loosen certain constraints to allow the user, in the simulation, to try to match or outperform performance metrics relative to the medical procedure. In the generative mode, the system can use generative artificial intelligence techniques to initiate a virtual simulation that is based on the frame of the scene in the medical procedure as well as other parameters. For example, in the generative mode, the system can change patient parameters, anatomical structure parameters (e.g., amount of bleeding, textures, or state), etc.
[0047] Thus, this technical solution can generate a virtual simulation that can be photorealistic using machine learning and generative Al, thereby providing an improved interactive virtual simulation that can represent an actual scene in a medical procedure.
[0048] Referring now to FIG. 1 A, among others, an example system 100 for providing a 3D view reconstruction of a surgical scene using a monocular image frame. System 100 can include one or more medical environments 102 in which medical procedures (e.g., robotic surgeries) can be performed. A medical environment 102 can include one or more sensors 104 for detection of various sensor data 174 and one or more data capture devices 108 (e.g., cameras) for capturing data, such as image frames 121 (e.g., monocular images). The medical environment 102 can include one or more visualization tools 114 for facilitating data visualization and one or more displays 116 for displaying data. The medical environment can include one or more robotic medical systems (RMSs) 110 used by users (e.g., surgeons) to perform medical surgeries on a patient. The users can utilize one or more head mounted devices (HMDs) 122 that can include integrated one or more displays 116 and sensors 104.
[0049] The devices of the medical environment 102 can communicate with a computing system 105 via a network 101. The computing system 105 can include one or more frame functions 120 configured to receive and process image frames 121 (e.g., 2D monocular94855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) images). Computing system 105 can include one or more segmentation functions 122 for generating masks 124 to identify anatomy segments 128 and provide labels 126. The computing system 105 can include one or more depth estimators 125 for generating depth values 132. The computing system 105 can include one or more mesh generators 130 for generating 3D meshes 136, using for example, one or more inpainting functions 134. The computing system 105 can include one or more trigger condition detectors 144 for detecting trigger conditions 146 (e.g., milestone events during the surgical procedure) based on one or more thresholds 148. Using the 3D mesh 136, a multi-view generator 140 can generate one or more view-angle frames 142 that can depict the scene of the frame 121 from different view angles. The computing system 105 can include or more 3D video generators 150 for generating videos 152 (e.g., 3D video multi-angle view simulations) from the 3D meshes 136 and the view-angle frames 142. The computing system 105 can include one or more prompt generators 154 for generating prompts 156 for using one or more machine learning (ML) models 182. The computing system 105 can include or more data repositories 160 storing surgeon data 172 and one or more data streams 162 that can include sensor data 174 and video data 178. The computing system 105 can include one or more machine learning (ML) frameworks 180. An ML framework 180 can include one or more of: ML models 182, ML trainers 184, label annotators 186 and encoder and decoder functions 190. The computing system 105 can include one or more user interfaces 192, such as graphical user interfaces for providing user presentations or outputs to the users of the RMS 110.
[0050] The RMS 110 can include any robotic system that can utilize or manipulate medical instruments 112 to perform medical procedures, such as a robotic surgery. Robotic medical system 110, also referred to as an RMS 110, can be deployed in any medical environment 102. A medical environment 102 can include any space or facility for performing medical procedures (e.g., robotic surgeries), including for example any surgical facility, or an operating room. A medical environment 102 can include any number of different medical instruments 112 that the RMS 110 can use for performing surgical patient procedures, whether invasive, non-invasive, in-patient, or out-patient procedures.
[0051] The medical environment 102 can include one or more data capture devices 108 (e.g., optical devices, such as cameras or sensors or other types of sensors or detectors) for capturing data streams 162, that can include video data 178 of images or a video stream of a surgery as well as any sensor data 174. The medical environment 102 can include one or more visualization tools 114 to gather the captured data streams 162 and process it for display104855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) to the user (e.g., a surgeon or other medical professional) at one or more displays 116. A display 116 can present data stream 162 (e.g., video frames, kinematics or sensor data) of an ongoing medical procedure (e.g., an ongoing surgery) performed using the robotic medical system 110 handling, manipulating, holding or otherwise utilizing medical instruments or tools 112 to perform surgical tasks at the surgical site.
[0052] Data repository 160 can include various data streams 162 generated by the RMS 110, including various sensor data 174 and video data 178 that can be collected, organized, stored and provided for use to various data processing system components (e.g., segmentation function 122, depth estimator 125, mesh generator 130, multi-view generator 140, trigger condition detector 144, 3D video generator 150, prompt generator 154, or any of the ML models 182 that can be utilized). For instance, ML framework 180 can use data streams 162 as inputs into one or more ML models 182, such as ML models 182 trained to generate masks 124 for segmenting anatomy segments 128 (e.g., anatomical structure) from image frames 121 or ML models 182 trained to generate view-angle frames 142 from a plurality of view angles. ML framework 180 can include one or more ML model trainers 184 for training the ML models along with attention mechanisms 188 that can be utilized by the ML models for detection of anatomies, instruments, various instrument to anatomy interactions, medical procedure tasks, anatomy segments 128, depth values 132, trigger conditions 146, view-angle frames 142, 3D meshes 136 or any other features represented or captured in the image frames 121. ML framework 180 can include one or more encoder and decoder functions 190 for processing and detection of image features and for processing data streams 162 for data relevant to ML framework determinations.
[0053] Machine learning (ML) framework 180 can include any combination of hardware and software for providing a system that integrates ML models 182 with attention mechanisms 188 or rule-based modeling to generate masks 124 for labeling of anatomy segments 128, generating view-angle frames 142, synthesize new pixels to add (e.g., for inpainting function 134) and generate prompts 156 for a prompt generator 154 to new pixel synthesis (e.g., inpainting). ML trainers 184 can be used for training, setting or configuring ML models 182 and their related functions or components, such as label annotators 186, attention mechanisms 188 or encoder and decoder functions 190.
[0054] ML framework 180 can include ML models 182 trained, set or configured to perform variety of tasks for 3D view reconstruction. ML models 182 can include any type and form of artificial intelligence (Al) models, such as neural network models, transformer-114855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) based mechanism models, any graph neural network. For example, ML models 182 can be trained or configured to generate masks 124 for labeling of anatomy segments 128 via labels 126 on behalf of a segmentation function 122. ML models 182 can be trained or configured to generate prompts 156 on behalf of a prompt generator 154 to cause other one or more ML models 182 to synthesize new pixels to match anatomy segments 128 of areas (e.g., holes) in a multi view-angle frame 142 generated by multi-view generator 140. The ML models 182 can be configured to generate view-angle frames 142 or synthesize new pixels to add into view-angle frames 142 (e.g., on behalf of an inpainting function 134). The ML models 182 can be configured to generate view angle frames 142 for a plurality of view angles (e.g., perspectives) surrounding the view-angle of the monocular image frame 121, allowing for generation of a 3D video 152, allowing the user to view a surgical scene from multiple viewangle perspectives.
[0055] ML models 182 can include any type of ML or Al architecture for processing images and generating multi-view angle images of a surgical scene, or for filling in missing data from 2D image. ML models 182 can include, for example, neural networks, such as Convolutional Neural Networks (CNNs) that can be configured to recognize and reconstruct spatial features (e.g., anatomy segments 128) from frames 121. ML models 182 can include self-attention mechanisms, such as those in transformer models, which can be configured to enhance image generation by focusing on different parts of the image frame 121. ML models 182 can include Generative Adversarial Networks (GANs) that can be configured to create realistic images (e.g., or portions of images) from incomplete data (e.g., synthesize missing pixels) using a generator and a discriminator. ML models 182 can include Variational Autoencoders (VAEs) that can generate new images, or portions of new images (e.g., fill in the holes in the view-angle frames 142) by learning the distribution of input data and sampling from it. For 3D reconstructions, 3D-CNNs and voxel-based functionalities can be used to process volumetric data to create accurate 3D models from 2D images. For instance, ML models 182 can include or utilize Neural Radiance Fields (NeRF) that can synthesize new regions of images (e.g., fill in the holes) by processing or optimizing a volumetric scene function (e.g., 3D mesh). ML model 182 can include or use a Reinforcement Learning (RL) that can improve model accuracy in predicting desired or optimal angles for multi-view reconstruction.
[0056] ML framework 180 can include attention mechanisms 188, implemented as neural networks, which enable the extraction of spatial and temporal features from the input data124855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) steams 162. Attention mechanisms 188 can facilitate or improve the capacity of the ML models to discern, detect or recognize specific details within the surgical context, thereby improving the accuracy of detection and recognition tasks. ML framework 180 can include and provide rule-based modeling to determine and quantify the consistency of anatomical features or segments along an area. ML framework 180 can include and provide a framework for generating a 3D mesh of various anatomy segments 128 that are labeled using labels 126 and for inpainting empty spaces (e.g., holes) generated by view-angle movement over obstructing objects, such as by using the tissue texture (e.g., color and shapes) of areas of the same anatomy (e.g., organ or tissue type) surrounding the hole. By integrating encoder and decoder functions 190 for extracting image features or time-series for using kinematics data and sensor measurements (including force data), the ML framework 180 improves the quality of the detection and recognition by the ML models. For example, the ML framework 180 can utilize attention mechanisms 188 to focus on the texture, color or shape patterns of the surround tissue of the same anatomy segment 128 that is to be inpainted or filled-in by synthesized pixels to be generated for the 3D mesh 136 or view-angle frames 142.
[0057] Encoder and decoder functions 190 can include any combination of algorithms and neural network architectures designed for transforming and reconstructing data representations. Encoder and decoder functions 190 can be utilized in by segmentation function 122 for anatomy segmentation and by depth estimator 125 for metric depth estimation. The encoder functionality of the encoder and decoder functions 190 can process the image frame 121, compressing its information into a lower-dimensional latent space, capturing features such as anatomical structures or depth information. The decoder function of the encoder and decoder functions 190 can take the compressed representation and reconstruct it back into a detailed output, such as a segmentation mask 124 or a metric depth mask 138. Encoder and decoder functions 190 can be combined with transformer units, allow for identifying relationships within the data.
[0058] Using the ML framework 180, the technical solutions can provide RMS 110 users with a spatial awareness intra-operatively, thereby facilitating a greater understanding than is achievable with 2D monocular image frames 121. By using 2D images (e.g., frames 121) to capture moments in real time, the technical solutions allow for the reconstruction of the 3D scenery as the 2D image recording continues. The computing system 105 can provide for rendering of a 2D monocular image into a plurality of view-angle frames 142 (e.g., based on user input device selections or movements in a user interface 192), or for a 3D view video134855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)152 that includes viewing angle change in multiple spatial directions around the original viewing angle (e.g., via changes to spatial Cartesian coordinates X-Y-Z) and attention point zoom-in / out. The 2D image of the scene can be used as an input to a modularized pipeline. The pipeline can include an input framing in which the system can render a single 2D image from surgery into 3D. The system can sample the selected frame from the video recording and pass through one or more modules (e.g., computing system 105 components).
[0059] Data repository 160 of the computing system 105 can include one or more data streams 162, such as video data 178 including a stream of video frames. Data streams 162 can include measurements from sensors, which can be referred to as sensor data 174 and which can include various force, torque or biometric data, haptic feedback data, pressure or temperature data, vibration, tension or compression data, endoscopic images or data, ultrasound images or videos or communication and command data streams. Data repository 160 can include installation data, such as system files or logs including time stamps and data on installation, activation, calibration or use of particular medical instruments 112.
[0060] The system 100 can include one or more data capture devices 108 (e.g., video cameras, sensors or any other detectors) for collecting any data stream 162. Data capture devices 108 can include cameras or other image capture devices for capturing video data 178 (e.g., videos or images) from a particular viewpoint within the medical environment 102. The data capture devices 108 can be positioned, mounted, or otherwise located to capture content from any viewpoint that facilitates the computing system 105 capturing various surgical tasks or actions.
[0061] Data capture devices 108 can include any of a variety of sensors, cameras, video imaging devices, infrared imaging devices, visible light imaging devices, intensity imaging devices (e.g., black, color, grayscale imaging devices, etc.), depth imaging devices (e.g., stereoscopic imaging devices, time-of-flight imaging devices, etc.), medical imaging devices such as endoscopic imaging devices, ultrasound imaging devices, etc., non-visible light imaging devices, any combination or sub-combination of the above mentioned imaging devices, or any other type of imaging devices that can be suitable for the purposes described herein. Data capture devices 108 can include cameras that a surgeon can use to perform a surgery and observe manipulation components within a purview of field of view suitable for the given task performance.144855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0062] Data capture devices 108 can capture, detect, or acquire sensor data, such as videos or images, including for example, still images (e.g., monocular 2D images from a single-view angle), video images, vector images, bitmap images, other types of images, or combinations thereof. The data capture devices 108 can capture the images at any suitable predetermined capture rate or frequency. Settings, such as zoom settings or resolution, of each of the data capture devices 108 can vary as desired to capture suitable images from any viewpoint. For instance, data capture devices 108 can have fixed viewpoints, locations, positions, or orientations. The data capture devices 108 can be portable, or otherwise configured to change orientation or telescope in various directions. The data capture devices 108 can be part of a multi-sensor architecture including multiple sensors, with each sensor being configured to detect, measure, or otherwise capture a particular parameter (e.g., sound, images, or pressure).
[0063] Data capture devices 108 can include any type and form of a sensor for providing sensor data 174, including a positioning sensor, a biometric sensor, a velocity sensor, an acceleration sensor, a vibration sensor, a motion sensor, a pressure sensor, a light sensor, a distance sensor, a current sensor, a focus sensor, a temperature sensor, a haptic or tactile sensor or any other type and form of sensor used for providing data on medical tools 112, or data capture devices (e.g., optical devices). For example, a data capture device 108 can include a location sensor, a distance sensor or a positioning sensor providing coordinate locations of a medical tool 112 or a data capture device 108. Data capture device 108 can include a sensor providing information or data on a location, position or spatial orientation of an object (e.g., medical tool 112 or a lens of data capture device 108) with respect to a reference point. The reference point can include any fixed, defined location used as the starting point for measuring distances and positions in a specific direction, serving as the origin from which all other points or locations can be determined.
[0064] Display 116 can show, illustrate or play data streams 162, including video data 178, in which medical tools 112 at or near surgical sites are shown. For example, display 116 can display a rectangular image (e.g., a frame of a video data 178) of a surgical site along with at least a portion of medical instruments 112 being used to perform surgical tasks. Display 116 can provide compiled or composite images generated by the visualization tool 114 from a plurality of data capture devices 108 to provide visual feedback from one or more points of view.154855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0065] The visualization tool 114 that can be configured or designed to receive any number of different data streams 162 from any number of data capture devices 108 and combine them into a single data stream displayed on a display 116. The visualization tool 114 can be configured to receive a plurality of data stream components and combine the plurality of data stream components into a single data stream 162. For instance, the visualization tool 114 can receive a visual sensor data from one or more medical tools 112, sensors or cameras with respect to a surgical site or an area in which a surgery is performed. The visualization tool 114 can incorporate, combine or utilize multiple types of data (e.g., positioning data of a medical tool 112 along sensor readings of pressure, temperature, vibration or any other data) to generate an output to present on a display 116. Visualization tool 114 can present locations of medical tools 112 along with locations of any reference points or surgical sites, including locations of anatomical parts of the patient (e.g., organs, glands or bones).
[0066] Medical instruments or tools 112 can be any type and form of tool or instrument used for surgery, medical procedures or a tool in an operating room or environment. Medical tool 112 can be imaged by, associated with or include an image capture device. For instance, a medical tool 112 can be a tool for making incisions, a tool for suturing a wound, an endoscope for visualizing organs or tissues, an imaging device, a needle and a thread for stitching a wound, a surgical scalpel, forceps, scissors, retractors, graspers, or any other tool or instrument to be used during a surgery. Medical tools 112 can include hemostats, trocars, surgical drills, suction devices or any instruments for use during a surgery. The medical tool 112 can include other or additional types of therapeutic or diagnostic medical imaging implements. The medical tool 112 can be configured to be installed in, coupled with, or manipulated by an RMS 110, such as by manipulator arms or other components for holding, using and manipulating the medical instruments or tools 112.
[0067] RMS 110 can be a computer-assisted system configured to perform a surgical or medical procedure or activity on a patient via or using or with the assistance of one or more robotic components or medical tools 112. RMS 110 can include any number of manipulator arms for grasping, holding or manipulating various medical tools 112 and performing computer-assisted medical tasks using medical tools 112 controlled by the manipulator arms.
[0068] Video data 178, including any images or videos captured by a medical tool 112 (e.g., endoscopic camera) can be sent to the visualization tool 114. The robotic medical system 110 can include one or more input ports to receive direct or indirect connection of one164855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) or more auxiliary devices. For example, the visualization tool 114 can be connected to the RMS 110 to receive the images from the medical instrument 112 when the medical instrument 112 is installed in the RMS 110 (e.g., on a manipulator arm of the RMS 110 that is used for moving, managing or otherwise handing medical instruments 112). The visualization tool 114 can combine the data streams 162 from the data capture devices 108 and the medical tool 112 into a single combined data stream 162 for use by the ML framework 180 (e.g., ML models 182 or associated attention mechanisms 188 and encoder and decoder functions 190).
[0069] Computing system 105 can be deployed in, communicatively coupled with, or otherwise associated with any of the components (e.g., devices) of the medical environment 102 directly or via a network 101. Computing system 105 can be provided by a remote server (e.g., connected to medical environment 102 via a network 101), or can be deployed or provide via a cloud-based service or a function. The computing system 105 can include an interface 192 designed, constructed and operational to communicate with one or more component of system 100 via network 101, including, for example, the robotic medical system 110. Computing system 105 can be implemented using instructions stored in memory locations and processed by one or more processors, controllers or integrated circuitry.Computing system 105 can include functionalities, computer codes or programs for executing or implementing any functionality of ML framework 180, including any ML models 182. The computing system 105, as well as any of its components can each be a part of or include a cloud computing environment functionality or features. The computing system 105 can include multiple, logically grouped servers and facilitate distributed computing techniques. The logical group of servers may be referred to as a data center, server farm or a machine farm. The servers can also be geographically dispersed. A data center or machine farm may be administered as a single entity, or the machine farm can include a plurality of machine farms. The servers within each machine farm can be heterogeneous - one or more of the servers or machines can operate according to one or more type of operating system platform.
[0070] The machine learning (ML) trainer 184 can include any combination of hardware and software for training ML models 182 to perform their designated or trained operations. For instance, the ML trainer 184 can include, train, configure, generate or adjust (e.g., retrain) any ML models 182. The ML trainer 184 can use the training datasets that can include any selection or collection of data streams 162 corresponding to various medical procedures using the RMS 110. ML trainer 184 can include a framework or functionality for training different174855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) machine learning models 182, such as a neural network spatial attention mechanisms models for determining spatial orientations of various components or features in the image frames 121. The ML trainer can train neural network spatial-temporal attention mechanism models 182, for determining features from image frames 121 having timestamps indicative of preceding or following image frames 121 (e.g., providing different view angles of the area). The ML trainer 184 can train the ML model designed for detecting medical instruments 112 as well as detecting anatomical parts of a patient (e.g., anatomy segments 128) using various data from data streams 162, including video data 178 or sensor data 174 (e.g., temperature, pressure, proximity or force data).
[0071] The ML trainer 184 can include, utilize, implement or provide an attention mechanism 188 that can be used to address the noise challenges in the data. Attention mechanism 188 can include a neural network with spatial attention or spatial -temporal attention that can be performed or implemented using an encoder and decoder function 190 to learn to identify various anatomical features or segments. Attention mechanism 188 can include a spatial attention or spatial-temporal attention that can be performed to identify tasks or phases of a medical procedure to detect trigger conditions 146 (e.g., milestone events) based on the type of medical instruments 112 and types of tissues (e.g., anatomy segments 128) involved in a particular order of operation. Attention mechanism 188 can utilize weights to emphasize different types of information in the data stream 162, such as movements in a region of an image frame that corresponds to a prior image frame in which a particular type of movement was detected. Such spatial and temporal weights used in the attention mechanism 188 can facilitate an improved or selective focus of the ML functions onto particular features in the data stream 162, assigning varying degrees of importance to each part of the input data during the learning process. For example, the attention mechanism 188 can include or utilize a neural network architecture that configures an ML model to focus selectively on relevant spatial features in the video data, thereby improving the accuracy of the detection. By assigning weights to selected segments of the input videos data 178, the attention mechanism 188 can allow the model to attenuate the impact less relevant portions of the data, emphasizing the importance of the more relevant cues (e.g., focus on a detected medical instrument 112 or a detected anatomical tissue of a patient) for more accurate determinations.
[0072] Interface 192 can include any combination of hardware and software for interfacing with a user of the robotic medical system 110. Interface 192 can include184855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) components or features designed, constructed and operational to communicate with one or more component of system 100 via network 101, including, for example, the RMS 110 or another device, such as a client’s personal computer. The interface 192 can include a network interface. The interface 192 can include or provide a user interface, such as a graphical user interface. The graphical user interface can include, for example, a window for displaying video data 178, or any indications or outputs to be provided or displayed with the end user. Interface 192 can provide data for presentation via a display, such as a display 116, and can depict, illustrate, render, present, or otherwise provide indications indicating determinations (e.g., outputs) of the ML models 182.
[0073] Interface 192 can be configured to utilize graphical user elements for utilization or operation of a computing system 105 functions or operations. Graphical user elements can include any graphical features, attributes, user selections or commands that can be utilized to receive user commands or controls for utilizing a multi view-angle frame 142 or a 3D presentation, such as a video 152. Graphical user elements can include user input device (e.g., mouse selections) that can be used to indicate movement from the original monocular image frame 121 view-angle to another view-angle frame 142 to be generated, responsive to the user selection of the graphical user element. For instance, interface 192 can receive, via a graphical user interface element, an input with a range of view angles (e.g., 142) for which to generate the video 152. The interface 192 can provide the output (e.g., of the desired viewangle for which to generate a view-angle frame 142) and trigger the computing system 105 to generate the video 152 (e.g., via 3D video generator 150) corresponding to the plurality of view angles and responsive to the input with the range of view angles. The interface 192 can display the video 152 or one or more generated view-angle frames 142 (e.g., generated responsive to the user selection) via a display 116. Interface 192 can operate on a client device or head mounted device 106 and can be used to display the view-angle frames 142 or the 3D videos 152 via a display 116 on the head mounted device 106.
[0074] System 100 can include a head mounted device 106, also referred to as an HMD 106, that can include any combination of hardware and software for using any features of the computing system 105 or the RMS 110. HMD 106 can include any combination of a display 116 and sensors 104, allowing a surgeon to perform or view surgery via a robotic medical system 110. The HMD 106 can include a display 116 for providing an immersive user interface 192, displaying multi-view angle frames 142 and 3D video reconstructions (e.g., videos 152) of the surgical scene. The sensors 104 can include hand tracking sensors or194855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) devices or eye tracking sensors or devices to detect user motions or selections. The display 116 can present high-resolution, real-time visuals. The integrated sensors 104 can track the surgeon's head movements and adjust the perspective accordingly. The HMD 106 can incorporate gesture and voice recognition functionalities, allowing the surgeon to control the user interface 192 of the HMD 106 hands-free.
[0075] The computing system 105 can interface with, communicate with, or otherwise receive or provide information with one or more component of system 100 via network 101. The computing system 105, RMS 110 and devices in the medical environment 102 can each include at least one logic device such as a computing device having a processor to communicate via the network 101. The computing system 105, any portion of the ML framework 180, the RMS 110 or a client device that can be communicatively coupled with the DPS or the RMS 110 via the network 101, can each include at least one computation resource, server, processor or memory for processing data. For example, the computing system 105 can include a plurality of computation resources or processors coupled with memory.
[0076] The network 101 can be any type or form of a medium for facilitating communication between devices or systems, such as the computing system 105 and the devices in a medical environment 102. The geographical scope of the network can vary widely and can include a body area network (BAN), a personal area network (PAN), a localarea network (LAN) (e.g., Intranet), a metropolitan area network (MAN), a wide area network (WAN), or the Internet. The topology of the network 101 can assume any form such as point-to-point, bus, star, ring, mesh, tree, etc. The network 101 can utilize different techniques and layers or stacks of protocols, including, for example, the Ethernet protocol, the internet protocol suite (TCP / IP), the ATM (Asynchronous Transfer Mode) technique, the SONET (Synchronous Optical Networking) protocol, the SDH (Synchronous Digital Hierarchy) protocol, etc. The TCP / IP internet protocol suite can include application layer, transport layer, internet layer (including, e.g., IPv6), or the link layer. The network 101 can be a type of a broadcast network, a telecommunications network, a data communication network, a computer network, a Bluetooth network, or other types of wired and wireless networks.
[0077] Frame function 120 can include any combination of hardware and software for receiving and processing image frames 121. Frames 121 can include any monocular images, such as 2D images from a single view-angle. A frame 121 can include an image of a surgical204855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) scene capturing or depicting one or more organs or tissues along with one or more medical instruments 112. The depicted surgical scene can be a 3D volumetric space having one or more objects (e.g., tissues, organs, glands or medical instruments 112) that are obstructing one another, providing one or more holes (e.g., obstructed areas that cannot be viewed from the given monocular image frame 121).
[0078] Frame function 120 can be configured to identify a frame 121 capturing a single view angle of a surgical scene in a medical procedure (e.g., surgical operation) that can be performed with a robotic medical system 110. The frame function 120 can receive a data stream 162 of the medical procedure captured via one or more sensors 104 of the robotic medical system 110. The frame function 120 can capture the data stream 162 from a repository 160. The frame function 120 can determine to generate one or more multi-view frames, such as view-angle frames 142 for a multi view-angle presentation, such as a 3D video. The frame function can determine to generate, or can trigger, command or start the generation of a multi view-angle presentation in which the scene captured by the frame(s) 108 can be displayed (e.g., via user interface 192) in a still format and from multiple viewangles that can be controlled or selected based on a selection (e.g., mouse movement or a click) of a user input device controlling the view-angles of the multi view-angle presentation.
[0079] Prompt generator 154 can include any combination of hardware and software for generating one or more prompts for one or more ML models 182. Prompts 156 can include any combination of parameters, values or instructions for directing or configuring one or more ML models 182 to implement particular (e.g., intended) functionality. Prompt generator 154 can generate any prompts 156 for any ML models 182. For instance, prompt generator 154 can generate prompts 156 to direct ML models 182 to generate view-angle frames 142 or synthesize new pixels to add into view-angle frames 142 (e.g., on behalf of an inpainting function 134). For instance, prompt generator 154 can generate prompts 156 to direct ML models 182 to generate view angle frames 142 for a plurality of view angles (e.g., perspectives) surrounding the view-angle of the monocular image frame 121, allowing for generation of a 3D video 152, allowing the user to view a surgical scene from multiple viewangle perspectives.
[0080] The prompt generator 154 can generate a prompt 156 based on the segments (e.g., anatomy segments 128) or based on labels 126 of a particular mask 124. The prompt generator 154 can input the prompt 156 into an ML model 182 that is a generative machine learning model configured to synthesize new pixels, to trigger the ML model 182 to fill the214855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) one or more holes (e.g., one or more areas with the new pixels synthesized by a generative ML model 182). The prompt generator 154 can generate the prompts 156 based on the masks 124 that can indicate the distance between pixels to be inpainted in the frame (e.g., viewangle frame 142) and the reference point. For example, when utilizing an inpainting function 134 to fill in the empty holes or areas of the view-angle frames 142 in which no pixel data exists (e.g., where there is no visibility in a new view-angle frame 142), the prompts generator 154 can generate the prompt 156 to cause the ML model 182 to fill in the pixels in the hole (e.g., empty area) of the view-angle frame 142 based (e.g., at least in part) on the distance between each of the pixels to be filled or inpainted and a reference point (e.g., one or more pixels in the original frame 121 or a view-angle frame 142).
[0081] Anatomy segments 128 can include any segments of an anatomy of a patient on whom a medical procedure is performed. Anatomy segments 128 can include, for example, segments of one or more body parts, organs, tissues, bones, blood vessels, nerves, muscles, joints, ligaments, tendons, cartilage, glands, and connective tissues. Anatomy segments 128 can include areas such as a cardiovascular system, respiratory system, digestive system, nervous system, musculoskeletal system, and endocrine system. The identification and visualization of these anatomy segments 128 can be done by the segmentation function or ML models in precise surgical planning, navigation, and execution during medical procedures.
[0082] Segmentation masks 124, also referred to as segment masks 124 or masks 124, can include any region, area or a map corresponding to, or identifying, any distinct anatomical area, segment or a region within an image frame. Mask 124 can be defined as binary or multi-class maps that delineates different anatomy segments 128 from each other, such as distinguishing a blood vessel from a fat tissue. Mask 124 can be used for labeling and distinguishing different regions within an image. For example, in a surgical scene, a segmentation mask 124 can identify areas corresponding to the liver, kidneys, blood vessels, and other critical organs or tissues. The areas indicated by the segmentation masks 124 can be labeled using labels 126, allowing ML models 182 to distinguish one anatomy segment 128 from another. For instance, each pixel, or a group of pixels, in the segmentation mask 124 can be assigned a class label 126, indicating the anatomical structure (e.g., anatomy segment 128) to which it belongs. For instance, pixels representing the liver can be assigned a first color, indication, or a label, while pixels representing the kidneys can be assigned a second color, indication or a label 126.224855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0083] Segmentation function 122 can include any combination of hardware and software for generating masks 124 for segmenting and labeling anatomy segments 128 using labels 126. Segmentation function 122 can generate one or more segmentation masks 124 for identifying, distinguishing or labeling different anatomy segments 128. For instance, a segmentation function 122 can utilize ML models 182 to segment, identify or distinguish different anatomy segments 128 from each other and label those segments based on their type. Segmentation function 122 can generate, using one or more ML models 182 trained with machine learning, one or more masks 124 that segment an anatomical structure in the frame 121 or view-angle frame 142. The segmentation function 122 can use one or more ML models 182 to label the anatomical structure in the frame (e.g., 108 or 142) using one or more labels 126. For instance, a segmentation function 122 can generate a mask 124 using a vision transformer based neural network (e.g., ML model 182) trained with a dataset comprising images of medical procedures that are annotated with labels 126. The mask 124 can include the anatomy segment 128 of the anatomical structure (e.g., a particular instance or a type of anatomy segment 128) and the label 126 of the anatomical structure (e.g., for the same anatomy segment 128).
[0084] Segmentation function 122 can include or utilize an ML model 182 having or using a vision transformer based neural network that can be configured and trained with human annotated clinical image dataset. For instance, an ML model 182 can be trained using labels 126 labeling various anatomy segments 128. Given an input image frame 121, an ML model 182 can generate a segmentation mask 124 along with any labels for each of one or more anatomy segments 128 within the image frame 121. The ML model 182 can label the segmentation mask 124 to generate a segmentation map. The segmentation map can include a spatial representation (e.g., a 2D representation) of locations of various anatomy segments 128 along with any labels 126 of such anatomy segments 128.
[0085] Label annotator 186 can include any function (e.g., application or process) for assigning labels to any segmentation masks 124 or depth masks 138. Label annotator 186 can include, for example, an ML model 182 trained to annotate or assign labels 126 to any segmentation masks 124 or depth masks 138. Label annotator 186 can include, for example, a generative Al model trained to generated information, metadata or otherwise labels 126 describing the type of anatomical structure and their depth or location with respect to the lens of the data capture device 108.234855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0086] Labels 126 can include any metadata providing or indicating information about a segmentation mask 124 corresponding to an anatomy segment 128 (e.g., name or type of tissue or an organ). Label 126 can include any metadata providing or indicating information corresponding to a depth mask 138, including any depth values 132 (e.g., distance of a given one or more pixels of a depth mask value 132 from a reference point, such as a location of a lens of the camera capturing the frame 121). Label annotator 186 can generate the labels 126 for any of the depth masks 138 or anatomy masks 124 using any functionality of a ML framework 180, including any ML models 182, attention mechanisms 188 or encoder and decoder functions 190. Labels can include “state”, such as a state of a liver being burned. The label 126 for a pixel can be used to generate a “burned” liver representation or a depiction (e.g., by an inpainting function 134). States can include various tissue or environment states, such as bleeding, burned or cut in a bleeding area.
[0087] Depth estimator 125 can include any combination of hardware and software for determining the actual depth or distance of objects or features within an image. The depth estimator 125 can include any functionality for generating depth masks 138 and their corresponding depth values 132. The depth estimator 125 can generate labels 126 for the depth masks 138. For instance, the depth estimator 125 can utilize ML models 182 (e.g., label annotator 186) to assign labels 126 to various depth values 132 of the depth masks 138. For instance, the depth estimator 125 can generate a depth mask 138 having a depth value 132 for each of the pixels of the 2D image frame 121.
[0088] Depth estimator 125 can use a 2D image frame 121 as its input into a neural network ML model 182, which can be configured as an encoder-decoder with a transformer unit, to map the image frame 121 to its metric depth values 132. The depth estimator 125 can generate a depth map for one or more depth masks 138 of each frame 121. Each pixel intensity in the resulting metric depth map can represent a real metric distance from the camera center to a point in the scene (e.g., the image frame 121). The ML model 182 can use camera intrinsic parameters for calibration during training, ensuring accurate depth estimation. The output depth map can be color-coded, with different colors or shades indicating various distances. This metric depth information can be used by the mesh generator 130 to generate a 3D mesh of the scene, allowing the creation of 3D representations (e.g., videos 152 or one or more depicted view-angle frames 142) along with new view angles and the filling (e.g., via inpanting) of any holes using generative Al.244855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0089] For example, a depth estimator 125 can provide a metric depth estimation that can include a distance for each of the one or more pixels of the 2D frame 121 from the lens of the camera to that one or more pixels. The depth estimator 125 can be configured as a neural network ML model 182 that maps a frame 121 to its metric depth. Each pixel intensity of the depth mask 138 can include a depth value 132 that represents a real metric distance between a reference point of the camera (e.g., the camera lens center) to the point of the scene. The neural network ML model 182 can include an encoder and decoder function 190 with a transformer unit. The output of the ML model 182 can include a metric depth map with depth values 132 for each of the pixels of the image frame 121. The camera intrinsic can be used during training for calibration purpose. The output from the ML model 182 can represent the distance using color coding, such that different distances are colored or shaded. Each coloring or shade intensity can represent a real distance from the shaded region (e.g., one or more pixels) to the lens of the data capture device 108 that has captured the 2D image frame 121. The metric depth information can be used to map or feed into a neural network ML model for generating the 3D mesh 136.
[0090] For example, a depth estimator 125 can generate a depth mask 138 using a neural network (e.g., ML model 182) trained to map an intensity of each pixel in the frame to the depth value 132. The depth value 132 can be a distance between a camera that captures the frame 121 and a location in the scene corresponding to the pixel of the frame 121. The depth estimator 125 can generate the depth mask 138 based on at least one of an optical center of the camera, a focal length of the camera, a scale factor of the camera, a principal point of the camera, a skew of the camera, or a geometric distortion of the camera.
[0091] Mesh generator 130 can include any combination of hardware and software for generating one or more 3D meshes 136. Mesh generator 130 can construct a 3D mesh 136 of a scene captured by an image frame 121 based on one or more segmentation masks 124 indicating or labeling anatomy segments 128 and depth masks 138 indicating or labeling depth values 132 of the various one or more pixels (e.g., individual pixels or groups of pixels) that can be color coded or shaded based on their metric distance from the camera. For example, the mesh generator 130 can create a 3D model of the patient's internal anatomy structures, allowing surgeons to visualize the exact locations and relationships of various anatomy segments 128 (e.g., organs or tissues). This 3D model representation in the 3D mesh 136 can be implemented respect to a frame of reference of a data capture device 108254855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)(e.g., the camera that captured the image frame 121). The 3D mesh 136 can reflect relative locations of various anatomy segments 128 from each other. This can greatly aid in preoperative planning, intraoperative navigation, and postoperative analysis. As 3D mesh 136 provides relative depth and spatial orientation of the internal structures in the frame 121, the 3D mesh 136 can be used by the multi -view generator 140 to generate view-angle frames 142 from various view-angles other than the original view-angle at which the 2D image frame 121 was captured.
[0092] 3D mesh 136 can include any 3D representation of a scene captured in an image frame 121. 3D mesh can be a graphic representation of a 3D area, indicating any anatomy segments 128 or depth values 132 with respect to such segments. 3D mesh can be or include a file having one or more vertices, edges and faces that provide a three-dimensional representation of the scene from the 2D image frame 121. The 3D mesh 136 can represent various anatomy segments 128 as well as their distances (e.g., depth values 132) from the lens location. The 3D mesh can be used to generate multiple view-angle frames 142 aside from the original view-angle of the 2D image frame 121. For instance, a user of the computing system 105 (e.g., a surgeon using a user interface 192) can rotate and view the 3D model of the imaged scene from different angles to gain a better understanding of complex anatomical relationships and to identify potential obstructions or satisfy trigger conditions 146.
[0093] For example, a mesh generator 130 can a mesh data to form the 3D mesh 136. The mesh data can include, provide or operate as a graphic representation of the 3D scene captured in the image frame 121. The mesh generator 130 can generate the 3D mesh 136 (e.g., a 3D mesh file) based on the depth map (e.g., depth mask 138 with the depth values 132), which can include a collection of vertices, edges, and faces that represent the shape and structure of a 3D scene. The 3D mesh 136 can allow for the representation of complex 3D shapes and surfaces by specifying the coordinates of vertices and the connectivity between them, which can be utilized by the multi -view generator 140 to generate view-angle frames 142 of the same scene from various view-angles. The 3D mesh 136 can be represented as a 2D color image according to its distance location.
[0094] Multi -view generator 140 can include any combination of hardware and software for generating view-angle frames 142. A view-angle frame 142 can include any generated image frame representing the scene captured by a frame 121 from a different view-angle264855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)(e.g., angular perspective) than the view-angle of the original frame 121. Multi-view generator 140 can utilize 3D mesh 136 and the image frame 121 information or data to generate any number of view-angle frames 142. The multi-view generator 140 can generate a plurality of view-angle frames 142 corresponding to a plurality of view-angles different from the original view-angle of the frame 121. Each of the generated view-angle frames 142 can correspond to its own view angle surrounding the view angle of the image frame 121.
[0095] For example, the multi -view generator 140 can generate, using the one or more ML models 182, a plurality of view-angle image frames 142 for the plurality of view angles of the frame 121. For instance, the multi-view generator 140 can utilize the 3D mesh 136 (e.g., the 3D model) to generate a plurality of view-angle frames 142 at various view-angles (e.g., other than the original view angle of the 2D image frame 121). The multi view generator 140 can identify one or more holes or empty spaces that can occur in the viewangle frames 142 due to their angular shift between the original view-angle of the image frame 121 and the view-angle frame 142. The multi-view generator 140 can update the plurality of images using an inpainting function 134 (e.g., an inpainting ML model 182) to fill in the one or more holes in the plurality of view-angle frames 142. For instance, the multiview generator 140 can utilize the inpainting function 134 to generate pixel coloring, texture or shape to conform to the pixel coloring, texture or shape in the anatomy segments 128 to which the hole (e.g., empty space) belongs or with which the hole overlaps. The multi -view image generation can be used to simulate the trajectory of a simulated moving camera.Along the trajectory, the image view can change and inpainting steps can be repeated for each sampled perspective (e.g., view-angle).
[0096] Inpainting function 134 can include any combination of hardware and software for reconstructing or filling in missing parts of an image frame 121. Inpainting function 134 can be used to restore any missing image areas that can appear or be generated while generating view-angle frames 142 from angles that reveal a region that was obstructed by another object or a feature in the original 2D image frame 121. For example, the inpainting function 134 can fill in gaps caused by obstructions with respect to any portions of anatomy segments 128 in a surgical scene captured by the frame 121, using generative Al (e.g., inpainting ML models 182) to predict and reconstruct the missing anatomical structures based on surrounding context.274855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0097] For example, the inpainting function 134 can determine, based on an ML model 182 trained on determining a full shape of an anatomy segment 128 that a portion of a hole (e.g., image area having missing pixel values in the view-angle frame 142) belongs to or is a portion of a particular anatomy structure (e.g., a particular tissue or an organ). In response to this determination, the inpainting function 134 can fill int the hole with the pixel values colored and arranged to match or correspond to the portion of the same anatomy segment 128 from the frame 121. For instance, the pixels of the hole can be filled with the same colored pixels as the pixels of the anatomy segments 128. For example, the hole can be filled in with an average color of the anatomy segment 128. For example, the hole can be filled in with a pattern of features of a plurality of pixels matching or corresponding to the pattern of features of the surrounding portions of the anatomy segments 128. The inpainting function 134 can achieve this through techniques such as deep learning, where neural network ML models 182 can be trained on large datasets to learn the patterns and features of different anatomical structures. The function can leverage convolutional neural networks (CNNs) for spatial inpainting or generative adversarial networks (GANs) for more complex and realistic reconstructions.
[0098] For example, the multi -view generator 140 can generate, using the one or more models 182, a plurality of view-angle image frames 142 for the plurality of view-angles of the frame 121. The inpainting function 134 can identify one or more holes (e.g., portions of the view-angle frames 142 whose pixels are not filled in with data) in the plurality of viewangle image frames 142. The inpainting function 134 can update the plurality of view-angle frames 142 using an inpainting ML model 182 to fill the one or more holes in the plurality of view-angle frames 142. For instance, the inpainting function 134 can synthesize, using the inpainting ML model 182, one or more new pixels based on the anatomy segments 128 (e.g., from the original frame 121) and the labels 126 in, or for, the one or more masks 124 and 138. The inpainting function 134 can fill the one or more holes with the new pixels and assign, to the new pixels, the label and the depth value based on the one or more masks (e.g., masks 124 and 138).
[0099] The inpainting function 134 can compute or calculate the pixel transformation (e.g., the pixel values of the holes) based on the depth map (e.g., depth mask 138 and depth values 132) and using the computational geometry. The inpainting function 134 can utilize an inpainting ML model 182, such as a generative Al model trained model with the capability284855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) to synthesize new pixels fitting in the holes (e.g., undetermined pixels) in the image. The ML model 182 can take a semantic segmentation map (e.g., segmentation mask 124) and labels 126 of the mask 124 to understand the categories of the infilling area. The semantic label can be used as, or included in, prompts 156 for the generative Al models 182 to better generate pixels as required. The inpainting function 134 can operate based on the geometrical constraints from the depth map-defined 3D structure for the synthesized pixels.
[0100] The pixels filled with an inpainting function 134 (e.g., during the generation of the view-angle frames 142 by the multi -view generator 140) can be generated to be smooth or continuous to their surroundings. The filled pixels can keep their semantic labels 126 unchanged. After inpainting the holes, the color and depth values can be assigned to the 3D mesh 136. The new viewing angle image can be generated based on the updated or adjusted 3D mesh 136 that can include the pixels filled in by the inpainting function 134. When inpainting various tissues, the inpainting function 134 can fill in the pixels based on the surrounding tissue or organ tissue. For instance, a gallbladder can be inpainted to a gallbladder, a fat tissue can be inpainted according to a view of a surrounding fat tissue. The distance of inpainted areas can be smooth to interpolate a 3D shape. Color and texture can be based on a semantic map. The technical solutions can use a standard interpolation function to smooth the shape. This can be done for other angles and such images can be generated and connected to play in 3D. The inpainting function 134 can be used to provide visibility or clarity for various visibility obstructions. For instance, a liquid or dirt on a camera can be addressed by the technical solutions that can utilize these techniques to clarify the image and remove the dirt.
[0101] The inpainting function 134 can be utilized to clean images in real-time by removing blurring or improving contrast and clarity of the image using masks (e.g., 124 or 138) and inpainting functionalities. For example, if the camera lens gets dirty during a procedure, the inpainting function 134 can clean the images using masks (e.g., 124 or 138) and generative Al models 182 in real-time to display a clean image or a video to the surgeon during the procedure, removing the artifacts from the dirty camera.
[0102] The inpainting function 134 can use anatomy states (e.g., burned, bloodied, damaged, cut, removed, connected or other) of various anatomy parts along with anatomy and medical instrument segmentation and metric depth map (e.g., depth mask 138) to compute pixel values for the image reconstruction and inpainting. Anatomy state detection294855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) can identify changes or states of the anatomy, such as the anatomy segment 128 being cut, bruised, burnt, bleeding or otherwise in other state. Anatomy and instrument segmentation can isolate relevant areas, allowing for focused inpainting. Metric depth maps provide depth information (e.g., depth values 132 across the image) for accurate spatial relationship determinations. By integrating the historic state of anatomy and detected changes, the inpainting can maintain consistency with previous states, allowing for accurate and reliable medical images for surgical planning and execution. For instance, the inpainting function 134 can detect bleeding in anatomy state detection and use historic state of anatomy to determine the changes to the state.
[0103] 3D video generator 150 can include any combination of hardware and software for generating a video 152. A video 152 can include a series of view-angle frames 142 or a video depicting a scene captured by an original frame 121 from one or more (e.g., a series of) view-angles other than the original view-angle of the image frame 121. The video 152 can include a stream of view-angle frames 142 along with the hole areas that are filled in, or inpainted, by the inpainting function 134. The 3D video generator 150 can generate the video 152 based on the 3D mesh 136 that can be adjusted or updated based on inpainted or filled in content (e.g., areas filled in by the inpainting function 134).
[0104] The 3D video generator 150 can generate any 3D representation of the scene captured in the image frame 121. The 3D video generator 150 can generate a 3D simulation of the scene, based on the 3D mesh 136, depth masks 138, segmentation masks 124 and any labels 126. The 3D video generator 150 can generate the video 152 corresponding to a plurality of view angles of the frame using the one or models and the 3D mesh 136 and the labels 126 of the anatomical structures from the one or more masks (e.g., 124 or 138). The 3D video generator 150 can generate a video 152, based on a set or predetermined frame rate. The video 152 can include multi -view images combined along with the view-angle of the original frame 121 as a 3D video and output to the frontend user interface 192. For instance, the 3d video generator 150 can receive a data stream 162 of the medical procedure captured via one or more sensors 104 of the robotic medical system 110 and determine, based on the data stream 162, a range of view angles in a portion of the video 152 that comprises the frame 121.
[0105] The 3D video generator 150 can allow for rendering of the surgical scene in a user-specified perspective. For instance, a user can specify a perspective of view at which304855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) the scene can be rendered (e.g., from a particular angle) by selecting with an input device (e.g., a computer mouse) an angle relative to the view-angle of the frame 121 (e.g., to the left or to the right of the view angle of the frame 121). Responsive to such a selection (e.g., via the user interface 192), the 3D video generator 150 can generate (e.g., from the view-angle frames 142) new viewing angles of previously invisible areas and contents. Those new contents can be photo-realistic and provide a same level of clarity as other features based on the filled in data generated or provided by the inpainting function 134 to render the previously invisible areas in color, texture and 3D shape of the anatomy segments 128 to which they correspond.
[0106] The 3D video generator 150 can determine the range of view angles for the video 152. The range of view-angles can correspond to the range of view-angles to generate for the view-angle frames 142. For instance, wider view-angle range can correspond to more viewangle frames 142 to be generated and more angular range for the video 152. The range of view angles can be determined based on various trigger conditions 146. For instance, the range of view-angles for the 3D video 152 can be based on a milestone event (e.g., a type of a milestone event) or a type of a medical task or a procedure. For instance, the range of viewangles for the 3D video 152 can be based on surgeon data 172 (e.g., a surgeon profile), such as a level of surgeon’s skill, surgeon’s historical performance or surgeon’s objective performance indicators (OPIs) indicating surgeon’s performance with respect to the particular task, phase or a type of a medical procedure. For instance, the range of view-angles for the 3D video 152 can be based on a data stream indicating a lack or absence of a sufficient view angles based on the type of tasks or procedure in the frame. For instance, the range of viewangles for the 3D video 152 can be based on a quality confidence level of the pixels or images created using the ML models 182.
[0107] The 3D video generator 150 can render a mask on 3D video 152 based on the confidence or quality of images at different view angles. For instance, the 3D video generator 150 can generate or render a mask on video 152 by providing different color or shade for pixels synthesized using the inpainting function 134, allowing for distinguishing pixels that are generated using ML models 182 from pixels from the frames 121. The 3D video generator 150 can generate a notification for the user interface 192 to notify the user on how many pixels are generated by the ML, versus how many are based on the frame 121.314855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0108] Trigger condition detector 144 can include any combination of hardware and software to detect trigger conditions 146 and trigger 3D view reconstruction based on one or more thresholds 148. The trigger condition 146 can include any milestones or features to be verified or affirmed as a part of a medical procedure process (e.g., a phase or a task). The trigger conditions 146 can occur based on any milestone events or based a level of surgeon’s skill, such as the surgeon’s historical performance or surgeon’s objective performance indicator (e.g., OPIs) or the surgeon’s demand. The trigger conditions 146 can occur based on the data stream 162 indicating a lack or absence of sufficient view angles based on the type of task or procedure in the image frame 121.
[0109] For example, a trigger condition 146 can include the verification that a particular first tissue (e.g., a first anatomy segment 128) is completely severed and separated form a second tissue (e.g., a second anatomy segment 128). The trigger condition detector 144 can utilize ML models 182 trained to determine or detect surgical procedures, their phases and individual tasks, to identify a particular trigger condition 146 present at a particular portion of the procedure (e.g., at an end of a surgical task). In response to the trigger condition detector 144 detecting a trigger condition 146, the trigger condition detector 144 can issue a request or an instruction (e.g., to the computing system 105) to trigger or start the 3D reconstruction of the surgical scenes. For instance, the trigger condition detector 144 can trigger any one or more of a mesh generator 130, the multi-view generator 140 and the 3D video generator 150 to implement their respective operations to provide a 3D representation of the surgical scene, to allow the user (e.g., the surgeon) to view the scene from a plurality of view-angles, facilitating a more convenient and more efficient and effective completion of the milestone event (e.g., the trigger condition 146).
[0110] The trigger condition detector 144 can utilize thresholds 148 to detect the presence of trigger conditions 146. A threshold 148 can include any value, parameter or an indication indicating the presence of the trigger condition 146. For instance, the threshold 148 can include a level of performance of a surgeon from a surgeon’s profile in the surgeon data 172. For instance, the threshold 148 can indicate a number of successfully performed surgeries for a different type of a phase or a task. In response to determining (e.g., from the surgeon’s data 172, such as the surgeon profile) that the surgeon has performed fewer than the threshold 148 number of successfully completed phases or tasks, the trigger condition detector 144 can identify the trigger condition 146. For instance, the threshold 148 can include a threshold 148 confidence score or a confidence level which can be provided by the ML model 182,324855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) which can be utilized by any one or more of the: segment function 122, depth estimator 125, mesh generator 130, multi-view generator 140 or 3D video generator 150. The threshold 148 can correspond to a minimal sufficient view angle of an area to view or verify given the task or a phase of a medical procedure. In response to the ML model 182 confidence level or score not satisfying the threshold 148, the trigger condition detector 144 can instruct the computing system 105 to begin the 3D view reconstruction process. For instance, in response to the trigger condition 146, the trigger condition detector 144 can instruct the computing system 105 and its components to perform the 3D view reconstruction to provide the viewangle frames 142 and the video 152 to the end user (e.g., the surgeon) to facilitate the view of the scene from multiple view angles.
[0111] For instance, a computing system 105 can select a frame 121 from the video stream 178 responsive to detection of the trigger condition 146. The trigger condition 146 can be detected, using for example the trigger condition detector 144 utilizing one or more ML models 182 trained to detect tasks of particular phases in the medical procedure at which the trigger condition 146 is set to occur (e.g., per the medical procedure performed). The computing system 105 can determine to generate the video 152 for the frame 121 based on the detection of the trigger condition 146.
[0112] For example, the computing system 105 can receive a data stream 162 of the medical procedure captured via one or more sensors 104 of the robotic medical system 110. The computing system 105 can determine, based on the data stream 162, a range of view angles in a portion of the video 152 that comprises the frame 121. A trigger condition detector 144 can detect the trigger condition 146 based on the occurrence of, or a temporal alignment with, a milestone event in the frame 121. A trigger condition detector 144 can detect the trigger condition 146 in response to the data stream 162 indicating that the view angle for the type of tasks or procedure in the frame 121 is below the threshold 148 corresponding to the minimum sufficient view angle based on the type of task or procedure. The trigger condition detector 144 can detect the trigger condition 146 the range of view angles in the portion of the video data 178 being less than or equal to a threshold. The threshold can include, for example, a view angle value sufficient to view the scene and verify a milestone event occurrence, based on the 2D image frame 121.
[0113] The trigger condition detector 144 can detect the trigger condition 146 based on a surgeon data 172. The surgeon data 172 can include any information about a surgeon, such334855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) as a surgeon profile providing information about the level of experience of the surgeon. The surgeon profile (e.g., the surgeon data 172) can indicate that the surgeon does not have a sufficient level of experience with the particular surgical task being performed. For instance, the trigger condition detector can utilize a threshold value for a number of prior performed surgical procedures or tasks. In response to determining that the threshold number of previously performed surgical procedures or tasks, the trigger condition detector 144 can detect the trigger condition 146 (e.g., using the surgeon profile) and trigger the computing system 105 to generate the 3D mesh 136, the view-angle frames 142 and 3D videos 152. The computing system 105 can include at least one simulation module 115, described in greater detail in FIG 2. The simulation module 115 can integrate or communicate with the other components of the computing system 105 and the system 100 to generate an interactive 3D simulation of a medical procedure to allow a medical practitioner to practice the medical procedure.
[0114] Referring now to FIG. IB, among others, depicts an example computing system 105 for performing 3D view reconstruction of a surgical scene in accordance with the discussion in connection with system 100 in FIG. 1 A. The example 200 system diagram can include a frame function 120 that can receive a 2D monocular frame 121, depicting a surgical scene. The surgical scene can include one or more anatomical structures and medical instruments 112. The frame 121 can be input into segmentation function 122 and a depth estimator 125. The segmentation function 122 can utilize ML models 182 to provide anatomy segmentation and generate segmentation masks 124 with labels 126 marking various anatomical structures according to their respective anatomy segments 128. The depth estimator 125 can utilize the frame 121 to implement metric depth estimation and generate the depth mask 138 along with any depth values 132.
[0115] The 3D mesh generator 130 can utilize the outputs from the segmentation function 122 and the depth estimator 125 (e.g., the segmentation mask 124 and the depth mask 138) to generate a 3D mesh 136 of the scene captured by the frame 121. The 3D mesh 136 can include a three-dimensional representation of the surgical scene. The 3D mesh 136 can be generated using ML models 182 trained to generate spatial relations between various features based on the metric distances (e.g., distance in meters, centimeters, millimeters, inches or any other unit of distance) between each of the portions of the image (e.g., one or more pixels) and the lens of the camera that captured the frame 121. The mesh generator 130 can provide344855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) the 3D mesh 136 to a simulation module 115. The simulation module 115 can generate an interactive 3D simulation of a medical procedure using the 3D mesh 136 to allow a medical practitioner to practice the medical procedure.
[0116] Referring now to FIG. 2, among others, an example computing system 105 to reconstruct a medical procedure environment in a 3D virtual scene 205. The simulation module 115 can receive a medical procedure video 178 from the robotic medical system 110. The medical procedure video 178 can be a video captured by a camera or endoscope of the robotic medical system 110. For example, the medical procedure video 178 can be a video, stream, frame, set of frames, or set of images of a medical procedure performed by the robotic medical system 110. The medical procedure can be a surgery, a therapy, a medical inspection, a biopsy, or any other type of medical procedure.
[0117] The simulation module 115 can include at least one simulator 220. The simulator 220 can receive the medical procedure video 178 from the robotic medical system 110, and generate, build, synthesize, or reconstruct the virtual scene 205 using the medical procedure video 178. The simulator 220 can include at least one frame identifier 225. The frame identifier 225 can select, determine, or identify at least one image or frame 121 of the medical procedure video 178 with which to generate, update, or initialize the virtual scene 205. The frame identifier 225 can identify a frame that captures a scene in a medical procedure of the medical procedure video 178 performed with the robotic medical system 110. For example, the frame identifier 225 can identify a frame 121 where an anatomical structure (e.g., a bone, a heart, an intestine, a cancerous growth, etc.) is visible or is not blocked or occluded (e.g., is completely visible or is occlude by less than a threshold amount). The frame 121 can be a monocular surgical or procedure image that can be used by the mesh generator 130 to produce a 3D mesh 136 and a texture file 255 to initialize a simulation with. The frame 121 can be a 2D image, such as a monocular image. The frame 121 may not be a stereo image or the video 178 may not be a stereo video. In some implementations, the video 178 is a stereo video, and the frame 121 can be a stereo frame (e.g., a set of multiple frames). The video 178 or the frame 121 may not have any depth or perspective data, and may only be a collection of raw pixel values.
[0118] The frame identifier 225 can analyze the entire procedure video 178 or a portion of the procedure video 178 (e.g., analyze all or part of the frames of the medical procedure video 178) to select a frame 121 from the analyzed frames to use to generate, build, update,354855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) or initialize the virtual scene 205. The frame identifier 225 can implement a model trained by machine learning (e.g., a convolutional neural network, a sequence to sequence neural network, a transformer, etc.) to select the frame 121 from the medical procedure video 178 with which to initialize or generate the virtual scene 205. The frame identifier 225 can identify the frame 121 from the medical procedure video 178 by selecting a frame 121 at a predefined timestamp or a frame 121 that occurs a predefined length of time into the medical procedure. The frame identifier 225 can sample the selected frame 121 from the video 178 and pass it to the depth estimator 125 or the anatomy segmentation function 120 to generate the depth mask 138, the segmentation masks 124, the 3D mesh 136, and / or the texture file 255.
[0119] The frame identifier 225 can identify an event, trigger, or triggering condition in or within the medical procedure video 178. For example, a trigger condition in the medical procedure video 178 can be a type of task performed with the robotic medical system 110 (e.g., cutting through tissue, dissecting an anatomical structure, removing a portion of the anatomical structure, etc.). The triggering condition can be an occurrence of a milestone event in the medical procedure video 178.
[0120] For example, the milestone event can be the beginning or end of a particular segment of the medical procedure video 178 (e.g., a step or phase of the medical procedure video 178). The medical procedure can include multiple phases, while at least some of the multiple phases can include multiple steps. The phases can include transection of an anatomical structure of a patient. For example, the phase can include transecting or dividing an anatomical structure, e.g., opening the anatomical structure. Transection can include cutting transversely across an anatomical structure. The phase can include extraction of an anatomical structure. For example, extraction can include cutting out or removing an anatomical structure from a patient or from an area or cavity of a patient. The phase can include reconstruction of an anatomical structure. For example, after opening, cutting, or exposing an anatomical structure, the robotic medical system 110 can reconstruct the anatomical structure. For example, the robotic medical system 110 can stitch up an opening in an anatomical structure, staple closed an opening in an anatomical structure, or otherwise reconstruct an anatomical structure after opening or cutting the structure. The phase can include dissection, such as cutting or dissecting an anatomical structure or patient. The phase can include exposure of an anatomical structure. For example, the phase can include peeling364855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) back tissue, cutting back tissue, moving other anatomical structures, or removing bones to expose or access an anatomical structure. The phases can include individual steps. The steps collectively can form or complete a particular phase. The steps can include steps for a cholecystectomy, for example specimen removal, litigation or division of a cystic artery, dissection of a gallbladder of a liver off a liver bed, etc. The steps can be steps of any type of medical procedure. The steps can be specific to a procedure type. The steps can be procedure type specific, while the phases can be universal or apply to multiple procedure types.
[0121] The triggering condition can be a particular metric or objective performance indication (OPI) occurring. The triggering condition can be the particular metric satisfying a threshold, e.g., reaching the threshold, exceeding the threshold, or falling below the threshold. The metric can be an amount of energy or power consumed by the robotic medical system 110. The various metric types of metrics can be a metric indicating an amount of energy or power consumed by the robotic medical system 110 to operate an individual instrument or endoscope. The metric can indicate a total duration of a segment of the medical procedure. For example, the metric can be a length of time of the particular segment of the medical procedure video 178, such as a length of time of the step or the phase. The metric can indicate a total linear distance of an instrument of the robotic medical system 110 during the segment. For example, the metric can indicate a total distance that the instrument traveled during the steps, the phases, or the entire video 178. The metric can be a total angular distance the instrument of the robotic medical system 110 traveled during the segment. The angular distance indicate an amount or distance that the instrument is rotated during manipulations. The metric 297 can indicate a total angular distance that the instrument traveled during the steps, the phases, or the entire video 178. The metric can indicate a total number of operations or a clutch or brake of the robotic medical system 110. For example, the metric can indicate a total number of actuators or activations of a clutch or brake for the instrument or the endoscope of the robotic medical system 110. The clutch or brake can be operated to float joints of the instrument or endoscope. The metric can indicate a total number of operations or a brake or clutch during the steps, the phases, or the entire video 178.
[0122] Responsive to detecting the triggering condition, the frame identifier 225 can identify at least one or a set of frames 235 that correspond to the triggering condition. For374855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) example, the at least one frame or set of frames 235 can have timestamps at (or within a length of time from) the time when the triggering condition is satisfied. The frame identifier 225 can generate or create a message, question, or prompt for user review. The frame identifier 225 can generate the prompt to identify the type of the triggering condition (e.g., an indication that a metric exceeded a threshold, an indication that a particular phase of the medical procedure was concluded, etc.). The prompt can include the at least one frame 121 associated with the triggering condition. The frame identifier 225 can cause the prompt to display in the interface or graphical user interface displayed on the client device 106. The frame identifier 225 can be or include a recommendation to establish the virtual scene 205 using the identified frame or frames. The user can provide input on the client device 106 accepting or declining to use the identified frames 235 to initiate the virtual scene 205. The input can selected one frame 121 (or multiple frames) from a set of different frames in the prompt to generate the virtual scene 205 with. Responsive to receiving the user approval and / or selected frames 235, the simulator 220 can initiate the virtual scene 205 or update the virtual scene 205 using the selected frames 235.
[0123] The frame identifier 225 can receive input via at least one client device 106. The frame identifier 225 can generate or cause the client device 106 to display an interface or graphical user interface that that the frame identifier 225 displays graphical data on, or receives user input via. The frame identifier 225 can display the medical procedure video 178 to the user of the client device 106. The frame identifier 225 can display all or some of the frames of the medical procedure video 178 to the user on the interface via the client device 106. The user can scan through the medical procedure video 178 via the client device 106, and provide input that selects or identifies one frame to initialize the virtual scene 205 with. The frame identifier 225 can receive, via the interface, a request to establish the virtual scene 205 in the interactive virtual environment 277 for the frame 121. The user, via the client device 106, can provide input that identifies a specific frame from which to initialize the virtual scene 205. Furthermore, the frame identifier 225 can receive another interaction or request to initialize or initiate the simulation using the selected frame 121. Responsive to receiving the request, the simulator 220 can generate the virtual scene 205 using the selected frame 121.
[0124] The frame identifier 225 can receive a requested view angle from the client device 106. For example, the frame identifier 225, via the interface, can receive a selection of a384855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) view angle to initialize the virtual scene 205 at from the client device 106. The frame identifier 225 can receive a pan, tilt, or roll angle measurements to the anatomical structure of the virtual scene 205 at. The frame identifier 225 can receive a position or orientation of the frame of view of the user for the virtual scene 205. The view angle that the user requests via the client device 106 can be different than the view angle of the selected frame for initializing the virtual scene 205. The simulator 220 can initialize the virtual scene 205 using the selected view angle. Furthermore, as the virtual scene 205 is viewed by the user on the client device 106, the user can provide input that changes the position of the user within the virtual scene 205, or changes the pan, tilt, or zoom of the user’s vision within the virtual scene 205. The simulator 220 can adapt the display of the virtual scene 205 to the user on the client device 106 according to the view navigation of the user. Thus, the user can view anatomical structures from a variety of angles in the virtual scene 205, and not only from the view angle of the frames used to initialize, generate, or update the virtual scene 205. The selected frame 121 can have or correspond to a single view of the scene in the medical procedure.
[0125] The simulator 220 can provide the selected frame 121 to at least one depth estimator 125. The depth estimator 125 can generate, using one or more models trained with machine learning, a depth map, a point cloud, or depth mesh 138 for the selected frame 121. The depth estimator 125 can output at least one depth value for at least one pixel of the selected frame 121. The depth estimator 125 can output a depth value for all, every, or most pixels of the selected frame 121. The depth estimator 125 can be a neural network that maps the selected frame 121 to a metric depth. Each pixel intensity in the depth mesh 138 produced by the depth estimator 125 can represent a real metric distance between a center of the camera that captured the selected frame 121 to a point of the scene. The neural network can be encoder-decoder structured with a transformer unit inside. The depth value can indicate a distance from the camera to a surface, anatomical structure, instrument, subject, etc. in the selected frame 121. The output of the model 125 can be a metric depth map rather than a relative depth mask. Then model 125 can output relative depth maps. The models of the depth estimator 125 can use a camera intrinsic during training for calibration purposes. In some implementations, the model of the depth estimator 125 can be the DEPTH ANYTHING model.
[0126] The simulator 220 can provide the selected frame 121 to the anatomy segmentation function 120. The anatomy segmentation function 120 can generate or produce394855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) anatomy or segmentation masks 245 that segment or define anatomical structures in the selected frame 121. Furthermore, the anatomy masks 245 can include labels (or names, identifiers, flags, or markers) or be linked to labels that identify the type of anatomical structure that the anatomy mask 124 represents. The anatomy segmentation function 120 can produce at least one state mask 124. The state masks 245 can identify the state of the anatomical structure or a state of a portion of the anatomical structure. The state can be bleeding, burned, bruised, cut, an abrasion, etc. The anatomy segmentation function 120 can be a hierarchical anatomy segmentation module (e.g., a module that implements a hierarchy of models that work together to generate segmentation masks 245 can state masks 245). The anatomy segmentation function 120 can implement a vision transformer-based neural network that is adopted and trained with human annotated clinical image dataset. Given the selected frame 121, the model can generate a segmentation mask that labels each anatomical structure with its spatial existence (e.g., produce a segmentation map or mask 124) and a name of the anatomical structure (e.g., text label) in a primary branch of the model. Within a detected anatomy area, a second level decoder can applied to segment out areas with defined states or conditions. The output of the module can include segmentation masks 245 and labels of the anatomy and structure in the view, and semantic masks 245 for each state within the anatomy / structure. The anatomy segmentation function 120 can implement the techniques and models described in U.S. Patent Application No. 63 / 566,531 filed March 18, 2024, the entirety of which is incorporated by reference herein.
[0127] The masks 245 can be frames that include values that indicate which pixels of the frame 121 correspond to the anatomical structure or states of the anatomical structures. For example, an anatomical structure mask for a heart can include a set of values (e.g., binary values or confidence values) that indicate whether one, multiple, each, or some pixels of the selected frame 121 are part of a heart of a patient. Similarly, the state masks 245 for bleeding of a heart can include a set of values (e.g., binary values or confidence values) that indicate whether one, multiple, each, or some pixels of the selected frame 121 indicate bleeding regions the heart of a patient.
[0128] The depth estimator 125 can provide the depth mesh 138 to the mesh generator 130. The anatomy segmentation function 120 can provide the anatomy masks 245 to the mesh generator 130. The mesh generator 130 can generate the 3D mesh 136 and the texture file 255 using the depth mesh 138 and the anatomy masks 245. For example, the mesh404855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) generator 130 can generate or construct the 3D mesh 136 and the texture file 255 using identified anatomical structures in the selected frame 121 (e.g., based on the anatomical structure masks 245). For example, the mesh generator 130 can generate or construct the 3D mesh 136 and the texture file 255 using identified states of anatomical structures (e.g., based on the state segmentation masks 245). The mesh generator 130 can execute at least one generative model or technique (or a combination of techniques) to generate the 3D mesh 136 (and / or texture file 255), such as Neural Radiance Fields (NeRF), Scene Representation Network (SRN), Local Light Field Fusion (LLFF), Stable Diffusion, Stable Diffusion XL (SDXL). The mesh generator 130 can be a model that is part of a generative Al tool, such as MESHY, STABLE ZERO 123, ALPHA3D, etc. The mesh generator 130 can transmit requests including a selected frame depth mesh 138 or masks 245 to a generative Al tool, which can reply with the 3D mesh 136 or texture file 255.
[0129] The 3D mesh 136 can be a graphic representation of the 3D scene or objects within the 3D scene. Based on the depth mesh 138, the mesh generator 130 can generate a 3D mesh file 250. The 3D mesh 136 can be or include a collection of vertices, edges, and faces that represent the shape and structure of a 3D scene or objects within the 3D scene (the 3D scene can be a medical scene where the video 178 is taken). The 3D mesh 136 can include or provide a representation of complex 3D shapes and surfaces by specifying the coordinates of vertices and the connections between vertices. The 3D mesh 136 can include vertices, edges, and faces. The faces can be formed by edges that extend between the vertices. The mesh generator 130 can output multiple 3D mesh files 250, e.g., on 3D mesh 136 for a first anatomical structure, a 3D mesh 136 for a second anatomical structure, a 3D mesh 136 for an environment shown in the frame 121.
[0130] The color and textures for the vertices, edges, or faces can be mapped directly (or indirectly) from the selected frame 121 to the mesh file 250 by the mesh generator 130. This mapping can generate the texture file 255. The mesh generator 130 can generate the texture file 255 based on the anatomy segmentation masks 245 and state masks 245. This can ensure the correctness of the semantics. The texture file 255 can store image data mapped to various different vertices, edges, or faces of the 3D mesh 136. The texture file 255 can store image data that is taken from, or generated using, the selected frame 121.
[0131] The simulator 220 can provide the virtual scene 205 in an interactive virtual environment 277 to the client device 106. For example, the simulator 220 can render the414855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) virtual scene 205 for display on the client device 106. The virtual scene 205 can be rendered to include one or multiple anatomical structures depicted in the selected frame 121, and include states for the anatomical structures. The simulator 220 can render the virtual scene 205 using the 3D mesh 136 and the texture file 255. For example, the simulator 220 can initialize the virtual scene 205 using the 3D mesh 136 and the texture file 255, and cause the initialized virtual scene 205 to be rendered on the client device 106. The virtual scene 205 can be a virtual space (e.g., such as an operating room, medical procedure room, doctor’s office, etc.) that includes various 3D meshes 250 and texture files 255 of anatomical structures, fluids, patients, cavities of a patient, etc. Furthermore, the virtual scene 205 can include 3D models 260 of objects such as medical instruments (e.g., a scalpel, a scissors, a monopolar curved scissors (MCS), a cautery hook tip, a cautery spatula tip, a needle driver, a forceps, a tooth retractor, a drill, or a clip applier), of a doctor or surgeon, an operating tables, instrument tables, operating lights, etc. The 3D models 260
[0132] The rendered virtual scene 205 can include the anatomical structures indicated by the anatomy masks 245. Furthermore, the rendered virtual scene 205 can include or represent the states of the anatomical structures using the state masks 245 or the simulator 220 can animate the states of the anatomical structures using the state masks 245. For example, the simulator 220 can animate bleeding, fluid leaking, burning, bruising, etc.
[0133] The virtual scene 205 can be displayed on the client device 106 within an interactive environment 277. The interactive environment 277 can allow a user to provide input via the client device 106 to move within the virtual scene, zoom in, zoom out, pan, tilt, or interact with the anatomical structures displayed in the virtual scene 205. The interactive environment 277 can simulate physics in the virtual scene 205, animate the virtual scene 205, provide lighting in the virtual scene 205, render the virtual scene 205 to a user on the client device 106, etc. For example, the user can provide input to the client device 106 to interact with the anatomical structures of the virtual scene 205 to perform a medical procedure on the anatomical structures (e.g., cutting, sawing, sewing, removing, cauterizing, or cleaning). The simulator 220 can persist the changes made by the user in the virtual scene 205 over time, and allow the user to provide input via the client device 106 to perform a portion of, or a complete medical procedure.
[0134] The simulator 220 can display the virtual scene 205 and the interactive virtual environment 277 on a display device of the client device 106. For example, the client device424855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)106 can be or include a virtual reality (VR) system or augmented reality (AR) system. For example, the client device 106 can be a head-mounted display, e.g., the client device 106 can be or include VR / AR glasses, VR / AR goggles, VR / AR smart contact lenses, or a heads-up display. The client device 106 can be part of, or integrated with, the robotic medical system 110. The client device 106, or the robotic medical system 110, can include manipulators (e.g., joysticks, buttons, wheels, switches, etc.) to move surgical instruments in the interactive virtual environment 277 and interact with the anatomical structures. For example, the simulator 220 can animate movement of instruments or endoscopes of the robotic medical system 110 within the interactive virtual environment 277 according to the inputs received from the client device 106 (e.g., according to the input that the user provides via the manipulators).
[0135] The simulator 220 can include or store at least one 3D model 260. The 3D models 260 can be or include general anatomical 3D models for use in building, maintaining, or animating the virtual scene 205 in the interactive virtual environment 277. The 3D models 260 can be or include anatomical structures, instruments, bodily fluids (e.g., blood, interstitial fluid, saliva, gastric juice, etc.), peripheral objects (e.g., a surgical needle, an ultra- sound probe, etc.), etc. The 3D models 260 can model the physical attributes, shape, geometry, mesh, texture, size, or other information of the various elements. The simulator 220 can maintain or update the virtual scene 205 over time to maintain simulation continuity of the various 3D objects or elements that appear within the virtual scene 205. For example, the simulator 220 can maintain continuity of the simulation by building virtual models or objects (e.g., the virtual scene 205 and its collection of meshes, texture, etc.) and persisting changes to the models or objects over time. The simulator 220 can initialize or maintain the virtual scene 205 at least on part on the 3D models 260. For example, if the virtual scene 205 can be initialized to include a bleeding pancreas, the simulator 220 can cause a blood model 260 to be used to animate or display blood bleeding from the pancreas rendered in the virtual scene 205.
[0136] The simulation module 115 can include at least one data repository 265. The data repository 265 can be a database, an SQL database, a knowledgebase, a noSQL database, a graph database, etc. The data repository 265 can store or include at least one subject parameter 270. The subject parameters 270 can include parameters that describe the condition, status, state, health, or characteristics of the patient, person, or animal on which the434855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) medical procedure is performed. For example, the subject parameters 270 can store the weight of the subject, the height of the subject, comorbidities of the subject, a history of past medical procedures performed on the subject, medical conditions of the patient (e.g., whether the patient is diabetic, whether the patient is on blood thinners, whether the patient is a cancer survivor, etc.). The mesh generator 130 can retrieve at least some subject parameters 270 from the data repository 265, and use the retrieved subject parameters 270 to generate the 3D mesh 136 and the texture file 255. The mesh generator 130 can generate the shape or size of the 3D mesh 136 or the texture file 255 using the retrieved subject parameters 270. For example, if the subject has a high BMI (e.g., above 30), the mesh generator 130 can generate a corresponding amount of fatty tissue on or around the anatomical structures of the selected frame 121. If the BMI of the subject is low, e.g., below 25, the mesh generator 130 can generate less fatty tissue around the anatomical structures of interest in the selected frame 121. Furthermore, the mesh generator 130 can generate the thickness of different or shape of different anatomical structures in the 3D mesh 136 according to the subject parameters 270. For example, if the subject parameters 270 indicate that the subject is an athlete, such as a runner or swimmer, the mesh generator 130 can generate a thickness of a wall of a heart to be thicker than normal or thicker than a nominal heart thickness for.
[0137] The mesh generator 130 can generate the texture file 255 at least in part on the subject parameters 270. For example, if the subject parameters 270 indicate that a subject has had open heart surgery before, the mesh generator 130 can generate a texture file 255 for a heart or chest of the subject to include scaring or scar tissue according to the subject parameters 270 indicating that the patient had open heart surgery.
[0138] The subject parameters 270 can indicate the likelihood of a state of the anatomical structures. The subject parameters 270 can indicate a likelihood that an anatomical structure will bleed after being cut, burn from heat, or tear from pressure. For example, the subject parameters 270 can indicate that a patient is on a blood thinner, and if the anatomical structure of the patient is cut, the patient will excessively bleed. The simulator 220 can initialize the anatomical structures in the virtual scene 205 according to the likelihood of the state, or modify the virtual scene 205 to according to the likelihood of the state. For example, if the user provides input via the client device 106 that cuts the anatomical structure in the virtual scene 205, the simulator 220 can animate the anatomical structure to bleed excessively if the subject parameters 270 indicate that the patient is taking a blood thinner.444855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0139] The simulator 220 can receive a data stream of information from the robotic medical system 110. The data stream can include data that indicates changes to the state of the subject. For example, the state of the subject can be changes in heart rate, changes in temperature, changes in blood pressure, etc. The mesh generator 130 can generate or update the 3D mesh 136 or the texture file 255 according to the updates in the state.
[0140] The robotic medical system 110 can generate, record, or store kinematics data 275. The robotic medical system 110 can provide, send, or transmit the kinematics data 275 to the simulation module 115, e.g., to the simulator 220. The kinematics data 275 can be or include information, data, data frames, or values collected by or from the robotic medical system 110 when performing the medical procedure. The kinematics data 275 can be time correlated data, e.g., data with timestamps such as a timeseries. The kinematics data 275 can be time correlated with the frames of the video 178. For example, at least one medical robotic system 110 can collect and store kinematics data 275 for a medical procedure, and then transmit the kinematics data 275 to the simulation module 115. The kinematics data 275 can be or include force, torque, acceleration, or velocity data of joints, links, arms, appendages, or manipulators of the robotic medical system 110. The kinematics data 275 can be captured or recorded from at least one sensor associated with the scene of the medical procedure. For example, the robotic medical system 110 can include sensors such as encoders, tachometers, current sensors, power meters, or force sensors.
[0141] The mesh generator 130 can generate the 3D mesh 136 or texture file 255 using the kinematics data 275. The kinematics data 275 can indicate an amount of force detected by a sensor between an instrument of the robotic medical system 110 and at least a portion of the anatomical structure in the scene. For example, the mesh generator 130 can generate the 3D mesh 136 or the texture file 255 using the kinematics data to accurately represent the thickness, toughness, or flexibility of anatomical structures according to the amount of force exerted by instruments of the robotic medical system 110 (e.g., indicated by the kinematics data 275) on the anatomical structures to cut, bend, or move the anatomical structures. Furthermore, the simulator 220 can generate or update the virtual scene 205 using the kinematics data 275. For example, after the 3D mesh 136 and the texture file 255 are generated by the mesh generator 130, and the simulator 220 initiates the virtual scene 205, the simulator 220 can update the virtual scene 205 according to the kinematics data 275 to454855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) accurately represent the thickness, toughness, or flexibility of anatomical structures of the virtual scene 205.
[0142] The simulation module 115 can include different modules or modes for interacting with the virtual scene 205. The simulation module 115 can establish a mode of interacting with the virtual scene 205 in the interactive virtual environment 277. For example, a user can provide input from the client device 106 to select between the different modes of the simulation module 115. The simulation module 115 can include at least one tutorial module 280 to implement a tutorial mode. The simulation module 115 can include at least one competitive module 285 to implement a competitive mode. The simulation module 115 can include at least one generative module 290 to implement a generative mode.
[0143] The tutorial module 280 can implement a tutorial mode in which parameters 270 of the subject are fixed. The tutorial module 280 can provide a mode that uses guardrails or provides guidance to a user when performing the simulated medical procedure. The guardrail can be or can be included by the fixed trajectory 295, for example, the guardrail can constrain the user to moving virtual instruments along the fixed trajectory 295. The tutorial module 280 can retrieve the subject parameters 270 from the data repository 265, and use the fixed subject parameters 270 to generate the 3D mesh 136, the texture file 255, or the virtual scene 205. In this regard, in the tutorial mode, the subject and the anatomical structures of the subject can be true to life or true to the medical procedure video 178. For example, by fixing the parameters 270, a user may not be able to adjust the subject parameters 270 via the client device 106 to explore performing the medical procedure on different patients of different characteristics. A surgeon can use the tutorial module 280 for training. For example, a surgeon can record a procedure, and use the tutorial mode for walking a trainee through the recorded procedure in the virtual scene 205. The simulation module 115 can provide a list of different cases (or videos 178) for tutorial mode, and can receive input via the client device 106 selecting a case from the list. The simulation module 115 can generate the virtual scene 205 for the interactive virtual environment 277 responsive to the selection. The simulation of the selected case can be performed locally on the simulator 220 or run locally on the client device 106.
[0144] The tutorial module 280 can include a fixed trajectory 295. The fixed trajectory 295 can be a guardrail for the user when moving the virtual surgical instruments, a constraint on a user’s input to move the virtual surgical instruments, or a component that provides464855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) guidance to the movements of the virtual surgical instrument controlled by the user. The trajectory 295 can represent the paths, motions, or movements of the instruments or endoscopes of the robotic medical system 110 took to perform the medical procedure. For example, the trajectories 295 can be based on, or can represent, the kinematics paths or data 275 recorded for each instrument of the robotic medical system 110 throughout the medical procedure video 178. The fixed trajectory 295 can represent different sets of coordinates or ranges of coordinates that an instrument of the robotic medical system 110 must move through in the virtual scene 205 to perform a particular act, segment, or step of the medical procedure. For example, for creating an incision, the fixed trajectory 295 can indicate a sequence of x, y, and z coordinates (or a sequence of ranges of x, y, and z coordinates for an instrument to stay within) in the virtual scene 205 for an instrument to move to create the incision. For example, the fixed trajectory 295 can indicate a sequence of coordinates that specify that the incision is made left to right across an anatomical structure at a particular position with a particular depth.
[0145] The tutorial module 280 can constrain or limit the movements of the virtual instruments in the virtual scene 205 to the sequence of coordinates. In this regard, a trainee or other surgeon can be constrained based on the trajectory 295. When the user provides input via the client device 106 to move the instruments in the virtual scene 205, the user may be limited to move the instruments according to the fixed trajectory 295 to follow the movements of the medical instruments of the robotic medical system 110 during the recorded medical procedure. If the user provides input that moves instrument outside or against the fixed trajectory 295, the tutorial module 280 can cause the simulator 220 to not move the instrument outside or against the fixed trajectory 295. In this regard, the trainee can be pulled through the motions of the surgery and with integrations within robotic medical system 110 can apply the same feelings that surgeon experienced while performing the surgery live.#
[0146] The tutorial module 280 can cause data to be displayed in the virtual scene 205 that represents the fixed trajectories 295 of the instruments of the robotic medical system 110. For example, the tutorial module 280 can display advanced views of trajectories or other visualization aids to show the context of maneuvers to a trainee. The tutorial module 280 can display arrows, lines, vectors, depth markers, dotted lines, etc. to guide the user in following the fixed trajectory 295 in the virtual scene 205.#474855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0147] The competitive module 285 can implement a virtual scene 205 of a particular case for multiple different surgeons or users to perform a medical procedure with. The competitive module 285 can implement a competition or competitive mode where the parameters 270 of the subject are fixed, but the trajectory of movement of instruments by the virtual robotic system (e.g., a virtual 3D model 260 or the robotic medical system 110 and virtual 3D instruments or endoscopes controlled by the user in the virtual scene 205) is unconstrained relative to the trajectory of movement of the instruments recorded when the medical procedure was first performed. In this regard, each user or surgeon who completes a virtual medical procedure in the virtual scene 205 can perform the procedure differently and achieve various results. This can allow surgeons to explore new techniques. Surgeons can perform different actions based on their discretion, e.g., some surgeons may create certain incisions, while other surgeons may not, or the length, size, or depth of the incisions can vary between surgeons. These individualized inputs of the surgeons can result in different changes to the virtual scene 205.
[0148] The simulator 220 can run at least one physics model or modeling techniques to adjust the anatomical structures displayed in the virtual scene 205 over time. This can allow the anatomical structures to be manipulated according to the inputs provided by the client device 106. For example, the simulator 220 can implement modeling techniques to simulate an instrument putting pressure on an anatomical structure, an instrument cutting or dissecting an anatomical structure, the anatomical structure bleeding, the anatomical structure tearing, the anatomical structure being moved by an instrument, etc. This can result in new virtual scenes 205 (or new versions of an existing virtual scene 205) being created different from the scenes that occurred in the video 178.
[0149] The competitive module 285 can generate metrics 297 for each surgeon (or for each session of one surgeon) to track the performance of the surgeon performing the particular case virtually in the simulator 220. The competitive module 285 can provide a virtual scenes 205 that are identical on various client devices 230 for various surgeons (or on various sessions implemented by one client device 106, one session for each surgeon). The virtual scenes 205 can be generated from the same selected frames 235, the same 3D mesh 136, or the same texture file 255. A user or set of users can select a particular case from a menu of cases generated by the simulation module 115, and generate the virtual scenes 205. For each user, the competitive module 285 can track OPIs or performance metrics 297 based484855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) on the input provided on the client device 106 to control the virtual robotic instruments to perform the medical procedure.
[0150] The individual metrics 297 can indicate a virtual amount of energy or power that a user consumed to perform the medical procedure in the virtual scene 205 by operating virtual instruments or virtual endoscopes. The metric 297 can indicate a total duration of a segment of the virtual medical procedure. For example, the metric 297 can be a length of time of the particular segment of the virtual medical procedure, such as a length of time of the step or the phase. The metric 297 can indicate a total linear distance a virtual instrument traveled in the virtual scene 205 according to the user’s inputs via the client device 106 during the segment. For example, the metric 297 can indicate a total distance that the virtual instrument traveled in the virtual scene 205 during the steps, the phases, or the entire video 178. The metric 297 can be a total angular distance the virtual instrument of the robotic medical system 110 traveled in the virtual scene 205 during the segment. The angular distance can indicate an amount or distance that the virtual instrument is rotated in the virtual scene 205 during manipulations. The metric 297 can indicate a total angular distance that the virtual instrument traveled in the virtual scene 205 during the steps, the phases, or the virtual medical procedure. The metric 297 can indicate a total number of virtual operations or a clutch or brake to manipulate the virtual instruments in the virtual scene 205. For example, the metric 297 can indicate a total number of actuations or activations of a clutch or brake for the virtual instrument or the virtual endoscope in the virtual scene 205 (e.g., to float virtual joints of the virtual instruments or endoscopes in the virtual scene 205). The metric 297 can indicate a total number of operations or a brake or clutch during the steps, the phases, or the entire medical procedure.
[0151] The competitive module 285 can include at least one metric comparator 293. The metric comparator 293 can compare individual metrics 297 against one another or against an expert surgeon’s metrics 297. For example, the expert surgeon’s metrics 297 can be the metrics 297 determined from the live medical procedure video 178. In this regard, the metrics 297 determined for virtually performing the medical procedure in the virtual scene 205 can be compared against the metric 297 determined from the live medical procedure video 178. The metric comparator 293 can identify a value of a particular type of metric (e.g., total instrument distance traveled) for the simulation, and identify another value for the same type of metric, but the actual total instrument distance traveled in the real medical494855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) procedure. The metric comparator 293 can cause the client device 106 to display both values of the metric type to the user, so that the user can gauge whether they performed better or worse than the live medical procedure. Similarly, the metric comparator 293 can identify values of a particular type of metric for different surgeons who performed the simulated medical procedure, and cause the client device 106 to display the values of the particular type of metric, so surgeons can compare their performance in performing the simulated medical procedure against other surgeons. The competitive module 285 can track a leaderboard of different surgeons, and display the leaderboard on the client device 106.
[0152] The metric comparator 293 can track metrics 297 of a user as they perform the simulated medical procedure. The metric comparator 293 can continuously or iteratively compare the tracked metrics 297 against thresholds. The thresholds can be predefined, or set for various segments of the virtual medical procedure. For example, the thresholds can be total distance that an instrument can travel during the simulated medical procedure, a total length of time of a segment, or an amount of power that can be consumed by the virtual robotic medical system in the virtual medical procedure. Responsive to a detection that an individual metric 297 satisfies, falls below, or exceeds a threshold, the metric comparator 293 can cause the simulator 220 to reset the virtual scene 205, e.g., return the virtual scene 205 to its initial or starting condition. For example, if a metric 297 indicates that the user has taken more than a threshold length of time to make an incision, the metric comparator 293 can reset the virtual scene 205.
[0153] The generative module 290 can implement a generative mode in which the parameters 287 of the subject are adjustable. For example, the generative module 290 can vary the parameters 287 of the subject (or the anatomical structures of the subject) in the virtual scene 205. Furthermore, the movements of the virtual instruments in the virtual scene 205 can be unconstrained relative to the trajectory of movements of the instruments of the robotic medical system 110 in the live medical procedure. For example, a user can provide any desired input to manipulate the virtual instruments in the virtual scene 205 from the client device 106. Furthermore, the generative module 290 can retrieve the subject parameters 270 from the data repository 265, and cause the subject parameters 270 to be displayed on the client device 106. A user can provide input on the client device 106 to adjust, change, or update the parameters 270. The resulting simulation parameters 287 can be used to run the simulation performed by the simulator 220. For example, a user can change the weight,504855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) height, BMI, age, comorbidities, drug prescriptions, etc. of the subject. The generative module 290 can retrieve patient health records, and a user can input adjustments or changes via the client device 106 to adjust the health records of the patient to influence the model (e.g., adjust BMI, age, previous surgeries, etc.). The adjustments to the parameters 287 can allow a surgeon to generate a realistic practice scenario for the surgeon to plan out an upcoming surgery and attempt techniques before performing the live scenario. The generative module 290 can store multiple videos 178 of historical cases. The historical cases can group the videos 178 into different categories enabling the formulating of a toolbox (surgical instruments) and library (anatomy library, anatomy states library). Such a library could support the creation of a model that can generate completely new cases of surgery with a variety of unique scenarios.
[0154] The generative module 290 can include at least one generative model 283. The generative model 283 can be a 3D generative Al model that outputs a three dimensional mesh 136, texture file 255, or other objects for the virtual scene 205. The generative model 283 can transform an image, such as the selected frame 121, into the 3D mesh 136 or an object for the virtual scene 205. In some implementations, pre-surgical images received from imaging or mapping technologies can be used to produce data for executing the generative model 283 on. The generative model 283 can receive text or parameters 287, and output 3D information using the text or parameters 287. For example, the generative model 283 can a model or technique (or a combination of techniques) such as Neural Radiance Fields (NeRF), Scene Representation Network (SRN), Local Light Field Fusion (LLFF), Stable Diffusion, Stable Diffusion XL (SDXL). The generative model 283 can be a model that is part of a generative Al tool, such as MESHY, STABLE ZERO 123, ALPHA3D, etc. The generative module 290 can transmit requests including a selected frame 121 and simulation parameters 287 to a generative Al tool, which can reply with a 3D object.
[0155] The simulation module 115 can receive a request from the client device 106 to operate in the generative mode. The request can include a request to initiate a generative simulation with the simulator 220 based on a particular selected frame 121. The user, via the client device 106, can specify which frame or frames of the video 178 to use to initiate or start the generative simulation. Responsive to receiving the request, the generative module 290 can identify parameters 287 for producing the 3D mesh 136, the texture file 255, or the virtual scene 205. The generative module 290 can identify parameters 287 that are different514855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) than the parameters of the live scene, e.g., a body weight of the simulated subject for the generative simulation can be different than an actual body weight of the patient. At least one of the identified parameters 287 can be different than the actual parameters of the live medical procedure of the actual scene. At least one of the identified parameters 287 can be the same as actual parameters of the live medical procedure.
[0156] The generative model 283 can execute on the selected frame 121 (e.g., on the depth mesh 138 or the anatomy masks 245) and the identified parameters 287. The generative model 283 can output the 3D mesh 136 and the texture file 255. The simulator 220 can generate, update, or initiate the virtual scene 205 using the 3D mesh 136 or the texture file 255 output by the generative model 283. The simulator 220 can receive input from the client device 106 to move instruments or endoscopes in the virtual scene 205, and interact with the anatomical structures of the subject in the virtual scene 205. The simulator 220 can detect interactions with the virtual scene 205 by a user, and generate updated or new virtual scenes 205. The updated or new virtual scenes 205 can be generated by the generative model 283.
[0157] The simulator 220 can cause the virtual scene 205 to be displayed in an interactive virtual environment 277. The simulator 220 can display both the selected frame 121 and the resulting virtual scene 205. The simulator 220 can cause different graphical user interface elements to be included in one graphical user interface of the client device 106. The elements can be windows, boxes, regions of pixels, etc. A first element can display or include the selected frame 121, while a second element can display or include the virtual scene 205 in an interactive virtual environment 277.
[0158] The simulation module 115 can automatically (e.g., without user input) or semi- automatically (e.g., first ask a user to confirm the switch) switch between different modes, e.g., running the tutorial module 280, the competitive module 285, or the generative module 290. For example, the simulation module 115 can switch from one mode to another responsive to detecting a condition or event. For example, the simulation module 115 can detect, based on data received from the client device 106, that a user is veering or moving a surgical instrument away from an established trajectory for the surgical instrument (e.g., a trajectory of surgical instrument during the actual medical procedure) when in the tutorial mode. For example, if the user navigates a predefined distance away from the established trajectory or navigates away from the trajectory for a predefined length of time, the524855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) simulation module 115 can switch from the tutorial mode of the tutorial module 280 to the competitive mode of the competitive module 285 or the generative mode of the generative module 290. For example, the simulation module 115 can cause the client device 106 to display a prompt asking the user if they wish to switch from the tutorial mode to the competitive mode or the generative mode.
[0159] As another example, the simulation module 115 can be operating in the tutorial mode, but determine based on metrics, such as OPIs, to increase the complexity of the simulated medical procedure by switching to the generative mode of the generative module 290. For example, the simulation module 115 can detect that an OPI of the simulation meets or exceeds a baseline OPI (e.g., an OPI of the actual medical procedure). The simulation module 115 can switch to the generative mode, where the generative module 290 can increase the complexity of the medical procedure, e.g., by increasing the BMI of the patient, increasing the patient’s predisposition to bleeding, etc. The simulation module 115 can automatically switch to the more difficult simulated medical procedure, or can first prompt a user to confirm that they wish to switch modes. Furthermore, the simulation module 115 can switch between modes based on a historical profile of a surgeon, e.g., based on whether historical simulated procedures are typically one type of mode or another, the time during the simulated procedure at which the surgeon changes mode, etc.
[0160] Referring now to FIG. 3, among others, an example 300 of a three dimensional virtual scene 205 initialized using a frame 121 of a medical procedure video 178. In FIG. 3, a video player 305 is shown. The video player 305 can be a graphic component included within a graphical user interface displayed on the client device 106. The video player 305 can play the medical procedure video 178 to the user, and allow the user to scan through frames of the video 178. The user can stop the video player 305 at a particular frame, and provide an input via the interface indicating to select the frame 121 for initiating the virtual scene 205 (or for updating the virtual scene 205).
[0161] The frame 121 selected by the video player 305 can be input into the depth estimator 125 to produce the depth mesh 138. The frame 121 selected by the video player 305 can be input into the anatomy segmentation function 120 to produce the anatomy masks 245. The 3D mesh 136 and the texture file 255 can be used to initialize the virtual scene 205 a first time. However, the virtual scene 205 can be progressively updated. For example, frames 235 can be sampled from the video 178 at an interval, according to rules, or534855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) responsive to user selections, and the sampled frames 235 can be used to update or recreate the virtual scene 205. For example, the models, 3D meshes 250, or texture files 255 of the virtual scene 205 can be updated over time with new frames 235. Each virtual scene 205 can be a new or updated set or collection of models, meshes, or texture files, and each virtual scene 205 can be generated for a different frame 121. For example, an initial frame 121 can be used to create a virtual scene 205 Mo, while an i-th frame 121 can be used to create a virtual scene Mi.
[0162] For example, as shown in FIG. 3, at a second point in time, a second frame 121 can be selected and used to update the virtual scene 205 (e.g., the second frame 121 can be used to create a second depth mesh 138, second anatomy masks 245, second 3D mesh 136, and second texture files 255). For example, as shown in FIG. 3, at a later third point in time, a third frame 121 can be selected and used to update the virtual scene 205 (e.g., the third frame 121 can be used to create a third depth mesh 138, third anatomy masks 245, third 3D mesh 136, and third texture files 255).
[0163] At a fourth point in time, a frame 121 can again be sampled and used to update the virtual scene 205. At the fourth point in time, force data 275 can be used to estimate physics parameters of an anatomical structure (e.g., a soft tissue model). A soft tissue model can be a model of soft tissue that can bend, flex, or deform under pressure, e.g., under pressure of a surgical instrument. The simulator 220 can use the force data 275 to update the 3D mesh 136 or the texture file 255 (or the virtual scene 205) to accurately represent the flexibility, softness, thickness, etc. of tissue of an anatomical structure.
[0164] Referring now to FIG. 4, among others, an example 400 of outcome virtual scenes 205 produced according to different actions performed in the virtual scene 205 initialized using a frame 121 of a medical procedure video 178 is shown. In the example 400, the virtual scene 205 is initialized from a selected frame 121. The user can interact with the virtual scene 205 by providing input via the client device 106 to perform a particular phase 410 (e.g., or a step, act, action, portion, or segment of a medical procedure). For example, the phase 410 can be a phase that occurred in the video 178. The phase 410 can extend from the selected frame 121 to a later frame 405 of the video 178.
[0165] The user (or multiple different users), can interact with the anatomical structures of the virtual scene 205 by providing one or more inputs via the client device 106 to control544855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) or move virtual instruments in the virtual scene 205 to interact with the anatomical structures (e.g., cut, pinch, hold, move, stitch, etc.). The user can attempt performing the phase 410 once, or multiple times. Multiple users can perform the phase 410. The output of each performance of the phase 410 can be a new configuration of the anatomical structures of the original virtual scene 205. The outcome of each virtual scene 205 can be different, and may differ from the actual outcome of the video 178. The multiple different simulation trials can have different results, different outcomes, different OPIs, different timeseries data, different noodle plots, etc. The results of the outcome virtual scenes 205 can be compared against each other, or against the results of the actual medical procedure video 178.
[0166] For example, the competitive module 285 can record metrics 297 for each outcome scene 205, and compare the metrics 297 against baseline metrics for the actual video 178 (e.g., baseline metrics determined from the actual video 178). The competitive module 285 can output, or provide, a result of the comparison to a graphical user interface on the client device 106. For example, the competitive module 285 can output a difference between the individual metrics 297 and the baseline metrics, or an indication of whether the simulated scenario was successful or unsuccessful, or any other comparison.
[0167] Referring now to FIG. 5, among others, an example method 500 of reconstructing a medical procedure environment in a three dimensional virtual scene is shown. At least a portion of the method 500 can be performed by the computing system 105, the simulation module 115, the depth estimator 125, the anatomy segmentation function 120, the mesh generator 130, or the client device 106. The method 500 can include an ACT 505 of identifying a frame capturing a medical procedure scene. The method 500 can include an ACT 510 of generating an anatomical mask and depth data. The method 500 can include an ACT 515 of constructing a 3D mesh and texture data. The method 500 can include an ACT 520 of providing a virtual scene using the 3D mesh and the texture data.
[0168] At ACT 505, the method 500 can include identifying, by the computing system 105, a frame capturing a medical procedure scene. The method 500 can include receiving, by the simulation module 115, the medical procedure video 178 from a robotic medical system 110. The method 500 can include receiving the video 178 capturing a medical scene, e.g., a subject or patient, an operating table, a set of robotic instruments or endoscopes, etc. The method 500 can include receiving the video 178 capturing images or views of anatomical structures, e.g., organs, bones, tissue, veins, arteries, etc. The method 500 can include554855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) identifying, by the frame identifier 225, one or multiple different frames to use to initialize the virtual scene 205. The method 500 can include receiving a selection of the frame 121 from the client device 106. The method 500 can include displaying the video 178 to the user on the client device 106, and the user can scan through the video, and provide an input that selects one or a set of frames 235 to use to construct the virtual scene 205.
[0169] At ACT 510, the method 500 can include generating, by the computing system 105, an anatomical mask and depth data. The method 500 can include generating a depth mesh 138 using the selected frame 121. The method 500 can include generating a depth mesh 138 that includes a set of pixel intensities that represent a real metric distance between a center of a camera to a point in the scene. The method 500 can include determining at least one depth value for at least one pixel of the selected frame 121. The method 500 can include determining a distance from the camera to a surface, anatomical structure, instrument, subject, etc. in the selected frame 121. The method 500 can include determining a depth value for all, every, or most pixels of the selected frame 121. The method 500 can execute one or multiple models on the selected frame 121 to produce the depth mesh 138. The method 500 can include executing a neural network that maps the selected frame 121 to a metric depth. The method 500 can include executing a neural network, such as an encoderdecoder structured with a transformer unit inside.
[0170] The method 500 can include generating, by the anatomy segmentation function 120, anatomy masks 245 for the selected frame 121. The method 500 can include executing one, or multiple models trained by machine learning. For example, the method 500 can include executing a hierarchy of models, one model to identify anatomical structures from the selected frame 121, and another model to identify states of the anatomical structures. The anatomical masks 245 can be masks that identify the pixels in the selected frame 121 that represent a particular anatomical structure. Each anatomical mask 124 can be or include a label that identifies the anatomical structure, e.g., the label can be a name (e.g., heart, lung, bone, etc.) or a code that maps to a name. The method 500 can include determining the state masks 245 at least in part on the anatomical masks 245. The method 500 can include determining the state masks 245 by executing a model with the anatomical masks 245 as at least one input.
[0171] At ACT 515, the method 500 can include constructing, by the computing system 105, a 3D mesh or texture data. The method 500 can include constructing the 3D mesh 136564855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) from at least the depth mesh 138 and the masks 245 (e.g., the anatomical masks 245 and / or the state masks 245. The method 500 can include executing at least one algorithm or technique that translates the depth mesh 138 and detected anatomical structures (e.g., the anatomical masks 245) into a 3D mesh 136 of vertices, edges, and faces. The method 500 can include executing at least one algorithm or technique that causes the vertices, edges, or faces of the 3D mesh anatomical structures to represent states using the state masks 245 (e.g., burning, bruising, charring, bleeding, etc.). Furthermore, the method 500 can map the selected frame 121 onto the 3D mesh 136. The mapping (or the mapped image 235) can be or be included in the texture file 255. The method 500 can include executing at least one generative Al model or technique (or a combination of techniques) such as Neural Radiance Fields (NeRF), Scene Representation Network (SRN), Local Light Field Fusion (LLFF), Stable Diffusion, Stable Diffusion XL (SDXL) using depth mesh 138 or the masks 245 as inputs to output the 3D mesh 136 and / or the texture file 255.
[0172] At ACT 520, the method 500 can include providing, by the computing system105, a virtual scene using the 3D mesh and the texture data. For example, the method 500 can include initializing the virtual scene 205 using the 3D mesh 136 or the texture file 255. The virtual scene 205 can be a collection or set of 3D meshes 250 in various places within a scene (e.g., a facility, an operating room, a doctor’s office). The virtual scene 205 can include various texture files 255 for rending on the 3D meshes 250. The virtual scene 205 can include one or multiple different anatomical structures that are being operated on.
[0173] The method 500 can include providing the virtual scene 205 in an interactive virtual environment 277, e.g., an environment or interface where a user can view the scene 205, and provide inputs to move virtual instruments about the virtual scene 205. The simulator 220 can simulate interactions between the virtual instruments and the anatomical structures. The simulator 220 can model physics in the virtual scene 205, and model cutting, slicing, splittingjoining, sewing, etc. based on user inputs via the client device 106 to move the virtual instruments about the virtual scene 205.
[0174] Referring now to FIG. 6, among others, an example block diagram of a computing system 105 is shown. The computing system 105 can include or be used to implement a data processing system or its components. The architecture described in FIG. 6 can be used to implement the computing system 105, the robotic medical system 110, or the client device106. The computing system 105 can include at least one bus 625 or other communication574855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) component for communicating information and at least one processor 630 or processing circuit coupled to the bus 625 for processing information. The computing system 105 can include one or more processors 630 or processing circuits coupled to the bus 625 for processing information. The computing system 105 can include at least one main memory 610, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 625 for storing information, and instructions to be executed by the processor 630. The main memory 610 can be used for storing information during execution of instructions by the processor 630. The computing system 105 can further include at least one read only memory (ROM) 615 or other static storage device coupled to the bus 625 for storing static information and instructions for the processor 630. A storage device 620, such as a solid state device, magnetic disk or optical disk, can be coupled to the bus 625 to persistently store information and instructions.
[0175] The computing system 105 can be coupled via the bus 625 to a display 600, such as a liquid crystal display, or active matrix display. The display 600 can display information to a user. An input device 605, such as a keyboard or voice interface can be coupled to the bus 625 for communicating information and commands to the processor 630. The input device 605 can include a touch screen of the display 600. The input device 605 can include a cursor control, such as a mouse, a trackball, or cursor direction keys, for communicating direction information and command selections to the processor 630 and for controlling cursor movement on the display 600.
[0176] The processes, systems and methods described herein can be implemented by the computing system 105 in response to the processor 630 executing an arrangement of instructions contained in main memory 610. Such instructions can be read into main memory 610 from another computer-readable medium, such as the storage device 620. Execution of the arrangement of instructions contained in main memory 610 causes the computing system 105 to perform the illustrative processes described herein. One or more processors in a multiprocessing arrangement can be employed to execute the instructions contained in main memory 610. Hard-wired circuitry can be used in place of or in combination with software instructions together with the systems and methods described herein. Systems and methods described herein are not limited to any specific combination of hardware circuitry and software.
[0177] Although an example computing system has been described in FIG. 6, the subject584855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) matter including the operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.
[0178] Some of the description herein emphasizes the structural independence of the aspects of the system components or groupings of operations and responsibilities of these system components. Other groupings that execute similar overall operations are within the scope of the present application. Modules can be implemented in hardware or as computer instructions on a non-transient computer readable storage medium, and modules can be distributed across various hardware or computer based components.
[0179] The systems described above can provide multiple ones of any or each of those components and these components can be provided on either a standalone system or on multiple instantiations in a distributed system. In addition, the systems and methods described above can be provided as one or more computer-readable programs or executable instructions embodied on or in one or more articles of manufacture. The article of manufacture can be cloud storage, a hard disk, a CD-ROM, a flash memory card, a PROM, a RAM, a ROM, or a magnetic tape. In general, the computer-readable programs can be implemented in any programming language, such as LISP, PERL, C, C++, C#, PROLOG, Python, or in any byte code language such as JAVA. The software programs or executable instructions can be stored on or in one or more articles of manufacture as object code.
[0180] Example and non-limiting module implementation elements include sensors providing any value determined herein, sensors providing any value that is a precursor to a value determined herein, datalink or network hardware including communication chips, oscillating crystals, communication links, cables, twisted pair wiring, coaxial wiring, shielded wiring, transmitters, receivers, or transceivers, logic circuits, hard-wired logic circuits, reconfigurable logic circuits in a particular non-transient state configured according to the module specification, any actuator including at least an electrical, hydraulic, or pneumatic actuator, a solenoid, an op-amp, analog control elements (springs, filters, integrators, adders, dividers, gain elements), or digital control elements.
[0181] The subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware,594855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. The subject matter described in this specification can be implemented as one or more computer programs, e.g., one or more circuits of computer program instructions, encoded on one or more computer storage media for execution by, or to control the operation of, data processing apparatuses. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. While a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be, or be included in, one or more separate components or media (e.g., multiple CDs, disks, or other storage devices including cloud storage). The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0182] The terms “computing device”, “component” or “data processing apparatus” or the like encompass various apparatuses, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.
[0183] A computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including604855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program can correspond to a file in a file system. A computer program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0184] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatuses can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Devices suitable for storing computer program instructions and data can include non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0185] The subject matter described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification, or a combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).614855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO)
[0186] While operations are depicted in the drawings in a particular order, such operations are not required to be performed in the particular order shown or in sequential order, and all illustrated operations are not required to be performed. Actions described herein can be performed in a different order.
[0187] Having now described some illustrative implementations, it is apparent that the foregoing is illustrative and not limiting, having been presented by way of example. In particular, although many of the examples presented herein involve specific combinations of method acts or system elements, those acts and those elements may be combined in other ways to accomplish the same objectives. ACTs, elements and features discussed in connection with one implementation are not intended to be excluded from a similar role in other implementations.
[0188] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of “including” “comprising” “having” “containing” “involving” “characterized by” “characterized in that” and variations thereof herein, is meant to encompass the items listed thereafter, equivalents thereof, and additional items, as well as alternate implementations consisting of the items listed thereafter exclusively. In one implementation, the systems and methods described herein consist of one, each combination of more than one, or all of the described elements, acts, or components.
[0189] Any references to implementations or elements or acts of the systems and methods herein referred to in the singular may also embrace implementations including a plurality of these elements, and any references in plural to any implementation or element or act herein may also embrace implementations including only a single element. References in the singular or plural form are not intended to limit the presently disclosed systems or methods, their components, acts, or elements to single or plural configurations. References to any ACT or element being based on any information, act or element may include implementations where the act or element is based at least in part on any information, act, or element.
[0190] Any implementation disclosed herein may be combined with any other implementation or example, and references to “an implementation,” “some implementations,” “one implementation” or the like are not necessarily mutually exclusive and are intended to624855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) indicate that a particular feature, structure, or characteristic described in connection with the implementation may be included in at least one implementation or example. Such terms as used herein are not necessarily all referring to the same implementation. Any implementation may be combined with any other implementation, inclusively or exclusively, in any manner consistent with the aspects and implementations disclosed herein.
[0191] References to “or” may be construed as inclusive so that any terms described using “or” may indicate any of a single, more than one, and all of the described terms. References to at least one of a conjunctive list of terms may be construed as an inclusive OR to indicate any of a single, more than one, and all of the described terms. For example, a reference to “at least one of ‘A’ and ‘B’” can include only ‘A’, only ‘B’, as well as both ‘A’ and ‘B’. Such references used in conjunction with “comprising” or other open terminology can include additional items.
[0192] Where technical features in the drawings, detailed description or any claim are followed by reference signs, the reference signs have been included to increase the intelligibility of the drawings, detailed description, and claims. Accordingly, neither the reference signs nor their absence have any limiting effect on the scope of any claim elements.
[0193] Modifications of described elements and acts such as variations in sizes, dimensions, structures, shapes and proportions of the various elements, values of parameters, mounting arrangements, use of materials, colors, orientations can occur without materially departing from the teachings and advantages of the subject matter disclosed herein. For example, elements shown as integrally formed can be constructed of multiple parts or elements, the position of elements can be reversed or otherwise varied, and the nature or number of discrete elements or positions can be altered or varied. Other substitutions, modifications, changes and omissions can also be made in the design, operating conditions and arrangement of the disclosed elements and operations without departing from the scope of the present disclosure.634855-6965-8828.1
Claims
Atty. Dkt: 135039-0493 (P06965-WO)CLAIMSWhat is claimed is:
1. A system, comprising: one or more processors, coupled with memory, to: identify a frame that captures a scene in a medical procedure performed with a robotic medical system; generate, using one or more models trained with machine learning, one or more masks that segment an anatomical structure in the frame, label the anatomical structure in the frame, and indicate a depth value for each pixel in the frame; construct 3D mesh and texture data based on the one or more masks; and provide, in an interactive virtual environment, a virtual scene rendered using the 3D mesh and texture data generated from the frame that captures the scene in the medical procedure.
2. The system of claim 1, wherein the one or more processors are further configured to: generate, using the one or more models, the one or more masks comprising segments and labels for a plurality of anatomical structures in the frame; construct the 3D mesh and texture data for the plurality of anatomical structures in the frame; and provide the virtual scene rendered using the 3D mesh and texture data for the plurality of anatomical structures in the frame.
3. The system of claim 1, wherein the one or more masks comprise at least one of a segmentation mask and a state mask.
4. The system of claim 1, wherein the frame corresponds to a single view of the scene in the medical procedure.
5. The system of claim 1, wherein the one or more processors are further configured to: determine, based on the frame and using the one or more models, a state of the644855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) anatomical structure in the frame; construct, based on the state, the 3D mesh and texture data; and provide, in the interactive virtual environment, the virtual scene rendered to represent the state of the anatomical structure.
6. The system of claim 5, wherein the state comprises at least one of burned, cut, or bleeding.
7. The system of claim 1, wherein the one or more processors are further configured to: display the interactive virtual environment via a display device.
8. The system of claim 7, wherein the display device is coupled with a head-mounted display.
9. The system of claim 1, wherein the one or more processors are further configured to: receive a data stream comprising information about a state of a subject that undergoes the medical procedure; and construct the 3D mesh and texture data based on the state of the subject.
10. The system of claim 9, wherein the one or more processors are further configured to: access a data repository storing parameters for the subject; and construct the 3D mesh and texture data using the parameters for the subject.
11. The system of claim 10, wherein the parameters for the subject provide an indication of a likelihood of a state of the anatomical structure during the medical procedure.
12. The system of claim 1, wherein the one or more processors are further configured to: receive a kinematics data stream captured via one or more sensors associated with the scene in the medical procedure; and construct the 3D mesh and texture data using the kinematics data stream.
13. The system of claim 12, wherein the kinematics data stream indicates an amount force detected between an instrument of the robotic medical system and at least a portion of the654855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) anatomical structure in the scene.
14. The system of claim 1, wherein the one or more processors are further configured to: display a video of the medical procedure, the video comprising a plurality frames that comprises the frame; and receive, via an interface, a request to establish the virtual scene in the interactive virtual environment for the frame of the plurality of frames.
15. The system of claim 14, wherein the one or more processors are further configured to: receive, via an interface, a request to initiate a simulation in the interactive virtual environment for the frame of the scene in the medical procedure.
16. The system of claim 14, wherein the one or more processors are further configured to: receive a request to establish the virtual scene with a view angle that is different from the view angle with which the frame is captured for the scene in the medical procedure.
17. The system of claim 14 or 16, wherein the one or more processors are further configured to: detect a trigger condition in the video; determine, responsive to the trigger condition, to generate a prompt comprising a recommendation to establish the virtual scene for the medical procedure; and receive the request to establish the virtual scene responsive to the prompt.
18. The system of claim 17, wherein the trigger condition is based on at least one of a type of task performed with the robotic medical system in the video, an occurrence of a milestone event in the frame, or a performance metric associated with the medical procedure.
19. The system of any one of claim 1 or 14-18, wherein the one or more processors are further configured to: establish a mode of interaction for the interactive virtual environment, wherein the mode of interaction comprises at least one of: i) a tutorial mode in which parameters of a subject associated with the medical664855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) procedure are fixed and a trajectory of movement of a virtual robotic medical system in the virtual scene is constrained based on the trajectory of movement of the robotic medical system during the medical procedure, ii) a competition mode in which the parameters of the subject associated with the medical procedure are fixed and the trajectory of movement of the virtual robotic system is unconstrained relative to the trajectory of movement of the robotic medical system during the medical procedure, or iii) a generative mode in which parameters for a virtual subject in the virtual scene are adjustable and the trajectory of movement of the virtual robotic system is unconstrained relative to the trajectory of movement of the robotic medical system during the medical procedure.
20. The system of claim 1, wherein the one or more processors are further configured to: receive, via an interface, a request to initiate a generative simulation based on the frame of the scene; identify a plurality of parameters for the generative simulation, wherein at least one of the plurality of parameters is different than a parameter of the scene in the medical procedure; and construct, using the one or more models, the virtual scene using the 3D mesh and texture data based on the one or more masks and the plurality of parameters for the generative simulation.
21. The system of claim 20, wherein the one or more processors are further configured to: detect, via the interface, an interaction in the virtual scene rendered in the interactive virtual environment; generate, using the one or more models, one or more subsequent virtual scenes responsive to the interaction; and provide, in the interactive virtual environment, the one or more subsequent virtual scenes.
22. The system of any one of claims 1, 20 or 21, wherein the one or more processors are further674855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) configured to: identify a first value of a performance metric associated with the medical procedure performed with the robotic medical system; identify a second value of the performance metric associated with an interaction in the virtual scene; and display the first value of the performance metric and the second value of the performance metric.
23. The system of claim 22, wherein the one or more processors are further configured to: determine, based on a comparison of the second value with a threshold, to reset the virtual scene in the interactive virtual environment.
24. The system of claim 1, wherein the one or more processors are further configured to: display the frame in a first graphical user interface element in the interactive virtual environment; and display the virtual scene in a second graphical user interface element in the interactive virtual environment.
25. A method, comprising: identifying, by one or more processors coupled with memory, a frame that captures a scene in a medical procedure performed with a robotic medical system; generating, by the one or more processors, using one or more models trained with machine learning, one or more masks that segment an anatomical structure in the frame, label the anatomical structure in the frame, and indicate a depth value for each pixel in the frame; constructing, by the one or more processors, 3D mesh and texture data based on the one or more masks; and providing, by the one or more processors, in an interactive virtual environment, a virtual scene rendered using the 3D mesh and texture data generated from the frame that captures the scene in the medical procedure.
26. A non-transitory computer-readable medium storing processor-executable instructions that,684855-6965-8828.1Atty. Dkt: 135039-0493 (P06965-WO) when executed by one or more processors, cause the one or more processors to: identify a frame that captures a scene in a medical procedure performed with a robotic medical system; generate, using one or more models trained with machine learning, one or more masks that segment an anatomical structure in the frame, label the anatomical structure in the frame, and indicate a depth value for each pixel in the frame; construct 3D mesh and texture data based on the one or more masks; and provide, in an interactive virtual environment, a virtual scene rendered using the 3D mesh and texture data generated from the frame that captures the scene in the medical procedure.694855-6965-8828.1
Citation Information
Patent Citations
Virtual reality training, simulation, and collaboration in a robotic surgical system
WO2019006202A1
US202463566531P
Cited By
Method and System for Generating Simulated Videos of Laparoscopic Robotic Operation
CN122289311A