3D view reconstruction of a surgical scene from an image

The system addresses the limitations of 2D monocular images in surgical robotics by generating a 3D mesh from a single image, enhancing surgical scene visualization and improving procedural efficiency through multiple view angles.

WO2026043786A1PCT designated stage Publication Date: 2026-02-26INTUITIVE SURGICAL OPERATIONS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/042385
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-19
Filing Date
2025-08-18
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Existing medical robotic systems using 2D monocular images for surgical procedures face limitations in providing multiple view angles, leading to obstructed views due to patient anatomy or instruments, making it difficult for surgeons to verify surgical milestones in real-time.

Method used

A system utilizing machine learning models to generate a 3D mesh from a single monocular image, segmenting anatomical structures, estimating depth, and reconstructing the scene to provide multiple view angles, allowing for more efficient and proactive visualization of surgical scenes.

Benefits of technology

Enables surgeons to visualize surgical scenes from multiple angles, improving procedural efficiency and facilitating better understanding of spatial arrangements, especially during real-time milestone events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025042385_26022026_PF_FP_ABST
    Figure US2025042385_26022026_PF_FP_ABST
Patent Text Reader

Abstract

Three-dimensional view reconstruction of a surgical scene using an image is provided. A system can identify a frame that captures a single view angle of a scene in a medical procedure performed with a robotic medical system. The system can generate, using one or more models trained with machine learning, one or more masks that segment an anatomical structure in the frame, label the anatomical structure in the frame, and indicate a depth value for each pixel in the frame. The system can construct a 3D mesh based on the one or more masks. The system can generate a video corresponding to a plurality of view angles of the frame using the one or models and the 3D mesh and the labels of the anatomical structures from the one or more masks.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Dkt: 135039-0420 (P06964-WO)3D VIEW RECONSTRUCTION OF A SURGICAL SCENE FROM AN IMAGECROSS-REFERENCES TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority under 35 U.S.C. § 119 to U.S. Provisional Patent Application No. 63 / 684,540, filed August 19, 2024, which is hereby incorporated by reference herein in its entirety.BACKGROUND

[0002] A medical robotic system can include an instrument for performing a medical session or procedure. For example, the instrument can be used to perform surgery, therapy, or a medical evaluation. The medical robotic system can include an endoscope that captures a video of the medical procedure.SUMMARY

[0003] Technical solutions disclosed herein can include a three dimensional (3D) view reconstruction of a surgical scene from a monocular image of a robotic medical procedure. Sometimes, surgical scenes can be confined within spaces more suitable for 2D image camera systems whose images can be limited to single view-angle views of the surgical area. A robotic medical system can utilize such a 2D image camera to allow a surgeon to perform the robotic surgery. However, as the monocular images of the surgical scenery can be limited to only a single view-angle provided by the 2D image, it may be difficult for surgeons observing the surgical procedure to complete visual verification of particular surgical milestones given the view-angle of the monocular image. Moreover, 2D images can sometimes include certain portions of the surgical scene being obstructed by patient’s anatomy structures or medical instruments, making the visibility even more difficult. As a result, it can be beneficial for the robotic medical system to provide the surgeon with more than one view-angle of the surgical scenery in a manner that is more compute and energy efficient than when the surgical scene is reconstructed using the sequence of images. The technical solutions of this disclosure overcome these challenges by providing ML-based 3D view reconstruction of the surgical scene using a 3D mesh of the surgical scene generated based on segmentation and depth masks constructed from a single monocular image.14855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)

[0004] An aspect of the present disclosure can be directed to a system. The system can include one or more processors, coupled with memory. The one or more processors can be configured to identify a frame that captures a single view angle of a scene in a medical procedure performed with a robotic medical system. The one or more processors can be configured to generate, using one or more models trained with machine learning, one or more masks that segment an anatomical structure in the frame, label the anatomical structure in the frame, and indicate a depth value for each pixel in the frame. The one or more processors can be configured to construct a 3D mesh based on the one or more masks. The one or more processors can be configured to generate a video corresponding to a plurality of view angles of the frame using the one or models and the 3D mesh and the labels of the anatomical structures from the one or more masks.

[0005] The one or more processors can be configured to detect a trigger condition in a video stream of the medical procedure. The one or more processors can be configured to select the frame from the video stream responsive to detection of the trigger condition. The one or more processors can be configured to determine to generate the video for the frame based on the detection of the trigger condition. The one or more processors can be configured to detect the trigger condition based on a type of task performed with the robotic medical system in the video stream.

[0006] The one or more processors can be configured to receive a data stream of the medical procedure captured via one or more sensors of the robotic medical system. The one or more processors can be configured to determine, based on the data stream, a range of view angles in a portion of the video that comprises the frame. The one or more processors can be configured to detect the trigger condition based on the occurrence of the milestone event in the frame and the range of view angles in the portion of the video stream being less than or equal to a threshold.

[0007] The one or more processors can be configured to detect the trigger condition based on a profile associated with a surgeon that performs the medical procedure with the robotic medical system. The one or more processors can be configured to generate a first mask of the one or more masks using a vision transformer based neural network trained with a dataset comprising images of medical procedures that are annotated with labels, wherein the first mask comprises the segment of the anatomical structure and the label of the anatomical structure.24855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)

[0008] The one or more processors can be configured to generate a second mask of the one or more masks using a neural network trained to map an intensity of each pixel in the frame to the depth value. The depth value can be a distance between a camera that captures the frame and a location in the scene corresponding to the pixel. The one or more processors can be configured to generate the second mask based on at least one of an optical center of the camera, a focal length of the camera, a scale factor of the camera, a principal point of the camera, a skew of the camera, or a geometric distortion of the camera.

[0009] The one or more processors can be configured to construct the 3D mesh comprising a plurality of vertices, edges and faces that provide a three-dimensional representation of the scene. The one or more processors can be configured to generate, using the one or more models, a plurality of images for the plurality of view angles of the frame. The one or more processors can be configured to identify one or more holes in the plurality of images. The one or more processors can be configured to update the plurality of images using an inpainting model to fill the one or more holes in the plurality of images. The one or more processors can be configured to generate the video with the updated plurality of images.

[0010] The one or more processors can be configured to synthesize, using the inpainting model, new pixels based on the segments and the labels in the one or more masks. The one or more processors can be configured to fill the one or more holes with the new pixels. The one or more processors can be configured to assign, to the new pixels, the label and the depth value based on the one or more masks. The inpainting model can include a generative machine learning model. The one or more processors can be configured to generate a prompt based on the segments and the labels in the first mask. The one or more processors can be configured to input the prompt into the generative machine learning model to synthesize new pixels. The one or more processors can be configured to fill the one or more holes with the new pixels synthesized by the generative machine learning model responsive to the prompt.

[0011] The inpainting model can include a generative machine learning model. The one or more processors can be configured to generate the prompt based on the second mask that indicates the distance between pixels in the frame and the reference point.

[0012] The one or more processors can be configured to receive, via a graphical user interface element, an input with a range of view angles for which to generate the video. The one or more processors can be configured to generate the video corresponding to the plurality34855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) of view angles responsive to the input with the range of view angles. The one or more processors can be configured to display the video via a head-mounted display device.

[0013] An aspect of this disclosure can be directed to a method. The method can be performed by one or more processors coupled with memory. The method can include the one or more processors identifying a frame that captures a scene in a medical procedure performed with a robotic medical system. The method can include one or more processors generating, using one or more models trained with machine learning, a first mask that segments an anatomical structure in the frame and labels the anatomical structure in the frame. The method can include one or more processors generating, using the one or more models, a second mask that indicates a depth value for each pixel in the frame. The method can include one or more processors constructing a 3D mesh based on the first mask and the second mask. The method can include one or more processors generating, by the one or more processors, a plurality of images corresponding to a plurality of view angles of the frame using the one or models and the 3D mesh and the labels of the anatomical structures.

[0014] The method can include one or more processors receiving a request to generate a three-dimensional video for the frame. The method can include one or more processors generating, responsive to the request, the plurality of images to create the three-dimensional video. The method can include detecting, a trigger condition in a video stream of the medical procedure. The method can include selecting the frame from the video stream responsive to detection of the trigger condition. The method can include determining to generate the video for the frame based on the detection of the trigger condition.

[0015] An aspect of the technical solutions is directed to a non-transitory computer- readable medium storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to identify a frame that captures a single view angle of a scene in a medical procedure performed with a robotic medical system. The instructions, when executed by the one or more processors, can cause the one or more processors to generate, using one or more models trained with machine learning, one or more masks that segment an anatomical structure in the frame, label the anatomical structure in the frame, and indicate a depth value for each pixel in the frame. The instructions, when executed by the one or more processors, can cause the one or more processors to construct a 3D mesh based on the one or more masks. The instructions, when executed by the one or more processors, can cause the one or more processors to generate a video corresponding to a44855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) plurality of view angles of the frame using the one or models and the 3D mesh and the labels of the anatomical structures from the one or more masks.

[0016] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of various aspects and implementations and provide an overview or framework for understanding the nature and character of the claimed aspects and implementations. The drawings provide illustration and a further understanding of the various aspects and implementations and are incorporated in and constitute a part of this specification. The foregoing information and the following detailed description and drawings include illustrative examples and should not be considered as limiting.BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings are not intended to be drawn to scale. Like reference numbers and designations in the various drawings indicate like elements. For purposes of clarity, not every component may be labeled in every drawing. In the drawings:

[0018] FIG. 1 depicts an example system for providing a 3D view reconstruction of a surgical scene using a 2D monocular image frame.

[0019] FIG. 2 depicts an example of a system diagram for performing 3D view reconstruction of a surgical scene.

[0020] FIG. 3 illustrates an example view of a surgical scene captured by a 2D monocular frame.

[0021] FIG. 4 illustrates an example view of a surgical scene in a view-angle frame before and after the inpainting process.

[0022] FIG. 5 depicts an example method of providing a 3D view reconstruction of a surgical scene using a 2D monocular image frame.

[0023] FIG. 6 depicts an example architecture of a computing system.54855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)DETAILED DESCRIPTION

[0024] Following below are more detailed descriptions of various concepts related to, and implementations of, methods, apparatuses, and systems for 3D view reconstruction of a surgical scene from an image. The various concepts introduced above and discussed in greater detail below may be implemented in any of numerous ways.

[0025] When recording of a surgical scene using 2D images, the user observing the surgical procedure can be limited to the single view angle of the 2D images provided. However, when aiming to view particular visual details in the surgical scene recorded by the 2D image, the user (e.g., the surgeon viewing the surgery) can find it more helpful to observe the scene form a perspective that is different than the perspective of the 2D image. For instance, during a surgical procedure, a surgeon may encounter a milestone event in which the surgeon is supposed to verify the presence or the absence of a particular object or feature in the scene, such as a completion of removal of a tissue or completion of a closing of a cut. However, as some surgical procedures may be confined to small spaces, or a narrow field of view or a single-view camera (e.g., a monocular camera) may be the only available option, the view of a particular object of interest may be obstructed or difficult to see in the 2D image. For example, a surgical scene can be such that a piece of a tissue or a medical instrument obstructs an object of interest at the view angle of the 2D monocular image provided, making it hard for the observer to complete the milestone event by confirming the viewing of the area of interest. The technical solutions of this disclosure allow for a multiview 3D scene generation (e.g., 3D perspective) based on the 2D monocular single-view image, facilitating multi-angle views that improve the visibility of the surgical scene, improving the surgeon’s performance both in intra-procedure and in post-procedure scenarios.

[0026] Implementation of stereoscopic depth information (e.g., continuous mapping and localizing based on image data) can be very compute and resource intensive and energy inefficient. As a result, 3D reconstruction of surgical scenes can be compute intensive and challenging to implement intra-operatively (e.g., in real time and during the ongoing procedure). Meanwhile, robotic medical surgeries can include tasks with real-time milestone moments in which the surgeon seeks to verify or check a task completion before proceeding to the next stage. For instance, in a cholecystectomy, prior to proceeding to the next stage of a procedure, a surgeon can seek to verify that a cystic duct is separated from the gallbladder or64855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) that a fat tissue on the cystic duct is removed such that a particular surgical area is cleared from any surrounding anatomies. In these scenarios, not having a clear visibility of a particular detail due to a 2D view limitations in real time, can make the procedure more complicated and challenging.

[0027] The technical solutions of this disclosure overcomes these challenges by reconstructing a 3D version of the surgical scene using a single monocular image and one or more machine learning models. Using masks indicative of anatomical segment features and the pixel depth from a camera reference point, the technical solutions can generate a 3D mesh of the imaged surgical scene. Using the 3D mesh, the technical solutions can generate various view-angle frame images from different view-angles, allowing for a 3D presentation (e.g., a video) that can be generated in a more compute efficient manner than may be the case with image sequencing. Using the 3D reconstructed scene, the technical solutions can allow the surgeons to visualize the depth within a surgical recording and explore the surgical workspace from multiple view angles that are not present in the monocular 2D image. In addition to facilitating a multi angle viewing in real-time, the solutions also allow for more proactive user experience. For instance, in student-teacher situations when an experienced surgeon reviews a procedure with a resident, both viewers can benefit from multi-angle viewing of the area. Also, in individual reviews of the surgical procedure in a post-operative analysis, it can be beneficial for the surgeons to look back at a scene at a given moment in time and understand the complete scenery they were experiencing live to consider other options or approaches that may be taken in the future. The technical solutions allow users to explore structures in images in which the spatial arrangements (e.g., obstructions) of the monocular camera may not allow the surgeons to view all of the structures or details clearly using the 2D camera’s viewing angle alone.

[0028] FIG. 1 depicts an example system 100 for providing a 3D view reconstruction of a surgical scene using a 2D monocular image frame. System 100 can include one or more medical environments 102 in which medical procedures (e.g., robotic surgeries) can be performed. A medical environment 102 can include one or more sensors 104 for detection of various sensor data 174 and one or more data capture devices 108 (e.g., cameras) for capturing data, such as image frames 121 (e.g., monocular images). The medical environment 102 can include one or more visualization tools 114 for facilitating data visualization and one or more displays 116 for displaying data. The medical environment can include one or more74855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) robotic medical systems (RMSs) 110 used by users (e.g., surgeons) to perform medical surgeries on a patient. The users can utilize one or more head mounted devices (HMDs) 122 that can include integrated one or more displays 116 and sensors 104.

[0029] The devices of the medical environment 102 can communicate with a data processing system 118 via a network 101. The data processing system 118 can include one or more frame functions 120 configured to receive and process image frames 121 (e.g., 2D monocular images). Data processing system 118 can include one or more segmentation functions 122 for generating masks 124 to identify anatomy segments 128 and provide labels 126. The data processing system 118 can include one or more depth estimators 125 for generating depth values 132. The data processing system 118 can include one or more mesh generators 130 for generating 3D meshes 136, using for example, one or more inpainting functions 134. The data processing system 118 can include one or more trigger condition detectors 144 for detecting trigger conditions 146 (e.g., milestone events during the surgical procedure) based on one or more thresholds 148. Using the 3D mesh 136, a multi-view generator 140 can generate one or more view-angle frames 142 that can depict the scene of the frame 121 from different view angles. The data processing system 118 can include or more 3D video generators 150 for generating videos 152 (e.g., 3D video multi-angle view simulations) from the 3D meshes 136 and the view-angle frames 142. The data processing system 118 can include one or more prompt generators 154 for generating prompts 156 for using one or more machine learning (ML) models 182. The data processing system 118 can include or more data repositories 160 storing surgeon data 172 and one or more data streams 162 that can include sensor data 174 and video data 178. The data processing system 118 can include one or more machine learning (ML) frameworks 180. An ML framework 180 can include one or more of: ML models 182, ML trainers 184, label annotators 186 and encoder and decoder functions 190. The data processing system 118 can include one or more user interfaces 192, such as graphical user interfaces for providing user presentations or outputs to the users of the RMS 110.

[0030] The RMS 110 can include any robotic system that can utilize or manipulate medical instruments 112 to perform medical procedures, such as a robotic surgery. Robotic medical system 110, also referred to as an RMS 110, can be deployed in any medical environment 102. A medical environment 102 can include any space or facility for performing medical procedures (e.g., robotic surgeries), including for example any surgical facility, or an operating room. A medical environment 102 can include any number of84855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) different medical instruments 112 that the RMS 110 can use for performing surgical patient procedures, whether invasive, non-invasive, in-patient, or out-patient procedures.

[0031] The medical environment 102 can include one or more data capture devices 108 (e.g., optical devices, such as cameras or sensors or other types of sensors or detectors) for capturing data streams 162, that can include video data 178 of images or a video stream of a surgery as well as any sensor data 174. The medical environment 102 can include one or more visualization tools 114 to gather the captured data streams 162 and process it for display to the user (e.g., a surgeon or other medical professional) at one or more displays 116. A display 116 can present data stream 162 (e.g., video frames, kinematics or sensor data) of an ongoing medical procedure (e.g., an ongoing surgery) performed using the robotic medical system 110 handling, manipulating, holding or otherwise utilizing medical instruments or tools 112 to perform surgical tasks at the surgical site.

[0032] Data repository 160 can include various data streams 162 generated by the RMS 110, including various sensor data 174 and video data 178 that can be collected, organized, stored and provided for use to various data processing system components (e.g., segmentation function 122, depth estimator 125, mesh generator 130, multi-view generator 140, trigger condition detector 144, 3D video generator 150, prompt generator 154, or any of the ML models 182 that can be utilized). For instance, ML framework 180 can use data streams 162 as inputs into one or more ML models 182, such as ML models 182 trained to generate masks 124 for segmenting anatomy segments 128 (e.g., anatomical structure) from image frames 121 or ML models 182 trained to generate view-angle frames 142 from a plurality of view angles. ML framework 180 can include one or more ML model trainers 184 for training the ML models along with attention mechanisms 188 that can be utilized by the ML models for detection of anatomies, instruments, various instrument to anatomy interactions, medical procedure tasks, anatomy segments 128, depth values 132, trigger conditions 146, view-angle frames 142, 3D meshes 136 or any other features represented or captured in the image frames 121. ML framework 180 can include one or more encoder and decoder functions 190 for processing and detection of image features and for processing data streams 162 for data relevant to ML framework determinations.

[0033] Machine learning (ML) framework 180 can include any combination of hardware and software for providing a system that integrates ML models 182 with attention mechanisms 188 or rule-based modeling to generate masks 124 for labeling of anatomy segments 128, generating view-angle frames 142, synthesize new pixels to add (e.g., for94855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) inpainting function 134) and generate prompts 156 for a prompt generator 154 to new pixel synthesis (e.g., inpainting). ML trainers 184 can be used for training, setting or configuring ML models 182 and their related functions or components, such as label annotators 186, attention mechanisms 188 or encoder and decoder functions 190.

[0034] ML framework 180 can include ML models 182 trained, set or configured to perform variety of tasks for 3D view reconstruction. ML models 182 can include any type and form of artificial intelligence (Al) models, such as neural network models, transformerbased mechanism models, any graph neural network. For example, ML models 182 can be trained or configured to generate masks 124 for labeling of anatomy segments 128 via labels 126 on behalf of a segmentation function 122. ML models 182 can be trained or configured to generate prompts 156 on behalf of a prompt generator 154 to cause other one or more ML models 182 to synthesize new pixels to match anatomy segments 128 of areas (e.g., holes) in a multi view-angle frame 142 generated by multi-view generator 140. The ML models 182 can be configured to generate view-angle frames 142 or synthesize new pixels to add into view-angle frames 142 (e.g., on behalf of an inpainting function 134). The ML models 182 can be configured to generate view angle frames 142 for a plurality of view angles (e.g., perspectives) surrounding the view-angle of the monocular image frame 121, allowing for generation of a 3D video 152, allowing the user to view a surgical scene from multiple viewangle perspectives.

[0035] ML models 182 can include any type of ML or Al architecture for processing images and generating multi-view angle images of a surgical scene, or for filling in missing data from 2D image. ML models 182 can include, for example, neural networks, such as Convolutional Neural Networks (CNNs) that can be configured to recognize and reconstruct spatial features (e.g., anatomy segments 128) from frames 121. ML models 182 can include self-attention mechanisms, such as those in transformer models, which can be configured to enhance image generation by focusing on different parts of the image frame 121. ML models 182 can include Generative Adversarial Networks (GANs) that can be configured to create realistic images (e.g., or portions of images) from incomplete data (e.g., synthesize missing pixels) using a generator and a discriminator. ML models 182 can include Variational Autoencoders (VAEs) that can generate new images, or portions of new images (e.g., fill in the holes in the view-angle frames 142) by learning the distribution of input data and sampling from it. For 3D reconstructions, 3D-CNNs and voxel -based functionalities can be used to process volumetric data to create accurate 3D models from 2D images. For instance,104855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)ML models 182 can include or utilize Neural Radiance Fields (NeRF) that can synthesize new regions of images (e.g., fill in the holes) by processing or optimizing a volumetric scene function (e.g., 3D mesh). ML model 182 can include or use a Reinforcement Learning (RL) that can improve model accuracy in predicting desired or optimal angles for multi-view reconstruction.

[0036] ML framework 180 can include attention mechanisms 188, implemented as neural networks, which enable the extraction of spatial and temporal features from the input data steams 162. Attention mechanisms 188 can facilitate or improve the capacity of the ML models to discern, detect or recognize specific details within the surgical context, thereby improving the accuracy of detection and recognition tasks. ML framework 180 can include and provide rule-based modeling to determine and quantify the consistency of anatomical features or segments along an area. ML framework 180 can include and provide a framework for generating a 3D mesh of various anatomy segments 128 that are labeled using labels 126 and for inpainting empty spaces (e.g., holes) generated by view-angle movement over obstructing objects, such as by using the tissue texture (e.g., color and shapes) of areas of the same anatomy (e.g., organ or tissue type) surrounding the hole. By integrating encoder and decoder functions 190 for extracting image features or time series for using kinematics data and sensor measurements (including force data), the ML framework 180 improves the quality of the detection and recognition by the ML models. For example, the ML framework 180 can utilize attention mechanisms 188 to focus on the texture, color or shape patterns of the surround tissue of the same anatomy segment 128 that is to be inpainted or filled-in by synthesized pixels to be generated for the 3D mesh 136 or view-angle frames 142.

[0037] Encoder and decoder functions 190 can include any combination of algorithms and neural network architectures designed for transforming and reconstructing data representations. Encoder and decoder functions 190 can be utilized in by segmentation function 122 for anatomy segmentation and by depth estimator 125 for metric depth estimation. The encoder functionality of the encoder and decoder functions 190 can process the image frame 121, compressing its information into a lower-dimensional latent space, capturing features such as anatomical structures or depth information. The decoder function of the encoder and decoder functions 190 can take the compressed representation and reconstruct it back into a detailed output, such as a segmentation mask 124 or a metric depth mask 138. Encoder and decoder functions 190 can be combined with transformer units, allow for identifying relationships within the data.114855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)

[0038] Using the ML framework 180, the technical solutions can provide RMS 110 users with a spatial awareness intra-operatively, thereby facilitating a greater understanding than is achievable with 2D monocular image frames 121. By using 2D images (e.g., frames 121) to capture moments in real time, the technical solutions allow for the reconstruction of the 3D scenery as the 2D image recording continues. The data processing system 118 can provide for rendering of a 2D monocular image into a plurality of view-angle frames 142 (e.g., based on user input device selections or movements in a user interface 192), or for a 3D view video 152 that includes viewing angle change in multiple spatial directions around the original viewing angle (e.g., via changes to spatial Cartesian coordinates X-Y-Z) and attention point zoom-in / out. The 2D image of the scene can be used as an input to a modularized pipeline. The pipeline can include an input framing in which the system can render a single 2D image from surgery into 3D. The system can sample the selected frame from the video recording and pass through one or more modules (e.g., data processing system 118 components).

[0039] Data repository 160 of the DPS 118 can include one or more data streams 162, such as video data 178 including a stream of video frames. Data streams 162 can include measurements from sensors, which can be referred to as sensor data 174 and which can include various force, torque or biometric data, haptic feedback data, pressure or temperature data, vibration, tension or compression data, endoscopic images or data, ultrasound images or videos or communication and command data streams. Data repository 160 can include installation data, such as system files or logs including time stamps and data on installation, activation, calibration or use of particular medical instruments 112.

[0040] The system 100 can include one or more data capture devices 108 (e.g., video cameras, sensors or any other detectors) for collecting any data stream 162. Data capture devices 108 can include cameras or other image capture devices for capturing video data 178 (e.g., videos or images) from a particular viewpoint within the medical environment 102. The data capture devices 108 can be positioned, mounted, or otherwise located to capture content from any viewpoint that facilitates the data processing system 118 capturing various surgical tasks or actions.

[0041] Data capture devices 108 can include any of a variety of sensors, cameras, video imaging devices, infrared imaging devices, visible light imaging devices, intensity imaging devices (e.g., black, color, grayscale imaging devices, etc.), depth imaging devices (e.g., stereoscopic imaging devices, time-of-flight imaging devices, etc.), medical imaging devices124855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) such as endoscopic imaging devices, ultrasound imaging devices, etc., non-visible light imaging devices, any combination or sub-combination of the above mentioned imaging devices, or any other type of imaging devices that can be suitable for the purposes described herein. Data capture devices 108 can include cameras that a surgeon can use to perform a surgery and observe manipulation components within a purview of field of view suitable for the given task performance.

[0042] Data capture devices 108 can capture, detect, or acquire sensor data, such as videos or images, including for example, still images (e.g., monocular 2D images from a single-view angle), video images, vector images, bitmap images, other types of images, or combinations thereof. The data capture devices 108 can capture the images at any suitable predetermined capture rate or frequency. Settings, such as zoom settings or resolution, of each of the data capture devices 108 can vary as desired to capture suitable images from any viewpoint. For instance, data capture devices 108 can have fixed viewpoints, locations, positions, or orientations. The data capture devices 108 can be portable, or otherwise configured to change orientation or telescope in various directions. The data capture devices 108 can be part of a multi-sensor architecture including multiple sensors, with each sensor being configured to detect, measure, or otherwise capture a particular parameter (e.g., sound, images, or pressure).

[0043] Data capture devices 108 can include any type and form of a sensor for providing sensor data 174, including a positioning sensor, a biometric sensor, a velocity sensor, an acceleration sensor, a vibration sensor, a motion sensor, a pressure sensor, a light sensor, a distance sensor, a current sensor, a focus sensor, a temperature sensor, a haptic or tactile sensor or any other type and form of sensor used for providing data on medical tools 112, or data capture devices (e.g., optical devices). For example, a data capture device 108 can include a location sensor, a distance sensor or a positioning sensor providing coordinate locations of a medical tool 112 or a data capture device 108. Data capture device 108 can include a sensor providing information or data on a location, position or spatial orientation of an object (e.g., medical tool 112 or a lens of data capture device 108) with respect to a reference point. The reference point can include any fixed, defined location used as the starting point for measuring distances and positions in a specific direction, serving as the origin from which all other points or locations can be determined.

[0044] Display 116 can show, illustrate or play data streams 162, including video data 178, in which medical tools 112 at or near surgical sites are shown. For example, display 116134855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) can display a rectangular image (e.g., a frame of a video data 178) of a surgical site along with at least a portion of medical instruments 112 being used to perform surgical tasks. Display 116 can provide compiled or composite images generated by the visualization tool 114 from a plurality of data capture devices 108 to provide visual feedback from one or more points of view.

[0045] The visualization tool 114 that can be configured or designed to receive any number of different data streams 162 from any number of data capture devices 108 and combine them into a single data stream displayed on a display 116. The visualization tool 114 can be configured to receive a plurality of data stream components and combine the plurality of data stream components into a single data stream 162. For instance, the visualization tool 114 can receive a visual sensor data from one or more medical tools 112, sensors or cameras with respect to a surgical site or an area in which a surgery is performed. The visualization tool 114 can incorporate, combine or utilize multiple types of data (e.g., positioning data of a medical tool 112 along sensor readings of pressure, temperature, vibration or any other data) to generate an output to present on a display 116. Visualization tool 114 can present locations of medical tools 112 along with locations of any reference points or surgical sites, including locations of anatomical parts of the patient (e.g., organs, glands or bones).

[0046] Medical instruments or tools 112 can be any type and form of tool or instrument used for surgery, medical procedures or a tool in an operating room or environment. Medical tool 112 can be imaged by, associated with or include an image capture device. For instance, a medical tool 112 can be a tool for making incisions, a tool for suturing a wound, an endoscope for visualizing organs or tissues, an imaging device, a needle and a thread for stitching a wound, a surgical scalpel, forceps, scissors, retractors, graspers, or any other tool or instrument to be used during a surgery. Medical tools 112 can include hemostats, trocars, surgical drills, suction devices or any instruments for use during a surgery. The medical tool 112 can include other or additional types of therapeutic or diagnostic medical imaging implements. The medical tool 112 can be configured to be installed in, coupled with, or manipulated by an RMS 110, such as by manipulator arms or other components for holding, using and manipulating the medical instruments or tools 112.

[0047] RMS 110 can be a computer-assisted system configured to perform a surgical or medical procedure or activity on a patient via or using or with the assistance of one or more robotic components or medical tools 112. RMS 110 can include any number of manipulator144855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) arms for grasping, holding or manipulating various medical tools 112 and performing computer-assisted medical tasks using medical tools 112 controlled by the manipulator arms.

[0048] Video data 178, including any images or videos captured by a medical tool 112 (e.g., endoscopic camera) can be sent to the visualization tool 114. The robotic medical system 110 can include one or more input ports to receive direct or indirect connection of one or more auxiliary devices. For example, the visualization tool 114 can be connected to the RMS 110 to receive the images from the medical instrument 112 when the medical instrument 112 is installed in the RMS 110 (e.g., on a manipulator arm of the RMS 110 that is used for moving, managing or otherwise handing medical instruments 112). The visualization tool 114 can combine the data streams 162 from the data capture devices 108 and the medical tool 112 into a single combined data stream 162 for use by the ML framework 180 (e.g., ML models 182 or associated attention mechanisms 188 and encoder and decoder functions 190).

[0049] Data processing system 118 can be deployed in, communicatively coupled with, or otherwise associated with any of the components (e.g., devices) of the medical environment 102 directly or via a network 101. Data processing system 118 can be provided by a remote server (e.g., connected to medical environment 102 via a network 101), or can be deployed or provide via a cloud-based service or a function. The data processing system 118 can include an interface 192 designed, constructed and operational to communicate with one or more component of system 100 via network 101, including, for example, the robotic medical system 110. Data processing system 118 can be implemented using instructions stored in memory locations and processed by one or more processors, controllers or integrated circuitry. Data processing system 118 can include functionalities, computer codes or programs for executing or implementing any functionality of ML framework 180, including any ML models 182. The data processing system 118, as well as any of its components can each be a part of or include a cloud computing environment functionality or features. The data processing system 118 can include multiple, logically grouped servers and facilitate distributed computing techniques. The logical group of servers may be referred to as a data center, server farm or a machine farm. The servers can also be geographically dispersed. A data center or machine farm may be administered as a single entity, or the machine farm can include a plurality of machine farms. The servers within each machine farm can be heterogeneous - one or more of the servers or machines can operate according to one or more type of operating system platform.154855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)

[0050] The machine learning (ML) trainer 184 can include any combination of hardware and software for training ML models 182 to perform their designated or trained operations. For instance, the ML trainer 184 can include, train, configure, generate or adjust (e.g., retrain) any ML models 182. The ML trainer 184 can use the training datasets that can include any selection or collection of data streams 162 corresponding to various medical procedures using the RMS 110. ML trainer 184 can include a framework or functionality for training different machine learning models 182, such as a neural network spatial attention mechanisms models for determining spatial orientations of various components or features in the image frames 121. The ML trainer can train neural network spatial-temporal attention mechanism models 182, for determining features from image frames 121 having timestamps indicative of preceding or following image frames 121 (e.g., providing different view angles of the area). The ML trainer 184 can train the ML model designed for detecting medical instruments 112 as well as detecting anatomical parts of a patient (e.g., anatomy segments 128) using various data from data streams 162, including video data 178 or sensor data 174 (e.g., temperature, pressure, proximity or force data).

[0051] The ML trainer 184 can include, utilize, implement or provide an attention mechanism 188 that can be used to address the noise challenges in the data. Attention mechanism 188 can include a neural network with spatial attention or spatial -temporal attention that can be performed or implemented using an encoder and decoder function 190 to learn to identify various anatomical features or segments. Attention mechanism 188 can include a spatial attention or spatial-temporal attention that can be performed to identify tasks or phases of a medical procedure to detect trigger conditions 146 (e.g., milestone events) based on the type of medical instruments 112 and types of tissues (e.g., anatomy segments 128) involved in a particular order of operation. Attention mechanism 188 can utilize weights to emphasize different types of information in the data stream 162, such as movements in a region of an image frame that corresponds to a prior image frame in which a particular type of movement was detected. Such spatial and temporal weights used in the attention mechanism 188 can facilitate an improved or selective focus of the ML functions onto particular features in the data stream 162, assigning varying degrees of importance to each part of the input data during the learning process. For example, the attention mechanism 188 can include or utilize a neural network architecture that configures an ML model to focus selectively on relevant spatial features in the video data, thereby improving the accuracy of the detection. By assigning weights to selected segments of the input videos data 178, the164855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) attention mechanism 188 can allow the model to attenuate the impact less relevant portions of the data, emphasizing the importance of the more relevant cues (e.g., focus on a detected medical instrument 112 or a detected anatomical tissue of a patient) for more accurate determinations.

[0052] Interface 192 can include any combination of hardware and software for interfacing with a user of the robotic medical system 110. Interface 192 can include components or features designed, constructed and operational to communicate with one or more component of system 100 via network 101, including, for example, the RMS 110 or another device, such as a client’s personal computer. The interface 192 can include a network interface. The interface 192 can include or provide a user interface, such as a graphical user interface. The graphical user interface can include, for example, a window for displaying video data 178, or any indications or outputs to be provided or displayed with the end user. Interface 192 can provide data for presentation via a display, such as a display 116, and can depict, illustrate, render, present, or otherwise provide indications indicating determinations (e.g., outputs) of the ML models 182.

[0053] Interface 192 can be configured to utilize graphical user elements for utilization or operation of a data processing system 118 functions or operations. Graphical user elements can include any graphical features, attributes, user selections or commands that can be utilized to receive user commands or controls for utilizing a multi view-angle frame 142 or a 3D presentation, such as a video 152. Graphical user elements can include user input device (e.g., mouse selections) that can be used to indicate movement from the original monocular image frame 121 view-angle to another view-angle frame 142 to be generated, responsive to the user selection of the graphical user element. For instance, interface 192 can receive, via a graphical user interface element, an input with a range of view angles (e.g., 142) for which to generate the video 152. The interface 192 can provide the output (e.g., of the desired viewangle for which to generate a view-angle frame 142) and trigger the data processing system 118 to generate the video 152 (e.g., via 3D video generator 150) corresponding to the plurality of view angles and responsive to the input with the range of view angles. The interface 192 can display the video 152 or one or more generated view-angle frames 142 (e.g., generated responsive to the user selection) via a display 116. Interface 192 can operate on a head mounted device 106 and can be used to display the view-angle frames 142 or the 3D videos 152 via a display 116 on the head mounted device 106.174855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)

[0054] System 100 can include a head mounted device 106, also referred to as an HMD 106, that can include any combination of hardware and software for using any features of the data processing system 118 or the RMS 110. HMD 106 can include any combination of a display 116 and sensors 104, allowing a surgeon to perform or view surgery via a robotic medical system 110. The HMD 106 can include a display 116 for providing an immersive user interface 192, displaying multi-view angle frames 142 and 3D video reconstructions (e.g., videos 152) of the surgical scene. The sensors 104 can include hand tracking sensors or devices or eye tracking sensors or devices to detect user motions or selections. The display 116 can present high-resolution, real-time visuals. The integrated sensors 104 can track the surgeon's head movements and adjust the perspective accordingly. The HMD 106 can incorporate gesture and voice recognition functionalities, allowing the surgeon to control the user interface 192 of the HMD 106 hands-free.

[0055] The data processing system 118 can interface with, communicate with, or otherwise receive or provide information with one or more component of system 100 via network 101. The data processing system 118, RMS 110 and devices in the medical environment 102 can each include at least one logic device such as a computing device having a processor to communicate via the network 101. The DPS 118, any portion of the ML framework 180, the RMS 110 or a client device that can be communicatively coupled with the DPS or the RMS 110 via the network 101, can each include at least one computation resource, server, processor or memory for processing data. For example, the data processing system 118 can include a plurality of computation resources or processors coupled with memory.

[0056] The network 101 can be any type or form of a medium for facilitating communication between devices or systems, such as the data processing system 118 and the devices in a medical environment 102. The geographical scope of the network can vary widely and can include a body area network (BAN), a personal area network (PAN), a localarea network (LAN) (e.g., Intranet), a metropolitan area network (MAN), a wide area network (WAN), or the Internet. The topology of the network 101 can assume any form such as point-to-point, bus, star, ring, mesh, tree, etc. The network 101 can utilize different techniques and layers or stacks of protocols, including, for example, the Ethernet protocol, the internet protocol suite (TCP / IP), the ATM (Asynchronous Transfer Mode) technique, the SONET (Synchronous Optical Networking) protocol, the SDH (Synchronous Digital Hierarchy) protocol, etc. The TCP / IP internet protocol suite can include application layer,184855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) transport layer, internet layer (including, e.g., IPv6), or the link layer. The network 101 can be a type of a broadcast network, a telecommunications network, a data communication network, a computer network, a Bluetooth network, or other types of wired and wireless networks.

[0057] Frame function 120 can include any combination of hardware and software for receiving and processing image frames 121. Frames 121 can include any monocular images, such as 2D images from a single view-angle. A frame 121 can include an image of a surgical scene capturing or depicting one or more organs or tissues along with one or more medical instruments 112. The depicted surgical scene can be a 3D volumetric space having one or more objects (e.g., tissues, organs, glands or medical instruments 112) that are obstructing one another, providing one or more holes (e.g., obstructed areas that cannot be viewed from the given monocular image frame 121).

[0058] Frame function 120 can be configured to identify a frame 121 capturing a single view angle of a surgical scene in a medical procedure (e.g., surgical operation) that can be performed with a robotic medical system 110. The frame function 120 can receive a data stream 162 of the medical procedure captured via one or more sensors 104 of the robotic medical system 110. The frame function 120 can capture the data stream 162 from a repository 160. The frame function 120 can determine to generate one or more multi -view frames, such as view-angle frames 142 for a multi view-angle presentation, such as a 3D video. The frame function can determine to generate, or can trigger, command or start the generation of a multi view-angle presentation in which the scene captured by the frame(s) 108 can be displayed (e.g., via user interface 192) in a still format and from multiple viewangles that can be controlled or selected based on a selection (e.g., mouse movement or a click) of a user input device controlling the view-angles of the multi view-angle presentation.

[0059] Prompt generator 154 can include any combination of hardware and software for generating one or more prompts for one or more ML models 182. Prompts 156 can include any combination of parameters, values or instructions for directing or configuring one or more ML models 182 to implement particular (e.g., intended) functionality. Prompt generator 154 can generate any prompts 156 for any ML models 182. For instance, prompt generator 154 can generate prompts 156 to direct ML models 182 to generate view-angle frames 142 or synthesize new pixels to add into view-angle frames 142 (e.g., on behalf of an inpainting function 134). For instance, prompt generator 154 can generate prompts 156 to direct ML models 182 to generate view angle frames 142 for a plurality of view angles (e.g.,194855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) perspectives) surrounding the view-angle of the monocular image frame 121, allowing for generation of a 3D video 152, allowing the user to view a surgical scene from multiple viewangle perspectives.

[0060] The prompt generator 154 can generate a prompt 156 based on the segments (e.g., anatomy segments 128) or based on labels 126 of a particular mask 124. The prompt generator 154 can input the prompt 156 into an ML model 182 that is a generative machine learning model configured to synthesize new pixels, to trigger the ML model 182 to fill the one or more holes (e.g., one or more areas with the new pixels synthesized by a generative ML model 182). The prompt generator 154 can generate the prompts 156 based on the masks 124 that can indicate the distance between pixels to be inpainted in the frame (e.g., viewangle frame 142) and the reference point. For example, when utilizing an inpainting function 134 to fill in the empty holes or areas of the view-angle frames 142 in which no pixel data exists (e.g., where there is no visibility in a new view-angle frame 142), the prompts generator 154 can generate the prompt 156 to cause the ML model 182 to fill in the pixels in the hole (e.g., empty area) of the view-angle frame 142 based (e.g., at least in part) on the distance between each of the pixels to be filled or inpainted and a reference point (e.g., one or more pixels in the original frame 121 or a view-angle frame 142).

[0061] Anatomy segments 128 can include any segments of an anatomy of a patient on whom a medical procedure is performed. Anatomy segments 128 can include, for example, segments of one or more body parts, organs, tissues, bones, blood vessels, nerves, muscles, joints, ligaments, tendons, cartilage, glands, and connective tissues. Anatomy segments 128 can include areas such as a cardiovascular system, respiratory system, digestive system, nervous system, musculoskeletal system, and endocrine system. The identification and visualization of these anatomy segments 128 can be done by the in precise surgical planning, navigation, and execution during medical procedures.

[0062] Segmentation masks 124, also referred to as segment masks 124 or masks 124, can include any region, area or a map corresponding to, or identifying, any distinct anatomical area, segment or a region within an image frame. Mask 124 can be defined as binary or multi-class maps that delineates different anatomy segments 128 from each other, such as distinguishing a blood vessel from a fat tissue. Mask 124 can be used for labeling and distinguishing different regions within an image. For example, in a surgical scene, a segmentation mask 124 can identify areas corresponding to the liver, kidneys, blood vessels,204855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) and other critical organs or tissues. The areas indicated by the segmentation masks 124 can be labeled using labels 126, allowing ML models 182 to distinguish one anatomy segment 128 from another. For instance, each pixel, or a group of pixels, in the segmentation mask 124 can be assigned a class label 126, indicating the anatomical structure (e.g., anatomy segment 128) to which it belongs. For instance, pixels representing the liver can be assigned a first color, indication, or a label, while pixels representing the kidneys can be assigned a second color, indication or a label 126.

[0063] Segmentation function 122 can include any combination of hardware and software for generating masks 124 for segmenting and labeling anatomy segments 128 using labels 126. Segmentation function 122 can generate one or more segmentation masks 124 for identifying, distinguishing or labeling different anatomy segments 128. For instance, a segmentation function 122 can utilize ML models 182 to segment, identify or distinguish different anatomy segments 128 from each other and label those segments based on their type. Segmentation function 122 can generate, using one or more ML models 182 trained with machine learning, one or more masks 124 that segment an anatomical structure in the frame 121 or view-angle frame 142. The segmentation function 122 can use one or more ML models 182 to label the anatomical structure in the frame (e.g., 108 or 142) using one or more labels 126. For instance, a segmentation function 122 can generate a mask 124 using a vision transformer based neural network (e.g., ML model 182) trained with a dataset comprising images of medical procedures that are annotated with labels 126. The mask 124 can include the anatomy segment 128 of the anatomical structure (e.g., a particular instance or a type of anatomy segment 128) and the label 126 of the anatomical structure (e.g., for the same anatomy segment 128).

[0064] Segmentation function 122 can include or utilize an ML model 182 having or using a vision transformer based neural network that can be configured and trained with human annotated clinical image dataset. For instance, an ML model 182 can be trained using labels 126 labeling various anatomy segments 128. Given an input image frame 121, an ML model 182 can generate a segmentation mask 124 along with any labels for each of one or more anatomy segments 128 within the image frame 121. The ML model 182 can label the segmentation mask 124 to generate a segmentation map. The segmentation map can include a spatial representation (e.g., a 2D representation) of locations of various anatomy segments 128 along with any labels 126 of such anatomy segments 128.214855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)

[0065] Label annotator 186 can include any function (e.g., application or process) for assigning labels to any segmentation masks 124 or depth masks 138. Label annotator 186 can include, for example, an ML model 182 trained to annotate or assign labels 126 to any segmentation masks 124 or depth masks 138. Label annotator 186 can include, for example, a generative Al model trained to generated information, metadata or otherwise labels 126 describing the type of anatomical structure and their depth or location with respect to the lens of the data capture device 108.

[0066] Labels 126 can include any metadata providing or indicating information about a segmentation mask 124 corresponding to an anatomy segment 128 (e.g., name or type of tissue or an organ). Label 126 can include any metadata providing or indicating information corresponding to a depth mask 138, including any depth values 132 (e.g., distance of a given one or more pixels of a depth mask value 132 from a reference point, such as a location of a lens of the camera capturing the frame 121). Label annotator 186 can generate the labels 126 for any of the depth masks 138 or anatomy masks 124 using any functionality of a ML framework 180, including any ML models 182, attention mechanisms 188 or encoder and decoder functions 190. Labels can include “state”, such as a state of a liver being burned. The label 126 for a pixel can be used to generate a “burned” liver representation or a depiction (e.g., by an inpainting function 134). States can include various tissue or environment states, such as bleeding, burned or cut in a bleeding area.

[0067] Depth estimator 125 can include any combination of hardware and software for determining the actual depth or distance of objects or features within an image. The depth estimator 125 can include any functionality for generating depth masks 138 and their corresponding depth values 132. The depth estimator 125 can generate labels 126 for the depth masks 138. For instance, the depth estimator 125 can utilize ML models 182 (e.g., label annotator 186) to assign labels 126 to various depth values 132 of the depth masks 138. For instance, the depth estimator 125 can generate a depth mask 138 having a depth value 132 for each of the pixels of the 2D image frame 121.

[0068] Depth estimator 125 can use a 2D image frame 121 as its input into a neural network ML model 182, which can be configured as an encoder-decoder with a transformer unit, to map the image frame 121 to its metric depth values 132. The depth estimator 125 can generate a depth map for one or more depth masks 138 of each frame 121. Each pixel intensity in the resulting metric depth map can represent a real metric distance from the224855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) camera center to a point in the scene (e.g., the image frame 121). The ML model 182 can use camera intrinsic parameters for calibration during training, ensuring accurate depth estimation. The output depth map can be color-coded, with different colors or shades indicating various distances. This metric depth information can be used by the mesh generator 130 to generate a 3D mesh of the scene, allowing the creation of 3D representations (e.g., videos 152 or one or more depicted view-angle frames 142) along with new view angles and the filling (e.g., via inpanting) of any holes using generative Al.

[0069] For example, a depth estimator 125 can provide a metric depth estimation that can include a distance for each of the one or more pixels of the 2D frame 121 from the lens of the camera to that one or more pixels. The depth estimator 125 can be configured as a neural network ML model 182 that maps a frame 121 to its metric depth. Each pixel intensity of the depth mask 138 can include a depth value 132 that represents a real metric distance between a reference point of the camera (e.g., the camera lens center) to the point of the scene. The neural network ML model 182 can include an encoder and decoder function 190 with a transformer unit. The output of the ML model 182 can include a metric depth map with depth values 132 for each of the pixels of the image frame 121. The camera intrinsic can be used during training for calibration purpose. The output from the ML model 182 can represent the distance using color coding, such that different distances are colored or shaded. Each coloring or shade intensity can represent a real distance from the shaded region (e.g., one or more pixels) to the lens of the data capture device 108 that has captured the 2D image frame 121. The metric depth information can be used to map or feed into a neural network ML model for generating the 3D mesh 136.

[0070] For example, a depth estimator 125 can generate a depth mask 138 using a neural network (e.g., ML model 182) trained to map an intensity of each pixel in the frame to the depth value 132. The depth value 132 can be a distance between a camera that captures the frame 121 and a location in the scene corresponding to the pixel of the frame 121. The depth estimator 125 can generate the depth mask 138 based on at least one of an optical center of the camera, a focal length of the camera, a scale factor of the camera, a principal point of the camera, a skew of the camera, or a geometric distortion of the camera.

[0071] Mesh generator 130 can include any combination of hardware and software for generating one or more 3D meshes 136. Mesh generator 130 can construct a 3D mesh 136 of a scene captured by an image frame 121 based on one or more segmentation masks 124234855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) indicating or labeling anatomy segments 128 and depth masks 138 indicating or labeling depth values 132 of the various one or more pixels (e.g., individual pixels or groups of pixels) that can be color coded or shaded based on their metric distance from the camera. For example, the mesh generator 130 can create a 3D model of the patient's internal anatomy structures, allowing surgeons to visualize the exact locations and relationships of various anatomy segments 128 (e.g., organs or tissues). This 3D model representation in the 3D mesh 136 can be implemented respect to a frame of reference of a data capture device 108 (e.g., the camera that captured the image frame 121). The 3D mesh 136 can reflect relative locations of various anatomy segments 128 from each other. This can greatly aid in preoperative planning, intraoperative navigation, and postoperative analysis. As 3D mesh 136 provides relative depth and spatial orientation of the internal structures in the frame 121, the 3D mesh 136 can be used by the multi -view generator 140 to generate view-angle frames 142 from various view-angles other than the original view-angle at which the 2D image frame 121 was captured.

[0072] 3D mesh 136 can include any 3D representation of a scene captured in an image frame 121. 3D mesh can be a graphic representation of a 3D area, indicating any anatomy segments 128 or depth values 132 with respect to such segments. 3D mesh can be or include a file having one or more vertices, edges and faces that provide a three-dimensional representation of the scene from the 2D image frame 121. The 3D mesh 136 can represent various anatomy segments 128 as well as their distances (e.g., depth values 132) from the lens location. The 3D mesh can be used to generate multiple view-angle frames 142 aside from the original view-angle of the 2D image frame 121. For instance, a user of the data processing system 118 (e.g., a surgeon using a user interface 192) can rotate and view the 3D model of the imaged scene from different angles to gain a better understanding of complex anatomical relationships and to identify potential obstructions or satisfy trigger conditions 146.

[0073] For example, a mesh generator 130 can a mesh data to form the 3D mesh 136. The mesh data can include, provide or operate as a graphic representation of the 3D scene captured in the image frame 121. The mesh generator 130 can generate the 3D mesh 136 (e.g., a 3D mesh file) based on the depth map (e.g., depth mask 138 with the depth values 132), which can include a collection of vertices, edges, and faces that represent the shape and structure of a 3D scene. The 3D mesh 136 can allow for the representation of complex 3D244855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) shapes and surfaces by specifying the coordinates of vertices and the connectivity between them, which can be utilized by the multi -view generator 140 to generate view-angle frames 142 of the same scene from various view-angles. The 3D mesh 136 can be represented as a 2D color image according to its distance location.

[0074] Multi -view generator 140 can include any combination of hardware and software for generating view-angle frames 142. A view-angle frame 142 can include any generated image frame representing the scene captured by a frame 121 from a different view-angle (e.g., angular perspective) than the view-angle of the original frame 121. Multi-view generator 140 can utilize 3D mesh 136 and the image frame 121 information or data to generate any number of view-angle frames 142. The multi-view generator 140 can generate a plurality of view-angle frames 142 corresponding to a plurality of view-angles different from the original view-angle of the frame 121. Each of the generated view-angle frames 142 can correspond to its own view angle surrounding the view angle of the image frame 121.

[0075] For example, the multi -view generator 140 can generate, using the one or more ML models 182, a plurality of view-angle image frames 142 for the plurality of view angles of the frame 121. For instance, the multi-view generator 140 can utilize the 3D mesh 136 (e.g., the 3D model) to generate a plurality of view-angle frames 142 at various view-angles (e.g., other than the original view angle of the 2D image frame 121). The multi view generator 140 can identify one or more holes or empty spaces that can occur in the viewangle frames 142 due to their angular shift between the original view-angle of the image frame 121 and the view-angle frame 142. The multi-view generator 140 can update the plurality of images using an inpainting function 134 (e.g., an inpainting ML model 182) to fill in the one or more holes in the plurality of view-angle frames 142. For instance, the multiview generator 140 can utilize the inpainting function 134 to generate pixel coloring, texture or shape to conform to the pixel coloring, texture or shape in the anatomy segments 128 to which the hole (e.g., empty space) belongs or with which the hole overlaps. The multi -view image generation can be used to simulate the trajectory of a simulated moving camera. Along the trajectory, the image view can change and inpainting steps can be repeated for each sampled perspective (e.g., view-angle).

[0076] Inpainting function 134 can include any combination of hardware and software for reconstructing or filling in missing parts of an image frame 121. Inpainting function 134 can be used to restore any missing image areas that can appear or be generated while generating254855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) view-angle frames 142 from angles that reveal a region that was obstructed by another object or a feature in the original 2D image frame 121. For example, the inpainting function 134 can fill in gaps caused by obstructions with respect to any portions of anatomy segments 128 in a surgical scene captured by the frame 121, using generative Al (e.g., inpainting ML models 182) to predict and reconstruct the missing anatomical structures based on surrounding context.

[0077] For example, the inpainting function 134 can determine, based on an ML model 182 trained on determining a full shape of an anatomy segment 128 that a portion of a hole (e.g., image area having missing pixel values in the view-angle frame 142) belongs to or is a portion of a particular anatomy structure (e.g., a particular tissue or an organ). In response to this determination, the inpainting function 134 can fill int the hole with the pixel values colored and arranged to match or correspond to the portion of the same anatomy segment 128 from the frame 121. For instance, the pixels of the hole can be filled with the same colored pixels as the pixels of the anatomy segments 128. For example, the hole can be filled in with an average color of the anatomy segment 128. For example, the hole can be filled in with a pattern of features of a plurality of pixels matching or corresponding to the pattern of features of the surrounding portions of the anatomy segments 128. The inpainting function 134 can achieve this through techniques such as deep learning, where neural network ML models 182 can be trained on large datasets to learn the patterns and features of different anatomical structures. The function can leverage convolutional neural networks (CNNs) for spatial inpainting or generative adversarial networks (GANs) for more complex and realistic reconstructions.

[0078] For example, the multi -view generator 140 can generate, using the one or more models 182, a plurality of view-angle image frames 142 for the plurality of view-angles of the frame 121. The inpainting function 134 can identify one or more holes (e.g., portions of the view-angle frames 142 whose pixels are not filled in with data) in the plurality of viewangle image frames 142. The inpainting function 134 can update the plurality of view-angle frames 142 using an inpainting ML model 182 to fill the one or more holes in the plurality of view-angle frames 142. For instance, the inpainting function 134 can synthesize, using the inpainting ML model 182, one or more new pixels based on the anatomy segments 128 (e.g., from the original frame 121) and the labels 126 in, or for, the one or more masks 124 and 138. The inpainting function 134 can fill the one or more holes with the new pixels and264855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) assign, to the new pixels, the label and the depth value based on the one or more masks (e.g., masks 124 and 138).

[0079] The inpainting function 134 can compute or calculate the pixel transformation (e.g., the pixel values of the holes) based on the depth map (e.g., depth mask 138 and depth values 132) and using the computational geometry. The inpainting function 134 can utilize an inpainting ML model 182, such as a generative Al model trained model with the capability to synthesize new pixels fitting in the holes (e.g., undetermined pixels) in the image. The ML model 182 can take a semantic segmentation map (e.g., segmentation mask 124) and labels 126 of the mask 124 to understand the categories of the infilling area. The semantic label can be used as, or included in, prompts 156 for the generative Al models 182 to better generate pixels as required. The inpainting function 134 can operate based on the geometrical constraints from the depth map-defined 3D structure for the synthesized pixels.

[0080] The pixels filled with an inpainting function 134 (e.g., during the generation of the view-angle frames 142 by the multi -view generator 140) can be generated to be smooth or continuous to their surroundings. The filled pixels can keep their semantic labels 126 unchanged. After inpainting the holes, the color and depth values can be assigned to the 3D mesh 136. The new viewing angle image can be generated based on the updated or adjusted 3D mesh 136 that can include the pixels filled in by the inpainting function 134. When inpainting various tissues, the inpainting function 134 can fill in the pixels based on the surrounding tissue or organ tissue. For instance, a gallbladder can be inpainted to a gallbladder, a fat tissue can be inpainted according to a view of a surrounding fat tissue. The distance of inpainted areas can be smooth to interpolate a 3D shape. Color and texture can be based on a semantic map. The technical solutions can use a standard interpolation function to smooth the shape. This can be done for other angles and such images can be generated and connected to play in 3D. The inpainting function 134 can be used to provide visibility or clarity for various visibility obstructions. For instance, a liquid or dirt on a camera can be addressed by the technical solutions that can utilize these techniques to clarify the image and remove the dirt.

[0081] The inpainting function 134 can be utilized to clean images in real-time by removing blurring or improving contrast and clarity of the image using masks (e.g., 124 or 138) and inpainting functionalities. For example, if the camera lens gets dirty during a procedure, the inpainting function 134 can clean the images using masks (e.g., 124 or 138)274855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) and generative Al models 182 in real-time to display a clean image or a video to the surgeon during the procedure, removing the artifacts from the dirty camera. Thus, an aspects of the technical solutions disclosed herein can be directed to an improved real-time display of a surgical scene in which the image or video presented is clearer or of higher quality due to removing, reducing, or otherwise eliminating blurry portions or artifacts from the image.

[0082] The inpainting function 134 can use anatomy states (e.g., burned, bloodied, damaged, cut, removed, connected or other) of various anatomy parts along with anatomy and medical instrument segmentation and metric depth map (e.g., depth mask 138) to compute pixel values for the image reconstruction and inpainting. Anatomy state detection can identify changes or states of the anatomy, such as the anatomy segment 128 being cut, bruised, burnt, bleeding or otherwise in other state. Anatomy and instrument segmentation can isolate relevant areas, allowing for focused inpainting. Metric depth maps provide depth information (e.g., depth values 132 across the image) for accurate spatial relationship determinations. By integrating the historic state of anatomy and detected changes, the inpainting can maintain consistency with previous states, allowing for accurate and reliable medical images for surgical planning and execution. For instance, the inpainting function 134 can detect bleeding in anatomy state detection and use historic state of anatomy to determine the changes to the state.

[0083] 3D video generator 150 can include any combination of hardware and software for generating a video 152. A video 152 can include a series of view-angle frames 142 or a video depicting a scene captured by an original frame 121 from one or more (e.g., a series of) viewangles other than the original view-angle of the image frame 121. The video 152 can include a stream of view-angle frames 142 along with the hole areas that are filled in, or inpainted, by the inpainting function 134. The 3D video generator 150 can generate the video 152 based on the 3D mesh 136 that can be adjusted or updated based on inpainted or filled in content (e.g., areas filled in by the inpainting function 134).

[0084] The 3D video generator 150 can generate any 3D representation of the scene captured in the image frame 121. The 3D video generator 150 can generate a 3D simulation of the scene, based on the 3D mesh 136, depth masks 138, segmentation masks 124 and any labels 126. The 3D video generator 150 can generate the video 152 corresponding to a plurality of view angles of the frame using the one or models and the 3D mesh 136 and the labels 126 of the anatomical structures from the one or more masks (e.g., 124 or 138). The284855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)3D video generator 150 can generate a video 150, based on a set or predetermined frame rate. The video 152 can include multi -view images combined along with the view-angle of the original frame 121 as a 3D video and output to the frontend user interface 192. For instance, the 3d video generator 150 can receive a data stream 162 of the medical procedure captured via one or more sensors 104 of the robotic medical system 110 and determine, based on the data stream 162, a range of view angles in a portion of the video 152 that comprises the frame 121.

[0085] The 3D video generator 150 can allow for rendering of the surgical scene in a user-specified perspective. For instance, a user can specify a perspective of view at which the scene can be rendered (e.g., from a particular angle) by selecting with an input device (e.g., a computer mouse) an angle relative to the view-angle of the frame 121 (e.g., to the left or to the right of the view angle of the frame 121). Responsive to such a selection (e.g., via the user interface 192), the 3D video generator 150 can generate (e.g., from the view-angle frames 142) new viewing angles of previously invisible areas and contents. Those new contents can be photo-realistic and provide a same level of clarity as other features based on the filled in data generated or provided by the inpainting function 134 to render the previously invisible areas in color, texture and 3D shape of the anatomy segments 128 to which they correspond.

[0086] The 3D video generator 150 can determine the range of view angles for the video 152. The range of view-angles can correspond to the range of view-angles to generate for the view-angle frames 142. For instance, a wider view-angle range can correspond to more viewangle frames 142 to be generated and more angular range for the video 152. The range of view angles can be determined based on various trigger conditions 146. For instance, the range of view-angles for the 3D video 152 can be based on a milestone event (e.g., a type of a milestone event) or a type of a medical task or a procedure. For instance, the range of viewangles for the 3D video 152 can be based on surgeon data 172 (e.g., a surgeon profile), such as a level of surgeon’s skill, surgeon’s historical performance or surgeon’s objective performance indicators (OPIs) indicating surgeon’s performance with respect to the particular task, phase or a type of a medical procedure. For instance, the range of view-angles for the 3D video 152 can be based on a data stream indicating a lack or absence of a sufficient view angles based on the type of tasks or procedure in the frame. For instance, the range of viewangles for the 3D video 152 can be based on a quality confidence level of the pixels or images created using the ML models 182.294855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)

[0087] The 3D video generator 150 can render a mask on 3D video 152 based on the confidence or quality of images at different view angles. For instance, the 3D video generator 150 can generate or render a mask on video 152 by providing different color or shade for pixels synthesized using the inpainting function 134, allowing for distinguishing pixels that are generated using ML models 182 from pixels from the frames 121. The 3D video generator 150 can generate a notification for the user interface 192 to notify the user on how many pixels are generated by the ML, versus how many are based on the frame 121.

[0088] Trigger condition detector 144 can include any combination of hardware and software to detect trigger conditions 146 and trigger 3D view reconstruction based on one or more thresholds 148. The trigger condition 146 can include any milestones or features to be verified or affirmed as a part of a medical procedure process (e.g., a phase or a task). The trigger conditions 146 can occur based on any milestone events or based a level of surgeon’s skill, such as the surgeon’s historical performance or surgeon’s objective performance indicator (e.g., OPIs) or the surgeon’s demand. The trigger conditions 146 can occur based on the data stream 162 indicating a lack or absence of sufficient view angles based on the type of task or procedure in the image frame 121.

[0089] For example, a trigger condition 146 can include the verification that a particular first tissue (e.g., a first anatomy segment 128) is completely severed and separated form a second tissue (e.g., a second anatomy segment 128). The trigger condition detector 144 can utilize ML models 182 trained to determine or detect surgical procedures, their phases and individual tasks, to identify a particular trigger condition 146 present at a particular portion of the procedure (e.g., at an end of a surgical task). In response to the trigger condition detector 144 detecting a trigger condition 146, the trigger condition detector 144 can issue a request or an instruction (e.g., to the data processing system 118) to trigger or start the 3D reconstruction of the surgical scenes. For instance, the trigger condition detector 144 can trigger any one or more of a mesh generator 130, the multi -view generator 140 and the 3d video generator 150 to implement their respective operations to provide a 3D representation of the surgical scene, to allow the user (e.g., the surgeon) to view the scene from a plurality of view-angles, facilitating a more convenient and more efficient and effective completion of the milestone event (e.g., the trigger condition 146).

[0090] The trigger condition detector 144 can utilize thresholds 148 to detect the presence of trigger conditions 146. A threshold 148 can include any value, parameter or an indication304855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) indicating the presence of the trigger condition 146. For instance, the threshold 148 can include a level of performance of a surgeon from a surgeon’s profile in the surgeon data 172. For instance, the threshold 148 can indicate a number of successfully performed surgeries for a different type of a phase or a task. In response to determining (e.g., from the surgeon’s data 172, such as the surgeon profile) that the surgeon has performed fewer than the threshold 148 number of successfully completed phases or tasks, the trigger condition detector 144 can identify the trigger condition 146. For instance, the threshold 148 can include a threshold 148 confidence score or a confidence level from the ML model 182 to use by any one or more of the: segment function 122, depth estimator 125, mesh generator 130, multi-view generator 140 or 3D video generator 150. The threshold 148 can correspond to a minimal sufficient view angle of an area to view or verify given the task or a phase of a medical procedure. In response to the ML model 182 confidence level or score not satisfying the threshold 148, the trigger condition detector 144 can instruct the data processing system 118 to begin the 3D view reconstruction process. For instance, in response to the trigger condition 146, the trigger condition detector 144 can instruct the data processing system 118 and its components to perform the 3D view reconstruction to provide the view-angle frames 142 and the video 152 to the end user (e.g., the surgeon) to facilitate the view of the scene from multiple view angles.

[0091] For instance, a data processing system 118 can select a frame 121 from the video stream 178 responsive to detection of the trigger condition 146. The trigger condition 146 can be detected, using for example the trigger condition detector 144 utilizing one or more ML models 182 trained to detect tasks of particular phases in the medical procedure at which the trigger condition 146 is set to occur (e.g., per the medical procedure performed). The data processing system 118 can determine to generate the video 152 for the frame 121 based on the detection of the trigger condition 146.

[0092] For example, the data processing system 118 can receive a data stream 162 of the medical procedure captured via one or more sensors 104 of the robotic medical system 110. The data processing system 118 can determine, based on the data stream 162, a range of view angles in a portion of the video 152 that comprises the frame 121. A trigger condition detector 144 can detect the trigger condition 146 based on the occurrence of, or a temporal alignment with, a milestone event in the frame 121. A trigger condition detector 144 can detect the trigger condition 146 in response to the data stream 162 indicating that the view angle for the type of tasks or procedure in the frame 121 is below the threshold 148314855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) corresponding to the minimum sufficient view angle based on the type of task or procedure. The trigger condition detector 144 can detect the trigger condition 146 responsive to an input of a user (e.g., a surgeon) into a user interface of the data processing system, such as a graphical user interface allowing the surgeon to select, click on, or otherwise identify the trigger condition 146. The trigger condition detector 144 can detect the trigger condition 146 the range of view angles in the portion of the video data 178 being less than or equal to a threshold. The threshold can include, for example, a view angle value sufficient to view the scene and verify a milestone event occurrence, based on the 2D image frame 121.

[0093] The trigger condition detector 144 can detect the trigger condition 146 based on a surgeon data 172. The surgeon data 172 can include any information about a surgeon, such as a surgeon profile providing information about the level of experience of the surgeon. The surgeon profile (e.g., the surgeon data 172) can indicate that the surgeon does not have a sufficient level of experience with the particular surgical task being performed. For instance, the trigger condition detector can utilize a threshold value for a number of prior performed surgical procedures or tasks. In response to determining that the threshold number of previously performed surgical procedures or tasks, the trigger condition detector 144 can detect the trigger condition 146 (e.g., using the surgeon profile) and trigger the data processing system 118 to generate the 3D mesh 136, the view-angle frames 142 and 3D videos 152.

[0094] FIG. 2 depicts an example 200 of a system diagram for performing 3D view reconstruction of a surgical scene using a system 100 of FIG. 1. The example 200 system diagram can include a frame function 120 that can receive a 2D monocular frame 121, depicting a surgical scene. The surgical scene can include one or more anatomical structures and medical instruments 112. The frame 121 can be input into segmentation function 122 and a depth estimator 125. The segmentation function 122 can utilize ML models 182 to provide anatomy segmentation and generate segmentation masks 124 with labels 126 marking various anatomical structures according to their respective anatomy segments 128. The depth estimator 125 can utilize the frame 121 to implement metric depth estimation and generate the depth mask 138 along with any depth values 132.

[0095] The 3D mesh generator 130 can utilize the outputs from the segmentation function 122 and the depth estimator 125 (e.g., the segmentation mask 124 and the depth mask 138) to generate a 3D mesh 136 of the scene captured by the frame 121. The 3D mesh 136 can324855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) include a three-dimensional representation of the surgical scene. The 3D mesh 136 can be generated using ML models 182 trained to generate spatial relations between various features based on the metric distances (e.g., distance in meters, centimeters, millimeters, inches or any other unit of distance) between each of the portions of the image (e.g., one or more pixels) and the lens of the camera that captured the frame 121.

[0096] The multi -view generator 140 can utilize the 3D mesh 136 to generate one or more view-angle frames 142 showing or depicting the scene from the frame 121 from different view-angles than the view-angle of the frame 121. The multi-view generator 140 can utilize the inpainting function 134 to fill in the pixels of the holes in the view-angle frames 142 that were not visible or filled in in the frame 121. The 3D video generator 150 can generate the 3D video 152 using one or more of the frame 121 and the view-angle frames 142, filling in the holes in the view-angle frames 142 to provide a complete image view from various view-angles and allowing for a 3D view experience to the end user.

[0097] FIG. 3 illustrates an example view 300 of a surgical scene 302 captured by a 2D monocular frame 108. The surgical scene 302 provided via a frame 121. The surgical scene 302 can include an anatomical structure of a cystic duct with a hepatic triangle, which can form a part of a first anatomy segment 128. The surgical scene 302 can also include a gallbladder, which can be a part of a second anatomy segment 128. In the medical procedure performed, a task or a phase of the procedure can dictate that the cystic duct be separated from the gallbladder and that a hepatic triangle be visually verified as the area should be clearly seen and dissected from the surrounding anatomies. Such a medical procedure dictate or a milestone can be a trigger condition 146 that can be used to trigger the data processing system 118 to implement the functionalities of the system 100 to generate 3D view reconstruction of the surgical scene 302.

[0098] FIG. 4 illustrates an example view 400 of a surgical scene 302 in a view-angle frame 142a before the start of the inpainting process and a view-angle frame 142b following the completion of the inpainting process. The view-angle frame 142a can include holes 402, including areas of the surgical scene 302 that are not filled with pixel data. The holes 402 are shown as gray shapes in the view-angle frame 142a. As shown in the view-angle frame 142b, upon the inpainting, the holes 402 are turned into inpainted regions 404 in which in inpainting function 134 fills the pixels of the holes 402 with the colors and textures of the surrounding anatomical structure.334855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)

[0099] Turning now to FIG. 5, an example flow diagram of a method 500 for providing a 3D view reconstruction of a surgical scene using a 2D monocular image frame. The method 500 can be performed by a system having one or more processors (e.g., 610) executing computer-readable instructions stored on a memory (e.g., 615, 620 or 625). The method 500 can be performed, for example, by any combination of features discussed in connection with example systems or features in examples 100-400 of FIGs. 1-4 and using a computing system 600 of FIG. 6. The method can include acts 505-535. At 505, the method can identify one or more trigger conditions. At 510, the method can identify a 2D image frame of a surgical scene. At 515, the method can determine if the 2D frame satisfies a threshold for a trigger condition. At 520, the method can generate anatomical masks with labels. At 525, the method can generate depth masks with depth values. At 530, the method can construct a 3D mesh. At 535, the method can generate a 3D video of the surgical scene.

[0100] At 505, the method can identify one or more trigger conditions. The method can include one or more processors coupled with memory identifying, detecting or otherwise determining one or more trigger conditions that apply to, or correspond to, the medical procedure being performed by the robotic medical system. The trigger conditions can include any conditions, such as events or occurrences, that can trigger 3D view reconstruction of a surgical scene using a 2D image frame (e.g., a monocular 2D image). The trigger conditions can be identified, based on the surgeon data, such as a profile of a surgeon including surgeon’s objective performance indicators from prior performed surgeries, level of experience of the surgeon with respect to a type of a procedure or a task, or the demand level for the surgeon.

[0101] The trigger condition can be identified based on a task of a phase of a particular medical procedure, such as a milestone event that can be scheduled or expected to occur at a particular stage of a medical procedure (e.g., at the end of a particular surgical task). The milestone event can include a visual verification of a particular portion of a surgical scene (e.g., a verification of a state of a surgical tissue). The trigger condition can include a detection or a determination that a view-angle of a particular 2D image frame does not satisfy a threshold for a view-angle sufficient to verify a milestone event.

[0102] At 510, the method can identify a 2D image frame of a surgical scene. The method can include the one or more processors identifying, detecting, receiving or capturing a two-dimensional image frame capturing or depicting a three-dimensional surgical scene344855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) from a single view-angle perspective. The method can include the frame function identifying a frame that captures a single view angle of a scene in a medical procedure performed with a robotic medical system. For instance, the frame function can receive the 2D frame as a part of a sequence of 2D frames received as a part of a video stream captured from a 2D monocular camera system. The frame can be identified, for example, responsive to determining that a trigger condition may occur or is expected within a particular time range. For instance, one or more ML models can be trained to monitor the tasks and phases of the ongoing surgical procedure and can determine that a milestone event is expected during an ongoing task.

[0103] The frame function can select the frame from the video stream responsive to detection of the trigger condition. For example, a trigger condition can be determined or detected and the frame function can identify, capture or store a 2D frame responsive to determining that the 2D frame corresponds to a time period during which the milestone event is expected to occur. The 2D frame can include a single view-angle view of the surgical scene. The 2D frame can include one or more obstructions to one or more portions of the scene that is desired to be viewed or observed.

[0104] At 515, the method can determine if the 2D frame satisfies a threshold for a trigger condition. The method can include on the one or more processors determining if the 2D frame identified at 510 satisfies at least one threshold condition of the one or more threshold conditions identified at 505. For instance, the method can include a threshold condition detector that can compare the 2D frame, or the data or contents from the 2D frame with one or more thresholds to determine if the trigger condition for triggering the 3D view reconstruction has occurred. For instance, the method can include the trigger condition detector detecting a trigger condition in a video stream of the medical procedure. The video stream can include a sequence of image frames that can be received as a part of a video stream that includes a sequence or a series of 2D image frames. Based on a determination or a detection of a type of task performed with the robotic medical system in the video stream, the trigger condition detector can detect an occurrence of a trigger condition.

[0105] The method can determine to generate one or more 3D videos or a 3D representations of a surgical scene using the 2D image frame based on the detection of the trigger condition. For example, the data processing system can receive a data stream of the medical procedure captured via one or more sensors of the robotic medical system. The data processing system can determine, based on the data stream, a range of view angles in a354855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) portion of the video that comprises the frame. The trigger condition detector can detect the trigger condition based on the occurrence of the milestone event in the frame and the range of view angles in the portion of the video stream being less than or equal to a threshold. The trigger condition detector can detect the trigger condition based on a profile associated with a surgeon that performs the medical procedure with the robotic medical system.

[0106] If the method determines that the 2D frame does not satisfy a threshold for a trigger condition, the method can move to act 510 and identify another 2 image frame of the surgical scene to check for trigger conditions. If the method at act 515 determines that the 2D frame satisfies the threshold for the trigger condition, the method can move to acts 520 or 525 to proceed with the 3D view reconstruction.

[0107] At 520, the method can generate anatomical masks with labels. The one or more processors can generate one or more anatomical masks and label one or more anatomy segments identified with the anatomical masks using one or more ML models. For example, the method can utilize a segmentation function to generate, using one or more models trained with machine learning, one or more segmentation masks. The one or more segmentation masks can segment one or more anatomical structures in the 2D frame. The segmentation function can utilize one or more ML models to label the anatomical structure in the frame. For instance, the segmentation function can generate a segmentation mask in which each of the plurality of segment structures is identified and labeled (e.g., via metadata) and provided in a segmentation map in which anatomy segments are indicated (e.g., via color coding or labeling).

[0108] The method can include generating a first mask of the one or more masks using an ML model configured to include a vision transformer based neural network. The ML model can be trained with a dataset comprising images of medical procedures that are annotated with labels. The first mask (e.g., a segmentation mask) can include the segment of the anatomical structure and the label of the anatomical structure.

[0109] At 525, the method can generate depth masks with depth values. The one or more processors can generate one or more depth masks with depth values corresponding to distance between each of the features or pixels in the 2D frame and a location of a point of a camera that captured the 2D image frame. For instance, the method can include a depth estimator determining a depth mask with depth values. The depth mask can include depth364855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) values for each of the pixels (e.g., to each one or more of the pixels) of the depth mask. The depth mask pixels can correspond to the pixels of the 2D frame (e.g., via one to one or more pixel correlations, such as 1 to 1, 1 to 2, 1 to 4 or any other relation or correlation). The depth values can correspond to distances from each of the respective pixels of the depth map to a camera-based reference point, such as a center point of a lens of the camera capturing the 2D frame.

[0110] The method can include indicating a depth value for each pixel in the frame. The data processing system can generate a second mask (e.g., a depth mask) of the one or more masks using an ML model configured to include or utilize a neural network. The neural network can be trained to map an intensity of each pixel in the frame to the depth value. The depth value can be a distance between a camera that captures the frame and a location in the scene corresponding to the pixel. The data processing system can generate the second mask based on at least one of an optical center of the camera, a focal length of the camera, a scale factor of the camera, a principal point of the camera, a skew of the camera, or a geometric distortion of the camera.

[0111] At 530, the method can construct a 3D mesh. The one or more processors can generate or construct a three-dimensional mesh representing a 3D model or a 3D reconstruction of the surgical scene. The 3D mesh can include vector and parameter values along with anatomy segments and depth values. The method can include the one or more processors utilizing at least the segmentation mask and a depth mask to generate a 3D mesh representing the three-dimensional surgical scene captured by the 2D image frame. The mesh generator can construct the 3D mesh comprising a plurality of vertices, edges and faces that provide a three-dimensional representation of the scene. The 3D mesh can include three- dimensional coordinates along with any vector values depicting, describing, defining or otherwise providing a 3D reconstruction of the surgical scene.

[0112] At 535, the method can generate a 3D video of the surgical scene. The one or more processors can generate, based on the 3D mesh, a 3D representation of the surgical scene having with at least two view-angle representations of the surgical scene. For instance, the one or more processors can generate a view-angle frame at a different angle than the original view-angle of the 2D frame, thereby providing at least two view-angles of the surgical scene. The method can include the data processing system generating a video corresponding to a plurality of view angles of the frame. The video can be a 3D video and374855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) can be generated using the one or models and the 3D mesh. The video can be generated using the labels of the anatomical structures from the one or more masks.

[0113] The data processing system can use the video (e.g., the 3D representation or reconstruction) to generate or provide, using the one or more models, a plurality of images for the plurality of view angles of the frame. The data processing system can identify one or more holes in the plurality of images. The data processing system can utilize the inpainting function to update the plurality of images using an inpainting model to fill the one or more holes in the plurality of images. The data processing system can utilize the 3D video generator to generate the video with the updated plurality of images.

[0114] The data processing system can utilize the inpainting function to synthesize, using the inpainting model, new pixels based on the segments and the labels in the one or more masks and fill the one or more holes with the new pixels. The inpainting ML model can assign, to the new pixels, the label and the depth value based on the one or more masks. The inpainting ML model can include a generative machine learning model, which can be used to generate a prompt based on the segments and the labels in the first mask. The data processing system can utilize the prompt function to input the prompt into the generative machine learning model to synthesize new pixels. The inpainting ML model can fill the one or more holes with the new pixels synthesized by the generative machine learning model responsive to the prompt.

[0115] The inpainting model can include a generative machine learning model that can generate the prompt based on the second mask that indicates the distance between pixels in the frame and the reference point. The user interface can receive, via a graphical user interface element, an input with a range of view angles for which to generate the video. The 3D video generator can generate the video corresponding to the plurality of view angles responsive to the input with the range of view angles.

[0116] The user interface can display the video via a head-mounted display device, a display of a medical environment or any other display. The displayed 3D video or 3D representation can include any number of view-angle frames that can depict the surgical scene from their corresponding view-angles, each of which is different from the view-angle of the 2D image frame. The user interface can continue receiving user input (e.g., via a user input device, such as a mouse) and can generate view-angle frames based on the user input device movement, allowing the user to control the view-angle of the depiction. The user384855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) interface can display or output information or data on the generated image quality, such as data on the number or percentage of pixels that were inpainted versus pixels that were captured in the 2D frame.

[0117] Referring now to FIG. 6, among others, an example block diagram 600 of a computing system or a computer system 602 is shown. The computing system 602 can include or be used to implement a data processing system or its components. The architecture described in FIG. 6 can be used to implement the computing system 602, the robotic medical system 110, or the data processing system 118. The computing system 602 can include at least one bus 625 or other communication component for communicating information and at least one processor 630 or processing circuit coupled to the bus 625 for processing information. The computing system 602 can include one or more processors 630 or processing circuits coupled to the bus 625 for processing information. The computing system 602 can include at least one main memory 610, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 625 for storing information, and instructions to be executed by the processor 630. The main memory 610 can be used for storing information during execution of instructions by the processor 630. The computing system 602 can further include at least one read only memory (ROM) 615 or other static storage device coupled to the bus 625 for storing static information and instructions for the processor 630. A storage device 620, such as a solid state device, magnetic disk or optical disk, can be coupled to the bus 625 to persistently store information and instructions.

[0118] The computing system 602 can be coupled via the bus 625 to a display 635, such as a liquid crystal display, or active matrix display. The display 635 can display information to a user. An input device 605, such as a keyboard or voice interface can be coupled to the bus 625 for communicating information and commands to the processor 630. The input device 605 can include a touch screen of the display 635. The input device 605 can include a cursor control, such as a mouse, a trackball, or cursor direction keys, for communicating direction information and command selections to the processor 630 and for controlling cursor movement on the display 635.

[0119] The processes, systems and methods described herein can be implemented by the computing system 602 in response to the processor 630 executing an arrangement of instructions contained in main memory 610. Such instructions can be read into main memory 610 from another computer-readable medium, such as the storage device 620. Execution of394855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) the arrangement of instructions contained in main memory 610 causes the computing system 602 to perform the illustrative processes described herein. One or more processors in a multiprocessing arrangement can be employed to execute the instructions contained in main memory 610. Hard-wired circuitry can be used in place of or in combination with software instructions together with the systems and methods described herein. Systems and methods described herein are not limited to any specific combination of hardware circuitry and software.

[0120] Although an example computing system has been described in FIG. 6, the subject matter including the operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.

[0121] Some of the description herein emphasizes the structural independence of the aspects of the system components or groupings of operations and responsibilities of these system components. Other groupings that execute similar overall operations are within the scope of the present application. Modules can be implemented in hardware or as computer instructions on a non-transient computer readable storage medium, and modules can be distributed across various hardware or computer based components.

[0122] The systems described above can provide multiple ones of any or each of those components and these components can be provided on either a standalone system or on multiple instantiations in a distributed system. In addition, the systems and methods described above can be provided as one or more computer-readable programs or executable instructions embodied on or in one or more articles of manufacture. The article of manufacture can be cloud storage, a hard disk, a CD-ROM, a flash memory card, a PROM, a RAM, a ROM, or a magnetic tape. In general, the computer-readable programs can be implemented in any programming language, such as LISP, PERL, C, C++, C#, PROLOG, Python, or in any byte code language such as JAVA. The software programs or executable instructions can be stored on or in one or more articles of manufacture as object code.

[0123] Example and non-limiting module implementation elements include sensors providing any value determined herein, sensors providing any value that is a precursor to a value determined herein, datalink or network hardware including communication chips,404855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) oscillating crystals, communication links, cables, twisted pair wiring, coaxial wiring, shielded wiring, transmitters, receivers, or transceivers, logic circuits, hard-wired logic circuits, reconfigurable logic circuits in a particular non-transient state configured according to the module specification, any actuator including at least an electrical, hydraulic, or pneumatic actuator, a solenoid, an op-amp, analog control elements (springs, filters, integrators, adders, dividers, gain elements), or digital control elements.

[0124] The subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. The subject matter described in this specification can be implemented as one or more computer programs, e.g., one or more circuits of computer program instructions, encoded on one or more computer storage media for execution by, or to control the operation of, data processing apparatuses. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. While a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be, or be included in, one or more separate components or media (e.g., multiple CDs, disks, or other storage devices including cloud storage). The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0125] The terms “computing device”, “component” or “data processing apparatus” or the like encompass various apparatuses, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates414855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.

[0126] A computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program can correspond to a file in a file system. A computer program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0127] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatuses can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Devices suitable for storing computer program instructions and data can include non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0128] The subject matter described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client424855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification, or a combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0129] While operations are depicted in the drawings in a particular order, such operations are not required to be performed in the particular order shown or in sequential order, and all illustrated operations are not required to be performed. Actions described herein can be performed in a different order.

[0130] Having now described some illustrative implementations, it is apparent that the foregoing is illustrative and not limiting, having been presented by way of example. In particular, although many of the examples presented herein involve specific combinations of method acts or system elements, those acts and those elements may be combined in other ways to accomplish the same objectives. ACTs, elements and features discussed in connection with one implementation are not intended to be excluded from a similar role in other implementations.

[0131] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of “including” “comprising” “having” “containing” “involving” “characterized by” “characterized in that” and variations thereof herein, is meant to encompass the items listed thereafter, equivalents thereof, and additional items, as well as alternate implementations consisting of the items listed thereafter exclusively. In one implementation, the systems and methods described herein consist of one, each combination of more than one, or all of the described elements, acts, or components.

[0132] Any references to implementations or elements or acts of the systems and methods herein referred to in the singular may also embrace implementations including a plurality of these elements, and any references in plural to any implementation or element or act herein may also embrace implementations including only a single element. References in the singular or plural form are not intended to limit the presently disclosed systems or methods,434855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) their components, acts, or elements to single or plural configurations. References to any ACT or element being based on any information, act or element may include implementations where the act or element is based at least in part on any information, act, or element.

[0133] Any implementation disclosed herein may be combined with any other implementation or example, and references to “an implementation,” “some implementations,” “one implementation” or the like are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic described in connection with the implementation may be included in at least one implementation or example. Such terms as used herein are not necessarily all referring to the same implementation. Any implementation may be combined with any other implementation, inclusively or exclusively, in any manner consistent with the aspects and implementations disclosed herein.

[0134] References to “or” may be construed as inclusive so that any terms described using “or” may indicate any of a single, more than one, and all of the described terms. References to at least one of a conjunctive list of terms may be construed as an inclusive OR to indicate any of a single, more than one, and all of the described terms. For example, a reference to “at least one of ‘A’ and ‘B’” can include only ‘A’, only ‘B’, as well as both ‘A’ and ‘B’. Such references used in conjunction with “comprising” or other open terminology can include additional items.

[0135] Where technical features in the drawings, detailed description or any claim are followed by reference signs, the reference signs have been included to increase the intelligibility of the drawings, detailed description, and claims. Accordingly, neither the reference signs nor their absence have any limiting effect on the scope of any claim elements.

[0136] Modifications of described elements and acts such as variations in sizes, dimensions, structures, shapes and proportions of the various elements, values of parameters, mounting arrangements, use of materials, colors, orientations can occur without materially departing from the teachings and advantages of the subject matter disclosed herein. For example, elements shown as integrally formed can be constructed of multiple parts or elements, the position of elements can be reversed or otherwise varied, and the nature or number of discrete elements or positions can be altered or varied. Other substitutions, modifications, changes and omissions can also be made in the design, operating conditions444855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) and arrangement of the disclosed elements and operations without departing from the scope of the present disclosure.454855-6965-8828.1

Claims

Atty. Dkt: 135039-0420 (P06964-WO)CLAIMSWhat is claimed is:

1. A system, comprising: one or more processors, coupled with memory, to: identify a frame that captures a single view angle of a scene in a medical procedure performed with a robotic medical system; generate, using one or more models trained with machine learning, one or more masks that segment an anatomical structure in the frame, label the anatomical structure in the frame, and indicate a depth value for each pixel in the frame; construct a 3D mesh based on the one or more masks; and generate a video corresponding to a plurality of view angles of the frame using the one or models and the 3D mesh and the labels of the anatomical structures from the one or more masks.

2. The system of claim 1, comprising the one or more processors to: detect a trigger condition in a video stream of the medical procedure; select the frame from the video stream responsive to at least one of a detection of the trigger condition or an input in a user interface; and determine to generate the video for the frame based on the detection of the trigger condition.

3. The system of claim 2, comprising the one or more processors to: detect the trigger condition based on a type of task performed with the robotic medical system in the video stream.

4. The system of claim 2, comprising the one or more processors to: receive a data stream of the medical procedure captured via one or more sensors of the robotic medical system; determine, based on the data stream, a range of view angles in a portion of the video that comprises the frame; and detect the trigger condition based on an occurrence of a milestone event in the frame and the range of view angles in the portion of the video stream being less than or equal to a464855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) threshold.

5. The system of claim 1, comprising the one or more processors to: detect the trigger condition based on at least one of a profile associated with a surgeon that performs the medical procedure with the robotic medical system or an input in a user interface.

6. The system of claim 1, comprising the one or more processors to: generate a first mask of the one or more masks using a vision transformer based neural network trained with a dataset comprising images of medical procedures that are annotated with labels, wherein the first mask comprises the segment of the anatomical structure and the label of the anatomical structure.

7. The system of claim 1, comprising the one or more processors to: generate a second mask of the one or more masks using a neural network trained to map an intensity of each pixel in the frame to the depth value, wherein the depth value is a distance between a camera that captures the frame and a location in the scene corresponding to the pixel.

8. The system of claim 7, comprising the one or more processors to: generate the second mask based on at least one of an optical center of the camera, a focal length of the camera, a scale factor of the camera, a principal point of the camera, a skew of the camera, or a geometric distortion of the camera.

9. The system of claim 1, comprising the one or more processors to: construct the 3D mesh comprising a plurality of vertices, edges and faces that provide a three-dimensional representation of the scene.

10. The system of claim 1, comprising the one or more processors to: generate, using the one or more models, a plurality of images for the plurality of view angles of the frame; identify one or more holes in the plurality of images; update the plurality of images using an inpainting model to fill the one or more holes in the plurality of images; and generate the video with the updated plurality of images.474855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO)11. The system of claim 10, comprising the one or more processors to: synthesize, using the inpainting model, new pixels based on the segments and the labels in the one or more masks; and fill the one or more holes with the new pixels.

12. The system of claim 11, comprising the one or more processors to: assign, to the new pixels, the label and the depth value based on the one or more masks.

13. The system of claim 10, wherein the inpainting model comprises a generative machine learning model, comprising the one or more processors to: generate a prompt based on the segments and the labels in the one or more masks; input the prompt into the generative machine learning model to synthesize new pixels; and fill the one or more holes with the new pixels synthesized by the generative machine learning model responsive to the prompt.

14. The system of claim 13, wherein the inpainting model comprises a generative machine learning model, comprising the one or more processors to: generate the prompt based on the one or more masks that indicate a distance between pixels in the frame and a reference point.

15. The system of claim 1, comprising the one or more processors to: receive, via a graphical user interface element, an input with a range of view angles for which to generate the video; and generate the video corresponding to the plurality of view angles responsive to the input with the range of view angles.

16. The system of claim 1, comprising the one or more processors to: display the video via a head-mounted display device.

17. A method, comprising: identifying, by one or more processors coupled with memory, a frame that captures a scene in a medical procedure performed with a robotic medical system;484855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) generating, by the one or more processors, using one or more models trained with machine learning, a first mask that segments an anatomical structure in the frame and labels the anatomical structure in the frame; generating, by the one or more processors, using the one or more models, a second mask that indicates a depth value for each pixel in the frame; constructing, by the one or more processors, a 3D mesh based on the first mask and the second mask; and generating, by the one or more processors, a plurality of images corresponding to a plurality of view angles of the frame using the one or models and the 3D mesh and the labels of the anatomical structures.

18. The method of claim 17, comprising: receiving, by the one or more processors, a request to generate a three-dimensional video for the frame; and generating, by the one or more processors, responsive to the request, the plurality of images to create the three-dimensional video.

19. The method of claim 17, comprising: detecting, by the one or more processors, a trigger condition in a video stream of the medical procedure; selecting, by the one or more processors, the frame from the video stream responsive to detection of the trigger condition; and determining, by the one or more processors, to generate the video for the frame based on the detection of the trigger condition.

20. A non-transitory computer-readable medium storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to: identify a frame that captures a single view angle of a scene in a medical procedure performed with a robotic medical system; generate, using one or more models trained with machine learning, one or more masks that segment an anatomical structure in the frame, label the anatomical structure in the frame, and indicate a depth value for each pixel in the frame; construct a 3D mesh based on the one or more masks; and generate a video corresponding to a plurality of view angles of the frame using the one494855-6965-8828.1Atty. Dkt: 135039-0420 (P06964-WO) or models and the 3D mesh and the labels of the anatomical structures from the one or more masks.504855-6965-8828.1

Citation Information

Patent Citations

  • Method and System for Calculating Resected Tissue Volume from 2D / 2.5D Intraoperative Image Data

    US20170105601A1

  • Multi-view segmentation and perceptual inpainting with neural radiance fields

    US20240153046A1

  • Apparatus, systems, and methods for intraoperative instrument tracking and information visualization

    WO2023166417A1