Systems and methods for sensor data generation in medical environments
By training a machine learning model with multi-modal data, the sparsity of 3D point clouds in medical environments is improved, enhancing reconstruction and robotic perception.
Patent Information
- Application Number
- PCT/US2024/053225
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-01
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-08
AI Technical Summary
Medical environments, such as operating rooms, face challenges with sparse 3D point clouds due to highly reflective and IR-absorbing objects, hindering full reconstruction and impacting robotic perception.
A machine learning model is trained using multi-modal data from sensors to enhance depth data, improving the sparsity of 3D point clouds by leveraging intensity and depth modalities.
The approach effectively increases the density of 3D point clouds, enabling better reconstruction of medical environments and improving robotic perception within these spaces.
Smart Images

Figure US2024053225_08052025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR SENSOR DATA GENERATION INMEDICAL ENVIRONMENTSCLAIM OF PRIORITY
[0001] This application claims priority to U.S. Patent Application No. 63 / 595,244, filedNovember 1, 2023, the full disclosure of which is incorporated herein in its entirety.TECHNICAL FIELD
[0002] Various of the disclosed embodiments relate to systems, apparatuses, methods, and non-transitory computer-readable media for generating three-dimensional point clouds for surgical and hospital processes using machine learning models.BACKGROUND
[0003] Medical environments such as operating rooms (ORs) typically feature various types of highly reflective objects (e.g., metallic medical tools) and infrared (IR) absorbing objects which do not reflect IR light well. Thus, 3D point clouds generated for the medical environments using multi-modal sensors are traditionally sparse. Sparse 3D point clouds hinder the capability to fully reconstruct the medical environments and layout for workflow analytic applications, and negatively impact 3D robot perception within medical environments.BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Various of the embodiments introduced herein may be better understood by referring to the following Detailed Description in conjunction with the accompanying drawings, in which like reference numerals indicate identical or functionally similar elements:
[0005] FIG. 1 A is a schematic view of various elements appearing in a surgical theater during a surgical operation, as may occur in relation to some embodiments.
[0006] FIG. IB is a schematic view of various elements appearing in a surgical theater during a surgical operation employing a robotic surgical system, as may occur in relation to some embodiments.
[0007] FIG. 2A is a schematic depth map rendering from an example theater-wide sensor perspective, as may be used in some embodiments.
[0008] FIG. 2B is a schematic top-down view of objects in the theater of FIG. 2A, with corresponding sensor locations.
[0009] FIG. 3 is a diagram illustrating an example method for training a machine learning model to improve sparsity of depth images, according to various embodiments.
[0010] FIG. 4 illustrates an input depth image, a patchified input depth image, and a patchified and masked input depth image, according to various embodiment.
[0011] FIG. 5 illustrates an input intensity image, a patchified input intensity image, and a patchified and masked input intensity image, according to various embodiment.
[0012] FIG. 6 is a flowchart diagram illustrating an example method for training a machine learning model to improve sparsity of a depth image captured for a medical procedure, according to various arrangements.
[0013] FIG. 7 is a flowchart diagram illustrating an example method for training a machine learning model to improve sparsity of a depth image captured for a medical procedure, according to various arrangements.
[0014] FIG. 8 is a flowchart diagram illustrating an example method for training a machine learning model to improve sparsity of a depth image captured for a medical procedure, according to various arrangements.
[0015] FIG. 9 is a flowchart diagram illustrating an example method for using a machine learning model to improve sparsity of a depth image captured during a medical procedure, according to various embodiments.
[0016] FIG. 10 is a flowchart diagram illustrating an example method for using a machine learning model to improve sparsity of a depth image captured during a medical procedure, according to various embodiments.
[0017] FIG. 11 is a flowchart diagram illustrating an example method for using a machine learning model to improve sparsity of a depth image captured during a medical procedure, according to various embodiments.
[0018] FIG. 12 is a block diagram of an example computer system as may be used in conjunction with some of the embodiments.
[0019] The specific examples depicted in the drawings have been selected to facilitate understanding. Consequently, the disclosed embodiments should not be restricted to the specific details in the drawings or the corresponding disclosure. For example, the drawings may not be drawn to scale, the dimensions of some elements in the figures may have been adjusted to facilitate understanding, and the operations of the embodiments associated with the flow diagrams may encompass additional, alternative, or fewer operations than those depicted here. Thus, some components and / or operations may be separated into different blocks or combined into a single block in a manner other than as depicted. The embodiments are intended to cover all modifications, equivalents, and alternatives falling within the scope of the disclosed examples, rather than limit the embodiments to the particular examples described or depicted.DETAILED DESCRIPTIONExample Surgical Theaters Overview
[0020] FIG. 1 A is a schematic view of various elements appearing in a surgical theater 100a during a surgical operation as may occur in relation to some embodiments. Particularly, FIG. 1A depicts a non-robotic surgical theater 100a, wherein a patient-side surgeon 105a performs an operation upon a patient 120 with the assistance of one or more assisting members 105b, who may themselves be surgeons, physician’s assistants, nurses, technicians, etc. The surgeon 105a may perform the operation using a variety of tools, e.g., a visualization tool 110b such as a laparoscopic ultrasound, visual image acquiring endoscope, etc., and a mechanical instrument 110a such as scissors, retractors, a dissector, etc.
[0021] The visualization tool 110b provides the surgeon 105a with an interior view of the patient 120, e.g., by displaying visualization output from an imaging device mechanicallyand electrically coupled with the visualization tool 110b. The surgeon may view the visualization output, e.g., through an eyepiece coupled with visualization tool 110b or upon a display 125 configured to receive the visualization output. For example, where the visualization tool 110b is a visual image acquiring endoscope, the visualization output may be a color or grayscale image. Display 125 may allow assisting member 105b to monitor surgeon 105a’s progress during the surgery. The visualization output from visualization tool 110b may be recorded and stored for future review, e.g., using hardware or software on the visualization tool 110b itself, capturing the visualization output in parallel as it is provided to display 125, or capturing the output from display 125 once it appears on-screen, etc. While two-dimensional video capture with visualization tool 110b may be discussed extensively herein, as when visualization tool 110b is a visual image endoscope, one will appreciate that, in some embodiments, visualization tool 110b may capture depth data instead of, or in addition to, two- dimensional image data (e.g., with a laser rangefinder, stereoscopy, etc.).
[0022] A single surgery may include the performance of several groups (e.g., phases or stages) of actions, each group of actions forming a discrete unit referred to herein as a task. For example, locating a tumor may constitute a first task, excising the tumor a second task, and closing the surgery site a third task. Each task may include multiple actions, e.g., a tumor excision task may require several cutting actions and several cauterization actions. While some surgeries require that tasks assume a specific order (e.g., excision occurs before closure), the order and presence of some tasks in some surgeries may be allowed to vary (e.g., the elimination of a precautionary task or a reordering of excision tasks where the order has no effect). Transitioning between tasks may require the surgeon 105a to remove tools from the patient, replace tools with different tools, or introduce new tools. Some tasks may require that the visualization tool 110b be removed and repositioned relative to its position in a previous task. While some assisting members 105b may assist with surgery -related tasks, such as administering anesthesia 115 to the patient 120, assisting members 105b may also assist with these task transitions, e.g., anticipating the need for a new tool 110c.
[0023] Advances in technology have enabled procedures such as that depicted in FIG. 1 A to also be performed with robotic systems, as well as the performance of procedures unable to be performed in non-robotic surgical theater 100a. Specifically, FIG. IB is a schematic viewof various elements appearing in a surgical theater 100b during a surgical operation employing a robotic surgical system, such as a da Vinci™ surgical system, as may occur in relation to some embodiments. Here, patient side cart 130 having tools 140a, 140b, 140c, and 140d attached to each of a plurality of arms 135a, 135b, 135c, and 135d, respectively, may take the position of patient-side surgeon 105a. As before, one or more of tools 140a, 140b, 140c, and 140d may include a visualization tool (here visualization tool 140d), such as a visual image endoscope, laparoscopic ultrasound, etc. An operator 105c, who may be a surgeon, may view the output of visualization tool 140d through a display 160a upon a surgeon console 155. By manipulating a hand-held input mechanism 160b and pedals 160c, the operator 105c may remotely communicate with tools 140a-d on patient side cart 130 so as to perform the surgical procedure on patient 120. Indeed, the operator 105c may or may not be in the same physical location as patient side cart 130 and patient 120 since the communication between surgeon console 155 and patient side cart 130 may occur across a telecommunication network in some embodiments. An electronics / control console 145 may also include a display 150 depicting patient vitals and / or the output of visualization tool 140d.
[0024] Similar to the task transitions of non-robotic surgical theater 100a, the surgical operation of theater 100b may require that tools 140a-d, including the visualization tool 140d, be removed or replaced for various tasks as well as new tools, e.g., new tool 165, be introduced. As before, one or more assisting members 105d may now anticipate such changes, working with operator 105c to make any necessary adjustments as the surgery progresses.
[0025] Also similar to the non-robotic surgical theater 100a, the output from the visualization tool 140d may here be recorded, e.g., at patient side cart 130, surgeon console 155, from display 150, etc. While some tools 110a, 110b, 110c in non-robotic surgical theater 100a may record additional data, such as temperature, motion, conductivity, energy levels, etc., the presence of surgeon console 155 and patient side cart 130 in theater 100b may facilitate the recordation of considerably more data than is only output from the visualization tool 140d. For example, operator 105c’s manipulation of hand-held input mechanism 160b, activation of pedals 160c, eye movement with respect to display 160a, etc., may all be recorded. Similarly, patient side cart 130 may record tool activations (e.g., the application of radiative energy, closing of scissors, etc.), movement of instruments, etc., throughout the surgery. In someembodiments, the data may have been recorded using an in-theater recording device, which may capture and store sensor data locally or at a networked location (e.g., software, firmware, or hardware configured to record surgeon kinematics data, console kinematics data, instrument kinematics data, system events data, patient state data, etc., during the surgery).
[0026] Within each of theaters 100a, 100b, or in network communication with the theaters from an external location, may be computer systems 190a and 190b, respectively (in some embodiments, computer system 190b may be integrated with the robotic surgical system, rather than serving as a standalone workstation). As will be discussed in greater detail herein, the computer systems 190a and 190b may facilitate, e.g., data collection, data processing, etc.
[0027] Similarly, many of theaters 100a, 100b may include sensors placed around the theater, such as sensors 170a and 170c, respectively, configured to record activity within the surgical theater from the perspectives of their respective fields of view 170b and 170d. Sensors 170a and 170c may be, e.g., visual image sensors (e.g., color, RGB, or grayscale image sensors), depth-acquiring sensors (e.g., via stereoscopically acquired visual image pairs, via Time-of-Flight (ToF) with a laser rangefinder, a Photonic Mixer Device (PMD), structural light, etc.), or a multi-modal sensor including a combination of a visual image sensor and a depth-acquiring sensor (e.g., a red green blue depth RGB-D sensor). In some embodiments, sensors 170a and 170c may also include audio acquisition sensors or sensors specifically dedicated to audio acquisition may be placed around the theater. A plurality of such sensors may be placed within theaters 100a, 100b, possibly with overlapping fields of view and sensing range, to achieve a more holistic assessment of the surgery. For example, depth-acquiring sensors may be strategically placed around the theater so that their resulting depth frames at each moment may be consolidated into a single three-dimensional virtual element model depicting objects in the surgical theater. Examples of a three-dimensional virtual element model include a three-dimensional point cloud (also referred to as three-dimensional point cloud data). Similarly, sensors may be strategically placed in the theater to focus upon regions of interest. For example, sensors may be attached to display 125, display 150, or patient side cart 130 with fields of view focusing upon the patient 120’s surgical site, attached to the walls or ceiling, etc. Similarly, sensors may be placed upon console 155 to monitor the operator105c. Sensors may likewise be placed upon movable platforms specifically designed to facilitate orienting of the sensors in various poses within the theater.
[0028] As used herein, a “pose” refers to a position or location and an orientation of a body. For example, a pose refers to the translational position and rotational orientation of a body. For example, in a three-dimensional space, one may represent a pose with six total degrees of freedom. One will readily appreciate that poses may be represented using a variety of data structures, e.g., with matrices, with quaternions, with vectors, with combinations thereof, etc. Thus, in some situations, when there is no rotation, a pose may include only a translational component. Conversely, when there is no translation, a pose may include only a rotational component.
[0029] Similarly, for clarity, “theater- wide” sensor data refers herein to data acquired from one or more sensors configured to monitor a specific region of the theater (the region encompassing all, or a portion, of the theater) exterior to the patient, to personnel, to equipment, or to any other objects in the theater, such that the sensor can perceive the presence within, or passage through, at least a portion of the region of the patient, personnel, equipment, or other objects, throughout the surgery. Sensors so configured to collect such “theater- wide” data are referred to herein as “theater-wide sensors.” For clarity, one will appreciate that the specific region need not be rigidly fixed throughout the procedure, as, e.g., some sensors may cyclically pan their field of view so as to augment the size of the specific region, even though this may result in temporal lacunae for portions of the region in the sensor’s data (lacunae which may be remedied by the coordinated panning or fields of view of other nearby sensors). Similarly, in some cases, personnel or robotics systems may be able to relocate theater-wide sensors, changing the specific region, throughout the procedure, e.g., to better capture different tasks. Accordingly, sensors 170a and 170c are theater- wide sensors configured to produce theaterwide data. “Visualization data” refers herein to visual image or depth image data captured from a sensor. Thus, visualization data may or may not be theater-wide data. For example, visualization data captured at sensors 170a and 170c is theater- wide data, whereas visualization data captured via visualization tool 140d would not be theater-wide data (for at least the reason that the data is not exterior to the patient).Example Theater-Wide Sensor Topologies
[0030] For further clarity regarding theater-wide sensor deployment, FIG. 2A is a schematic depth map rendering from an example theater-wide sensor perspective 205 as may be used in some embodiments. Specifically, this example depicts depth values corresponding to an electronics / control console 205a (e.g., the electronics / control console 145) and a nearby tray 205b, and cabinet 205c. Also within the field of view are depth values associated with a first technician 205d, presently adjusting a robotic arm (associated with depth values 205f) upon a robotic surgical system (associated with depth values 205e). Team members, with corresponding depth values 205g, 205h, and 205i, likewise appear in the field of view, as does a portion of the surgical table 205j . Depth values 2051 corresponding to a movable dolly and a boom with a lighting system’s depth values 205k also appear within the field of view.
[0031] The theater-wide sensor capturing the perspective 205 may be only one of several sensors placed throughout the theater. For example, FIG. 2B is a schematic top- down view of objects in the theater at a given moment during the surgical operation. Specifically, the perspective 205 may have been captured via a theater-wide sensor 220a with corresponding field of view 225a. Thus, for clarity, cabinet depth values 205c may correspond to cabinet 210c, electronics / control console depth values 205a may correspond to electronics / control console 210a, and tray depth values 205b may correspond to tray 210b. Robotic system 210e may correspond to depth values 205e, and each of the individual team members 210d, 210g, 21 Oh, and 210i may correspond to depth values 205d, 205g, 205h, and 205i, respectively. Similarly, dolly 2101 may correspond to depth values 2051. Depth values 205j may correspond to table 21 Oj (with an outline of a patient shown here for clarity, though the patient has not yet been placed upon the table corresponding to depth values 205j in the example perspective 205). A top-down representation of the boom corresponding to depth values 205k is not shown for clarity, though one will appreciate that the boom may likewise be considered in various embodiments.
[0032] As indicated, each of the sensors 220a, 220b, 220c is associated with different fields of view 225a, 225b, and 225c, respectively. The fields of view 225a-c may sometimes have complementary characters, providing different perspectives of the same object, or providing a view of an object from one perspective when it is outside, or occluded within, another perspective. Complementarity between the perspectives may be dynamic bothspatially and temporally. Such dynamic character may result from movement of an object being tracked, but also from movement of intervening occluding objects (and, in some cases, movement of the sensors themselves). For example, at the moment depicted in FIGs. 2A and 2B, the field of view 225a has only a limited view of the table 21 Oj , as the electronics / control console 210a substantially occludes that portion of the field of view 225a. Consequently, in the depicted moment, the field of view 225b is better able to view the surgical table 210j . However, neither field of view 225b nor 225a has an adequate view of the operator 21 On in console 210k. To observe the operator 210n (e.g., when they remove their head in accordance with “head out” events), field of view 225c may be more suitable. However, over the course of the data capture, these complementary relationships may change. For example, before the procedure begins, electronics / control console 210a may be removed and the robotic system 210e moved into the position 210m. In this configuration, field of view 225a may instead be much better suited for viewing the patient table 21 Oj than the field of view 225b. As another example, movement of the console 210k to the presently depicted pose of electronics / control console 210a may render field of view 225a more suitable for viewing operator 21 On, than field of view 225c. Suitability of a field of view may thus depend upon the number and duration of occlusions, quality of the field of view (e.g., how close the object of interest is to the sensor), and movement of the object of interest within the theater. Such changes may be transitory and short in duration, as when a team member moving in the theater briefly occludes a sensor, or they may be chronic or sustained, as when equipment is moved into a fixed position throughout the duration of the procedure.
[0033] As mentioned, the theater-wide sensors may take a variety of forms and may, e.g., be configured to acquire visual image data, depth data, both visual and depth data, etc. One will appreciate that visual and depth image captures may likewise take on a variety of forms, e.g., to afford increased visibility of different portions of the theater. For example, FIG. 2C is a pair of images 250b, 255b depicting a grid-like pattern of orthogonal rows and columns in perspective, as captured from a theater-wide sensor having a rectilinear view and a theaterwide sensor having a fisheye view, respectively. More specifically, some theater-wide sensors may capture rectilinear visual images or rectilinear depth frames, e.g., via appropriate lenses, post-processing, combinations of lenses and post-processing, etc. while other theater-wide sensors may instead, e.g., acquire fisheye or distorted visual images or rectilinear depth frames,via appropriate lenses, post-processing, combinations of lenses and post-processing, etc. For clarity, image 250b depicts a checkboard pattern in perspective from a rectilinear theater wide sensor. Accordingly, the orthogonal rows and columns 250a shown here in perspective, retain linear relations with their vanishing points. In contrast, image 255b depicts the same checkboard pattern in the same perspective, but from a fish-eye theater-wise sensor perspective. Accordingly, the orthogonal rows and columns 255a, while in reality retaining a linear relationship with their vanishing points (as they appear in image 250b) appear here from the sensor data as having curved relations with their vanishing points. Thus, each type of sensor, and other sensor types, may be used alone, or in some instances, in combination, in connection with various embodiments.
[0034] Similarly, one will appreciate that not all sensors may acquire perfectly rectilinear, fisheye, or other desired mappings. Accordingly, checkered patterns, or other calibration fiducials (such as known shapes for depth systems), may facilitate determination of a given theater-wide sensor’s intrinsic parameters. For example, the focal point of the fisheye lens, and other details of the theater-wide sensor (principal points, distortion coefficients, etc.), may vary between devices and even across the same device over time. Thus, it may be necessary to recalibrate various processing methods for the particular device at issue, anticipating the device variation when training and configuring a system for machine learning tasks. Additionally, one will appreciate that the rectilinear view may be achieved by undistorting the fisheye view once the intrinsic parameters of the camera are known (which may be useful, e.g., to normalize disparate sensor systems to a similar form recognized by a machine learning architecture). Thus, while a fisheye view may allow the system and users to more readily perceive a wider field of view than in the case of the rectilinear perspective, when a processing system is considering data from some sensors acquiring undistorted perspectives and other sensors acquiring distorted perspectives, the differing perspectives may be normalized to a common perspective form (e.g., mapping all the rectilinear data to a fisheye representation or vice versa).Point Cloud Generation
[0035] Systems, methods, apparatuses, and non-transitory computer-readable media are provided for enhancing point cloud and depth maps captured using sensors such as thetheater-wide sensors. At least one sensor can provide streams of data in multiple modalities. For example, a first sensor can provide a stream of data referred to as first input data in a first modality (e.g., a three-dimensional depth modality or a three-dimensional point cloud modality), which can include a stream of input depth images (or frames). The first sensor can include a depth-acquiring sensor, such as a ToF sensor, a PMD, a structural light sensor, and so on. A second sensor can provide a stream of data referred to as second input data in a second modality, which includes one or more of an intensity modality, vision modality, a heatmap modality, wireless signal modality, and so on. For example, in the intensity modality, the second input data includes input intensity image (or frames), and the second sensor can include a ToF sensor, a PMD, a structural light sensor, and so on. For example, in the vision modality, the second input data includes input RGB images (or frames), and the second sensor can include a visual image sensor (e.g., a color, RGB, or grayscale image sensor). For example, in the heatmap modality, the second input data includes input heatmap images (or frames), and the second sensor can include an infrared sensor. For example, in the wireless signal modality, the second input data includes wireless signal sensed images (or frames), and the second sensor can include a wireless signal sensor (with wireless transmitters, receivers, and so on). The sights gained from the second input data can be used to improve data quality (e.g., density, sparsity, and so on) of the first input data. Depth and intensity modalities are well suited for capturing information and data regarding medical environments and for medical procedures given that the depth and intensity modalities can provide data that can accurately and efficiently capture information and data regarding medical environments and for medical procedures without collecting protected health information (PHI).
[0036] In some embodiments, the first and second sensors can be provided as separate sensors to separately provide the streams of the first input data and the second input data. That is, each of the first sensor and the second sensor can sense and output data in a different modality. In other embodiments, a multi-modal sensor that combines the functionalities of the first and second sensors can output both the first input data and the second input data. In some examples, a multi-modal sensor emits Infrared (IR) light (e.g., pulse-based or continuous waves) toward an environment (e.g., the within theaters 100a, 100b) and outputs the first input data including a depth image and the second input data corresponding to an intensity image and a depth image of the environment using the reflections of the emitted IR light. Examplesof the multi-modal sensor include a ToF sensor, a PMD, a structured light sensor, and so on. In some examples, a multi-modal sensor can include a combination of a depth acquiring sensor and a visual image sensor. In some examples, a multi-modal sensor can include a combination of a depth acquiring sensor and an infrared sensor. In some examples, a multi-modal sensor can include a combination of a depth acquiring sensor and a wireless signal sensor.
[0037] In some embodiments, each of the one or more each theater-wide sensor described herein can be a first sensor, a second sensor, or a multi-modal sensor that can be disposed within or around an indoor environment such as medical environments (e.g., ORs, within the theaters 100a, 100b) to create a three-dimensional perception of the medical environments, used for optimization of the medical environment layout, workflow, and robotic perception. The combination of the first input data and the second input data can be referred to as multi-modal data.
[0038] The medical environment typically feature various types of highly reflective objects (e.g., metallic medical tools) and IR absorbing objects (which do not reflect IR light well). Thus, 3D point clouds generated using the first input data (e.g., the input depth image) can traditionally be sparse in some examples. Sparse three-dimensional point clouds hinder the capability to fully reconstruct the medical environment scene and layout for workflow analytic applications, and negatively impact three-dimensional robot perception.
[0039] In some embodiments, a machine learning model can leverage the multi-modal data that can be time synchronized to enhance depth data and therefore enhancing the generation of three-dimensional point cloud data. The machine learning model is trained to learn to in-fill the first input data such as the depth data (e.g., a portion of a depth image) using information of the second input data. A database of dataset obtained by deploying first sensors, second sensors, and / or multi-modal sensors in multiple medical environments can be used as training data. This dataset contains multi-modal data from various medical settings, procedures, medical environments, etc. The machine learning model can be trained using unsupervised training (without labels).
[0040] FIG. 3 is a diagram illustrating an example training system 300 architecture for training a machine learning model to improve density / sparsity of a first data modality (e.g.,depth data, depth images, 3D point cloud data) using a second data modality (e.g., intensity data, visual images, RGB images, heatmaps, etc.), according to various embodiments. The machine learning model can be implemented using a first encoder 320, the second encoder 330, and a multi-modal decoder 340. That is, the machine learning model updates the weights, parameters, coefficients, constants, and other information in the first encoder 320, the second encoder 330, and the multi-modal decoder 340. The training system 300 has suitable processing and memory capabilities. Examples of the first encoder 320 and the second encoder 330 include a transformer.
[0041] The training system 300 can receive first input data 301 of a first modality and a second input data 302 of a second modality. Examples of the first input data 301 includes an input depth image of a frame of an input depth video. The first modality includes a three- dimensional depth modality, a three-dimensional point cloud modality, depth modality, and so on. In some examples, the first input data 301 can be referred to as three-dimensional data as it provides information regarding an object in the theater in three dimensions. Examples of the second input data 302 includes at least one of an input intensity image or a frame of input intensity video (for the intensity modality), an input RGB image or a frame of input RGB video (for the vision modality), an input heatmap image or a frame of input RGB video (for the depth modality), an input wireless signal sensed images or a frame of input wireless signal sensed video (for the wireless signal modality). In some examples, the second input data 302 can be referred to as two-dimensional data as it provides information regarding an object in the theater in two dimensions.
[0042] The combination of the first input data 301 and the second input data 302 can be referred to as multi-modal input data. In some examples, the first sensor outputs the first input data 301, and the second sensor outputs the second input data 302. In some examples, a multi-modal sensor can output both the first input data 301 and the second input data 302.
[0043] In some examples, the first input data 301 and the second input data 302 are time synchronized. That is, the first input data 301 and the second input data 302 are frames of their respective video streams at a same time as defined by a same frame sequence or a same timestamp. For example, the first input data 301 and the second input data 302 can use a same timestamp or frame sequence system or authority. In the examples in which the first sensorand the second sensor are separate sensors, the first sensor and the second sensor use a same timestamp or frame sequence system or authority. In the example in which the multi-modal sensor outputs both the first input data 301 and the second input data 302, the multi-modal sensor aligns the first input data 301 and the second input data 302 to a same timestamp or frame sequence. In some examples, the timestamps and the frame sequence of one of the first sensor or the second sensor can be adjusted (e.g., based on offsets) to align with the timestamps and the frame sequence of the other one of the first sensor or the second sensor.
[0044] In some examples, the first input data 301 and the second input data 302 are asynchronous in the time domain. That is, the first input data 301 and the second input data 302 are frames of their respective video streams at different times as defined by different frame sequence numbers or different timestamps. In other words, the machine learning model can be trained using multi-modal data that is both synchronous and asynchronous. This allows both historic information contained in the second input data 302 (a frame of the second input data 302 that is earlier than a frame of the first input data 301) and future information contained in the second input data 302 (a frame of the second input data 302 that is later than a frame of the first input data 301) to improve the quality of the first input data 301.
[0045] In some examples, the first input data 301 and the second input data 302 are data collected from a same sensor pose, in the examples in which both the first input data 301 and the second input data 302 can provided by a same multi-modal sensor or by the first and second sensors with the same pose within the environment. Improving the sparsity of the first input data 301 using multimodal data collected from a same sensor pose allow the machine learning model to learn from information contained in the second input data 302 that has significant correlation with the first input data 301. In some examples, the first input data 301 and the second input data 302 are data collected from different sensor poses, in the example in which the first sensor and the second sensor have different poses within the environment. Improving the sparsity of the first input data 301 using multimodal data collected from different sensor poses allow the machine learning model to learn from information that may be absent from the second input data 302 collected from a same sensor pose, such as information about occluded aspects of objects that would otherwise be absent from a particular sensor pose.
[0046] The training system 300 can run multiple iterations for multiple frames of the first input data 301 and the second input data 302. In some iterations, the first input data 301 and the second input data 302 are time synchronized. In some iterations, the first input data 301 and the second input data 302 are asynchronous in the time domain. In some iterations, the first input data 301 and the second input data 302 have the same sensor pose. In some iterations, the first input data 301 and the second input data 302 have different sensor poses. In different iterations, a same type of the second input data 302 or modality thereof (e.g., intensity modality, vision modality, a heatmap modality, or wireless signal modality) can be used. In different iterations, different types of the second input data 302 or modalities thereof (e.g., different ones of intensity modality, vision modality, a heatmap modality, or wireless signal modality) can be used.
[0047] The first input data 301 can be patchified and masked at 310a. The first input data 301 can be segmented (e.g., partitioned, divided, or patchified) into first patches, and one or more patches of the first patches can be masked. FIG. 4 illustrates an example input depth image 410, an example patchified input depth image 420, and an example patchified and masked input depth image 430, according to various embodiments. The input depth image 410 is an example of the first input data 301. As shown, the input depth image 410 can be segmented into multiple patches, each having a square or rectangular shape to obtain the patchified input depth image 420. While 16 square or rectangular patches are shown in the patchified input depth image 420, the number, shape, size, and arrangement of the patches can differ. The position of each patch can be defined by a location (e.g., a set of coordinates, a location index, and so on) of that patch within the patchified input depth image 420. One or more of the patches of the patchified input depth image 420 can be masked to obtain the masked input depth image 430. The masked patches are shown in grey. The masked input depth image 430 is an example of the masked / patchified first input data 311.
[0048] The images 410, 420, and 430 are visualizations of depth data. For example, a particular sensed depth value at a location in each of the images 410, 420, and 430 (e.g., pixel location) can be graphically illustrated with color. A warmer color represents a closer distance between the sensor and a point on an object (e.g., lesser depth), and a cooler color represents a farther distance between the sensor and a point on an object (e.g., greater depth).
[0049] The second input data 302 can be patchified and masked at 310b. The second input data 302 can be segmented (e.g., partitioned, divided, or patchified) into second patches, and one or more patches of the second patches can be masked. FIG. 5 illustrates an example input intensity image 510, an example patchified input intensity image 520, and an example patchified and masked input intensity image 530, according to various embodiments. The input intensity image 510 is an example of the second input data 302. As shown, the input intensity image 510 can be segmented into multiple patches, each having a square or rectangular shape to obtain the patchified input depth image 520. While 16 square or rectangular patches are shown in the patchified input intensity image 520, the number, shape, size, and arrangement of the patches can differ. The position of each patch can be defined by a location (e.g., a set of coordinates, a location index, and so on) of that patch within the patchified input intensity image 520. One or more of the patches of the patchified input intensity image 520 can be masked to obtain the masked input intensity image 530. The masked patches are shown in grey. The masked input intensity image 530 is an example of the masked / patchified second input data 312.
[0050] The images 510, 520, and 530 are visualizations of intensity data. For example, a particular sensed intensity value at a location in each of the images 510, 520, and 530 (e.g., pixel location) can be graphically illustrated with grayscale color. A brighter color represents greater measured intensity value, and a cooler color represents a lesser measured intensity value.
[0051] Other types of the second input data 302, such as two-dimensional RGB images, heatmap images, wireless signal sensed images, and so on can be likewise patchified and masked using similar methods.
[0052] In some examples, the first encoder 320 includes a depth encoder. The masked / patchified first input data 311 is applied to the first encoder 320 as an input to determine the first encoder output 321. The first encoder output 321 includes one or more features for the masked / patchified first input data 311. The first encoder output 321 can be at least one depth feature or a three-dimensional feature. For example, the first encoder output 321 includes at least one feature (e.g., one-dimensional feature), at least one tensor, at least one value, or so on for each patch of the masked / patchified first input data 311. For any maskedpatches of the masked / patchified first input data 311, the first encoder 320 determines masked tokens. The first encoder output 321 includes features of the unmasked patches as well as masked tokens corresponding to the masked patches of the masked / patchified first input data 311.
[0053] In some examples, the second encoder 330 includes an intensity encoder. The masked / patchified second input data 312 is applied to the second encoder 330 as an input to determine the second encoder output 331. Examples of the second encoder output 331 includes one or more of an intensity encoder output correspondingly generated by the second encoder 330 for the input intensity image, an RGB encoder output corresponding generated by the second encoder 330 for the input RGB image, a heatmap encoder output correspondingly generated by the second encoder 330 for the input heatmap image, or a wireless signal encoder output correspondingly generated by the second encoder 330 for the input wireless signal image, and so on.
[0054] The second encoder output 311 can be at least one two-dimensional feature or a two-dimensional pixel feature. For example, the second encoder output 331 includes at least one feature (e.g., one-dimensional feature), at least one tensor, at least one value, or so on for each patch of the masked / patchified second input data 312. For any masked patches of the masked / patchified second input data 312, the second encoder 330 determines masked tokens. The second encoder output 331 includes both features of the unmasked patches as well as masked tokens corresponding to the masked patches of the masked / patchified second input data 312.
[0055] A first loss 325 can be determined between the first encoder output 321 and the second encoder output 331. The first loss 325 (as a part of a final loss or by itself) is intended to train the first encoder 320 and the second encoder 330. The first loss 325 includes contrastive loss between the first encoder output 321 and the second encoder output 331 such as a pixel- to-depth knowledge transfer loss and a modality invariance loss. The features for the different patches of the first encoder output 321 and the features for the different patches of the second encoder output 331 can be pooled. The first loss 325 can be determined for two corresponding patches (e.g., patches having the same location in the respective images of the first modality input data and the second modality input data).
[0056] Taking into consideration the contrastive loss minimizes the relative distance between the representations of the first encoder output 321 (e.g., depth representation, 3D representation, etc.) and the second encoder output 331 (e.g., intensity representation, pixel representation, RGB representation, 2D representation, etc.). The loss function of the contrastive loss is to form a feature space by attracting a feature (e.g., depth feature) of the first encoder output 321 and its corresponding feature (e.g., a two-dimensional feature) of the second encoder output 321 while separating the feature of the first encoder output 321 from other features (e.g., other two-dimensional features) of the second encoder output 331 at the same time. For example, if a depth value and an intensity or RGB value share the same coordinate or location within a frame, they are a positive pairs. On the other hand, if a depth value and an intensity or RGB value do not share a same coordinate within a same frame, they are a negative pair. In some examples in which the first input data 301 is depth data and the second input data 302 is intensity data, features corresponding to the first input data 301 and features corresponding to the second input data 302 are separately computed using the encoders 320 and 330 respectively, followed by a global average pooling layer to pool the computed features (e.g., the first encoder output 321 and the second encoder output 331), and contrastive loss is determined on those pooled features. The contrastive loss £ccan be determined using expression (1) below:where r is the temperature, z\ is the feature of the second encoder output 331 for a feature vector z, and zdis the feature of the first encoder output 321 for a feature vector i.
[0057] The first encoder output 321 and the second encoder output 331 are applied as inputs into the multi-modal decoder 340, which outputs the decoder output 345 in response. In some examples, the decoder output 345 is of the first modality and includes patches of an output depth image.
[0058] The decoder output 345 can be combined (e.g., assembled or unpatchified) to generate the unpatchified output 355, which is the output depth image. For example, thepatches of the decoder output 345 can be combined according to the respective locations mapped to the patches of the decoder output 345.
[0059] A second loss 360 can be determined between the first input data 301 and the unpatchified output 355, between the second input data 302 and the unpatchified output 355, and / or for the unpatchified output 355. The second loss 360 (as a part of a final loss or by itself) is intended to train the multi-modal decoder 340. The second loss 360 includes a reconstruction loss, depth regularization loss, and an occlusion regularization term. The depth regularization loss enforces smoothness in low-textured regions. The occlusion regularization term minimizes the shadow areas generated in the decoder output 345 or the unpatchified output 355.
[0060] In some embodiments, the reconstruction loss (also referred to as a depth reconstruction loss) can be determined between the first input data 301 and the unpatchified output 355. For the first input data 301, which can include an input depth image of a depth modality, can be referred to as D. The input depth image has a size defined by Q x H x W, where H is the height of the input depth image, W is the width of the input depth image, and Q (e.g., 1) corresponds to the number of channels in the input depth image. The reconstruction loss -Cdepthcanbe determined using expression (2) below:where m is a token index, is the set of masked tokens. In addition, D corresponds to the reconstruction or the prediction of the machine learning model, which is the unpatchified output 355.
[0061] In some embodiments, the depth regularization loss can be determined between the second input data 302 and the unpatchified output 355. The depth regularization loss is used to address the local blurring of the borders of the unpatchified output 355 (e.g., the output depth image). To smooth out prediction discontinuities and preserve sharp details, edge-aware smoothness regularization term is applied to enforce smoothness in depths in order to regularize the disparities in texture-less low-image gradient regions. The depth regularization loss ■^smoothcanbe determined using expression (3) below:where & and dyare the I gradient of the second input data 302 (e.g., an input intensity image) along horizontal and vertical axes of the image, respectively. D corresponds to the reconstruction or the prediction of the machine learning model, which is the unpatchified output 355 (e.g., the output depth image). In some examples, 3x and dy are the gradient along horizontal and vertical axes respectively while D is the prediction. The combination of the 3x and dy refers to the gradients of depth predicted map D along x-axis and y-axis, respectively. In some examples, cyl and ex I are the gradients for the input intensity image. In some examples, the functions 3x, dy, cyl and 3x1 are used as evaluation metrics to update the model.
[0062] In some embodiments, the occlusion regularization term can be determined between the first input data 301 and the unpatchified output 355. In some cases, application of ■^smooth may generate a shadow area where values gradually change from foreground to background due to occlusion. In some examples, the occlusion regularization term can be applied to minimize the shadow areas generated in the depth map, especially across high gradient disparity regions. Some implementations that introduce an LI loss function (e.g., least absolute deviations to minimize error determined based on sum of all absolute differences between truth and reconstruction / prediction) over the reconstruction / prediction can be improved in terms of background depths by penalizing the total sum of absolute depth in the reconstruction / prediction. The occlusion regularization term Looc( / ) for the / gradient can be determined using expression (4) below:-Cooc(0 = |D| (4), where D corresponds to the reconstruction or the prediction of the machine learning model, which is the unpatchified output 355 (e.g., the output depth image).
[0063] In some embodiments, the machine learning model including the first encoder 320, the second encoder 330, and the decoder 340 can be trained using a final loss -Ctotai, which is the final pre-training loss of the three-dimensional point cloud enhancement method performed by the training system 300. The final loss Ltotaican be determined using expression (5) below:where a, >, y and / are the hyper-parameters tuned during pre-training. Thus, the final loss is the combination of the first loss 325 and the second loss 360. The machine learning model can thus be updated using unsupervised learning (e.g., the loss functions), without labels.
[0064] For example, the training system 300 can minimize the final loss by minimizing at least one of a mean absolute error (MAE), root mean squared error (RMSE), mean absolute error of an inverse depth (iMAE), and root mean squared error of an inverse depth (iRMSA). For example, MAE can be defined using expression (6) below:For example, RMSE can be defined using expression (7) below:For example, iMAE can be defined using expression (8) below:For example, iRMSA can be defined using expression (9) below:
[0065] In some examples, the machine learning model including the first encoder 320, the second encoder 330, and the decoder 340 can be trained using the final loss in a one-stage training method. In some examples, the machine learning model including the first encoder 320, the second encoder 330, and the decoder 340 can be pre-trained in two stages. In a first stage, the encoders 320 and 330 can be trained with the contrastive loss without any decoder and masking. The result of the first stage is that the encoders 320 and 330 are defied by stage- 1 weights. In the second stage, the encoder is initialized with stage- 1 weights, the decoder 340 is added to the training pipeline, average pool layer is dropped, and the pre-training processes using remaining losses (e.g., the second loss 360) only. In the second stage, the first loss 325is not considered as the first loss 325 is already considered in the first stage, which includes cross-modal learning.
[0066] FIG. 6 is a flowchart diagram illustrating an example method 600 for training a machine learning model to improve sparsity of a depth image captured in a medical environment, according to various embodiments. The method 600 can be performed by the training system which includes the first encoder 320, the second encoder 330, and the decoder 340. The training system 300 can implement or perform the method 600.
[0067] At 610, the training system receives first input data 301 of a first modality and second input data 302 of a second modality. At 620, the training system determines, using the first encoder 320, first encoder output 321 with the first input data 301 applied as an input into the first encoder 320. At 630, the training system determines, using the second encoder 330, second encoder output 331 using the second input data 302 applied as an input into the second encoder 330. At 640, the training system determines a first loss 325 between the first encoder output 321 and the second encoder output 331.
[0068] At 650, the training system determines, using a multi-modal decoder 340, a decoder output 345 using the first encoder output 321 and the second encoder output 331 applied as inputs into the multi-modal decoder 340. The multi-modal decoder output 345 is of the first modality and includes a reconstruction of at least one of the first input data 301 and the second input data 302. At 660, the training system determines a second loss 360 for the decoder output 345. At 670, the training system update a machine learning model using the first loss and the second loss. In some examples, the training system updating (training) the machine learning model includes updating the weights, parameters, coefficients, constants, and other information in one or more of the first encoder 320, the second encoder 330, and the multi-modal decoder 340.
[0069] In some examples, the first loss includes contrastive loss, an example of which is defined using expression (1). In some examples, the second loss includes decoder loss. In some examples, the second loss or the decoder loss includes at least one of a reconstruction loss between the first input data and the decoder output (e.g., defined using expression (2)), a depth regularization loss between the second input data and the decoder output (e.g., definedusing expression (3)), or an occlusion regularization term for the decoder output (e.g., defined using expression (4)). In some examples, the machine learning model is updated using a combination of the first loss and the second loss, for example, according to expression (5)
[0070] In some examples, determining using the first encoder 320 the first encoder output 321 with the first input data 301 applied as the input into the first encoder 320 includes generating first masked data (e.g., the masked / patchified first input data 311) based on the first input data 301 and applying the first masked data as the input to the first encoder 320. In some examples, determining using the second encoder 330 the second encoder output 331 with the second input data 302 applied as the input into the second encoder 330 includes generating second masked data (e.g., the masked / patchified second input data 312) based on the second input data 302 and applying the second masked data as the input to the second encoder 330.
[0071] In some examples, determining the decoder output 345 using the first encoder output 321 and the second encoder output 331 as the inputs includes applying a first set of masked tokens representing masked portions of the first input data 301 and a second set of masked tokens representing masked portions of the second input data 302 as the inputs to the multi-modal decoder 340 to generate the decoder output 345.
[0072] In some examples, the training system segments the first input data 301 into a plurality of first patches and masks one or more patches of the plurality of first patches, for example, at 310a. In some examples, the training system segments the second input data 302 into a plurality of second patches and masks one or more patches of the plurality of second patches, for example, at 310b.
[0073] In some examples, locations of the one or more masked patches of the plurality of first patches and locations the one or more masked patches of the plurality of second patches are same. That is, a patch for the first input data 301 and a patch for the second input data 302 with the same location are masked. In some examples, a number of the one or more masked patches of the plurality of first patches and a number of the one or more masked patches of the plurality of second patches are same. In some examples, a size or shape of the one or more masked patches of the plurality of first patches and a size or shape of the one or more masked patches of the plurality of second patches are same.
[0074] In some examples, a location of a masked patch of the one or more masked patches of the plurality of first patches and a location of any of the one or more masked patches of the plurality of second patches are different. In some examples, a masked patch of the first input data 301 and an unmasked patch of the second input data 302 have a same location. In some examples, a number of the one or more masked patches of the plurality of first patches and a number of the one or more masked patches of the plurality of second patches are different. In some examples, a size or shape of the one or more masked patches of the plurality of first patches and a size or shape of the one or more masked patches of the plurality of second patches are different.
[0075] In some examples, locations the one or more masked patches of the plurality of first patches and locations the one or more masked patches of the plurality of second patches are exclusive. That is, patch locations that are masked for the first input data 301 are all unmasked for the second input data 302, and patch locations that are masked for the second input data 302 are all unmasked for the first input data 301. In some examples, a location of each of the one or more masked patches of the plurality of first patches is different from any location of any of the one or more masked patches of the plurality of second patches.
[0076] In some examples, locations of the one or more masked patches of the plurality of first patches are random. In some examples, locations of the one or more masked patches of the plurality of second patches are random. In some examples, a number of the one or more masked patches of the plurality of first patches is random. In some examples, a number of the one or more masked patches of the plurality of second patches are random. In some examples, a sampling rate for sampling the one or more masked patches of the plurality of first patches is same as a sampling rate for sampling the one or more masked patches of the plurality of second patches. In some examples, a sampling rate for sampling the one or more masked patches of the plurality of first patches is different from a sampling rate for sampling the one or more masked patches of the plurality of second patches. Greater sampling rate yields a greater number of masked patches.
[0077] In some examples, a ratio of a number of the one or more masked patch of the plurality of first patches to a number of unmasked patches of the plurality of first patches is between 75%-90% inclusive. In some examples, a ratio of a number of the one or more maskedpatch of the plurality of second patches to a number of unmasked patches of the plurality of second patches is between 75%-90% inclusive.
[0078] In some examples, the training system segments the first input data 301 into a plurality of first patches and determines that a quality of a first patch of the plurality of first patches is below a threshold. A quality of a patch of an input depth image can be evaluated using an average (mean) depth value for each location or pixel of the patch, a sum of all depth values for all locations or pixels of the patch, sparsity of the patch, and so on. In response, the training system masks the first patch of the plurality of first patches. In some examples, the training system segments the second input data into a plurality of second patches and masks a second patch of the plurality of second patches. A location of the first patch can be different from a location of any second patch.
[0079] In some examples, the training system segments the first input data 301 into a plurality of first patches and determines that a quality of a first patch of the plurality of first patches is above a threshold. The training system masks the first patch of the plurality of first patches. In some examples, the training system segments the second input data into a plurality of second patches and masks a second patch of the plurality of second patches. A location of the first patch is same as a location of the second patch.
[0080] The method 600 can be rerun in multiple iterations for multiple frames of the first input data 301 and the second input data 302. That is, after 670 is performed, the method 600 returns to 610 with one or more of the same first input data 301, different first input data 301 of the same modality, different first input data 301 of different modalities, different methods for processing the first input data 301, the same second input data 302, different second input data 302of the same modality, different second input data 302 of different modalities, or different methods for processing the second input data 302.
[0081] In some iterations, the first input data 301 and the second input data 302 are time synchronized. In some iterations, the first input data 301 and the second input data 302 are asynchronous in the time domain. In some iterations, the first input data 301 and the second input data 302 have the same sensor pose. In some iterations, the first input data 301 and the second input data 302 have different sensor poses. In different iterations, a same typeof the second input data 302 or modality thereof (e.g., intensity modality, vision modality, a heatmap modality, or wireless signal modality) can be used. In different iterations, different types of the second input data 302 or modalities thereof (e.g., different ones of intensity modality, vision modality, a heatmap modality, or wireless signal modality) can be used.
[0082] With respect to masking, in some examples, locations of the one or more masked patches of the plurality of first patches in a first iteration of updating the machine learning model is same as locations of the one or more masked patches of the plurality of first patches in a second iteration of updating the machine learning model. In some examples, locations of the one or more masked patches of the plurality of second patches in the first iteration of updating the machine learning model is same as locations of the one or more masked patches of the plurality of second patches in the second iteration of updating the machine learning model.
[0083] In some examples, a number of the one or more masked patches of the plurality of first patches in a first iteration of updating the machine learning model is different from a number of the one or more masked patches of the plurality of first patches in a second iteration of updating the machine learning model. In some examples, a number of the one or more masked patches of the plurality of second patches in the first iteration of updating the machine learning model is different from a number of the one or more masked patches of the plurality of second patches in the second iteration of updating the machine learning model.
[0084] In some examples, a size or shape of each of the one or more masked patches of the plurality of first patches in a first iteration of updating the machine learning model is same as a size or shape of each of the one or more masked patches of the plurality of first patches in a second iteration of updating the machine learning model. In some examples, a size or shape of each of the one or more masked patches of the plurality of second patches in the first iteration of updating the machine learning model is same as a size or shape of each of the one or more masked patches of the plurality of second patches in the second iteration of updating the machine learning model.
[0085] FIG. 7 is a flowchart diagram illustrating an example method 700 for training a machine learning model to improve sparsity of a depth image captured for a medical procedure,according to various embodiments. The method 700 can be performed by the training system which includes the first encoder 320, the second encoder 330, and the decoder 340. The training system 300 can implement or perform the method 700.
[0086] At 710, the training system receives multi-modal data. The multi-modal data includes at least a first input data 301 of a first modality and a second input data 302 of a second modality. At 720, the training system generate, using a first encoder 320, a first encoder output 321 using the first input data 301 as input to the first encoder 320. At 730, the training system generates, using a second decoder 330, a second encoder output 331 using the second input data 302 as input to the second encoder 330. At 740, the training system generates, using a multi-modal decoder 340, a decoder output 345 using the first encoder output 321 and the second encoder output 331 as inputs to the multi-modal decoder 340. The decoder output 345 is of the first modality. At 750, the training system updates a machine-learning model based at least in part on the decoder output 345. In response to performing 750, the method 700 returns to 710.
[0087] In some examples, the first modality includes a three-dimensional depth modality or a three-dimensional point cloud modality. In some examples, the second modality comprises one or more of an intensity modality, a vision modality, a heatmap modality, or a wireless signal sensing modality. In some examples, updating the machine-learning model includes updating at least one of the first encoder 320, the second encoder 330, or the decoder 340.
[0088] In some examples, updating the machine-learning model based at least in part on the decoder output includes determining a first loss 325 between the first encoder output 321 and the second encoder output 331, determining a second loss 36 between the first input data 301 and the decoder output 345 or 355, between the second input data 302 and the decoder output 345 or 355, or for the decoder output 345 or 355. The training system updates the machine-learning model based on the first loss and the second loss. In some examples, the first loss includes contrastive loss and the second loss includes decoder loss.
[0089] In some examples, generating the first encoder output using the first input data includes generating first masked data based on the first input data and providing the firstmasked data as the input to the first encoder. In some examples, generating the second encoder output using the second input data comprises generating second masked data based on the second input data and providing the second masked data as the input to the second encoder.
[0090] In some examples, generating the decoder output using the first encoder output and the second encoder output as the inputs includes applying a first set of masked tokens representing masked portions of the first input data and a second set of masked tokens representing masked portions of the second input data as the inputs to the multi-modal decoder 340 to generate the decoder output. In some examples, the decoder output includes reconstructed portions of the first input data corresponding to the first set of masked tokens.
[0091] In some examples, the training system segments the first input data into a plurality of first patches, masks at least one patch of the plurality of patches, segments the second input data into a plurality of second patches, and masks at least one patch of the plurality of second patches.
[0092] FIG. 8 is a flowchart diagram illustrating an example method 800 for training a machine learning model to improve sparsity of a depth image captured for a medical procedure, according to various embodiments. The method 800 can be performed by the training system which includes the first encoder 320, the second encoder 330, and the decoder 340. The method 300 can implement or perform the method 800.
[0093] At 810, a multi-modal decoder 340 receives a first input (e.g., the first encoder output 321) and a second input (e.g., the second encoder output 331). The first input is determined based on first input data 301 of a first modality. The second input is determined based on second input data 302 of a second modality. At 820, the multi-modal decoder 340 generates a decoder output 345 by applying the first input and the second input as inputs to the multi-modal decoder 340. The decoder output 345 is of the first modality. At 830, the training system updates a machine-learning model based at least in part on the decoder output. In response to performing 830, the method 800 returns to 810.
[0094] In some examples, the first input includes at least one first masked tokens and at least one portion of the first input data 301. Each of the at least one first masked token corresponds to a masked portion of the first input data 301. Each of the at least one portion ofthe first input data corresponds to an unmasked portion of the first input data. In some examples, the second input includes at least one second masked tokens and at least one portion of the second input data. Each of the at least one second masked token corresponds to a masked portion of the second input data. Each of the at least one portion of the second input data corresponds to an unmasked portion of the second input data. In some examples, the at least one first masked token is generated by a first encoder 320 of the first modality based on the first input data 301. The at least one second masked token is generated by a second encoder 331 of the second modality based on the second input data 302.
[0095] FIG. 9 is a flowchart diagram illustrating an example method 900 for using a machine learning model to improve sparsity of a depth image captured during a medical procedure, according to various embodiments. The method 900 can be performed by a deployment system that may reside within the computer system 190a and 190b, which include the first encoder 320, the second encoder 330, and the decoder 340. That is, the machine learning model including the first encoder 320, the second encoder 330, and the decoder 340, as trained using the training system, can be downloaded, installed, or otherwise provided to the deployment system. The method 900 is similar to the training methods described herein, except that no loss is determined to update the machine learning models. In other deployment scenarios, however, the deployment method can be the same as a training method as losses are continuously being calculated, and the machine learning model is continuously updated for the downstream deployment task, during deployment.
[0096] At 910, the deployment system receives first input data 301 of a first modality and second input data 302 of a second modality. At 920, the deployment system determines, using a first encoder 320, first encoder output 321 with the first input data 301 applied as an input into the first encoder 320. At 930, the deployment system determines using a second encoder 330, second encoder output 331 using the second input data 302 applied as an input into the second encoder 330. At 940, the deployment system determines, using a multi-modal decoder 340, a decoder output 345 using the first encoder output 321 and the second encoder output 331 applied as inputs into the multi-modal decoder 340. The multi-modal decoder output 345 is of the first modality. The decoder output (e.g., the unpatchified version thereof, which is the unpatchified output 355) has a density that is greater than the first input data (e.g.,a sparsity that is less than the first input data). Density refers to a number of subdivisions or units (e.g., pixels) of a frame of data that has useful values (e.g., values above a threshold, values below a threshold, values within a range, or so on). Sparsity refers to a number of subdivisions or units of a frame of data that has null or useless values (e.g., values above a threshold, values below a threshold, values within a range). Null and useless values can be results of reflection from highly reflective objects.
[0097] FIG. 10 is a flowchart diagram illustrating an example method 1000 for using a machine learning model to improve sparsity of a depth image captured during a medical procedure, according to various embodiments. The method 1000 can be performed by a deployment system that may reside within the computer system 190a and 190b, which include the first encoder 320, the second encoder 330, and the decoder 340, as described. The method 1000 is similar to the training methods described herein, except that no loss is determined to update the machine learning models. In other deployment scenarios, however, the deployment method can be the same as a training method as losses are continuously being calculated, and the machine learning model is continuously updated for the downstream deployment task, during deployment.
[0098] At 1010, the deployment system receives first input data 301 of a first modality and second input data 302 of a second modality. The first input data 301 and the second input data 302 are outputted from at least one sensor. At 1020, the deployment system generates first encoder output 321 by applying a plurality of patches of the first input data 301 to a first encoder 320. At least one of the plurality of patches of the first input data 301 is masked. At 1030, the deployment system generates second encoder output 331 by applying a plurality of patches of the second input data 302 to a second encoder 330. At least one of the plurality of patches of the second input data 302 is masked. At 1040, the deployment system generates a plurality of decoder output patches of the first modality by applying the first encoder output 321 and the second encoder output 331 to a multi-modal decoder 340. At 1050, the deployment system generates output data (e.g., the unpatchified output 355) of the first modality by combining the plurality of decoder output patches of the decoder output 345 via the unpatchified function 350. The output data 355 has greater density than the first input data
[0099] In some examples, the at least one sensor includes a first sensor configured to output the first input data 301 and a second sensor configured to output the second input data 302. In some examples, the at least one sensor includes a multi-modal sensor configured to output both the first input data 301 and the second input data 302.
[0100] FIG. 11 is a flowchart diagram illustrating an example method 1100 for using a machine learning model to improve sparsity of a depth image captured during a medical procedure, according to various embodiments. The method 10100 can be performed by a deployment system that may reside within the computer system 190a and 190b, which include the first encoder 320, the second encoder 330, and the decoder 340, as described. The method 1100 is similar to the training methods described herein, except that no loss is determined to update the machine learning models. In other deployment scenarios, however, the deployment method can be the same as a training method as losses are continuously being calculated, and the machine learning model is continuously updated for the downstream deployment task, during deployment.
[0101] At 1110, the deployment system receives first input data 301 of a first modality and second input data 302 of a second modality. The first input data 301 and the second input data 302 are outputted from at least one sensor. At 1120, the deployment system masks one or more patches of a plurality of first patches of the first input data 301 based on a quality for each of the plurality of first patches. For example, the deployment system segments the first input data 301 into a plurality of first patches and determines that a quality of a first patch of the plurality of first patches is below a threshold. A quality of a patch of the first input data 301 (e.g., an input depth image) can be evaluated using an average (mean) depth value for each location or pixel of the patch, a sum of all depth values for all locations or pixels of the patch, sparsity of the patch, and so on. In response, the deployment system masks the patch of the plurality of first patches. At 1130, the deployment system generates first encoder output 321 by applying the plurality of first patches of the first input data 301 to a first encoder 320.
[0102] At 1140, the deployment system masks one or more patches of a plurality of second patches of the second input data 302 based on at least one of a quality for each of the plurality of second patches or the one or more masked patches of the plurality of first patches. For example, the deployment system segments the second input data 302 into a plurality ofsecond patches and determines that a quality of a second patch of the plurality of second patches is below a threshold. A quality of a patch of the second input data 302 can be evaluated using an average (mean) pixel value for each location or pixel of the patch, a sum of all pixel values for all locations or pixels of the patch, sparsity of the patch, and so on. In response, the deployment system masks the patch of the plurality of second patches. At 1150, the deployment system generates second encoder output 331 by applying the plurality of patches of the second input data 302 to a second encoder 330. At 1160, the deployment system generates a plurality of decoder output patches (e.g., the decoder output 345) of the first modality by applying the first encoder output 321 and the second encoder output 331 to a multimodal decoder 340. At 1170, the deployment system generates output data (e.g., the unpatchified output 355) of the first modality by combining the plurality of decoder output patches. The output data has greater density than the first input data.
[0103] In some examples, the deployment system segments the first input data 301 into the plurality of first patches and determines that a quality of a first patch of the plurality of first patches is below a threshold. In response, the deployment system masks at least a portion of the first patch of the plurality of first patches. In some examples, the deployment system segments the second input data 302 into a plurality of second patches and masks a second patch of the plurality of second patches. A location of the first patch is different from a location of any second patch. The patch in the second input data 302 having the same location as the location of the masked first patch is unmasked to provide information about the masked first patch of the first input data 301. A quality of a patch of the first input data 301 (e.g., an input depth image) can be evaluated using an average (mean) depth value for each location or pixel of the patch, a sum of all depth values for all locations or pixels of the patch, sparsity of the patch, and so on.
[0104] In some examples, locations the one or more masked patches of the first input data 301 and locations the one or more masked patches of the second input data 302 are exclusive. That is, patch locations that are masked for the first input data 301 are all unmasked for the second input data 302, and patch locations that are masked for the second input data 302 are all unmasked for the first input data 301.Computer System
[0105] FIG. 12 is a block diagram of an example computer system 1200 as may be used in conjunction with some of the embodiments. The training system that performs the methods 120, 600, 700, 800 and the deployment system that performs the methods 900, 1000, and 1100 can be implemented using the computer system 1200. For example, the first encoder 320, the second encoder 330, and the multi-modal decoder 340 can be implemented using the same computer system 1200 or different computer systems each of which can be the computer system 1200. The computing system 1200 may include an interconnect 1205, connecting several components, such as, e.g., one or more processors 1210, one or more memory components 1215, one or more input / output systems 1220, one or more storage systems 1225, one or more network adaptors 1230, etc. The interconnect 1205 may be, e.g., one or more bridges, traces, busses (e.g., an ISA, SCSI, PCI, I2C, Firewire bus, etc.), wires, adapters, or controllers.
[0106] The one or more processors 1210 may include, e.g., an general-purpose processor (e.g., an x86 processor, an RISC processor, etc.), a math coprocessor, a graphics processor, etc. The one or more memory components 1215 may include, e.g., a volatile memory (RAM, SRAM, DRAM, etc.), a non-volatile memory (EPROM, ROM, Flash memory, etc.), or similar devices. The one or more input / output devices 1220 may include, e.g., display devices, keyboards, pointing devices, touchscreen devices, etc. The one or more storage devices 1225 may include, e.g., cloud-based storages, removable Universal Serial Bus (USB) storage, disk drives, etc. In some systems memory components 1215 and storage devices 1225 may be the same components. Network adapters 1230 may include, e.g., wired network interfaces, wireless interfaces, Bluetooth™ adapters, line-of-sight interfaces, etc.
[0107] One will recognize that only some of the components, alternative components, or additional components than those depicted in FIG. 12 may be present in some embodiments. Similarly, the components may be combined or serve dual-purposes in some systems. The components may be implemented using special-purpose hardwired circuitry such as, for example, one or more ASICs, PLDs, FPGAs, etc. Thus, some embodiments may be implemented in, for example, programmable circuitry (e.g., one or more microprocessors) programmed with software and / or firmware, or entirely in special-purpose hardwired (nonprogrammable) circuitry, or in a combination of such forms.
[0108] In some embodiments, data structures and message structures may be stored or transmitted via a data transmission medium, e.g., a signal on a communications link, via the network adapters 1230. Transmission may occur across a variety of mediums, e.g., the Internet, a local area network, a wide area network, or a point-to-point dial-up connection, etc. Thus, “computer readable media” can include computer-readable storage media (e.g., “non- transitory” computer-readable media) and computer-readable transmission media.
[0109] The one or more memory components 1215 and one or more storage devices 1225 may be computer-readable storage media. In some embodiments, the one or more memory components 1215 or one or more storage devices 1225 may store instructions, which may perform or cause to be performed various of the operations discussed herein. In some embodiments, the instructions stored in memory 1215 can be implemented as software and / or firmware. These instructions may be used to perform operations on the one or more processors 1210 to carry out processes described herein. In some embodiments, such instructions may be provided to the one or more processors 1210 by downloading the instructions from another system, e.g., via network adapter 1230.
[0110] For clarity, one will appreciate that while a computer system may be a single machine, residing at a single location, having one or more of the components of FIG. 12, this need not be the case. For example, distributed network computer systems may include multiple individual processing workstations, each workstation having some, or all, of the components depicted in FIG. 12. Processing and various operations described herein may accordingly be spread across the one or more workstations of such a computer system. For example, one will appreciate that a process amenable to being run in a single thread upon a single workstation may instead be separated into an arbitrary number of sub-threads across one or more workstations, such sub-threads then run in serial or in parallel to achieve a same, or substantially similar, result as the process run within the single thread. Similarly, one will appreciate that while a non-transitory computer readable medium may stand alone (e.g., in a single USB storage device), or reside within a single workstation (e.g., in the workstation’s random access memory or disk storage), such a medium need not reside at a single geographic location, but may include, e.g., multiple memory storage units residing across geographicallyseparated workstations of a computer system in network communication with one another or across geographically separated storage devices.Remarks[OHl] The drawings and description herein are illustrative. Consequently, neither the description nor the drawings should be construed so as to limit the disclosure. For example, titles or subtitles have been provided simply for the reader’s convenience and to facilitate understanding. Thus, the titles or subtitles should not be construed so as to limit the scope of the disclosure, e.g., by grouping features which were presented in a particular order or together simply to facilitate understanding. Unless otherwise defined herein, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In the case of conflict, this document, including any definitions provided herein, will control. A recital of one or more synonyms herein does not exclude the use of other synonyms. The use of examples anywhere in this specification including examples of any term discussed herein is illustrative only and is not intended to further limit the scope and meaning of the disclosure or of any exemplified term.
[0112] Similarly, despite the particular presentation in the figures herein, one skilled in the art will appreciate that actual data structures used to store information may differ from what is shown. For example, the data structures may be organized in a different manner, may contain more or less information than shown, may be compressed and / or encrypted, etc. The drawings and disclosure may omit common or well-known details in order to avoid confusion. Similarly, the figures may depict a particular series of operations to facilitate understanding, which are simply exemplary of a wider class of such collection of operations. Accordingly, one will readily recognize that additional, alternative, or fewer operations may often be used to achieve the same purpose or effect depicted in some of the flow diagrams. For example, data may be encrypted, though not presented as such in the figures, items may be considered in different looping patterns (“for” loop, “while” loop, etc.), or sorted in a different manner, to achieve the same or similar effect, etc.
[0113] Reference herein to “an embodiment” or “one embodiment” means that at least one embodiment of the disclosure includes a particular feature, structure, or characteristicdescribed in connection with the embodiment. Thus, the phrase “in one embodiment” in various places herein is not necessarily referring to the same embodiment in each of those various places. Separate or alternative embodiments may not be mutually exclusive of other embodiments. One will recognize that various modifications may be made without deviating from the scope of the embodiments.
Claims
What is claimed is:
1. A system, comprising: one or more processors, coupled with memory, to: receive first input data of a first modality and second input data of a second modality; determine, using a first encoder, first encoder output with the first input data applied as an input into the first encoder; determine, using a second encoder, second encoder output using the second input data applied as an input into the second encoder; determine a first loss between the first encoder output and the second encoder output; determine, using a multi-modal decoder, a decoder output using the first encoder output and the second encoder output applied as inputs into the multi-modal decoder, wherein the decoder output is of the first modality; determine a second loss for the decoder output; and update a machine learning model using the first loss and the second loss.
2. The system of claim 1, wherein, the first input data comprises an input depth image; the second input data comprise at least one of an input intensity image, an input RGB image, or a heatmap image; the decoder output comprises an output depth image.
3. The system of claim 1, wherein the first input data comprises an input depth image; the first modality comprises a three-dimensional depth modality or a three-dimensional point cloud modality; the first encoder output comprises depth encoder output.
4. The system of claim 1, wherein the second input data comprise an input intensity image; the second modality comprises an intensity modality; andthe second encoder output comprises intensity encoder output;5. The system of claim 1, wherein the second input data comprise an input RGB image; the second modality comprises a vision modality; and the second encoder output comprises RGB encoder output.
6. The system of claim 1, wherein the second input data comprise an input heatmap image; the second modality comprises a heatmap modality; and the second encoder output comprises heatmap encoder output.
7. The system of claim 1, wherein the second input data comprise an input wireless signal sensed image; the second modality comprises a wireless signal modality; and the second encoder output comprises wireless signal encoder output.
8. The system of claim 1, wherein updating the machine learning model comprises updating at least one of the first encoder, the second encoder, or the multi-modal decoder.
9. The system of claim 1, wherein the first loss comprises contrastive loss; and the second loss comprises decoder loss of the decoder output.
10. The system of claim 1, wherein the second loss comprises at least one of: a reconstruction loss between the first input data and the decoder output; a depth regularization loss between the second input data and the decoder output; or an occlusion regularization term for the decoder output.
11. The system of claim 1, wherein the machine learning model is updated using a combination of the first loss and the second loss.
12. The system of claim 1, wherein: determining using the first encoder the first encoder output with the first input data applied as the input into the first encoder comprises: generating first masked data based on the first input data; and applying the first masked data as the input to the first encoder; and determining using the second encoder the second encoder output with the second input data applied as the input into the second encoder comprises: generating second masked data based on the second input data; and applying the second masked data as the input to the second encoder.
13. The system of claim 1, wherein determining the decoder output using the first encoder output and the second encoder output as the inputs comprises applying a first set of masked tokens representing masked portions of the first input data and a second set of masked tokens representing masked portions of the second input data as the inputs to the multi-modal decoder to generate the decoder output.
14. The system of claim 1, further comprising: segmenting the first input data into a plurality of first patches; masking one or more patches of the plurality of first patches; segmenting the second input data into a plurality of second patches; and masking one or more patch of the plurality of second patches.
15. The system of claim 14, wherein at least one of: locations of the one or more masked patches of the plurality of first patches and locations the one or more masked patches of the plurality of second patches are same; a number of the one or more masked patches of the plurality of first patches and a number of the one or more masked patches of the plurality of second patches are same; or a size or shape of the one or more masked patches of the plurality of first patches and a size or shape of the one or more masked patches of the plurality of second patches are same.
16. The system of claim 14, wherein at least one of: a location of a masked patch of the one or more masked patches of the plurality of first patches and a location of an unmasked patch of the one or more masked patches of the plurality of second patches are same; a number of the one or more masked patches of the plurality of first patches and a number of the one or more masked patches of the plurality of second patches are different; or a size or shape of the one or more masked patches of the plurality of first patches and a size or shape of the one or more masked patches of the plurality of second patches are different.
17. The system of claim 14, wherein locations the one or more masked patches of the plurality of first patches and locations the one or more masked patches of the plurality of second patches are exclusive.
18. The system of claim 14, wherein a location of each of the one or more masked patches of the plurality of first patches is different from location of any of the one or more masked patches of the plurality of second patches.
19. The system of claim 14, wherein at least one of: locations of the one or more masked patches of the plurality of first patches are random; locations of the one or more masked patches of the plurality of second patches are random; a number of the one or more masked patches of the plurality of first patches is random; or a number of the one or more masked patches of the plurality of second patches are random.
20. The system of claim 14, wherein a sampling rate for sampling the one or more masked patches of the plurality of first patches is same as a sampling rate for sampling the one or more masked patches of the plurality of second patches.
21. The system of claim 14, wherein a sampling rate for sampling the one or more masked patches of the plurality of first patches is different from a sampling rate for sampling the one or more masked patches of the plurality of second patches.
22. The system of claim 14, wherein at least one of: locations of the one or more masked patches of the plurality of first patches in a first iteration of updating the machine learning model is same as locations of the one or more masked patches of the plurality of first patches in a second iteration of updating the machine learning model; or locations of the one or more masked patches of the plurality of second patches in the first iteration of updating the machine learning model is same as locations of the one or more masked patches of the plurality of second patches in the second iteration of updating the machine learning model.
23. The system of claim 14, wherein at least one of: a number of the one or more masked patches of the plurality of first patches in a first iteration of updating the machine learning model is different from a number of the one or more masked patches of the plurality of first patches in a second iteration of updating the machine learning model; or a number of the one or more masked patches of the plurality of second patches in the first iteration of updating the machine learning model is different from a number of the one or more masked patches of the plurality of second patches in the second iteration of updating the machine learning model.
24. The system of claim 14, wherein at least one of: a size or shape of each of the one or more masked patches of the plurality of first patches in a first iteration of updating the machine learning model is same as a size or shapeof each of the one or more masked patches of the plurality of first patches in a second iteration of updating the machine learning model; or a size or shape of each of the one or more masked patches of the plurality of second patches in the first iteration of updating the machine learning model is same as a size or shape of each of the one or more masked patches of the plurality of second patches in the second iteration of updating the machine learning model.
25. The system of claim 14, wherein at least one of: a ratio of a number of the one or more masked patch of the plurality of first patches to a number of unmasked patches of the plurality of first patches is between 75%-90% inclusive; or a ratio of a number of the one or more masked patch of the plurality of second patches to a number of unmasked patches of the plurality of second patches is between 75%-90% inclusive.
26. The system of claim 1, further comprising: segmenting the first input data into a plurality of first patches; determining that a quality of a first patch of the plurality of first patches is below a threshold; masking the first patch of the plurality of first patches.
27. The system of claim 26, further comprising: segmenting the second input data into a plurality of second patches; and masking at least one second patch of the plurality of second patches, wherein a location of the first patch is different from a location of any of the at least one second patch.
28. The system of claim 1, further comprising: segmenting the first input data into a plurality of first patches; determining that a quality of a first patch of the plurality of first patches is above a threshold; masking the first patch of the plurality of first patches.
29. The system of claim 28, further comprising: segmenting the second input data into a plurality of second patches; and masking a second patch of the plurality of second patches, wherein a location of the first patch is same as a location of the second patch.
30. The system of claim 1, wherein the machine learning model is updated using unsupervised learning, without labels.
31. A non-transitory computer-readable medium comprising instructions configured to cause the one or more processors of the system of claim 1-30 to perform the operations of the one or more processors in claims 1-30.
32. A system comprising: one or more processors, coupled with memory, to: receive multi-modal data, the multi-modal data comprising at least a first input data of a first modality and a second input data of a second modality; generate, using a first encoder, a first encoder output using the first input data as input to the first encoder; generate, using a second encoder, a second encoder output using the second input data as input to the second encoder; generate, using a multi-modal decoder, a decoder output using the first encoder output and the second encoder output as inputs to the multi-modal decoder, wherein the decoder output is of the first modality; and update a machine-learning model based at least in part on the decoder output.
33. The system of claim 32, wherein the first modality comprises a three-dimensional depth modality or a three-dimensional point cloud modality.
34. The system of claim 32, wherein the second modality comprises one or more of an intensity modality, a vision modality, a heatmap modality, or a wireless signal sensing modality.
35. The system of claim 32, wherein updating the machine-learning model comprises updating at least one of the first encoder, the second encoder, or the multi-modal decoder.
36. The system of claim 32, wherein updating the machine-learning model based at least in part on the decoder output comprises: determining a first loss between the first encoder output and the second encoder output; determining a second loss between the first input data and the decoder output, between the second input data and the decoder output, or for the decoder output; updating the machine-learning model based on the first loss and the second loss.
37. The system of claim 36, wherein the first loss comprises contrastive loss; and the second loss comprises decoder loss.
38. The system of claim 32, wherein: generating the first encoder output using the first input data comprises: generating first masked data based on the first input data; and providing the first masked data as the input to the first encoder; and generating the second encoder output using the second input data comprises: generating second masked data based on the second input data; and providing the second masked data as the input to the second encoder.
39. The system of claim 32, wherein generating the decoder output using the first encoder output and the second encoder output as the inputs comprises applying a first set of masked tokens representing masked portions of the first input data and a second set of masked tokens representing masked portions of the second input data as the inputs to the multi-modal decoder to generate the decoder output.
40. The system of claim 39, wherein the decoder output comprises reconstructed portions of the first input data corresponding to the first set of masked tokens.
41. The system of claim 32, further comprising: segmenting the first input data into a plurality of first patches; masking at least one patch of the plurality of patches; segmenting the second input data into a plurality of second patches; and masking at least one patch of the plurality of second patches.
42. A non-transitory computer-readable medium comprising instructions configured to cause the one or more processors of the system of claims 32-41 to perform operations of the one or more processors in claims 32-41.
43. A system, comprising: one or more processors, coupled with memory, to: receive first input data of a first modality and second input data of a second modality; determine, using a first encoder, first encoder output with the first input data applied as an input into the first encoder; determine, using a second encoder, second encoder output using the second input data applied as an input into the second encoder; determine, using a multi-modal decoder, a decoder output using the first encoder output and the second encoder output applied as inputs into the multi-modal decoder, wherein the decoder output is of the first modality, and wherein the decoder output has a density that is greater than the first input data.
44. The system of claim 43, wherein, the first input data comprises an input depth image; the second input data comprise at least one of an input intensity image, an input RGB image, or a heatmap image; the decoder output comprises an output depth image.
45. The system of claim 43, wherein the first input data comprises an input depth image;the first modality comprises a three-dimensional depth modality or a three-dimensional point cloud modality; the first encoder output comprises depth encoder output.
46. The system of claim 43, wherein the second input data comprise an input intensity image; the second modality comprises an intensity modality; and the second encoder output comprises intensity encoder output;47. The system of claim 43, wherein the second input data comprise an input RGB image; the second modality comprises a vision modality; and the second encoder output comprises RGB encoder output.
48. The system of claim 43, wherein the second input data comprise an input heatmap image; the second modality comprises a heatmap modality; and the second encoder output comprises heatmap encoder output.
49. The system of claim 43, wherein the second input data comprise an input wireless signal sensed image; the second modality comprises a wireless signal modality; and the second encoder output comprises wireless signal encoder output.+50. A non-transitory computer-readable medium comprising instructions configured to cause the one or more processors of the system of claims 43-49 to perform operations of the one or more processors in claims 43-49.
51. A system, comprising: one or more processors, coupled with memory, to:receive first input data of a first modality and second input data of a second modality, wherein the first input data and the second input data are outputted from at least one sensor; generate first encoder output by applying a plurality of patches of the first input data to a first encoder, at least one of the plurality of patches of the first input data is masked; generate second encoder output by applying a plurality of patches of the second input data to a second encoder, at least one of the plurality of patches of the second input data is masked; generate a plurality of decoder output patches of the first modality by applying the first encoder output and the second encoder output to a multi-modal decoder; and generate output data of the first modality by combining the plurality of decoder output patches, wherein the output data has greater density than the first input data.
52. The system of claim 51, wherein the at least one sensor comprises one of a first sensor configured to output the first input data and a second sensor configured to output the second input data; a multi-modal sensor configured to output both the first input data and the second input data.
53. A non-transitory computer-readable medium comprising instructions configured to cause the one or more processors of the system of claims 51-52to perform operations of the one or more processors in claims 51-52.
54. A system, comprising: one or more processors, coupled with memory, to: receive first input data of a first modality and second input data of a second modality, wherein the first input data and the second input data are outputted from at least one sensor; masking one or more patches of a plurality of first patches of the first input data based on a quality for each of the plurality of first patches; generate first encoder output by applying the plurality of first patches of the first input data to a first encoder;masking one or more patches of a plurality of second patches of the second input data based on at least one of a quality for each of the plurality of second patches or the one or more masked patches of the plurality of first patches; generate second encoder output by applying the plurality of patches of the second input data to a second encoder; generate a plurality of decoder output patches of the first modality by applying the first encoder output and the second encoder output to a multi-modal decoder; and generate output data of the first modality by combining the plurality of decoder output patches, wherein the output data has greater density than the first input data.
55. The system of claim 54, further comprising: segmenting the first input data into the plurality of first patches; determining that a quality of a first patch of the plurality of first patches is below a threshold; masking the first patch of the plurality of first patches.
56. The system of claim 55, further comprising: segmenting the second input data into a plurality of second patches; and masking at least one second patch of the plurality of second patches, wherein a location of the first patch is different from a location of any of the at least one second patch.
57. A non-transitory computer-readable medium comprising instructions configured to cause the one or more processors of the system of claims 54-56 to perform operations of the one or more processors in claims 54-56.
Citation Information
Patent Citations
US202363595244P