Systems and methods for low-compute high-resolution depth map generation using low-resolution cameras
By using a low-resolution camera to generate a low-resolution depth map and combining it with high-resolution texture information, the cost and computational complexity issues brought about by high-resolution cameras in mixed reality systems are solved, achieving efficient parallax correction and pass-through experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MICROSOFT TECHNOLOGY LICENSING LLC
- Filing Date
- 2022-04-15
- Publication Date
- 2026-07-03
AI Technical Summary
In existing mixed reality systems, using high-resolution stereo cameras to generate depth maps has problems such as high cost, large device size, high battery consumption and computational complexity, resulting in delays in passthrough experiences.
Low-resolution depth maps are generated using a low-resolution camera, and high-resolution depth maps are generated through stereo matching and reprojection. Combined with high-resolution texture information, computational complexity is reduced and frame rate is increased.
It reduces device cost, size, and power consumption, while improving the efficiency of parallax correction and the frame rate of passthrough experience, providing high-quality parallax-corrected images.
Smart Images

Figure CN117203662B_ABST
Abstract
Description
Background Technology
[0001] Mixed reality (MR) systems, including virtual reality and augmented reality systems, have garnered significant attention for their ability to create truly unique experiences for users. For reference, conventional virtual reality (VR) systems create fully immersive experiences by restricting the user's view to a virtual environment only. In VR systems, this is often achieved using a head-mounted display (HMD) that completely blocks any view of the real world. Therefore, the user is fully immersed in the virtual environment. In contrast, conventional augmented reality (AR) systems create augmented reality experiences by visually presenting virtual objects that are placed in or interact with in the real world.
[0002] As used herein, VR and AR systems are described and referenced interchangeably. Unless otherwise stated, the descriptions herein also apply to all types of mixed reality systems (as detailed above), including AR systems, VR reality systems, and / or any other similar systems capable of displaying virtual objects.
[0003] Many mixed reality systems include depth reconstruction systems (e.g., time-of-flight cameras, rangefinders, stereo depth cameras, etc.). The depth reconstruction system provides depth information about the real-world environment surrounding the mixed reality system, enabling the system to accurately render mixed reality content (e.g., holograms) about real-world objects. As an illustrative example, the depth reconstruction system can obtain depth information about a real-world table located within a real-world environment. The mixed reality system can then render and display a virtual statue accurately positioned on the real-world table, allowing the user to perceive the virtual statue as part of the user's real-world environment.
[0004] Some mixed reality systems employ stereo cameras for depth detection or for other purposes besides depth detection. For example, a mixed reality system can use images obtained from a stereo camera to provide the user with a pass-through view of the user's environment. This pass-through view can help users avoid disorientation and / or safety hazards when transitioning to and / or navigating within an immersive mixed reality environment.
[0005] Some mixed reality systems are also equipped with cameras of different modalities to enhance the user's view in low-visibility environments. For example, mixed reality systems equipped with long-wavelength thermal imaging cameras improve visibility in smoke, haze, fog, and / or dust. Similarly, mixed reality systems equipped with low-light imaging cameras improve visibility in dark environments where ambient light levels are below those required for human vision. In some cases, low-light and thermal images can be fused or combined to provide the user with visualizations from multiple camera modalities simultaneously.
[0006] Mixed reality systems can present users with views captured by stereo cameras in a variety of ways. The process of providing users with a three-dimensional view of a real-world environment using images captured by a world-facing camera presents many challenges.
[0007] Initially, the physical location of the stereo camera is physically separate from the physical location of the user's eyes. Therefore, presenting the image captured by the stereo camera directly to the user's eyes will cause the user to perceive the real-world environment incorrectly. For example, a vertical offset between the user's eye location and the stereo camera's location will cause the user to perceive real-world objects as being vertically offset relative to the user from their actual position. In another example, the difference between the distance between the user's eyes and the distance between the stereo cameras will cause the user to perceive real-world objects with incorrect depth.
[0008] The perceptual difference between how a camera observes an object and how a user's eye observes an object is often referred to as the "parallax problem" or "parallax error". Figure 1 The illustration depicts a conceptual representation of the parallax problem, where stereo cameras 105A and 105B are physically separated from the user's eyes 110A and 110B. Sensor region 115A conceptually depicts the image sensing areas of camera 105A (e.g., a pixel grid) and the user's eye 110A (e.g., the retina). Similarly, sensor region 115B conceptually depicts the image sensing areas of camera 105B and the user's eye 110B.
[0009] Cameras 105A and 105B, and the user's eyes 110A and 110B, perceive object 130, such as in Figure 1 The lines extending from object 130 to cameras 105A and 105B and the user's eyes 110A and 110B are respectively indicated. Figure 1 The illustration shows cameras 105A and 105B sensing object 130 at different locations on their respective sensor regions 115A and 115B. Similarly, Figure 1 The illustration shows the user's eyes 110A and 110B perceiving the object 130 at different locations on their respective sensor regions 115A and 115B. Furthermore, the user's eye 110A perceives the object 130 at a different location on sensor region 115A than on camera 105A, and the user's eye 110B perceives the object 130 at a different location on sensor region 115B than on camera 105B.
[0010] Some schemes for correcting parallax involve performing a camera reprojection from the stereo camera's viewpoint to the user's eye's viewpoint. For example, some schemes involve performing a calibration step to determine the difference in physical positioning between the stereo camera and the user's eyes. Then, after capturing stereo image pairs using the stereo camera, a step of calculating depth information (e.g., a depth map) based on the stereo image pairs is performed (e.g., by performing stereo matching). The system is then able to reproject the calculated depth information to correspond to the user's left and right eye viewpoints.
[0011] However, computing depth information (e.g., depth maps) based on stereo image pairs (e.g., to address parallax problems) is associated with many challenges. For example, as mentioned above, stereo cameras are typically used to capture stereo image pairs. To provide an improved user experience, high-resolution stereo cameras are often used to capture stereo images for generating pass-through images. High-resolution stereo cameras are expensive and increase device size, weight, and battery consumption. Furthermore, mixed reality systems that provide pass-through imaging for multiple camera modalities (e.g., both low-light and thermal) typically require high-resolution cameras for each different camera modality, further increasing device cost, size, battery consumption, and weight. Additionally, computing depth information using high-resolution stereo image pairs is computationally expensive, which can cause latency in the delivered pass-through experience.
[0012] For at least the reasons mentioned above, there is a continued need and expectation for improved techniques and systems for generating high-resolution depth maps using low-resolution cameras.
[0013] The subject matter addressed herein is not limited to addressing any shortcomings or to embodiments that operate only in environments such as those described above. Rather, this background is provided merely to illustrate an exemplary technical field in which some of the embodiments described herein can be practiced. Summary of the Invention
[0014] The disclosed embodiments include systems and methods for generating low-computational-resolution depth maps using low-resolution cameras.
[0015] Some publicly disclosed systems are configured to acquire stereo image pairs of an environment and generate a depth map of the environment by performing stereo matching on the stereo image pairs. The depth map includes depth information for the environment. These systems are also configured to acquire a first image including first texture information for the environment. The first image has a first image resolution higher than the image resolution of the stereo image pair. These systems are also configured to generate a reprojected first image by reprojecting the first image to correspond to an image capture viewpoint associated with the depth map. The reprojection of the first image is based on depth information from the depth map, and the reprojected first image includes reprojected first texture information for the environment. These systems are also configured to generate an upsampled depth map based on the depth map and the reprojected first texture information.
[0016] Some publicly available systems are configured to acquire a first image of the environment and a second image of the environment. The second image captures the environment synchronously with the first image in time. The second image has a higher image resolution than the first image. Such systems are also configured to generate an upsampled first image. The upsampled first image has the same image resolution as the second image. The system is also configured to generate a depth map of the environment by performing stereo matching on the upsampled first image and the second image.
[0017] Some embodiments include systems configured to acquire a first image of an environment and a second image including texture information for the environment. The second image captures the environment synchronously with the first image in time. The second image has a higher image resolution than the first image. These systems are also configured to: generate a downsampled second image having the same image resolution as the first image; generate a depth map of the environment by performing stereo matching on the downsampled second image and the first image; and generate an upsampled depth map based on the depth map and the texture information of the second image.
[0018] This summary is provided to introduce a selection of concepts in a simplified form, which are further described in the detailed description below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to determine the scope of the claimed subject matter.
[0019] Additional features and advantages will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practice of the teachings herein. The features and advantages of the invention can be realized and obtained by means of the instruments and combinations particularly pointed out in the appended claims. The features of the invention will become more apparent from the description which follows and the appended claims, or may be learned by practice of the invention as described below. Attached Figure Description
[0020] To describe how the above and other advantages and features can be obtained, a more specific description of the subject matter briefly described above will be presented with reference to specific embodiments shown in the accompanying drawings. It is understood that these drawings depict only typical embodiments and therefore should not be considered limiting in scope. The embodiments will be described and explained with additional specificity and detail using the drawings, in which:
[0021] Figure 1 The illustration shows an example of parallax problems that occur when the camera has a different perspective than the user's eyes;
[0022] Figure 2 The illustrations depict exemplary systems that may include or be used to implement the disclosed embodiments;
[0023] Figure 3 The illustration shows an exemplary structural configuration of components of an exemplary mixed reality system and an example of parallax correction operations;
[0024] Figure 4 An exemplary head-mounted display (HMD) including various cameras that can facilitate the disclosed embodiments is illustrated, including a pair of low-resolution stereo cameras;
[0025] Figure 5A and Figure 5B The illustration shows examples of using various cameras to capture images of the environment using an HMD and generating low-resolution depth maps based at least on low-resolution stereo images;
[0026] Figure 6A and Figure 6B The illustration shows a conceptual representation of reprojecting a high-resolution image to correspond to the capture viewpoint of a low-resolution depth map;
[0027] Figure 7 An example of a high-resolution image reprojected and spatially aligned with a low-resolution depth map is illustrated.
[0028] Figure 8 An example is illustrated of upsampling a low-resolution depth map to generate an upsampled depth map 804 that is spatially aligned with the reprojected high-resolution image.
[0029] Figure 9The illustration illustrates a conceptual representation of reprojecting additional high-resolution images from different camera modalities to correspond to the capture viewpoint of a low-resolution depth map;
[0030] Figure 10 An example of upsampling a low-resolution depth map to generate an upsampled depth map 804 spatially aligned with an additional reprojection of a high-resolution image from a different camera modality is illustrated.
[0031] Figure 11 An alternative embodiment of an HMD, comprising a single low-resolution camera instead of a pair of stereo low-resolution cameras, is illustrated.
[0032] Figure 12 The illustration shows examples of using various cameras in an HMD, including a single low-resolution camera, to capture images of the environment;
[0033] Figure 13 The illustration shows a conceptual representation of upsampling a captured low-resolution image to generate a high-resolution depth map;
[0034] Figure 14 The illustrations depict conceptual representations of downsampling a captured high-resolution image to generate a low-resolution depth map and upsampling the low-resolution depth map; and
[0035] Figure 15-17 An exemplary flowchart depicts the actions associated with low computational depth map generation to provide parallax correction for an image. Detailed Implementation
[0036] The disclosed embodiments include systems and methods for facilitating the computationally low generation of high-resolution depth maps using low-resolution images.
[0037] Examples of technological benefits, improvements, and practical applications
[0038] In view of this disclosure, those skilled in the art will recognize that at least some of the embodiments disclosed can address various drawbacks associated with conventional schemes, devices, and / or techniques for calculating depth information (e.g., depth maps). The following sections outline some exemplary improvements and / or practical applications provided by the disclosed embodiments. However, it will be appreciated that the following are merely examples, and the embodiments described herein are by no means limited to the exemplary improvements discussed herein.
[0039] As described herein, a high-resolution parallax-corrected image can be generated by: generating a low-resolution depth map by performing depth calculations on a pair of low-resolution stereo images; generating a high-resolution depth map by upsampling the low-resolution depth map; and reprojecting the high-resolution image to correspond to the viewpoint associated with the high-resolution depth map. In this regard, an HMD can implement a pair of low-resolution stereo cameras for capturing low-resolution images (for generating low-resolution depth maps) and a single high-resolution camera for capturing high-resolution texture information to generate the parallax-corrected image (or a single high-resolution camera can be implemented for each desired camera modality (e.g., thermal, low-light, etc.), while omitting the stereo camera pair for each desired camera modality).
[0040] The HMD disclosed herein utilizes a stereo high-resolution camera pair to capture depth information to generate parallax-corrected images. Since low-resolution cameras are generally cheaper, smaller, lighter, and require less power than their corresponding high-resolution cameras, the implementation of this disclosure can reduce the cost, size, weight, and / or power consumption of devices that capture images to generate a parallax-corrected view of the captured environment. Furthermore, performing depth calculations on low-resolution images is often computationally more efficient than performing depth calculations on high-resolution images, allowing reprojection algorithms to run at higher frame rates while consuming less computational power.
[0041] In light of this disclosure, it will be appreciated that at least some of the principles described herein can enhance applications that rely on accurate depth maps, such as performing disparity error correction to provide disparity-corrected images (e.g., pass-through images). While this disclosure focuses in some aspects on depth map generation for performing disparity error correction, it should be noted that at least some of the principles described herein can be applied to other implementations involving and / or relying on depth map generation. By way of non-limiting examples, at least some of the principles disclosed herein can be used for hand tracking (or tracking other real-world objects), stereoscopic video streaming, building surface reconstruction meshes, and / or other applications.
[0042] Having just described some of the various advanced features and advantages of the disclosed embodiments, we will now focus on... Figures 2 to 17 These figures provide various conceptual representations, systems, architectures, methods, and supporting illustrations related to the disclosed embodiments.
[0043] Exemplary System
[0044] Follow us now Figure 2 The illustration depicts an exemplary system 200 that may include or be used to implement one or more of the disclosed embodiments. Figure 2 System 200 is described as a head-mounted display (HMD) configured to be placed on a user's head to display virtual content for the user's eyes to view. Such an HMD may include an augmented reality (AR) system, a virtual reality (VR) system, and / or any other type of HMD. Although this disclosure focuses at least in some aspects on system 200 implemented as an HMD, it should be noted that the principles described herein can be implemented using other types of systems.
[0045] Figure 2 Various exemplary components of system 200 are illustrated. For example, Figure 2 The diagram illustrates an implementation of a system including one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, and one or more communication systems 214. Although Figure 2 The illustration shows a system 200 including specific components, but in light of this disclosure, one will appreciate that system 200 may include any number of additional or alternative components.
[0046] One or more processors 202 may include one or more sets of electronic circuitry, including any number of logic units, registers, and / or control units, to facilitate the execution of computer-readable instructions (e.g., instructions forming a computer program). Such computer-readable instructions may be stored within storage device 204. Storage device 204 may include physical system memory and may be volatile, non-volatile, or a combination thereof. Furthermore, storage device 204 may include local storage, remote storage (e.g., accessible via one or more communication systems 214 or otherwise), or a combination thereof. Additional details relating to the processor (e.g., one or more processors 202) and computer storage media (e.g., storage device 204) are provided below.
[0047] In some implementations, processor(s)202 may include or be configured to execute any combination of software and / or hardware components that are operable to facilitate processing using machine learning models or other artificial intelligence-based architectures. For example, processor(s) 202 may include and / or utilize hardware components or computer-executable instructions to perform function blocks and / or processing layers, which may be configured, by non-limiting examples, in the form of single-layer neural networks, feedforward neural networks, radial basis function networks, deep feedforward networks, recurrent neural networks, long short-term memory (LSTM) networks, gated recurrent units, autoencoder neural networks, variational autoencoders, denoising autoencoders, sparse autoencoders, Markov chains, Hopfield neural networks, Boltzmann machine networks, restricted Boltzmann machine networks, deep belief networks, deep convolutional networks (or convolutional neural networks), deconvolutional neural networks, deep convolutional inverse graph networks, generative adversarial networks, liquid state machines, extreme learning machines, echo state networks, deep residual networks, Kohonen networks, support vector machines, neural Turing machines, and / or others.
[0048] As will be described in more detail, processor(s) 202 may be configured to execute instructions 206 stored within storage device 204 to perform specific actions associated with generating a high-resolution depth map using low-resolution images from a low-resolution camera. These actions may depend at least in part on data 208 stored on storage device 204 in a volatile or non-volatile manner.
[0049] In some instances, the operation may rely at least in part on one or more communication systems 214 for receiving data from one or more remote systems 216, which may include, for example, separate systems or computing devices, sensors, and / or others. The one or more communication systems 214 may include any combination of software or hardware components operable to facilitate communication between components / devices on the system and / or with components / devices outside the system. For example, the one or more communication systems 214 may include ports, buses, or other physical connection means for communicating with other devices / components. Additionally or alternatively, the one or more communication systems 214 may include systems / components operable to wirelessly communicate with external systems and / or devices via any suitable communication channel, such as Bluetooth, ultra-wideband, WLAN, infrared communication, and / or others by non-limiting examples.
[0050] Figure 2The system 200 illustrated may include or communicate with one or more sensors 210. The one or more sensors 210 may include any device for capturing or measuring data representing a perceptible phenomenon. By way of non-limiting example, the one or more sensors 210 may include one or more image sensors, microphones, thermometers, barometers, magnetometers, accelerometers, gyroscopes, tracking systems (e.g., GPS), and / or others.
[0051] also, Figure 2 The illustration shows that system 200 may include or communicate with one or more I / O systems 212. The one or more I / O systems 212 may include any type of input or output device, such as, by way of non-limiting example, a touchscreen, mouse, keyboard, controller, and / or others, but are not limited thereto. For example, the one or more I / O systems 212 may include a display system, which may include any number of display panels, optics, laser scanning display components, and / or other components.
[0052] Figure 3 The illustration shows an exemplary HMD 300, which is from Figure 2 An exemplary implementation of system 200. HMD 300 is shown as including multiple different cameras (e.g., implementations of one or more sensors 210), including cameras 305, 310, 315, 320, and 325. Cameras 305-325 may include any type of camera mode, such as one or more visible light cameras, one or more low-light cameras (e.g., configured with large pixels for image sensing in environments with low ambient light, such as starlight conditions (e.g., about 10 lux or lower)), one or more thermal imaging cameras (e.g., long-wave infrared cameras for detecting thermal radiation), one or more UV cameras, and / or others. Although in Figure 3 The diagram shows five cameras, but the HMD 300 may include more or fewer than five cameras.
[0053] In some cases, the camera can be positioned at a specific location on the HMD 300. For example, in some cases, the first camera (e.g., camera 320) is positioned on the HMD 300 at a location above the designated left eye position of any user wearing the HMD 300 in the height direction relative to the HMD 300. For example, camera 320 is positioned above the pupil 330. As another example, the first camera (e.g., camera 320) is additionally positioned above the designated left eye position in the width direction relative to the HMD. That is, camera 320 is not only positioned above the pupil 330, but also in a straight line relative to the pupil 330. When using a VR system, the camera may be placed directly in front of the designated left eye position. For example, see reference... Figure 3 The camera can be physically positioned on the HMD300 in front of the pupil 330 along the z-axis.
[0054] When a second camera (e.g., camera 310) is provided, the second camera can be positioned on the HMD above the designated right eye position of any user wearing the HMD in the height direction relative to the HMD. For example, camera 310 is above the pupil 335. In some cases, the second camera is additionally positioned above the designated right eye position in the width direction relative to the HMD. When using a VR system, the camera can be placed directly in front of the designated right eye position. For example, see reference... Figure 3 The camera can be physically positioned on the HMD 300 in front of the pupil 335 in the Z-axis direction.
[0055] When a user wears the HMD 300, the HMD 300 is fitted to the user's head, and the HMD 300's display is positioned in front of the user's pupils (such as pupils 330 and 335). Typically, cameras 305-325 will be physically offset from the user's pupils 330 and 335 by a certain distance. For example, a vertical offset may exist in the HMD's height direction (i.e., the "Y" axis), as shown by offset 340 (representing the vertical offset between the user's eyes and camera 325). Similarly, a horizontal offset may exist in the HMD's width direction (i.e., the "X" axis), as shown by offset 345 (representing the horizontal offset between the user's eyes and camera 325). Each camera may be associated with a different amount of offset.
[0056] In some implementations, the HMD 300 can be used to generate parallax-corrected pass-through visualizations of the user's environment. "Pass-through" visualization refers to a visualization reflecting what the user would see without wearing the HMD 300, regardless of whether the HMD 300 is included as part of an AR, VR, or other type of system. To generate this pass-through visualization, the HMD 300 can utilize one or more of its cameras 305-325 to capture its surrounding environment, including any objects within it, and transmit this data to the user for viewing. In many cases, the pass-through data is modified to reflect or correspond to the user's pupil's field of view. The field of view can be determined using any type of eye-tracking technology. In some instances, because the camera module is not telecentric with the user's eye, the difference in field of view between the user's eye and the camera module can be corrected to provide parallax-corrected pass-through visualizations.
[0057] To convert raw images into pass-through images, depth information can be determined from the raw images (or using a separate depth detection system). This depth information can specify the distance from the sensor to any object captured by the raw images (e.g., z-axis range or measurement result). Once these raw images are obtained, a depth map can be calculated based on depth data embedded in or contained in the raw images, and pass-through images (e.g., one for each pupil) can be generated using the depth information used for any reprojection.
[0058] As used herein, a "depth map" details the positional relationships and depth of objects relative to their environment. Therefore, it is possible to determine the positioning, location, geometry, outline, and depth of objects relative to each other. Based on the depth map (and possibly the original image), a 3D representation of the environment can be generated.
[0059] Accordingly, through-visualization will enable users to perceive the content currently in their environment without having to remove or reposition the HMD 300. Furthermore, as will be described in more detail later, the disclosed through-visualization can also enhance a user's ability to view objects in their environment (e.g., by displaying additional environmental conditions that may be imperceptible to the human eye).
[0060] It should be noted that although part of this disclosure focuses on generating "one" pass-through image, the implementation described herein can generate a separate pass-through image for each eye in the user's eye. That is, two pass-through images can be generated concurrently with each other. Therefore, although it is often mentioned that the generation appears to be a single pass-through image, the implementation described herein is actually capable of generating multiple pass-through images simultaneously.
[0061] In some cases, the pass-through image may have various levels of processing performed on the sensor, including denoising, tone mapping, and / or other processing steps to produce a high-quality image. Additionally, a camera reprojection step (e.g., parallax correction) may or may not be performed to correct for the offset between the user's viewpoint and the camera position.
[0062] As in Figure 3 As shown, none of the cameras 305-325 are directly aligned with pupils 330 and 335. Offsets 340 and 345 introduce a viewing angle difference (i.e., parallax) between cameras 305-325 and pupils 330 and 335. As mentioned above, due to the parallax resulting from offsets 340 and 345, in some cases, the original image produced by cameras 305-325 cannot be immediately used as (one or more) through-images 350. Therefore, parallax correction 355 (also known as image synthesis or reprojection) can be performed on the original image to transform (or reproject) those viewing angles represented within the original image to correspond to the viewing angles of the user's pupils 330 and 335. Parallax correction 355 may include any number of distortion corrections 360 (e.g., for correcting concave or convex wide-angle or narrow-angle camera lenses), epipolar transformations 365 (e.g., for making the optical axis of the camera parallel), and / or reprojection transformations 370 (e.g., for repositioning the optical axis so that it is substantially in front of or in line with the user's pupil).
[0063] As mentioned above, parallax correction 355 may include performing depth calculations to determine the depth of the environment, and then reprojecting the image to a determined location or with a determined viewpoint. As used herein, the phrases “parallax correction” and “image synthesis” are interchangeable and may include performing stereoscopic pass-through parallax correction and / or image reprojection parallax correction.
[0064] In some cases, reprojection is based on the current pose 375 of the HMD 300 relative to its surroundings (e.g., determined via visual-inertial SLAM). Based on the generated pose 375 and depth map, the HMD 300 and / or other systems can correct for parallax errors by reprojecting the viewpoints represented by the original image to align with the viewpoints of the user's pupils 330 and 335.
[0065] By performing these different transformations, the HMD 300 is able to perform three-dimensional (3D) geometric transformations on the raw camera image to change the perspective of the raw image in a manner related to the user's pupils 330 and 335. Additionally, the 3D geometric transformations rely on depth calculations, in which objects in the HMD 300's environment are mapped to determine their depth and pose 375. Based on these depth calculations and pose 375, the HMD 300 can perform a three-dimensional reprojection or three-dimensional warping of the raw image in such a way that the appearance of object depth is preserved in (one or more) through-images 350, wherein the preserved object depth substantially matches, corresponds to, or visualizes the actual depth of objects in the real world. Therefore, the degree or amount of parallax correction 355 depends at least in part on the degree or amount of offsets 340 and 345.
[0066] By performing parallax correction 355, the HMD 300 effectively creates a “virtual” camera positioned in front of the user’s pupils 330 and 335. For further clarification, consider the position of camera 305, which is currently above and to the left of pupil 335. By performing parallax correction 355, the embodiment programmatically transforms the images generated by camera 305, or transforms the viewing angle of those images, so that the viewing angle appears as if camera 305 is actually positioned directly in front of pupil 335. That is, even if camera 305 does not actually move, the embodiment is able to transform the images generated by camera 305 so that those images have the appearance of camera 305 being coaxially aligned with pupil 335, and in some cases, at the precise location of pupil 335.
[0067] Generating high-resolution depth maps using low-computation cameras
[0068] Figure 4 An exemplary head-mounted display 400 (HMD 400) is illustrated, including various cameras that can facilitate the disclosed embodiments. In at least some aspects, HMD 400 may correspond to HMD 300 and / or system 200 discussed above. Figure 4 As illustrated, the HMD includes a high-resolution low-light camera 402, a high-resolution thermal camera 404, and two low-resolution thermal cameras 406A and 406B.
[0069] As mentioned above, a low-light camera may include image-sensing pixels configured to detect small amounts of electrons at a sufficiently high frame rate to facilitate image capture in environments including low ambient light (e.g., under starlight conditions, approximately 10 lux or less). Furthermore, as mentioned above, a thermal camera may be configured to detect infrared light to provide an image representing the thermal radiation from objects within the capture environment.
[0070] The image resolution of images captured by the high-resolution low-light camera 402 and / or the high-resolution thermal camera 404 is higher than that of images captured by the low-resolution thermal cameras 406A and 406B. For example, in some instances, the image resolution of the high-resolution low-light camera 402 and / or the high-resolution thermal camera 404 is sufficiently high (e.g., 1920 × 1080, or another value or aspect ratio) so that the image pixels do not appear indistinguishable to the user during presentation to the user for various applications (such as through-pass imaging as discussed above), as discussed above. In contrast, the image resolution of the low-resolution thermal cameras 406A and 406B is lower than that of the high-resolution low-light camera 402 and / or the high-resolution thermal camera (e.g., 480 × 270, or another value or aspect ratio). For example, if presented to the user under normal use conditions (e.g., through-pass imaging), the images captured by the low-resolution thermal cameras 406A and 406B may appear pixelated.
[0071] Low-resolution thermal cameras 406A and 406B form a stereo thermal camera pair (i.e., stereo thermal camera 406), which can be configured to capture time-synchronized thermal images of an environment that are substantially identical in image resolution, aspect ratio, etc. Although in many cases, low-resolution thermal cameras 406A and 406B may not be able to capture images of sufficient fidelity to provide the desired user experience, low-resolution thermal cameras can facilitate low-computational depth map computation and can avoid the need to implement a stereo high-resolution camera pair (e.g., a second high-resolution thermal camera or a second high-resolution low-light camera) to facilitate pass-through imaging of the environment, as described in more detail below.
[0072] As mentioned above, the low-resolution thermal cameras 406A and 406B operate by detecting heat radiated within the captured scene. In some implementations, the thermal cameras are advantageously capable of operating in no-light (e.g., in pitch-black environments) and / or in low-visibility environments (e.g., in areas with smoke or fog). Therefore, the stereo thermal camera 406 is capable of capturing low-resolution images that can be used to obtain depth information in a variety of environments (as described in more detail below), which may be beneficial for users of the HMD 400 in various environments.
[0073] Although this example focuses in at least some aspects on a particular camera modality (e.g., thermal and low-light) and / or a particular number of cameras (e.g., two high-resolution cameras of different modalities and a stereo low-resolution camera pair of the same modality), it will be appreciated in light of this disclosure that the principles described herein are not limited to the specific configuration of this example. For example, according to this disclosure, a system for facilitating low-computation high-resolution depth map computation using a low-resolution camera can include any combination of visible light cameras, infrared cameras, ultraviolet cameras, low-light cameras, and / or cameras of any modality. For example, in some implementations, the HMD can be implemented using a stereo low-resolution low-light camera instead of a stereo low-resolution thermal camera. In some instances, a stereo low-resolution low-light camera can provide high contrast and / or high fidelity (e.g., compared to a low-resolution thermal camera), but may not operate as expected in completely dark environments, or in environments with smoke or fog, etc.
[0074] Furthermore, the system may include only a single high-resolution camera or two or more high-resolution cameras, and the high-resolution cameras (one or more) may have the same or different camera modes as the stereo low-resolution camera pair.
[0075] Figure 4 The illustration also shows that the HMD 400 may include any number of additional cameras 410 to facilitate various functions associated with providing a mixed reality (MR) experience. For example, the additional cameras 410 may facilitate simultaneous localization and mapping (SLAM), object tracking (e.g., hand tracking), and / or other functions. Furthermore, Figure 4 The illustration shows an HMD 400 including a display 408 for displaying virtual content to a user wearing the HMD 400. For example, in some instances, the display 408 may include a laser diode, a mirror (e.g., a microelectromechanical system (MEMS) mirror), a waveguide, a diffraction grating, an LCD display (for VR systems), and / or other elements or associated therewith for displaying images to the user's eyes. In some instances, the display 408 includes a separate optical / display system for displaying an image to each eye of the user.
[0076] As in Figure 5A As depicted in the document, HMDs can use their cameras to capture images of objects within their environment. Figure 5A The illustration shows the HMD 400 worn by user 504 when the HMD 400 captures an image of object 506 in the environment. Figure 5AThe illustration shows a high-resolution thermal image 508 (e.g., captured by a high-resolution thermal camera 404 of the HMD 400), a high-resolution low-light image 512 (e.g., captured by a high-resolution low-light camera 402 of the HMD 400), and low-resolution thermal images 516A and 516B (e.g., captured by low-resolution thermal cameras 406A and 406B of the HMD 400). The low-resolution thermal images 516A and 516B form a stereo image pair 516 to which depth calculations (e.g., stereo matching) can be performed.
[0077] High-resolution thermal image 508 captures texture information 510 describing the thermal radiation properties of object 506 at the time of capture. Figure 5A The representation of object 506 in the high-resolution thermal image 508 is filled with a pattern of intersecting lines. The high-resolution low-light image 512 captures different texture information 514 describing the texture of object 506 that can be observed in the visible spectrum. Figure 5A The representation of object 506 in the high-resolution low-light image 512 is filled with a dashed pattern. As described below, texture information 510 and / or 514 can provide the basis for generating a pass-through view of object 506 (e.g., for rendering on display 408 of HMD 400).
[0078] Figure 5A The various images include a center line depicted as a dashed line extending horizontally and vertically across the images. The center line is contained within... Figure 5A (And in other figures) to more clearly depict the spatial differences between the various images captured by the cameras of the HMD 400. For example, high-resolution thermal image 508 depicts object 506 with its top handle intersecting the vertical center line, most of which is located to the right of the vertical center line. Conversely, high-resolution low-light image 512 depicts object 506 with its top handle intersecting the vertical center line, most of which is located to the left of the vertical center line. These spatial differences between high-resolution thermal image 508 and high-resolution low-light image 512 are caused by the physical displacement between high-resolution thermal camera 404 and high-resolution low-light camera 402 on the HMD 400.
[0079] Figure 5A The spatial differences between and between the low-resolution thermal images 516A and 516B and (i) the high-resolution thermal image 508 and the high-resolution low-light image 512 are also illustrated. For example, the low-resolution thermal image 516B depicts an object 506 with its handle completely to the right of the vertical center line, while the low-resolution thermal image 516A depicts an object 506 with its handle completely to the left of the vertical center line.
[0080] exist Figure 5AAt least some of the images depicted may be captured or acquired by a system (e.g., system 200, HMD 400, etc.) to facilitate the generation of low-computational high-resolution depth maps using low-resolution images. To facilitate the generation of low-computational high-resolution depth maps using low-resolution images, the system may generate low-resolution depth maps using low-resolution thermal images 516A and 516B. Figure 5B The illustration shows low-resolution thermal images 516A and 516B provided as input to depth processing 518 for generating depth information of objects captured in the low-resolution thermal images 516A and 516B (e.g., describing the depth information of object 506 relative to low-resolution thermal cameras 406A and 406B when capturing stereo image pair 516).
[0081] Depth processing 518 for calculating depth information can be performed in various ways, including stereo matching. To perform stereo matching, a pair of images (e.g., low-resolution thermal images 516A and 516B) is obtained. Typically, a correction process is performed, thereby aligning corresponding pixels in different images of this pair of images, representing common 3D points in the environment, along scan lines (e.g., horizontal scan lines, vertical scan lines, epipolar lines, etc.). For the corrected images, the coordinates of corresponding pixels in the different images differ only in one dimension (e.g., the dimension of the scan lines). The stereo matching algorithm can then search along the scan lines to identify pixels in the different images that correspond to each other (e.g., by performing pixel patch matching to identify pixels representing common 3D points in the environment) and identify disparity values for the corresponding pixels. The disparity values can be based on the differences in pixel positions between corresponding pixels in different images describing the same part of the environment. Depth per pixel can be determined based on the disparity value per pixel, thus providing a depth map.
[0082] Figure 5B The output of depth processing 518 is illustrated as a depth map 520, which includes depth information 522. As mentioned above, the depth map 520 is capable of describing the per-pixel distance between the object captured in the stereo image pair and one or more cameras in the cameras that captured the stereo image pair. Figure 5B The diagram illustrates a depth map 520 within the geometry of a low-resolution thermal image 516A. In other words, the objects represented in the low-resolution thermal image 516A and the depth map 520 are spatially aligned. As mentioned above, the system can generate depth maps within the geometry of both images of the stereo image pair 516, and can perform any of the processing described herein without loss of generality to generate multiple parallax-corrected views (e.g., one for the user's right eye and one for the user's left eye).
[0083] As discussed above, performing depth processing 518 on a low-resolution image is computationally much less expensive than performing depth processing on a high-resolution image, and using a low-resolution image to capture stereo image pairs for depth processing allows the HMD 400 to omit stereo image pairs from a high-resolution camera. However, as indicated above, there is a spatial difference between the depth map 520 (which is within the geometry of the low-resolution thermal image 516A) and the high-resolution thermal image 508 and the high-resolution low-light image 512. Furthermore, the depth map 520 has an image resolution similar to that of the low-resolution thermal images 516A and 516B, and therefore has a lower image resolution than the high-resolution thermal image 508 and / or the high-resolution low-light image 512.
[0084] These spatial and image resolution differences pose problems for generating a parallax-corrected view using depth information 522 and texture information 510 or 514 from depth map 520. However, these obstacles can be overcome by utilizing reprojection and upsampling operations, as described below.
[0085] Figure 6A and Figure 6B The illustration depicts a conceptual representation of a capture viewpoint that reprojects a high-resolution image to correspond to a low-resolution depth map. Figure 6A and Figure 6B The illustration shows how texture information 510 from a high-resolution thermal image is reprojected to spatially align with depth information 522 from a depth map 520. Figure 6A The illustration shows a depth map 520, in which a deprojection ray 604 extends from the principal point 602 through the individual pixels of the depth information 522 of the depth map 520. For illustrative purposes, using pinhole camera terminology, the principal point 602 corresponds to the optical center or camera center of the low-resolution thermal camera 406B (which is the spatially aligned camera of the depth map 520), while the low-resolution thermal camera 406B captures a low-resolution thermal image 516B for forming the depth map 520.
[0086] When the pixels of the depth information are located on a forward image plane positioned around the principal point 602, the deprojection ray 604 is illustrated as being projected from the principal point 602 through the pixels of the depth information 522 represented in the depth map 520. Each deprojection ray 604 extends through the corresponding pixel of the depth information 522 to a distance corresponding to the depth value of the corresponding pixel of the depth information 522. These deprojection rays 604 provide a plurality of 3D points or coordinates that depict a 3D representation 606 of the object 506 captured in the depth map 520.
[0087] Each 3D point or coordinate of the 3D representation 606 of object 506 can be associated with a specific pixel through which the corresponding deprojection ray 604 is projected to provide depth information 522 for the 3D point or coordinate. In this way, if the pixels of the texture information 510 of the high-resolution thermal image 508 can be associated with the 3D points or coordinates of the 3D representation 606, then the pixels of the texture information 510 can be associated with and / or aligned with the depth information 522 of the depth map 520.
[0088] Figure 6B A high-resolution thermal image 508 is also illustrated, wherein a deprojection ray 610 extends from the principal point 608 of the high-resolution thermal image 508 through the individual pixels of the texture information 510 of the high-resolution thermal image 508. When the high-resolution thermal camera 404 captures the high-resolution thermal image 508, the principal point 608 corresponds to the optical center or camera center of the high-resolution thermal camera 404.
[0089] When the pixels of texture information 510 are located on a front image plane positioned around principal point 608, deprojection ray 610 is illustrated as being projected from principal point 608 through the pixels of texture information 510. At least some deprojection rays 610 extend through the corresponding pixels of texture information 510 until the deprojection ray 610 intersects a 3D point of 3D representation 606. Each pixel of texture information 510 through which the deprojection ray 610 intersects a specific 3D point of 3D representation 606 can be associated with the pixels of depth information 522 of depth map 520 through which the deprojection ray 604 passes, to generate the specific 3D point of 3D representation 606.
[0090] By using 3D points of 3D representation 606 as an intermediary, pixels of texture information 510 can be associated with pixels of depth information 522 of depth map 520, even if both are captured from different camera perspectives. In other words, the system can reproject texture information 510 to correspond to the perspective associated with depth map 520 (e.g., spatially align texture information 510 with depth information 522 of depth map 520) by deprojecting texture information onto 3D representation 606 and onto depth map 520 (or toward principal point 602 of depth map 520).
[0091] Figure 7 A high-resolution thermal image 508 is depicted as input to reprojection 702, which can perform as reference Figure 6A and Figure 6B The operation is conceptually described and can be performed using any reprojection technique known in the art. As discussed above, reprojection 702 is based at least in part on depth information 522 from depth map 520. (As in...) Figure 7As shown, the output of reprojection 702 includes a reprojected high-resolution thermal image 704, which includes reprojected texture information 706. Figure 7 The reprojected high-resolution thermal image 704 is shown spatially aligned with the depth map 520, as discussed above. For example, both the reprojected high-resolution thermal image 704 and the depth map 520 depict the object 506, with its top handle positioned entirely to the right of the vertical centerline.
[0092] Although the high-resolution thermal image 704 and the depth map 520 are spatially aligned, the two images have different image resolutions, with the high-resolution thermal image 704 having a higher image resolution than the depth map 520 (as shown in the image). Figure 7 (It is obvious). The resolution difference between depth information from a depth map (e.g., depth map 520) and texture information from a texture image (e.g., a reprojected high-resolution thermal image 704) can reduce the quality of a parallax-corrected image generated using both depth and texture information.
[0093] therefore, Figure 8 The illustration shows the generation of an upsampled depth map 804, which may include an image resolution matching the image resolution of the reprojected high-resolution thermal image 704. Specifically, Figure 8 A depth map 520 is shown as input to upsampling 802 for generating an upsampled depth map 804, the upsampled depth map 804 having upsampled depth information 806 spatially aligned with the reprojected texture information 706 of the reprojected high-resolution thermal image 704.
[0094] Upsampling 802 for generating high-resolution images from low-resolution images can employ techniques such as spatial domain schemes (e.g., sampling transforms using sampling theorems and Nyquist theorems), frequency domain schemes (e.g., image registration using the properties of discrete Fourier transform), learning-based techniques (e.g., adaptive regularization, pair matching, etc.), iterative reconstruction and interpolation-based techniques (e.g., iterative backprojection, pixel duplication, nearest neighbor interpolation, bilinear or bicubic interpolation, etc.), resolution techniques based on dynamic trees and wavelets (e.g., mean-field methods), and / or other techniques.
[0095] In some instances, upsampling 802 includes or utilizes filtering algorithms, such as edge-preserving filtering operations that optionally utilize a guiding image to improve the algorithm's output. By way of non-limiting examples, such edge-preserving filters may include joint bilateral filters, guided filters, bilateral solvers, etc. Figure 8An exemplary implementation is illustrated in which a reprojected high-resolution thermal image 704 is provided as a guide 808 for upsampling 802 to facilitate improved alignment between the reprojected high-resolution thermal image 704 and the upsampled depth map 804 (i.e., the output of upsampling 802).
[0096] As mentioned above, Figure 8 The reprojected texture information 706 of the reprojected high-resolution thermal image 704 is illustrated as spatially aligned with the upsampled depth information 806 of the upsampled depth map 804. Furthermore, as in... Figure 8 As illustrated, both the reprojected high-resolution thermal image 704 and the upsampled depth map 804 include the same image resolution. Therefore, the reprojected texture information 706 and the upsampled depth information 806 can be combined to form a parallax-corrected image for display to the user, as discussed above (e.g., see reference 1). Figure 3 ).
[0097] For example, the system can use the upsampled depth information 806 to reproject the already reprojected texture information 706 to correspond to the viewpoint of one or more user eyes (e.g., one or more eyes of user 504). By way of non-limiting example, such reprojection may include deprojecting each pixel of the reprojected texture information 706 to a distance indicated by the corresponding pixel of the upsampled depth information 806 having the same pixel coordinates. Deprojection can provide 3D points in 3D space, and these 3D points can be projected onto a principal point associated with the user's viewpoint and onto a forward-facing image plane to form a parallax-corrected image. The parallax-corrected image can be displayed on one or more portions of the display 408 of the HMD 400 (see...). Figure 4 This provides users with a pass-through image of their environment (e.g., a pass-through thermal image of the environment).
[0098] Even if the high-resolution thermal image 508 and the high-resolution low-light image 512 are spatially misaligned, operations similar to those discussed above for generating a reprojected high-resolution low-light image 704 can be performed without loss of generality to generate a reprojected high-resolution low-light image (or a reprojected high-resolution image with any camera mode different from the stereo image pair 516 used to form the low-resolution depth map 520).
[0099] Figure 9 The illustration shows a conceptual representation of reprojecting a high-resolution low-light image 512 to correspond to a capture viewpoint associated with a low-resolution depth map 520. Similar to the above... Figure 6A and 6B , Figure 9 A depth map 520 is illustrated, in which deprojection rays 604 extend from a principal point 602 through the individual pixels of depth information 522 in the depth map 520. The deprojection ray 604 is illustrated as being projected from the principal point 602 through the pixels of depth information 522 represented in the depth map 520 when the pixels of the depth information image are located on a forward image plane positioned around the principal point 602. Each deprojection ray 604 extends through the corresponding pixel of depth information 522 a distance corresponding to the depth value of that corresponding pixel in depth information 522. These deprojection rays 604 provide multiple 3D points or coordinates that depict a 3D representation 606 of an object 506 captured in the depth map 520.
[0100] Each 3D point or coordinate of the 3D representation 606 of object 506 can be associated with a specific pixel through which the corresponding deprojection ray 604 is projected to provide depth information 522 for the 3D point or coordinate. In this way, if the pixels of the texture information 514 of the high-resolution low-light image 512 can be associated with the 3D points or coordinates of the 3D representation 606, then the pixels of the texture information 514 can be associated with and / or aligned with the depth information 522 of the depth map 520.
[0101] Figure 9 The diagram also illustrates a high-resolution low-light image 512, in which a deprojection ray 904 extends from the principal point 902 of the high-resolution low-light image 512 through the individual pixels of the texture information 514 of the high-resolution low-light image 512. When the high-resolution low-light camera 402 captures the high-resolution low-light image 512, the principal point 902 corresponds to the optical center or camera center of the high-resolution low-light camera 402.
[0102] When the pixels of texture information 514 are located on a front image plane positioned around principal point 902, deprojection ray 904 is illustrated as a projection from principal point 902 through the pixels of texture information 514. At least some of the deprojection rays 904 extend through the corresponding pixels of texture information 514 until the deprojection ray 904 intersects a 3D point of 3D representation 606. Each pixel of texture information 514 through which the deprojection ray 904 intersects a specific 3D point of 3D representation 606 can be associated with a pixel of depth information 522 of depth map 520 through which the deprojection ray 904 intersects to generate that specific 3D point of 3D representation 606.
[0103] By using 3D points of 3D representation 606 as an intermediary, pixels of texture information 514 can be associated with pixels of depth information 522 of depth map 520, even if the two are captured from different camera perspectives. In other words, by deprojecting texture information 514 onto 3D representation 606 and instead onto depth map 520 (or toward principal point 602 of depth map 520), the system can reproject texture information 514 to correspond to the perspective associated with depth map 520 (e.g., spatially aligning texture information 514 with depth information 522 of depth map 520).
[0104] Figure 10 The illustration shows a reprojected high-resolution low-light image 1002, which can be generated using a reprojection operation based on depth information 522 from depth map 520, similar to those operations discussed above (e.g., from...). Figure 7 (Reprojection 702). The reprojected high-resolution low-light image 1002 includes reprojected texture information 1004. Figure 10 The diagram also illustrates the generation of an upsampled depth map 1008, which may include an image resolution matching the image resolution of the reprojected high-resolution low-light image 1002. Specifically, Figure 10 A depth map 520 provided as input to upsampling 1006 is shown to generate an upsampled depth map 1008 having upsampled depth information 1010 spatially aligned with the texture information 1004 of the reprojected high-resolution low-light image 1002. In some instances, upsampling 1006 utilizes one or more aspects of the reprojected high-resolution low-light image 1002 as input for generating the upsampled depth map 1008 (e.g., guide 1012).
[0105] As mentioned above, Figure 10 The reprojected texture information 1004 of the reprojected high-resolution low-light image 1002 is illustrated as spatially aligned with the upsampled depth information 1010 of the upsampled depth map 1008. Furthermore, as in... Figure 10 As illustrated, both the reprojected high-resolution low-light image 1002 and the upsampled depth map 1008 include the same image resolution. Therefore, the system can utilize the upsampled depth information 1010 to generate a parallax-corrected image by reprojecting the already reprojected texture information 1004 to correspond to the viewing angle of one or more eyes in the user's eye. The parallax-corrected image can be displayed on one or more portions of the display 408 of the HMD 400 (see [link to HMD 400]). Figure 4 Such parallax-corrected images can include captured through images of the environment (e.g., through low-light images of the environment).
[0106] In light of this disclosure, it will be appreciated that the system can perform operations to generate two parallax-corrected images using two different high-resolution images captured by high-resolution cameras of different camera modalities. For example, the system can generate a parallax-corrected low-light image based on a high-resolution low-light image 512, and can also generate a parallax-corrected thermal image based on a high-resolution thermal image 508. These images can be fused or combined to generate a synthetic pass-through image for presentation to a user, which captures information about the environment obtained from multiple camera modalities. It should be noted that when generating multiple parallax-corrected images, it is not necessary to repeatedly generate upsampled depth maps, and depth information from the same upsampled depth map can be used to generate multiple parallax-corrected images based on high-resolution texture information captured by different high-resolution cameras (e.g., generating reprojected texture information by first reprojecting other high-resolution texture information to correspond to the viewpoint of the upsampled depth map, and then reprojecting the already reprojected texture information again to correspond to the viewpoint of the user's eye).
[0107] Figure 11 An alternative embodiment of the HMD 1100, which can be used to generate high-resolution depth maps from low-resolution images, is illustrated. The HMD 1100 is similar in many ways to... Figure 4 The HMD 1100 is similar to the HMD 400. For example, the HMD 1100 includes a high-resolution low-light camera 1102, a high-resolution thermal camera 1104, a display 1108, and one or more other cameras 1110. The key difference between the HMD 1100 and the HMD 400 is that the HMD 1100 includes only a single low-resolution thermal camera 1106, and does not include a stereo low-resolution thermal camera pair.
[0108] Figure 12 The illustration shows an HMD 1100 worn by a user 1204 when the HMD 1100 captures an image of an object 1206 in the environment. Figure 12 The illustration shows a high-resolution thermal image 1208 (e.g., captured by a high-resolution thermal camera 1104), a high-resolution low-light image 1212 (e.g., captured by a high-resolution low-light camera 1102), and a low-resolution thermal image 1216 (e.g., captured by a low-resolution thermal camera 1106). The high-resolution thermal image 1208 captures texture information 1210 describing the thermal radiation properties of the object 506 at the time of capture, and the high-resolution low-light image 1212 captures texture information 1214 describing the texture of the object 506 observable in the visible spectrum. Similar to the reference above. Figure 5A The spatial misalignment discussed exists in... Figure 12 Among the various images shown.
[0109] As will be described herein, in some cases, high-resolution depth maps and high-resolution parallax-corrected images can be generated even in the absence of any stereo camera pair or stereo image pair captured by a stereo camera. Such functionality can be facilitated in a variety of ways according to this disclosure. Figure 13 An exemplary technique is provided for generating a high-resolution depth map in the absence of stereo image pairs captured by a stereo camera. Figure 14 An alternative technique is provided for generating high-resolution depth maps in the absence of stereo image pairs captured by a stereo camera.
[0110] Figure 13 The illustration illustrates a conceptual representation of providing a low-resolution thermal image 1216 as input to upsampling 1302 to generate an upsampled thermal image 1304. Upsampling 1302 can utilize any technique described herein or known in the art to generate a high-resolution image from an initial image input. In some instances, the upsampled thermal image 1304 is configured to have the same image resolution as the high-resolution thermal image 1208, as shown in... Figure 13 As shown in the diagram. Figure 13 It is also shown that there is a parallax between the capture viewpoints associated with the high-resolution thermal image 1208 and the upsampled thermal image 1304 (e.g., the representations of object 1206 in the two images are offset vertically and horizontally from each other).
[0111] Because both the high-resolution thermal image 1208 and the upsampled thermal image 1304 have the same high image resolution but different associated capture viewpoints, these two images can be used as input to perform depth processing 1306, as in... Figure 13 As described herein, depth processing 1306 can utilize any techniques described herein or known in the art to generate a depth map from an image input. Figure 13 The diagram illustrates the output of depth processing 1306 as a high-resolution depth map 1308, which includes depth information 1310.
[0112] Figure 13 The diagram also illustrates a depth map 1308 within the geometry of the high-resolution thermal image 1208, such that the depth information 1310 of the high-resolution depth map 1308 and the texture information 1210 of the high-resolution thermal image 1208 are spatially aligned. Considering this spatial alignment, the system can use the depth information 1310 to reproject the texture information 1210 to correspond to the user's eye's perspective, as shown in... Figure 13The diagram illustrates the use of a high-resolution depth map 1308 and a high-resolution thermal image 1208 as inputs to the reprojection 1312 (as indicated by arrows extending from the high-resolution depth map 1308 and the high-resolution thermal image 1208 to the reprojection 1312). The reprojection 1312 can provide one or more parallax-corrected images 1314, such as parallax-corrected thermal images that can be displayed to a user (e.g., a display 1108 using an HMD 1100).
[0113] Depth information 1310 can also be used to generate a parallax-corrected view based on texture information captured by a high-resolution camera of other modalities, such as a high-resolution low-light image 1212. For example, the system can use depth information 1310 to generate reprojected low-light texture information to reproject the texture information 1214 of the high-resolution low-light image 1212 to become spatially aligned with the depth information 1310 of the high-resolution depth map 1308. The system can then use depth information 1310 to reproject the already reprojected low-light texture information again to correspond to the view of the user's eye, thereby forming a parallax-corrected low-light image that can be displayed to the user (e.g., using the display 1108 of the HMD 1100). Such operation is... Figure 13 The image is depicted by arrows extending from the high-resolution low-light image 1212 to the reprojection 1312, which may contribute to one or more parallax-corrected images 1314. As discussed above, the parallax-corrected thermal image and the parallax-corrected low-light image can be combined to form a composite parallax-corrected image to be presented to the user.
[0114] Figure 14 An alternative method for generating high-resolution depth maps is illustrated when a pair of stereo images captured by a stereo camera is not available. Figure 14 A conceptual representation is provided that provides a high-resolution thermal image 1208 as input to downsampling 1402 to generate a downsampled thermal image 1404.
[0115] In some implementations, downsampling 1402 includes reducing a portion of pixels in the original image (e.g., high-resolution thermal image 1208) to a single pixel in the downsampled image (e.g., downsampled thermal image 1404). For example, in some cases, each pixel in the downsampled image is defined by pixels in the original image:
[0116] pa(m,n)=p(Km,Kn)
[0117] Among them, P dHere, p is the pixel in the downsampled image, K is the scaling factor, m is the pixel coordinate on the horizontal axis, and n is the pixel coordinate on the vertical axis. In some cases, downsampling 1402 also includes pre-filtering functions for defining the pixels of the downsampled image, such as anti-aliasing pre-filtering to prevent aliasing artifacts.
[0118] In some implementations, downsampling 1402 utilizes an averaging filter to define the pixels of the downsampled image based on the average value of a pixel portion in the original image. In one example of downsampling by a factor of 2 along each axis, each pixel in the downsampled image is defined by the average value of a 2x2 pixel portion in the original image:
[0119]
[0120] Where, p d is the pixel in the downsampled image, p is the pixel in the original image, m is the pixel coordinate on the horizontal axis, and n is the pixel coordinate on the vertical axis.
[0121] Downsampling 1402 may include iterative downsampling operations performed iteratively to obtain a downsampled image at the desired final image resolution. Figure 14 A downsampled thermal image 1404 with the same image resolution as the low-resolution thermal image 1216 is illustrated, but there is parallax between the capture viewpoints associated with the downsampled thermal image 1404 and the low-resolution thermal image 1216. Therefore, depth processing 1406 can be performed using both the downsampled thermal image 1404 and the low-resolution thermal image 1216 as input. Figure 13 Compared to the high-resolution thermal image 1208 and the upsampled thermal image 1304, the depth processing 1406 can be more computationally efficient, considering the smaller image size of the downsampled thermal image 1404 and the low-resolution thermal image 1216.
[0122] Figure 14 The illustration shows the output of depth processing 1406 as depth map 1408, which has a lower image resolution than high-resolution thermal image 1208. Figure 14 The depth map 1408 in the geometry of the downsampled thermal image 1404 is also illustrated. Figure 14 The illustration shows a conceptual representation of a depth map 1408 provided as input to upsampling 1408 to generate an upsampled depth map 1412 including upsampled depth information 1414. Upsampling 1408 can utilize any technique described herein or known in the art to generate a high-resolution image from an initial image input. For example, upsampling 1410 can utilize guidance based on a high-resolution thermal image 1208 (although in...) Figure 14 (Not explicitly shown in the text).
[0123] In some instances, the upsampled depth map 1412 is configured to have the same image resolution as the high-resolution thermal image 1208, as in Figure 14 As shown in the diagram. Figure 13 It is also shown that the upsampled depth information 1414 of the upsampled depth map 1412 is spatially aligned with the texture information 1210 of the high-resolution thermal image 1208 (e.g., given that the low-resolution depth map 1408 is within the geometry of the downsampled thermal image 1404, and / or given that the high-resolution thermal image 1208 is used as a guide for the upsampled 1410). Due to the same image resolution and the spatial alignment between the upsampled depth map 1412 and the high-resolution thermal image 1208, the system can use the upsampled depth information 1414 to reproject the texture information 1210 to correspond to the user's eye's viewing angle, as shown in... Figure 14 As illustrated, a high-resolution thermal image 1208 and an upsampled depth map 1412 are provided as input to reprojection 1416 (as indicated by the arrows extending from the high-resolution thermal image 1208 and the upsampled depth map 1412 to reprojection 1416). Reprojection 1416 can provide one or more parallax-corrected images 1418, such as a parallax-corrected thermal image that can be displayed to a user (e.g., a display 1108 using HMD 1100).
[0124] The upsampled depth information 1414 can also be used to generate a parallax-corrected view based on texture information captured by a high-resolution camera of other modalities, such as a high-resolution low-light image 1212. For example, the system can use the upsampled depth information 1414 to generate reprojected low-light texture information to reproject the texture information 1214 of the high-resolution low-light image 1212 so that it becomes spatially aligned with the upsampled depth information 1414 of the upsampled depth map 1412. The system can then use the upsampled depth information 1414 to reproject the already reprojected low-light texture information again to correspond to the user's eye's viewing angle, thereby forming a parallax-corrected low-light image that can be displayed to the user (e.g., using a display 1108 of the HMD 1100). Such operation is... Figure 14 The image 1418, depicted by arrows extending from the high-resolution low-light image 1212 to the reprojection 1416, can contribute to one or more parallax-corrected images 1418. As discussed above, the parallax-corrected thermal image and the parallax-corrected low-light image can be combined to form a composite parallax-corrected image for presentation to the user.
[0125] An exemplary method for dense depth computation assisted by sparse feature matching
[0126] The following discussion now concerns the many methods and method actions that can be performed. Although method actions may be discussed in a specific order or shown in a flowchart in a specific order, a specific order is not required unless specifically stated otherwise, or because an action depends on another action that is completed before that action is performed.
[0127] Figure 15-17 Exemplary flowcharts 1500, 1600, and 1700 are illustrated respectively, depicting actions associated with low-computational depth map generation to provide disparity correction for images. Discussions of the various actions represented in the flowcharts include references... Figure 2 , 4 And 11, which provides a more detailed description of various hardware components.
[0128] Figure 15 Action 1502 of flowchart 1500 includes acquiring a stereoscopic image pair of the environment. In some implementations, action 1502 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others. In some instances, system 200 includes a head-mounted display (HMD), and system 200 may include a pair of stereoscopic cameras for capturing stereoscopic image pairs of the environment.
[0129] Action 1504 of flowchart 1500 includes generating a depth map of the environment by performing stereo matching on stereo image pairs, the depth map including depth information for the environment. In some implementations, action 1504 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0130] Action 1506 of flowchart 1500 includes acquiring a first image including first texture information for the environment, the first image including a first image resolution higher than the image resolution of the stereo image pair. In some implementations, action 1506 is performed using one or more components of system 200, such as one or more processors 202, storage device 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others. In some implementations, system 200 includes a first camera that captures the first image. The first camera may have a different modality than the stereo camera pair that captures the stereo image pair. For example, the first camera may be a low-light camera, while the camera of the stereo camera pair may be a thermal camera. In another example, the first camera is a thermal camera, while the camera of the stereo camera pair is a low-light camera.
[0131] In some cases, the first camera may have the same modality as the stereo camera pair that captures the stereo images. For example, the camera in the first camera and the stereo camera pair may be a thermal camera, or the camera in the first camera and the stereo camera pair may be a low-light camera.
[0132] Action 1508 of flowchart 1500 includes generating a reprojected first image by reprojecting the first image to correspond to an image capture viewpoint associated with a depth map. The reprojection of the first image is based on depth information from the depth map, and the reprojected first image includes first texture information reprojected for the environment. In some implementations, action 1508 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0133] Action 1510 of flowchart 1500 includes generating an upsampled depth map based on the depth map. In some implementations, action 1510 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others. In some cases, the upsampled depth map and the reprojected first image have the same image resolution. In some cases, the generation of the upsampled depth map is based on the reprojected first texture information. Furthermore, in some cases, generating the upsampled depth map includes utilizing an edge-preserving filter, such as a joint bilateral filter.
[0134] Action 1512 of flowchart 1500 includes generating a parallax-corrected image by reprojecting the reprojected first image to correspond to the user's viewpoint. In some implementations, action 1512 is performed using one or more components of system 200, such as one or more processors 202, storage device 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others. Reprojecting the reprojected first image to generate the parallax-corrected image is based on upsampled depth information from an upsampled depth map. The parallax-corrected image can be displayed on a display of system 200 (e.g., display 408).
[0135] Action 1514 of flowchart 1500 includes acquiring a second image that includes second texture information specific to the environment. In some implementations, action 1514 is performed using one or more components of system 200, such as one or more processors 202, storage device 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others. The second image has a second image resolution higher than the image resolution of the stereo image pair. In some cases, the second image is associated with a camera modality different from the first image. For example, system 200 may include a second camera that captures the second image, and the second camera may have a different modality than the first camera and the stereo camera pair (e.g., the second camera may be a low-light camera while the other cameras are not low-light cameras; or the second camera may be a thermal camera while the other cameras are not thermal cameras).
[0136] Action 1516 of flowchart 1500 includes generating a reprojected second image by reprojecting a second image to correspond to an image capture viewpoint associated with the depth map, the reprojection of the second image being based on depth information from the depth map. In some implementations, action 1516 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others. The reprojected second image may include second texture information for reprojection of the environment.
[0137] Figure 16Action 1602 of flowchart 1600 includes obtaining a first image of the environment. In some implementations, action 1602 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0138] Action 1604 of flowchart 1600 includes acquiring a second image of the environment, the second image being captured synchronously with the first image in time, the second image including a higher image resolution than the first image. In some implementations, action 1604 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0139] Action 1606 of flowchart 1600 includes generating an upsampled first image, wherein the upsampled first image has the same image resolution as the second image. In some implementations, action 1606 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0140] Action 1608 of flowchart 1600 includes generating a depth map of the environment by performing stereo matching on the upsampled first and second images. In some implementations, action 1608 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0141] Action 1610 of flowchart 1600 includes generating a parallax-corrected image by reprojecting the second image to correspond to the user's viewpoint, wherein the reprojection of the second image to generate the parallax-corrected image is based on depth information from a depth map. In some implementations, action 1610 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0142] Action 1612 of flowchart 1600 includes acquiring an additional image that includes additional texture information specific to the environment. In some implementations, action 1612 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others. The additional image includes an image resolution higher than the first image. The additional image may be associated with a camera modality different from the second image discussed in reference action 1604 above.
[0143] Action 1614 of flowchart 1600 includes generating a reprojected additional image by reprojecting an additional image to correspond to the image capture viewpoint associated with the depth map. The reprojected additional image includes additional texture information for the environment. In some implementations, action 1614 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0144] Action 1616 of flowchart 1600 includes generating a parallax-corrected additional image by reprojecting the reprojected additional image to correspond to the user's viewpoint, wherein the reprojection of the reprojected additional image to generate the parallax-corrected additional image is based on depth information from a depth map. In some implementations, action 1616 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0145] Figure 17 Action 1702 of flowchart 1700 includes obtaining a first image of the environment. In some implementations, action 1702 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0146] Action 1704 of flowchart 1700 includes acquiring a second image that includes texture information about the environment, the second image capturing the environment synchronously with the first image in time, and the second image having a higher image resolution than the first image. In some implementations, action 1704 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others. In some cases, the first image and the second image are associated with the same camera modality.
[0147] Action 1706 of flowchart 1700 includes generating a downsampled second image, wherein the downsampled second image has the same image resolution as the first image. In some implementations, action 1706 is performed using one or more components of system 200, such as one or more processors 202, storage device 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0148] Action 1708 of flowchart 1700 includes generating a depth map of the environment by performing stereo matching on a downsampled second image and a first image. In some implementations, action 1708 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0149] Action 1710 of flowchart 1700 includes generating an upsampled depth map based on texture information from the depth map and the second image. In some implementations, action 1710 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0150] Action 1712 of flowchart 1700 includes generating a parallax-corrected image by reprojecting the second image to correspond to the user's viewpoint, wherein the reprojection of the second image to generate the parallax-corrected image is based on upsampled depth information from an upsampled depth map. In some implementations, action 1712 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0151] Action 1714 of flowchart 1700 includes acquiring an additional image that includes additional texture information specific to the environment. In some implementations, action 1714 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others. The additional image includes an image resolution higher than that of the first image. The additional image may be associated with a camera modality different from the first and second images.
[0152] Action 1716 of flowchart 1700 includes generating a reprojected additional image by reprojecting the additional image to correspond to an image capture viewpoint associated with the upsampled depth map. The reprojected additional image includes additional texture information for the environment. In some implementations, action 1716 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0153] Action 1718 of flowchart 1700 includes generating a parallax-corrected additional image by reprojecting the reprojected additional image to correspond to the user's viewpoint, wherein the reprojection of the reprojected additional image to generate the parallax-corrected additional image is based on upsampled depth information from an upsampled depth map. In some implementations, action 1718 is performed using one or more components of system 200, such as one or more processors 202, storage devices 204, one or more sensors 210, one or more I / O systems 212, one or more communication systems 214, and / or others.
[0154] The disclosed embodiments may include or utilize dedicated or general-purpose computers including computer hardware, as discussed in more detail below. The disclosed embodiments also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media may be any available media accessible by a general-purpose or dedicated computer system. A computer-readable medium storing computer-executable instructions in data form is one or more “physical computer storage media” or “hardware storage devices.” A computer-readable medium that carries only computer-executable instructions and does not store computer-executable instructions is a “transmission medium.” Therefore, by way of example and not limitation, the present embodiments may include at least two distinct types of computer-readable media: computer storage media and transmission media.
[0155] Computer storage media (also known as “hardware storage devices”) are computer-readable hardware storage devices, such as RAM, ROM, EEPROM, CD-ROM, RAM-based solid-state drives (“SSDs”), flash memory, phase-change memory (“PCM”), or other types of memory, or other optical disc storage, disk storage, or other magnetic storage devices, or any other medium that can be used to store desired program code in hardware in the form of computer-executable instructions, data, or data structures and that can be accessed by a general-purpose or special-purpose computer.
[0156] A “network” is defined as one or more data links that enable the transmission of electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer via a network or other communication connection (hardwired, wireless, or a combination of hardwired and wireless), the computer correctly regards the connection as a transmission medium. Transmission media may include networks and / or data links that can be used to carry program code in the form of computer-executable instructions or data structures and are accessible by general-purpose or special-purpose computers. Combinations of the above are also included within the scope of computer-readable media.
[0157] Furthermore, upon arrival at various computer system components, program code units in the form of computer-executable instructions or data structures can be automatically transferred from the transmission computer-readable medium to the physical computer-readable storage medium (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., a "NIC") and then eventually transferred to the computer system RAM and / or to a less volatile computer-readable physical storage medium at the computer system location. Therefore, computer-readable physical storage media can be contained within computer system components that also (or even primarily) utilize the transmission medium.
[0158] Computer-executable instructions include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a function or a set of functions. Computer-executable instructions can be, for example, binary files, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions described above. Rather, the described features and actions are disclosed as exemplary forms for implementing the claims.
[0159] The disclosed embodiments may include or utilize cloud computing. The cloud model may consist of various features (e.g., on-demand self-service, broad network access, resource pooling, rapid elasticity, metric services, etc.), service models (e.g., Software as a Service (“SaaS”), Platform as a Service (“PaaS”), Infrastructure as a Service (“IaaS”), and deployment models (e.g., private cloud, community cloud, public cloud, hybrid cloud, etc.).
[0160] Those skilled in the art will understand that this invention can be practiced in networked computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronic devices, network PCs, minicomputers, mainframes, mobile phones, PDAs, pagers, routers, switches, wearable devices, etc. This invention can also be practiced in distributed system environments, where multiple computer systems (e.g., local and remote systems) linked via a network (via hardwired data links, wireless data links, or a combination of hardwired and wireless data links) perform tasks. In a distributed system environment, program modules can reside in local and / or remote memory storage devices.
[0161] Alternatively or additionally, the functions described herein may be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), central processing units (CPUs), graphics processing units (GPUs), and / or others.
[0162] As used herein, the terms “executable module,” “executable component,” “component,” “module,” or “engine” can refer to a hardware processing unit or a software object, routine, or method that can be executed on one or more computer systems. The various components, modules, engines, and services described herein can be implemented as objects or processors that execute on one or more computer systems (e.g., as separate threads).
[0163] Those who wish to understand will also understand how any feature or operation disclosed herein can be combined with any or a combination of other features and operations disclosed herein. Furthermore, the content or feature in any of the figures can be combined or used in conjunction with any content or feature used in any other figure. In this regard, the content disclosed in any of the figures is not mutually exclusive, but can be combined with content from any other figure.
[0164] This invention may be embodied in other specific forms without departing from its spirit or characteristics. The described embodiments are to be considered illustrative rather than restrictive in all respects. Therefore, the scope of the invention is indicated by the appended claims rather than by the foregoing description. All modifications within the equivalent meaning and scope of the claims should be included within their scope.
Claims
1. A system for facilitating the generation of low-computational depth maps to provide disparity-corrected images, the system comprising: One or more processors; as well as One or more hardware storage devices storing instructions executable by the one or more processors to configure the system to facilitate the generation of low-computational depth maps to provide disparity-corrected images by configuring the system for the following operations: Obtain stereoscopic image pairs of the environment; A depth map of the environment is generated by performing stereo matching on the stereo image pairs, the depth map including depth information for the environment; A first image is obtained, including first texture information for the environment, the first image having a first image resolution higher than that of the stereo image pair, wherein there is a parallax between the image capture viewpoint associated with the depth map and the image capture viewpoint associated with the first image, and wherein the stereo image pair and the first image are captured at the same capture time; A reprojected first image is generated by reprojecting the first image such that the image capture viewpoint associated with the first image corresponds to the image capture viewpoint associated with the depth map. The reprojection of the first image is based on the depth information from the depth map. The reprojected first image includes first texture information reprojected for the environment. The first image reprojected is used to generate an upsampled depth map based on the depth map.
2. The system according to claim 1, wherein, The upsampled depth map and the reprojected first image have the same image resolution, and the generation of the upsampled depth map is based on the reprojected first texture information.
3. The system according to claim 1, wherein, Generating the upsampled depth map involves using an edge-preserving filter.
4. The system according to claim 3, wherein, The edge-preserving filter includes a joint bilateral filter.
5. The system according to claim 1, wherein, The instructions can be executed by the one or more processors to further configure the system to: generate a parallax-corrected image by reprojecting the reprojected first image to correspond to the user's viewpoint, wherein the reprojection of the reprojected first image to generate the parallax-corrected image is based on upsampled depth information from the upsampled depth map.
6. The system according to claim 5, wherein, The system also includes a display, and wherein the instructions are executable by the one or more processors to further configure the system to display the parallax-corrected image on the display.
7. The system according to claim 1, wherein, The system includes a pair of stereo cameras and a first camera, wherein the pair of stereo images is captured by the pair of stereo cameras, and wherein the first image is captured by the first camera.
8. The system according to claim 7, wherein, The first camera has a different mode than the stereo camera.
9. The system according to claim 7, wherein, The stereo camera pair has the same camera mode as the first camera.
10. The system according to claim 9, wherein, The instructions can be executed by the one or more processors to further configure the system as follows: Obtain a second image including second texture information for the environment, the second image including a second image resolution higher than the image resolution of the stereo image pair, wherein the second image is associated with a camera modality different from the first image; and The reprojected second image is generated by reprojecting the second image to correspond to the image capture viewpoint associated with the depth map. The reprojection of the second image is based on the depth information from the depth map. The reprojected second image includes second texture information reprojected for the environment.
11. The system according to claim 10, wherein, The system includes a second camera that captures the second image, the second camera having a mode different from the first camera and the stereo camera pair.