Aligning images from a separation camera using 6DOF pose information

By generating 3D feature maps and determining 6-DOF poses, the problem of image alignment difficulties in mixed reality systems was solved, enabling precise image alignment and display between a remote camera and a head-mounted device.

CN116194866BActive Publication Date: 2025-12-16MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180060798.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-17
Filing Date
2021-04-20
Publication Date
2025-12-16
Estimated Expiration
2041-04-20

AI Technical Summary

Technical Problem

In mixed reality systems, when using multiple cameras, the lack or misalignment of timestamp data during image alignment makes it difficult to align image content, especially between single-camera systems and separate camera systems, affecting the accuracy of image generation and display.

Method used

By generating a 3D feature map of the environment, the 6-DOF poses of the integrated and separate cameras are determined, and the pose information is used to reproject the image from the separate camera to align with the viewpoint of the integrated camera, generating a superimposed image, with optional parallax correction.

Benefits of technology

This enables precise alignment of image content between a remote camera system and a head-mounted device without relying on timestamp data, improving image quality and display effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116194866B_ABST
    Figure CN116194866B_ABST
Patent Text Reader

Abstract

Techniques for aligning images generated by an integrated camera physically mounted to an HMD with images generated by a separate camera physically unmounted from the HMD are disclosed. A 3D feature map is generated and shared with the separate camera. Both the integrated camera and the separate camera use the 3D feature map to reposition themselves and determine their respective 6DOF poses. The HMD receives an image of the environment of the separate camera and the 6DOF pose of the separate camera. A depth map of the environment is accessed. An overlay image is generated by reprojecting the perspective of the image of the separate camera to align with the perspective of the integrated camera and by superimposing the reprojected image of the separate camera onto the image of the integrated camera.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Mixed reality (MR) systems, including virtual reality (VR) and augmented reality (AR) systems, have received widespread attention for their ability to create truly unique experiences for users. By way of reference, conventional VR systems create a fully immersive experience by limiting the field of view of its users to only a virtual environment. This is often accomplished through the use of a head-mounted device (HMD) that completely blocks any view of the real world. Thus, the user is fully immersed within the virtual environment. In contrast, conventional AR systems create an augmented reality experience by visually presenting virtual objects that are placed in or interact with the real world.

[0002] As used herein, VR and AR systems are described and referenced interchangeably. Unless otherwise noted, the description herein applies equally to all types of MR systems, which (as described in detail above) include AR systems, VR reality systems, and / or any other similar system capable of displaying virtual content.

[0003] MR systems can also employ different types of cameras in order to display content to a user, such as in the form of pass-through images. Pass-through images or views can assist a user in avoiding disorientation and / or safety hazards when transitioning into and / or navigating within an MR environment. MR systems can present views captured by cameras in a variety of ways. However, the process of using images captured by world-facing cameras to provide views of the real-world environment presents a number of challenges.

[0004] Some of these challenges arise when attempting to align image content from multiple cameras. Typically, this alignment process requires detailed timestamp information in order to perform the alignment process. However, at times, timestamp data is not available because different cameras can operate in different time domains, such that they have a time offset. Additionally, at times, timestamp data is not available at all because the cameras can operate remotely from one another and timestamp data is not transmitted. Another issue occurs due to having left and right HMD cameras (i.e., a dual camera system) but only a single detached camera. Aligning image content between the image of the detached camera and the image of the left camera presents a number of problems in terms of computational efficiency and image alignment in addition to aligning image content between the image of the detached camera and the image of the right camera. In other words, aligning image content provides substantial benefits, particularly in terms of placement and generation of holograms, and thus these problems present a serious obstacle to the art. Accordingly, there is a substantial need in the art to improve how images are aligned with one another.

[0005] The subject matter claimed herein is not limited to implementations that solve any disadvantages or that operate only in environments such as those described above. Rather, this background is provided only to illustrate one exemplary technology area where some embodiments described herein can be practiced. SUMMARY

[0006] Embodiments disclosed herein relate to systems, devices (e.g., hardware storage devices, wearable devices, etc.) and methods for aligning and stabilizing images generated by an integrated camera physically mounted to a head-mounted device (HMD) with images generated by a separate camera physically unmounted from the HMD.

[0007] In some embodiments, a three-dimensional (3D) feature map of an environment in which both the HMD and the separate camera operate is generated. The 3D feature map is then shared with the separate camera. The 3D feature map is used to relocalize a position frame of the integrated camera based on a first image generated by the integrated camera. Thus, a 6 degrees of freedom (6DOF) pose of the integrated camera is determined. Further, the separate camera uses the 3D feature map to relocalize a position frame of the separate camera based on a second image generated by the separate camera. Thus, a 6DOF pose of the separate camera is also determined. Embodiments then receive from the separate camera (i) the second image of the environment and (ii) the 6DOF pose of the separate camera. A depth map of the environment is accessed. Additionally, an overlay image is generated by reprojecting a perspective of the second image to align or match a perspective of the first image and then by superimposing at least a portion of the reprojected second image onto the first image. Note that (i) the 6DOF pose of the integrated camera, (ii) the 6DOF pose of the separate camera, and (iii) the depth map are used to perform the reprojecting process.

[0008] Optionally, some embodiments additionally perform a parallax correction on the overlay image to modify the perspective of the overlay image to correspond to a new perspective. In some cases, the new perspective is a perspective of a pupil of a user wearing the HMD. An additional option is to display the overlay image for viewing by the user.

[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to determine the scope of the claimed subject matter.

[0010] Additional features and advantages will be set forth in the description that follows, and in part will be apparent from the description, or can be learned by practice of the teachings herein. Features and advantages of the application will be realized and attained by the instrumentalities and combinations particularly pointed out in the appended claims. The features and advantages of the application will become more fully apparent from the following description and appended claims, or can be learned by practice of the application as described in the following description. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to describe the manner in which the above-recited and other advantages and features can be obtained, a more particular description will be rendered by reference to specific embodiments thereof, which are illustrated in the appended drawings. It is appreciated that these drawings depict only typical embodiments and are therefore not to be considered limiting of its scope, the application will be described and explained with additional specificity and detail by the use of the accompanying drawings in which:

[0012] FIG. 1 An example scenario involving an integrated camera and a separate camera is illustrated.

[0013] FIG. 2 An example head mounted device (HMD) is illustrated.

[0014] FIG. 3 An example implementation or configuration of an HMD is illustrated.

[0015] FIG. 4A And FIG. 4B A flowchart of an example method for aligning images from a separate camera and an integrated camera using 6DOF pose information from both the separate camera and the integrated camera is illustrated.

[0016] FIG. 5 An example scenario in which an integrated camera is generating images of an environment is illustrated.

[0017] FIG. 6 An example resulting 3D feature map that can be generated is illustrated, where the 3D feature map identifies feature points located in the environment.

[0018] FIG. 7 How the HMD can share or transmit the 3D feature map with a separate camera is illustrated.

[0019] FIG. 8 An example scenario in which an integrated camera and a separate camera are generating images of an environment is illustrated.

[0020] FIG. 9 A repositioning process that can be performed by the integrated camera and the separate camera in order to enable both cameras to use the same coordinate system in the physical space of the HMD is illustrated.

[0021] FIG. 10 Figure illustrates how the separate camera can transmit its 6DOF pose information as well as its generated images to the HMD.

[0022] FIG. 11 Figure illustrates how the HMD maintains information about the images and 6DOF pose of the integrated camera; about the images and 6DOF pose of the separate camera; and about the depth map of the environment. The HMD can also update the 6DOF pose information based on inertial measurement unit (IMU) data obtained from IMUs associated with the integrated camera and the separate camera.

[0023] FIG. 12 Figure illustrates how a depth map can be generated based on different source information.

[0024] FIG. 13 Figure illustrates an example re-projection operation that can be performed to re-project the perspective of the image of the separate camera to a new perspective that matches the perspective of the image of the integrated camera so that the re-projected image can subsequently be overlaid onto the image of the integrated camera.

[0025] FIG. 14 Figure illustrates inputs that can be used to perform the re-projection operation.

[0026] FIG. 15 Figure illustrates an example overlay operation that can be performed to overlay the re-projected image (i.e., the image of the separate camera) onto the image of the integrated camera to generate an overlaid image.

[0027] FIG. 16 Figure illustrates how a parallax correction operation can be performed on the overlaid image to correct for parallax.

[0028] FIG. 17 Figure illustrates an example computer system that can perform any of the disclosed operations. DETAILED DESCRIPTION

[0029] Embodiments disclosed herein relate to systems, devices (e.g., hardware storage devices, wearable devices, etc.) and methods for aligning and stabilizing images generated by an integrated camera that is physically mounted to a head-mounted device (HMD) with images generated by a separate camera that is physically unmounted from the HMD.

[0030] In some embodiments, a 3D feature map of the environment is generated and then shared with the detached camera. The 3D feature map is used to re-localize the integrated camera so that its 6DOF pose is determined. The detached camera also re-localizes itself based on the 3D feature map so that its 6DOF pose is also determined. Then, the embodiments receive (i) an image of the environment by the detached camera and (ii) the 6DOF pose of the detached camera. A depth map of the environment is accessed. An overlay image is generated by re-projecting the perspective of the image of the detached camera to align with the perspective of the image of the integrated camera and by superimposing at least a portion of the re-projected image of the detached camera onto the image of the integrated camera. Note that (i) the 6DOF pose of the integrated camera, (ii) the 6DOF pose of the detached camera, and (iii) the depth map are used to perform the re-projection process. Optionally, some embodiments additionally perform a parallax correction on the overlay image and then display the overlay image.

[0031] Examples of technical benefits, improvements, and practical applications

[0032] The following sections outline some exemplary improvements and practical applications provided by the disclosed embodiments. However, it will be appreciated that these are merely examples and that the embodiments are not limited to only these improvements.

[0033] The disclosed embodiments provide substantial improvements, benefits and practical applications to the technical field. By way of example, the disclosed embodiments improve how images are generated and displayed and how image content is aligned.

[0034] In other words, the embodiments address the problem of aligning image content from a remote or detached camera image with image content from an integrated camera image to create a single composite or overlay image. Note that the overlay image is generated without the need to use timestamp data, but rather by determining the 6DOF poses of both camera systems using a 3D feature map. By having the 6DOF poses from both the remote camera system and the HMD, and with an understanding of the scene geometry, the disclosed embodiments are able to provide accurate image overlay between the remote camera system and the HMD, while accounting for the physical separation and different orientations. Once the poses are determined, the embodiments are able to beneficially re-project the image of the detached camera in a manner that aligns its perspective with the perspective of the image of the integrated camera. After the re-projection occurs, the image of the detached camera can then be overlaid onto the image of the integrated camera to form an overlay image. In this regard, the disclosed embodiments address problems related to image alignment when images are generated by separate cameras, and when both left and right pass-through images are desired despite only a single detached camera image being generated. By performing the disclosed operations, the embodiments are able to significantly improve image quality and image display.

[0035] Integrated camera and detached camera

[0036] FIG. 1 An example environment 100 in which the HMD 105 is operating is shown. As in FIG. 2 and FIG. 3 The HMD 105 can be configured in a variety of different ways, as illustrated in

[0037] By way of example, FIG. 1 The HMD 105 of FIG. 2 The HMD 200. The HMD 200 can be any type of MR system 200A, including a VR system 200B or an AR system 200C. It should be noted that while much of this disclosure is focused on the use of HMDs, embodiments are not limited to practicing using only HMDs. In other words, any type of scanning system can be used, even systems that are completely removed or separate from an HMD. As such, the disclosed principles should be broadly interpreted to encompass any type of scanning scenario or device. Some embodiments can not even actively use the scanning device itself, and can simply use data generated by the scanning device. For example, some embodiments can be practiced at least partially in a cloud computing environment.

[0038] The HMD 200 is shown as including scanning sensor(s) 205 (i.e., a type of scanning or camera system), and the HMD 200 can use the scanning sensor(s) 205 to scan the environment, map the environment, capture environment data, and / or generate any kind of image of the environment (e.g., by generating a 3D representation of the environment or by generating a "pass-through" visualization). The scanning sensor(s) 205 can include any number or type of scanning device without limitation.

[0039] According to the disclosed embodiments, the HMD 200 can be used to generate a parallax-corrected pass-through visualization of the user's environment. In some cases, the "pass-through" visualization refers to a visualization that reflects what the user would see if the user were not wearing the HMD 200, regardless of whether the HMD 200 is included as part of an AR system or a VR system. In other cases, the pass-through visualization reflects a different or novel perspective.

[0040] To generate this pass-through visualization, the HMD 200 can use its scanning sensor(s) 205 to scan, map, or otherwise record its surroundings, including any objects in the environment, and pass this data to the user for viewing. In many cases, the pass-through data is modified to reflect or correspond to the perspective of the user's pupils, although other perspectives can also be reflected by the image. The perspective can be determined by any type of eye tracking technology or other data.

[0041] To convert raw images into pass-through images, one or more scanning sensors 205 typically rely on their cameras (e.g., head-tracking cameras, hand-tracking cameras, depth cameras, or any other type of camera) to acquire one or more raw images (aka texture images) of the environment. In addition to generating pass-through images, these raw images can also be used to determine depth data, which details the distances from the sensor to any objects captured by the raw images (e.g., z-axis ranges or measurements). Once these raw images are acquired, depth maps can be computed (e.g., based on pixel differences) from the depth data embedded in or included in the raw images, and pass-through images (e.g., one per pupil) can be generated using the depth maps for any reprojection. In some cases, depth maps can be evaluated using 3D sensing systems, including time-of-flight, stereo, active stereo, or structured light systems. Furthermore, head-tracking cameras can be used to perform evaluations of visual maps of the surrounding environment, and these head-tracking cameras typically have stereo overlay regions to evaluate 3D geometry and generate environment maps. It is also worth noting that remote camera systems often have similar “head-tracking cameras” for identifying their position in 3D space.

[0042] As used herein, a "depth map" details the positional relationships and depth of objects relative to each other. Therefore, it is possible to determine the positional arrangement, orientation, geometry, contours, and depth of objects relative to each other. Based on the depth map, a 3D representation of the environment can be generated.

[0043] Accordingly, based on the described pass-through visualization, the user will be able to perceive the content currently in his / her environment without having to remove or reposition the HMD 200. Furthermore, as will be described in more detail later, the disclosed pass-through visualization will also enhance the user's ability to view objects within his / her environment (e.g., by displaying additional environmental conditions or image data that may be imperceptible to the human eye).

[0044] It should be noted that although much of this disclosure focuses on generating a “single” pass-through (or overlay) image, embodiments can generate separate pass-through images for each eye in a user’s eye. In other words, two pass-through images are typically generated simultaneously with each other. Therefore, although it is often mentioned that the generation appears to be a single pass-through image, embodiments are actually capable of generating multiple pass-through images simultaneously.

[0045] In some embodiments, the scanning sensor(s) 205 includes one or more visible light cameras 210, one or more low-light cameras 215, and one or more thermal imaging cameras 220, potentially (although not necessarily, as in...) FIG. 2The dashed box in FIG. 2B illustrates the ultraviolet (UV) camera(s) 225 and potentially (though not necessarily) the point illuminator (not shown). The ellipsis 230 illustrates how any other type of camera or camera system (e.g., depth camera, time-of-flight camera, virtual camera, depth laser, etc.) can be included in the scanning sensor(s) 205.

[0046] As an example, a camera configured to detect mid-infrared wavelengths can be included within the scanning sensor(s) 205. As another example, any number of virtual cameras that are re-projecting from actual cameras can be included in the scanning sensor(s) 205 and can be used to generate stereoscopic image pairs. In this way and as will be discussed in greater detail later, the scanning sensor(s) 205 can be used to generate stereoscopic image pairs. In some cases, the stereoscopic image pairs can be obtained or generated as a result of performing any one or more of the following: active stereoscopic image generation via the use of two cameras and one point illuminator; passive stereoscopic image generation by using two cameras; image generation via the use of one actual camera, one virtual camera, and one point illuminator using structured light; or image generation using a time-of-flight (TOF) sensor, where a baseline exists between a depth laser and a corresponding camera, and where the field of view (FOV) of the corresponding camera is offset relative to the illumination field of the depth laser.

[0047] Generally speaking, the human eye is capable of perceiving light within the so-called "visible spectrum," which includes light (or electromagnetic radiation) having wavelengths from about 380 nanometers (nm) to about 740 nm. As used herein, the visible light camera(s) 210 include two or more black-and-white cameras configured to capture photons within the visible spectrum. Typically, these black-and-white cameras are complementary metal-oxide-semiconductor (CMOS) type cameras, although other camera types (e.g., charge-coupled device, CCD) can be used. These black-and-white cameras can also be extended to the NIR range (up to 1100 nm).

[0048] The black and white cameras are typically stereo cameras, meaning that the fields of view of two or more black and white cameras at least partially overlap one another. With this overlap region, images generated by the visible light camera(s) 210 can be used to identify differences between certain pixels that collectively represent objects captured by the two images. Based on these pixel differences, embodiments can determine the depth of objects located within the overlap region (i.e., "stereo depth matching" or "stereo depth matching"). As such, the visible light camera(s) 210 can be used not only to generate a pass-through visualization, but also to determine object depths. In some embodiments, the visible light camera(s) 210 can capture both visible light and IR light.

[0049] The low light camera(s) 215 are configured to capture visible light and IR light. IR light is often segmented into three different classifications including near IR, mid-IR, and far IR (e.g., thermal IR). The classifications are determined based on the energy of the IR light. For example, near IR has a relatively high amount of energy due to having a relatively short wavelength (e.g., between about 750 nm and about 1,000 nm). In contrast, far IR has a relatively low amount of energy due to having a relatively long wavelength (e.g., up to about 30,000 nm). Mid-IR has an energy value that is between or intermediate to the near IR and far IR ranges. The low light camera(s) 215 are configured to detect or be sensitive to IR light at least within the near IR range.

[0050] In some embodiments, the visible light camera(s) 210 and the low light camera(s) 215 (a.k.a., low light night vision cameras) operate within approximately the same overlapping wavelength range. In some cases, this overlapping wavelength range is between about 400 nanometers to about 1,000 nanometers. Additionally, in some embodiments, both types of cameras are silicon detectors.

[0051] One distinguishing feature between the two types of cameras relates to the illumination conditions or illumination range(s) in which they are actively operated. In some cases, the visible light camera(s) 210 are low power cameras and operate in environments where the illumination is between about dusk illumination (e.g., about 10 lux) and bright midday sun illumination (e.g., about 100,000 lux), or, the illumination range starts at about 10 lux and increases beyond 10 lux. In contrast, the low light camera(s) 215 consume more power and operate in environments where the illumination range is between about starlight illumination (e.g., about 1 milli-lux) and dusk illumination (e.g., about 10 lux).

[0052] On the other hand, the thermal imaging camera(s) 220 are configured to detect electromagnetic radiation or IR light in the far-IR (i.e., thermal-IR) range, although some embodiments also enable the thermal imaging camera(s) 220 to detect radiation in the mid-IR range. To clarify, the thermal imaging camera(s) 220 can be long-wave infrared imaging cameras configured to detect electromagnetic radiation by measuring long-wave infrared wavelengths. Typically, the thermal imaging camera(s) 220 detect IR radiation with wavelengths between about 8 microns and 14 microns to detect blackbody radiation from the environment and people in the camera’s field of view. Because the thermal imaging camera(s) 220 detect far-IR radiation, the thermal imaging camera(s) 220 are able to operate under any illumination conditions, without limitation.

[0053] In some cases (although not all), the thermal imaging camera(s) 220 include a non-cooled thermal imaging sensor. Non-cooled thermal imaging sensors use a particular type of detector design based on a microbolometer array, which is a device that measures the amplitude or power of incident electromagnetic waves / radiation. To measure the radiation, the microbolometer uses a thin layer of an absorbing material (e.g., metal) that is connected to a thermal reservoir through a thermal chain. Incident waves strike and heat the material. In response to heating the material, the microbolometer detects a temperature-dependent electrical resistance. Changes in ambient temperature cause changes in the bolometer temperature, and these changes can be converted into an electrical signal, thereby producing a thermal image of the environment. According to at least some of the disclosed embodiments, a non-cooled thermal imaging sensor is used to generate any number of thermal images. The bolometers of a non-cooled thermal imaging sensor are able to detect electromagnetic radiation across a wide spectrum, spanning the mid-IR spectrum, the far-IR spectrum, and even up to millimeter-sized waves.

[0054] The UV camera(s) 225 are configured to capture light in the UV range. The UV range includes electromagnetic radiation with wavelengths between about 150 nm and about 400 nm. The disclosed UV camera(s) 225 should be interpreted broadly and can operate in a manner that includes both reflective UV photography and UV-induced fluorescence photography.

[0055] Thus, as used herein, a "visible light camera" (including a "head tracking camera") is a camera that is primarily used for computer vision to perform head tracking. These cameras are capable of detecting visible light, or even a combination of visible light and IR light (e.g., IR light in the range including wavelengths of about 850 nm). In some cases, these cameras are global shutter devices with pixel sizes of about 3 pm. On the other hand, low light cameras are cameras that are sensitive to visible light and near IR. These cameras are larger, and can have pixel sizes of about 8 pm or larger. These cameras are also sensitive to the wavelengths of silicon sensors, which are between about 350 nm and 1100 nm. These sensors can also be manufactured with III-V materials to have optical sensitivity to NIR wavelengths. Thermal / long wavelength IR devices (i.e., thermal imaging cameras) have pixel sizes of about 10 pm or larger, and detect heat radiating from the environment. These cameras are sensitive to wavelengths in the range of 8 pm to 14 pm. Some embodiments also include mid-IR cameras that are configured to detect at least mid-IR light. These cameras often include non-silicon materials (e.g., InP or InGaAs) that detect light in the 800 nm to 2 pm wavelength range.

[0056] Thus, the disclosed embodiments can be structured to utilize many different camera types. Different camera types include, but are not limited to: visible light cameras, low light cameras, thermal imaging cameras, and UV cameras. Stereoscopic depth matching can be performed using images generated from any of the above-listed camera types, or a combination of types.

[0057] Generally, the low light camera(s) 215, the thermal imaging camera(s) 220, and the UV camera(s) 225 (if present) consume more power than the visible light camera(s) 210. Thus, when not in use, the low light camera(s) 215, the thermal imaging camera(s) 220, and the UV camera(s) 225 are generally in a powered-off state, where those cameras are either turned off (and thus do not consume power), or in a reduced operability mode (and thus consume much less power than when those cameras are fully operational). In contrast, the visible light camera(s) 210 are generally in a powered-on state, where those cameras are fully operational by default.

[0058] It should be noted that any number of cameras can be provided on the HMD 200 for each of the different camera types. In other words, the visible light camera(s) 210 can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 cameras. However, the number of cameras is typically at least 2, so the HMD 200 can perform stereo depth matching, as described previously. Similarly, the low-light camera(s) 215, the thermal camera(s) 220, and the UV camera(s) 225 can each include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 corresponding cameras.

[0059] FIG. 3 FIGURE 1 illustrates an example HMD 300, which represents the HMD 200 from FIG. 2 FIGURE 1 illustrates an example HMD 300, which represents the HMD 200 from FIG. 2 FIGURE 1 illustrates an example HMD 300, which represents the HMD 200 from FIG. 3 Although only 5 cameras are illustrated in FIGURE 1, the HMD 300 can include more or less than 5 cameras.

[0060] In some cases, the cameras can be located at particular positions on the HMD 300. For example, in some cases, a first camera (e.g., which can be the camera 320) is disposed on the HMD 300 at a position above a specified left eye position of any user wearing the HMD 300 relative to a height direction of the HMD. For example, the camera 320 is located above the pupil 330. As another example, the first camera (e.g., the camera 320) is additionally located above the specified left eye position relative to a width direction of the HMD. In other words, the camera 320 is not only located above the pupil 330, but is also in line with the pupil 330. The camera can be placed directly in front of the specified left eye position when using the VR system. For example, with reference to FIGURE 1, the camera can be physically disposed at a position on the HMD 300 in front of the pupil 330 in the z-axis direction. FIG. 3

[0061] ​When a second camera (e.g., possibly camera 310) is provided, the second camera can be disposed on the HMD at a position above a specified right eye position of any user wearing the HMD in the height direction of the HMD. For example, camera 310 is above pupil 335. In some cases, the second camera is additionally above the specified right eye position in the width direction of the HMD. When using the VR system, the camera can be placed directly in front of the specified right eye position. For example, referring to FIG. 3 , the camera can be physically disposed on the HMD 300 at a position in front of pupil 335 in the z-axis direction.

[0062] When a user wears HMD 300, HMD 300 fits on the user's head, and the display of HMD 300 is in front of the user's pupils, such as pupil 330 and pupil 335. Typically, cameras 305-325 will be physically offset some distance from the user's pupils 330 and 335. For example, there can be a vertical offset in the HMD height direction (i.e., the "Y" axis), as shown by offset 340. Similarly, there can be a horizontal offset in the HMD width direction (i.e., the "X" axis), as shown by offset 345.

[0063] As previously described, HMD 300 is configured to provide pass-through image(s) for viewing by a user of HMD 300. In this way, HMD 300 is able to provide a visualization of the real world without requiring the user to remove or reposition HMD 300. These pass-through image(s) effectively represent the same view that the user would see if the user were not wearing HMD 300. Cameras 305-325 are used to provide these pass-through image(s).

[0064] However, none of cameras 305-325 are telecentrically aligned with pupils 330 and 335. Offsets 340 and 345 actually introduce differences in perspective between cameras 305-325 and pupils 330 and 335. These differences in perspective are referred to as "parallax."

[0065] As a result of offsets 340 and 345, the raw images (a.k.a., texture images) produced by cameras 305-325 can not be immediately usable as pass-through images. Instead, it is beneficial to perform parallax correction (a.k.a., image compositing) on the raw images to transform the perspectives contained within those raw images to correspond to the perspectives of user pupils 330 and 335. Parallax correction includes any number of corrections, which will be discussed in more detail later.

[0066] Returning to FIG. 1HMD 105 is shown to include a set of stereo camera pair 110, which includes a first camera 115 and a second camera 120, which represent the cameras mentioned in FIG. 2 and FIG. 3 Additionally, both the first camera 115 and the second camera 120 are integral parts of the HMD 105, so the first camera 115 can be considered an integrated camera, and the second camera 120 can also be considered an integrated camera.

[0067] FIG. 1 A detached camera 125 is also shown. Note that the detached camera 125 is physically unattached from the HMD 105, such that it can move independently of any motion of the HMD 105. Furthermore, the detached camera 125 is separated from the HMD 105 by a distance 130. This distance 130 can be any distance, but is typically less than 1.5 meters (i.e., the distance 130 is at most 1.5 meters).

[0068] In this example, the various different cameras are used in a scenario where objects in the environment 100 are relatively far away from the HMD 105 (as shown by a distance 135). The relationship between the distance 135 and the distance 130 will be discussed in more detail later. However, typically the distance 135 is at least 3 meters.

[0069] In any case, the first camera 115 captures images of the environment 100 from a first perspective 140. Similarly, the second camera 120 captures images of the environment 100 from a second perspective 145, and the detached camera 125 captures images of the environment 100 from a third perspective 150.

[0070] In cases involving the use of both integrated and detached cameras, it can be beneficial to superimpose the images of the detached camera onto the images of the integrated camera in order to generate a superimposed image. In order to provide a highly accurate superimposition between the two images, it is beneficial to first determine the 6 degrees of freedom (6DOF) pose of each camera, and then use this pose information (along with depth information) to re-project the images of the detached camera to a perspective that matches or is consistent with the perspective of the integrated camera. After the perspectives are aligned with each other, the images (or at least a portion thereof) of the detached camera can be superimposed onto the images of the integrated camera to generate a superimposed pass-through image. Thus, the remainder of this disclosure will present various techniques for using 6DOF pose information to align and stabilize image content between two separate cameras.

[0071] Exemplary method

[0072] The following discussion now will focus on numerous methods and method acts concerning the performance of the disclosed methods. Although the method acts can be discussed in a certain order, flowchart diagrams may be construed to represent certain methodologies or other possibilities by virtue of their nature(s). Particularly, unless otherwise noted, acts described herein can occur in any order. Furthermore, those skilled in the art will recognize that symbolic representations of acts as used in the flowcharts can be labeled using virtually any description, terminology, number, and / or letters.

[0073] Attention is now directed towards FIG. 4A and FIG. 4B which illustrates a flowchart of an example method 400 for aligning and stabilizing images generated by an integrated camera physically mounted to a head-mounted device (HMD) with images generated by a detached camera physically unmounted from the HMD. For example, the HMD in method 400 can be any of the HMDs discussed thus far (e.g., the HMD 105 from FIG. 1 , such that the method 400 can be performed by the HMD 105. Similarly, the so-called integrated camera can be either of the first camera 115 or the second camera 120 from FIG. 1 , and the detached camera can be the detached camera 125.

[0074] In some cases, the integrated camera is one camera selected from a group of cameras including a visible light camera, a low-light camera, or a thermal imaging camera, while the detached camera is also one camera selected from the group of cameras. In some cases, both the detached camera and the integrated camera have the same modality (e.g., both are thermal imaging cameras, or both are low-light cameras, etc.).

[0075] Initially, the method 400 includes an act of generating a three-dimensional (3D) feature map of an environment in which both the HMD and the detached camera operate (act 405). For example, the environment 100 from FIG. 1 may be the environment referred to in act 405. To generate the 3D feature map referred to in act 405, embodiments first perform a scan of the environment, as shown in FIG. 5

[0076] FIG. 5 An environment 500 is shown, which represents the environment 100 from FIG. 1 FIG. 5 ​​Also shown is HMD 505, which represents the HMDs discussed thus far, and in particular, the HMD referred to in act 405. In this example scenario, HMD 505 is performing a scan of environment 500 using its camera (e.g., possibly an integrated camera or possibly any one or more other cameras included on HMD 505), as shown by scan 510, scan 515, and scan 520 (e.g., HMD 505 is aligning different areas of environment 500). For example, HMD 505 can utilize its head tracking camera in order to perform the scans.

[0077] As a result of performing the scans, HMD 505 is able to generate a 3D feature map of environment 500, as shown in FIG. 6 . In particular, FIG. 6 shown is 3D feature map 600, which represents the 3D feature map referred to in act 405 of FIG. 4A . In FIG. 6 each of the black circles illustrated in . In

[0078] Generally, a "feature point" (e.g., any of feature points 605-615) refers to a discrete and identifiable point included within an object or image. Examples of feature points include corners, edges, or other geometric outlines that stand in sharp contrast to other areas of the environment. The black circles shown in FIG. 6 correspond to corners where walls meet or corners that form the corners of a table, which are considered feature points. Although only a few feature points are illustrated in FIG. 6 , one will appreciate that embodiments are able to identify any number of feature points in an image.

[0079] Identifying feature points can be performed using any type of image analysis, image segmentation, or even machine learning (ML). Any type of ML algorithm, model, or machine learning can be used to identify feature points. As used herein, a reference to a "machine learning" or ML model can include any type of machine learning algorithm or device, neural network (e.g., convolutional neural network(s), multi-layer neural network(s), recurrent neural network(s), deep neural network(s), dynamic neural network(s), etc.), decision tree model(s) (e.g., decision trees, random forests, and gradient boosted trees), linear regression model(s) or logistic regression model(s), support vector machine(s) ("SVMs"), artificial intelligence device(s), or any other type of intelligent computing system. Any amount of training data can be used (and possibly later refined) to train the machine learning algorithm to dynamically perform the disclosed operations.

[0080] In a general sense, the 3D feature map 600 is an assembly of a set of fused sparse depth maps that have been acquired over time. These depth maps identify 3D depth information as well as feature points. The set or fusion of these depth maps constitutes the 3D feature map 600.

[0081] Sharing and using 3D feature maps

[0082] Returning to FIG. 4A After the 3D feature map has been generated, the method 400 then includes an act of sharing the 3D feature map with a separate camera (act 410). FIG. 7 is a representation of this method act 410.

[0083] In particular, FIG. 7 An HMD 700 and a separate camera 705 are shown. The HMD 700 represents the HMD 105 from FIG. 1 , and the separate camera 705 represents the separate camera 125. A wideband radio connection 710 exists between the HMD 700 and the separate camera 705 to enable information to be quickly transmitted back and forth between the HMD 700 and the separate camera 705. The wideband radio connection 710 is a high-speed connection with high bandwidth availability.

[0084] As described in the method act 410, the HMD 700 can use the wideband radio connection 710 to transmit a 3D feature map 715 representing the 3D feature map 600 of FIG. 6 to the separate camera 705. In this regard, the separate camera 705 receives the 3D feature map 715 from the HMD 700. In other words, the process of sharing the 3D feature map with the separate camera can be performed by transmitting the 3D feature map to the separate camera via the wideband radio connection 710.

[0085] Before, during, or even possibly after the HMD shares the 3D feature map with the detached camera, both the integrated camera and the detached camera generate images of the environment, as shown in FIG. 8 Specifically, FIG. 8 An environment 800, an integrated camera 805, and a detached camera 810 are shown, all of which represent the environment, integrated camera, and detached camera, respectively, discussed herein.

[0086] FIG. 8 The integrated camera 805 is shown as having a field of view (FOV) 815 and as performing image capture in order to generate a first image 820. Similarly, the detached camera 810 has a FOV 825 and is performing image capture in order to generate a second image 830. The FOV of a camera generally refers to the area that the camera can observe. Here, the size of the FOV 815 is different than the size of the FOV 825. In some cases, the FOVs can be the same. In this case, the FOV 815 is larger than the FOV 825. In other cases, the FOV 825 can be larger than the FOV 815. Despite the difference in size of the FOVs, the resulting images can have the same resolution. This aspect will be discussed in more detail later.

[0087] In some implementations, the overall architecture includes computer vision (CV) visible light (VL) cameras on the detached system. These CV VL cameras are used to identify markers in the scene and reposition the location of the device in the shared map from the HMD. The FOV of the detached camera CV VL cameras is typically much larger than the main imaging cameras used in the remote camera system.

[0088] The two image capture processes can be performed simultaneously with respect to each other, or alternatively, there can be no temporal correlation. In some cases, the image capture process of the integrated camera at least partially overlaps in time with the image capture process of the detached camera, while in other cases there can be no overlap in time. Regardless, the integrated camera 805 generates a first image 820 and the detached camera 810 generates a second image 830. Notably, at least a portion of the FOVs of the two cameras overlap, such that at least a portion of the second image 830 overlaps with at least a portion of the first image 820.

[0089] By way of additional illustration, the dashed circle illustrated in FIG. 8 corresponds to the FOV 825 of the detached camera, while the rounded dashed rectangle corresponds to the FOV 815 of the integrated camera. In this exemplary scenario, the FOV 815 of the integrated camera completely consumes or encloses the FOV 825 of the detached camera.

[0090] Returning to FIG. 4A, method 400 includes an act (act 415) of relocalizing a position frame of the integrated camera using the 3D feature map based on a first image (e.g., FIG. 8 first image 820) generated by the integrated camera, thereby determining a 6 degrees of freedom (6DOF) pose of the integrated camera. Act 415 can be performed before, during, or even after act 410 (i.e., the act of sharing the 3D feature map).

[0091] In addition, method 400 includes an act (act 420) of relocalizing a position frame of the detached camera using the 3D feature map based on a second image (e.g., FIG. 8 second image 830) generated by the detached camera, thereby determining a 6DOF pose of the detached camera. Method act 420 is performed after act 410, but act 420 can be performed before, during, or even after act 415. In other words, in some cases, the detached camera and the integrated camera can (i) perform the relocalization process at the same time, (ii) during overlapping time periods, or (iii) during non-overlapping time periods. FIG. 9 The meaning of relocalization is more fully elucidated.

[0092] In general, relocalization refers to a process of determining a 6DOF pose of a camera relative to an environment so as to enable the camera to rely on a baseline coordinate system for that environment. In the context of the detached camera and the integrated camera, the detached camera is able to receive the 3D feature map from the HMD. Based on images of the detached camera (i.e., from FIG. 8 second image 830), the detached camera is able to identify feature points within the second image and correlate those feature points with feature points identified in the 3D feature map. Once those correlations are identified, the detached camera obtains or generates an understanding of the scene or environment geometry. The detached camera then determines or computes a geometric transform (e.g., a rotational transform) to determine the position of the detached camera relative to where the detected feature points are physically located (e.g., by determining a full 6 degrees of freedom (6DOF) pose).

[0093] In other words, relocalization refers to a process of matching feature points between a 3D feature map and an image, and then computing a geometric translation or transform to determine where the camera is physically located relative to the environment based on the 3D feature map and the current image. Performing relocalization enables both the detached camera and the integrated camera to rely on the same coordinate system. FIG. 9 The relocalization process performed by both the integrated camera and the detached camera is illustrated.

[0094] In particular, FIG. 9 A 3D feature map 900 and an image frame 905 are illustrated. The 3D feature map 900 represents FIG. 6the 3D feature map 600 and the other 3D feature maps discussed thus far. If the integrated camera is performing the relocalization process, the image frame 905 corresponds to FIG. 8 the first image 820. On the other hand, if the detached camera is performing the relocalization process, the image frame 905 corresponds to the second image 830. Note that the integrated camera and the detached camera independently perform their respective relocalization processes, which are typically the same, and which are shown in FIG. 9 .

[0095] The 3D feature map 900 and the image frame 905 are fed as inputs into a relocalization 910 operation. The relocalization 910 operation relocalizes the position frame 915 of the camera (e.g., the integrated camera or the detached camera) based on the correspondence between the feature points detected in the image frame 905 and the feature points contained in the 3D feature map 900. Simultaneous localization and mapping (SLAM) (e.g., SLAM 920) techniques can also be used to relocalize the camera system within the same physical space (i.e., the HMD space). SLAM techniques use images from a camera to make a map as a frame of reference for the physical system.

[0096] Historically, SLAM techniques have been used to allow multiple users to visualize holographic content in a scene. The disclosed embodiments can be configured to use SLAM to relocalize the position of the remote camera (i.e., the detached camera) relative to the camera mounted with the HMD (i.e., the integrated camera). By using SLAM from the remote camera system and the HMD-mounted system, the embodiments are able to determine the relative and absolute positions of the two camera systems. Thus, the result of the relocalization 910 operation is the 6DOF pose 925 of the camera (e.g., the detached camera and the separate integrated camera). By determining the 6DOF pose 925, the embodiments enable the two camera systems to effectively operate using the same coordinate system 930. By 6DOF pose 925, it is meant that the embodiments are able to determine the angular displacement (e.g., yaw, pitch, roll) and the translational displacement (e.g., forward / backward, left / right, and up / down) of the camera in the environment.

[0097] Thus, the embodiments are able to use the 3D feature map to relocalize the position frame of the integrated camera into the HMD physical space. This relocalization process is performed by identifying the feature points in the first image and the feature points in the 3D feature map. The embodiments then attempt to establish a correlation or match between these two sets of feature points. Once a sufficient number of matches are made, the embodiments are able to use this information to determine the 6DOF pose of the integrated camera.

[0098] Similarly, the detached camera can use the 3D feature map to reposition its positional frame into the HMD space. This repositioning process is performed in the same manner. In other words, the detached camera identifies the feature points in the second image and the feature points in the 3D feature map. The detached camera then attempts to establish a correlation or match between these two sets of feature points. Because the FOV of the detached camera at least partially overlaps the FOV of the integrated camera, the second image should include at least several of the same feature points included in the first image. Thus, the detached camera can identify matches between the feature points, some of which are the same as detected in the first image, thereby enabling it to determine its 6DOF pose as well. In this regard, the detached camera can determine its 6DOF pose based at least in part on some of the same identified feature points used by the integrated camera to determine its 6DOF pose. As a result of having the detached camera use the 3D feature map to reposition the detached camera's positional frame (e.g., into the HMD space), the detached camera will then be able to use the same coordinate system as the integrated camera.

[0099] In other words, the detached camera and the integrated camera compute a rotation base matrix that details the angular and translational differences between the perspectives embodied in the respective images with respect to the environment (e.g., the feature points detected in the environment) and with respect to each other. In this regard, the rotation base matrix provides a mapping with respect to translational or angular movement to map the feature points detected in the images to the feature points included in the 3D feature map. The mapping enables the system to determine what translational and angular translations are needed to transition from the perspective of the first image to the perspective of the second image, and vice versa. The process of having the detached camera and the integrated camera use the 3D feature map to reposition their positional frames (e.g., into the HMD space) can include performing a simultaneous localization and mapping (SLAM) operation to determine the relative position between the detached camera and the integrated camera.

[0100] Returning to FIG. 4A , the method 400 then includes an act (act 425) in which the HMD receives (i) a second image of the environment from the detached camera and (ii) a 6DOF pose of the detached camera from the detached camera. Thus, the HMD now includes data detailing the 6DOF pose of the detached camera, the image of the detached camera (i.e., the second image), the 6DOF pose of the integrated camera, and the image of the integrated camera (i.e., the first image). FIG. 10 is an illustration of method act 425.

[0101] FIG. 10 The HMD 1000 and the detached camera 1005 are shown, each of which represents the corresponding item referred to herein. As in FIG. 7As described in the middle, there is a wideband radio connection 1010 between the HMD 1000 and the separate camera 1005. In this case, the separate camera 1005 is transmitting a 6DOF pose 1015 and a second image 1020 to the HMD 1000. Here, the 6DOF pose 1015 corresponds to the 6DOF pose 925 (when computed for the separate camera), and the second image 1020 corresponds to the second image 830 from FIG. 8 In some cases, the 6DOF pose 1015 and the second image 1020 can be transmitted using the same transmission burst, while in other cases, these two pieces of information can be transmitted in separate and independent transmission bursts.

[0102] Depth map

[0103] As a result of performing the method actions 405-425, the HMD now includes the information described in detail in FIG. 11 . In particular, the HMD 1100 includes the first image 1105, the 6DOF pose of the integrated camera 1110, the second image 1115, and the 6DOF pose of the separate camera 1120. Each of these elements corresponds to its respective element discussed herein. Additionally, the HMD 1100 is able to generate, access, or obtain a depth map 1125 of the environment. To clarify, as described in the method action 430 illustrated in FIG. 4B The method 400 includes an action of accessing a depth map of the environment (e.g., the depth map 1125) (action 430).

[0104] As used herein, a "depth map" describes in detail the positional relationships and depths relative to objects in an environment. Thus, the positional arrangement, positioning, geometry, contours, and depths of objects relative to each other can be determined. As shown in FIG. 12 , the depth map 1125 can be computed in different ways.

[0105] In particular, FIG. 12 The depth map 1200 of FIG. 12 represents the depth map 1125. In some cases, the depth map 1200 can be computed using a rangefinder 1205. In some cases, the depth map 1200 can be computed by performing stereo depth matching 1210. The ellipsis 1215 shows how the depth map 1200 can be computed using other techniques, and is not limited to the techniques described intwo of the two illustrated in FIG. 12. In some implementations, the depth map 1200 can be a complete and full depth map in which a corresponding depth value is assigned to each pixel in the depth map. In some implementations, the depth map 1200 can be a single-pixel depth map. In some implementations, the depth map can be a planar depth map in which each pixel in the depth map is assigned the same depth value. In any case, FIG. 11 The depth map 1125 of the second image 1115 represents one or more depths of objects located in the environment. Note that the depth of the center of the secondary camera can also be determined by the rangefinder / single-pixel measurement system. Embodiments can superimpose the two camera images based on the 6DOF pose plus the single-pixel depth information.

[0106] Returning to FIG. 11 If the first image 1105, the 6DOF pose 1110, the second image 1115, and the 6DOF pose 1120 are computed prior to subsequent movement of the integrated camera and / or the detached camera, embodiments can update those pieces of data using inertial measurement unit (IMU) data 1130 obtained from an IMU 1135. To clarify, the integrated camera can be associated with its own corresponding IMU, and the detached camera can be associated with its own corresponding IMU. Both IMUs can generate IMU data, as represented by the IMU data 1130. The detached camera can transmit its IMU data to the HMD.

[0107] If the previously described rotation / rotation fundamental matrix (computed during the relocalization process) is computed prior to subsequent movement of any integrated camera or detached camera, embodiments can update the respective rotation fundamental matrix with the IMU data 1130 to account for the new movement. For example, by multiplying the rotation fundamental matrix of the integrated camera with matrix data generated based on the IMU data 1130, embodiments can eliminate the effects of movement of the integrated camera. Similarly, by multiplying the rotation fundamental matrix of the detached camera with matrix data generated based on its corresponding IMU data, embodiments can eliminate the effects of movement of the detached camera. Thus, the 6DOF pose 1110 and the 6DOF pose 1120 can be updated based on subsequently obtained IMU data. In other words, embodiments can update the 6DOF pose of the integrated camera (or the detached camera) based on detected movement of the integrated camera (or the detached camera). The detected movement can be detected based on IMU data obtained from the IMU of the integrated camera (or the detached camera).

[0108] Generating overlay images

[0109] Returning to FIG. 4BThe method 400 then includes an act of generating an overlaid image by re-projecting the perspective of the second image to align with the perspective of the first image (act 435) (e.g., using the two 6DOF poses and the depth map discussed previously to perform the re-projection). After aligning the perspectives, embodiments overlay at least a portion (and possibly all) of the re-projected second image onto the first image. To clarify, the 6DOF pose of the integrated camera, the 6DOF pose of the detached camera, and the depth map are used to perform the re-projection operation. Of course, the image of the detached camera (i.e., the second image) is also used to perform the re-projection operation. FIG. 13 The re-projection operation is illustrated, where the perspective of the second image is re-projected so as to align, match, or coincide with the perspective of the first image. By performing this alignment, embodiments are then able to selectively overlay portions of the second image onto the first image while ensuring accurate alignment between the content of the two images.

[0110] FIG. 13 A second image 1300 is shown, which represents the second image discussed thus far. The second image includes a 2D keypoint 1305A and a corresponding 3D point 1310 for this 2D keypoint 1305A. After determining the intrinsic camera parameters 1315 (e.g., the focal length, principal point, and lens distortion of the camera) and the extrinsic camera parameters 1320A (e.g., the position and orientation of the camera, or the 6DOF pose of the camera), embodiments are able to perform a re-projection 1325 operation on the second image 1300 to re-project the perspective embodied by this image to a new perspective, where the new perspective matches the perspective of the first image (thus, the second image can then be accurately overlaid onto the first image).

[0111] For example, as a result of performing the re-projection 1325 operation, a re-projected image 1330 is generated, where the re-projected image 1330 includes a 2D keypoint 1305B that corresponds to the 2D keypoint 1305A. In effect, the re-projection 1325 operation results in a synthetic camera having new extrinsic camera parameters 1320B so as to give the appearance that the re-projected image 1330 was captured by the synthetic camera at the new perspective (e.g., at the same position as the integrated camera). In this regard, re-projecting the second image (or at least a portion of the second image) compensates for the distance (e.g., distance 130 from FIG. 1 the detached camera to the integrated camera) and also compensates for the pose or perspective difference between the two cameras.

[0112] Thus, embodiments re-project the second image to a new perspective so as to align the perspective of the second image with the perspective of the first image. Further details are illustrated in FIG. 14

[0113] ​FIG. 14 how the separate camera's 6DOF pose 1400, the second image 1405, the integrated camera's 6DOF pose 1410, and the depth map 1415 (i.e., the depth map 1125) are fed as inputs into the re-projection 1420 operation (i.e., the re-projection 1325 from FIG. 13 to produce the re-projected image 1425 (i.e., the re-projected image 1330 from FIG. 13 As a result of performing the re-projection 1420 operation, the perspective embodied by the re-projected image 1425 matches the perspective of the integrated camera (e.g., the first camera 115 or the second camera 120 from FIG. 1 In some cases, the disclosed operations are performed twice, where the operations are performed once for the first camera 115 and a second time for the second camera 120 in order to produce two separate pass-through images.

[0114] FIG. 14 the re-projected image 1425 is also illustrated as a re-projected image 1500 in FIG. 15 Now that the re-projected image 1500 has a perspective that corresponds to the perspective of the first image, embodiments are able to perform an overlay 1505 operation to generate an overlaid image 1510. To clarify, embodiments generate the overlaid image 1510 by merging or blending pixels from the first image (i.e., first image pixels 1515) with pixels from the re-projected image 1500 (i.e., second image pixels 1520). In other words, one or more portions from the re-projected image 1500 are overlaid onto the first image to form the overlaid image 1510. As a result of performing the earlier re-projection operation on the second image, the second image pixels 1520 are properly aligned with the first image pixels 1515 below.

[0115] For example, the re-projected image 1500 shows a man wearing a baseball cap and the back of a woman. The first image (see, e.g., the first image 1105 illustrated in FIG. 11 includes the same content. For many reasons, it is beneficial to overlay the second image content onto the first image content.

[0116] For example, because the size of the FOV of different cameras can be different, the size of the resulting images can also be different. Despite the size difference, the resolution can still be the same. For example, FIG. 11 illustrates how the second image 1115 is smaller than the first image 1105. Despite the size difference, the resolution can still be the same. Thus, each pixel included in the second image 1115 is smaller compared to each pixel in the first image 1105 and provides an increased level of detail.

[0117] Accordingly, in some embodiments, the resolution of the second image 1115 can be the same as the resolution of the first image 1105, such that as a result of the FOV of the second image 1115 being smaller than the FOV of the first image 1105, each pixel in the second image 1115 is smaller than each pixel in the first image 1105. Accordingly, the pixels of the second image 1115 will give a sharper, clearer, or crisper visualization of the content than the pixels of the first image 1105. Thus, by superimposing the second image content onto the first image content, the portion of the superimposed image 1510 contained within the bounds 1525 (corresponding to the second image content) can appear sharper or higher in detail than other portions of the superimposed image 1510 (e.g., those pixels corresponding to the first image content). Thus, by superimposing the content, an enhanced image can be generated. FIG. 15

[0118] Parallax correction

[0119] Returning to FIG. 4B , the method 400 includes an optional (as indicated by the dashed box) action of performing a parallax correction on the superimposed image to modify the perspective of the superimposed image to correspond to a new perspective (act 440). In some implementations (although not all), the new perspective is the perspective of the pupil (e.g., the pupil 330 or 335) of the user wearing the HMD. The method 400 includes another optional action of displaying the superimposed image for viewing by the user (act 445). FIG. 3

[0120] The computer system that implements the disclosed operations (including the method 400) can be a head mounted device (HMD) worn by a user. The new perspective can correspond to one of the left eye pupil or the right eye pupil. If a second superimposed image is generated, the second superimposed image can also be parallax corrected to a second new perspective, where the second new perspective can correspond to the other of the left eye pupil or the right eye pupil. FIG. 16 Some additional explanation is provided regarding the parallax correction operation.

[0121] FIG. 16 A superimposed image 1600 is shown, which can be the superimposed image 1510 from FIG. 15 , and which can be the superimposed image discussed in the method 400. Here, the superimposed image 1600 is shown as having an original perspective 1605. In accordance with the disclosed principles, embodiments are able to perform a parallax correction 1610 to transform the original perspective 1605 of the superimposed image 1600 to a new or novel perspective. It should be noted how the two separate re-projection operations are subsequently performed on the pixels taken from the split camera images, one involving modifying the perspective of the split camera images to coincide with the perspective of the integrated camera, and one involving modifying the perspective of the superimposed image to coincide with the perspective of the user's pupil. ​​

[0122] Performing parallax correction 1610 involves using a depth map to reproject image content to a new viewpoint. This depth map may be the same as or different from the previously mentioned depth map. In some cases, the depth map is an updated version of a previous depth map to reflect the current localization and pose of the HMD. In other cases, the depth map is a new depth map generated specifically to perform parallax correction.

[0123] Parallax correction 1610 is shown as including any one or more of a plurality of different operations. For example, parallax correction 1610 may involve distortion correction 1615 (e.g., correcting a concave or convex wide-angle or narrow-angle camera lens), epipolar transformation 1620 (e.g., parallelizing the optical axis of a camera), and / or reprojection transformation 1625 (e.g., repositioning the optical axis to be substantially in front of or in line with the user's pupil). Parallax correction 1610 includes performing depth calculations to determine the depth of the environment and then reprojecting the image to the determined location or with a determined viewpoint. As used herein, the phrases “parallax correction” and “image synthesis” are interchangeable and may include performing stereo through-parallax correction and / or image reprojection parallax correction.

[0124] The reprojection is based on the original viewpoint 1605 of the overlaid image 1600 relative to the surrounding environment. Based on the original viewpoint 1605 and the generated depth map, embodiments can correct parallax by reprojecting the viewpoint represented by the overlaid image to match the new viewpoint, as shown with parallax-corrected image 1630 and new viewpoint 1635. In some embodiments, the new viewpoint 1635 may be from... FIG. 3 One of the user's pupil sizes is 330 or 335.

[0125] Some embodiments perform three-dimensional (3D) geometric transformations on the overlay image to change the viewing angle of the overlay image in a manner related to the viewing angles of the user's pupils 330 and 335. Furthermore, the 3D geometric transformation relies on depth calculations, where objects in the HMD environment are mapped to determine their depth and viewing angle. Based on these depth calculations and viewing angles, embodiments are able to perform 3D reprojection or 3D warping of the overlay image in a manner that preserves the appearance of object depth in the parallax-corrected image 1630 (i.e., a type of pass-through image), wherein the preserved object depth substantially matches, corresponds to, or visualizes the actual depth of objects in the real world. Therefore, the degree or amount of parallax correction 1610 depends at least in part on the depth of the object in the parallax correction image 1610. FIG. 3 The degree or amount of offsets of 340 and 345.

[0126] By performing parallax correction 1610, the embodiment effectively creates a "virtual" camera positioned in front of the user's pupils 330 and 335. With further clarification, consider the... FIG. 3the camera 305 is currently positioned above and to the left of the pupil 335. By performing disparity correction, embodiments programmatically transform the images generated by the camera 305, or the perspective of those images, so that the perspective appears as if the camera 305 is actually positioned directly in front of the pupil 335. In other words, even though the camera 305 does not actually move, embodiments are able to transform the images generated by the camera 305 so that those images have the appearance that the camera 305 is located in front of the pupil 335.

[0127] In some cases, the disparity correction 1610 relies on a full depth map to perform the re-projection, while in other cases the disparity correction 1610 relies on a planar depth map to perform the re-projection. In some embodiments, the disparity correction 1610 relies on a one-pixel depth map (e.g., one-pixel depth measurements for each camera frame), such as a depth map generated by a one-pixel rangefinder.

[0128] When performing re-projection using a full depth map on the superimposed image, it is sometimes beneficial to attribute a single depth to all of the pixels defined by the dot circle in the disparity corrected image 1630. Failing to do so can result in a skew or warping of the disparity corrected region corresponding to the defined pixels. For example, instead of producing a circle of pixels, failing to use a single common depth for the pixels in the circle can result in an elliptical shape or other skewed effect. Thus, some embodiments determine a depth corresponding to the depth of a particular pixel (e.g., which can be the center pixel of the circle), and then attribute that single depth to all of the pixels defined by the circle. To clarify, the same depth value is given to all of the pixels defined by the circle.

[0129] The full depth map is then used to perform the re-projection involved in the disparity correction operations discussed earlier. By attributing the same depth to all of the pixels defined by the circle in the superimposed image, embodiments prevent skewing of the image content from occurring as a result of performing disparity correction.

[0130] While most embodiments select the depth corresponding to the center pixel, some embodiments can be configured to select the depth of a different pixel defined by the circle. As such, using the depth of the center pixel is merely one example implementation, but it is not the only implementation. Some embodiments select multiple pixels that are located at the center, and then use the average depth of those pixels. Some embodiments select the depth of a pixel or group of pixels that are offset from the center.

[0131] Instead of using a full depth map to perform the re-projection, some embodiments use a fixed depth map to perform a fixed depth map re-projection. In this case, embodiments select a depth for a particular pixel from the pixels defined by the circle (e.g., possibly again the center pixel). Based on the selected depth, embodiments then attribute that single depth to all pixels of the depth map to generate a fixed depth map. To clarify, all depth pixels in the fixed depth map are assigned or attributed to the same depth, which is the depth of the selected pixel (e.g., possibly the center pixel or possibly some other selected pixel).

[0132] Once the fixed depth map is generated, the depth map can then be used to perform a re-projection (e.g., a planar re-projection) on the overlay image using the fixed depth map. In this regard, the re-projection overlay image (e.g., overlay image 1600 in FIG. 16) can be performed using either a full depth map or a fixed depth map to generate a disparity corrected image 1630. FIG. 16

[0133] Accordingly, the disclosed embodiments are able to align images by performing a re-projection using 6DOF poses in order to align the images to have matching perspectives. Embodiments then perform a disparity correction on the aligned overlay images in order to generate a pass-through image having a new perspective. Such operations significantly enhance the quality of the images by enabling the display of new and dynamic image content.

[0134] Exemplary computer / computer system

[0135] Attention is now directed towards FIG. 17 , FIG. 17 FIG. 16 illustrates an exemplary computer system 1700 that can include and / or be used for performing any of the operations described herein. The computer system 1700 can take various different forms. For example, the computer system 1300 can be embodied as a tablet computer 1700A, a desktop or laptop computer 1700B, a wearable device 1700C (e.g., any of the disclosed HMDs), a mobile device, a standalone device, or any other embodiment as represented by the ellipsis 1700D. The computer system 1700 can also be a distributed system including one or more connected computing components / devices in communication with the computer system 1700.

[0136] In its most basic configuration, the computer system 1700 includes various different components. FIG. 17 FIG. 16 illustrates that the computer system 1700 includes one or more processors 1305 (also referred to as "hardware processing units"), scanning sensors 1710 (e.g., scanning sensors 205 of FIG. 2), an image processing engine 1715, and a storage device 1720. FIG. 2 FIG. 16 illustrates that the computer system 1700 includes one or more processors 1305 (also referred to as "hardware processing units"), scanning sensors 1710 (e.g., scanning sensors 205 of FIG. 2), an image processing engine 1715, and a storage device 1720. ​

[0137] With respect to the processor 1705, it is to be understood that the functions described herein can be performed, at least in part, by one or more hardware logic components (e.g., processor 1705). For example, and without limitation, illustrative types of hardware logic components / processors that can be used include field programmable gate arrays (“FPGAs”), program- or application-specific integrated circuits (“ASICs”), program- specific standard products (“ASSPs”), system-on-a-chip (“SOCs”), complex programmable logic devices (“CPLDs”), central processing units (“CPUs”), graphics processing units (“GPUs”), or any other type of programmable hardware.

[0138] The computer system 1700 and the scanning sensor 1710 can use any type of depth detection. Examples include, but are not limited to, stereo depth detection (active illumination (e.g., using a point illuminator), structured light illumination (e.g., 1 real camera, 1 virtual camera, and 1 point illuminator), and passive (i.e., no illumination)), time-of-flight depth detection (with a baseline between a laser and a camera, where the camera’s field of view does not completely overlap the laser’s illuminated area), rangefinder depth detection, or any other type of distance or depth detection.

[0139] The image processing engine 1315 can be configured to perform any of the method actions discussed in connection with FIG. 4A River FIG. 4B The image processing engine 1715, in some cases, includes ML algorithms. In other words, ML can also be leveraged by the disclosed embodiments, as previously described. The ML can be implemented as a specific processing unit (e.g., a special-purpose processing unit, as previously described) that is configured to perform one or more specialized operations of the computer system 1700. As used herein, the terms “executable module,” “executable component,” “component,” “module,” “model,” or “engine” can refer to a hardware processing unit or a software object, routine, or method that can be executed on the computer system 1700. The different components, modules, engines, models, and services described herein can be implemented as objects or processors (e.g., as separate threads) executing on the computer system 1700. The ML model and / or processor 1705 can be configured to perform one or more disclosed method actions or other functions.

[0140] The storage device 1720 can be physical system memory, which can be volatile, non-volatile, or some combination of the two. The term “memory” can also be used herein to refer to non-volatile mass storage devices such as physical storage media. If the computer system 1700 is distributed, the processing, memory, and / or storage capability can also be distributed.

[0141] The storage device 1720 is shown to include executable instructions (i.e., code 1325). The executable instructions represent instructions that are executable by the processor 1705 (or even the image processing engine 1715) of the computer system 1700 to perform the disclosed operations such as those described in the various methods.

[0142] The disclosed embodiments can include or utilize special-purpose or general-purpose computer(s) including computer hardware, such as one or more processors (e.g., processor 1705) and system memory (e.g., storage device 1720), as discussed in more detail below. Embodiments also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media that is accessible by a general -purpose or special-purpose computer system. Computer-readable media that store computer- executable instructions are “physical computer storage media.” A computer-readable medium that carries computer-executable instructions is a “transmission medium.” Thus, by way of example, and not limitation, current embodiments can comprise at least two distinctly different kinds of computer-readable media: computer storage media and transmission media.

[0143] Computer storage media (also called “hardware storage devices”) are computer- readable hardware storage devices, such as RAM, ROM, EEPROM, CD-ROM, solid state drives (“SSDs”) that are based on RAM, Flash memory, phase-change memory (“PCM”), or other types of memory, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other hardware storage devices that can be used to store desired program code means in the form of computer-executable instructions, data, or data structures and that can be accessed by a general -purpose or special-purpose computer.

[0144] The computer system 1700 can also be connected to external sensors (e.g., one or more remote cameras) or devices via a network 1730 (via wired or wireless connections). For example, the computer system 1700 can communicate with any number of devices or cloud services to obtain or process data. In some cases, the network 1730 itself can be a cloud network. Moreover, the computer system 1700 can also be connected to remote / separate computer systems over one or more wired or wireless networks 1730 that are configured to perform any of the processing described with respect to the computer system 1700.

[0145] A "network" as used herein is defined as one or more data links and / or data switches that enable the transport of electronic data between computer systems, modules, and / or other electronic devices. When information is transferred or provided over a network (either hardwired, wireless, or a combination of hardwired and wireless) to a computer, the computer properly views the connection as a transmission medium. The computer system 1700 will include one or more communication channels that are used to communicate with the network 1730. Transmission media include a network that can carry data in the form of computer-executable instructions or data structures, and that can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.

[0146] Program code portions can be downloaded to the computer system 1700 from the network 1730 and / or another computer system via the communication interface 1710 (for example, on a computer- readable medium). The other computer system can provide program code portions in the form of computer-executable instructions or data structures, or both. These computer- executable instructions and / or data structures can be accessed from computer- storage media and executed by a processor of the computer system 1700. Alternatively, the computer system 1700 can access one or more computer- storage media 1704 that can include computer-executable instructions or data structures, or both, used to program, software, or firmware, installed on or used in one or more computer systems. The computer system 1700 can be programmed to provide the functionality described herein using one or more of the computer-executable instructions or data structures.

[0147] Computer-executable (or computer-interpretable) instructions comprise, for example, instructions which cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. The computer-executable instructions can be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0148] Those skilled in the art will realize that the embodiments can be practiced in a network computing environment with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, pagers, routers, switches, and the like. The embodiments can also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, each perform tasks (e.g., cloud computing, cloud services, etc.). In a distributed system environment, program modules can be located in both local and remote memory storage devices.

[0149] The application can be implemented in other specific forms without departing from the spirit or essential characteristics thereof. The described embodiments are to be considered in all respects only as illustrative and not restrictive. The scope of the application is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning of and equivalency of the claims are to be embraced within their scope.

Claims

1. A head-mounted device (HMD) configured to align and stabilize images generated by an integrated camera physically mounted to the HMD with images generated by a separate camera physically unmounted from the HMD, the HMD comprising: one or more processors; and one or more computer-readable hardware storage devices storing instructions executable by the one or more processors to cause the HMD to at least: generate a three-dimensional (3D) feature map of an environment in which both the HMD and the separate camera operate; share the 3D feature map with the separate camera; use the 3D feature map to relocalize a position frame of the integrated camera based on a first image generated by the integrated camera, thereby determining a 6 degrees of freedom (6DOF) pose of the integrated camera; cause the separate camera to use the 3D feature map to relocalize a position frame of the separate camera based on a second image generated by the separate camera, thereby determining a 6DOF pose of the separate camera; receive, from the separate camera, (i) the second image of the environment and (ii) the 6DOF pose of the separate camera; access a depth map of the environment; and generate an overlay image by reprojecting a perspective of the second image to align with a perspective of the first image and by superimposing at least a portion of the reprojected second image onto the first image, wherein (i) the 6DOF pose of the integrated camera, (ii) the 6DOF pose of the separate camera, and (iii) the depth map are used to perform the reprojecting. the instructions are executable to further cause the HMD to display the overlay image, wherein one or more of the integrated camera and the separate camera are head tracking cameras configured to perform relocalization. the depth map is one pixel depth measurement per camera frame.

2. The HMD of claim 1, wherein, reprojecting the perspective of the second image to align with the perspective of the first image compensates for a distance the separate camera is separated from the integrated camera.

3. The HMD of claim 1, wherein, the integrated camera is one camera selected from a group of cameras comprising a visible light camera, a low light camera, or a thermal imaging camera, and wherein the separate camera is also one camera selected from the group of cameras.

4. The HMD of claim 1, wherein, the separate camera is separated from the integrated camera by a distance of at most 1.5 meters.

5. The HMD of claim 1, wherein, causing the separate camera to use the 3D feature map to relocalize the position frame of the separate camera comprises performing a simultaneous localization and mapping (SLAM) operation to determine a relative position between the separate camera and the integrated camera.

6. The HMD of claim 1, wherein, both the separate camera and the integrated camera are thermal imaging cameras.

7. The HMD of claim 1, wherein, the instructions are executable to further cause the HMD to apply a parallax correction to the overlay image to modify a perspective of the overlay image to correspond to a new perspective.

8. The HMD of claim 1, wherein, the instructions are executable to further cause the HMD to update the 6DOF pose of the integrated camera based on a detected movement of the integrated camera, the detected movement detected based on inertial measurement unit (IMU) data obtained from an IMU of the integrated camera.

9. The HMD of claim 1, wherein, ​ 10. The HMD of claim 1, wherein, ​ 11. A method for aligning and stabilizing images generated by an integrated camera physically mounted to a head-mounted device (HMD) with images generated by a separate camera physically unmounted from the HMD, the method comprising: generating a three-dimensional (3D) feature map of an environment in which both the HMD and the separate camera operate; sharing the 3D feature map with the separate camera; using the 3D feature map to relocalize a positional frame of the integrated camera based on a first image generated by the integrated camera, thereby determining a 6 degrees of freedom (6DOF) pose of the integrated camera; causing the separate camera to use the 3D feature map to relocalize a positional frame of the separate camera based on a second image generated by the separate camera, thereby determining a 6DOF pose of the separate camera; receiving, from the separate camera, (i) the second image of the environment and (ii) the 6DOF pose of the separate camera; accessing a depth map of the environment; and generating an overlay image by reprojecting a perspective of the second image to align with a perspective of the first image and by superimposing at least a portion of the reprojected second image onto the first image, wherein (i) the 6DOF pose of the integrated camera, (ii) the 6DOF pose of the separate camera, and (iii) the depth map are used to perform the reprojecting. The method further comprises updating the 6DOF pose of the separate camera based on a detected movement of the separate camera, the detected movement being detected based on inertial measurement unit (IMU) data obtained from an IMU of the separate camera.

12. The method of claim 11, wherein, Sharing the 3D feature map with the separate camera is performed by transmitting the 3D feature map to the separate camera via a wideband radio connection.

13. The method of claim 11, wherein, Using the 3D feature map to relocalize the positional frame of the integrated camera is performed by identifying feature points contained in the 3D feature map and by determining the 6DOF pose of the integrated camera based on the identified feature points.

14. The method of claim 11, wherein, Both the separate camera and the integrated camera are thermal imaging cameras.

15. The method of claim 14, wherein, ​

Citation Information

Patent Citations

  • Large-scale surface reconstruction that is robust against tracking and mapping errors

    CN105765631A

  • Camera pose estimation for mobile devices

    CN107646126A