Image alignment using interleaved feature extraction

By using interlaced feature extraction technology in mixed reality systems, re-using previously identified feature sets to align images, solving the problems of image alignment delay and inefficiency in multi-camera systems, achieving more efficient image alignment and enhanced user experience.

CN120112944APending Publication Date: 2025-06-06MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380074879.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-01
Filing Date
2023-09-26
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In mixed reality systems, there are challenges in the alignment of images generated using multiple cameras, especially when the images have less than a sufficient number of detectable or extractable features, resulting in alignment failure and delays.

Method used

By interleaved feature extraction technology, the previously identified feature set is reused to align the second image with the first image, avoiding delays in waiting for new image generation and feature extraction.

Benefits of technology

Improves the success rate and efficiency of image alignment, reduces latency, enhances the system's frame rate analysis capabilities, and reduces power usage and processor utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112944A_ABST
    Figure CN120112944A_ABST
Patent Text Reader

Abstract

Techniques for performing image alignment between a first image generated by a first camera and a second image generated by a second camera are disclosed. Image alignment is performed using interleaved feature extraction, wherein the set of features is reused to align the second image with the first image. A first set of features is identified from within the first image and a second set of features previously detected within the second image is accessed. The second set of features is previously used at least once to perform a previous image alignment operation. A current image alignment operation is performed to align the first image with the second image by using the first set of features and by re-using the second set of features.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Mixed reality (MR) systems, including virtual reality (VR) and augmented reality (AR) systems, have received a great deal of attention due to their ability to create truly unique experiences for their users. For reference, conventional VR systems create a fully immersive experience by limiting their users' views to only the virtual environment. This is often achieved by using a head mounted device (HMD) that completely blocks any view of the real world. As a result, the user is completely immersed in the virtual environment. In contrast, conventional AR systems create an augmented reality experience by visually presenting virtual objects that are placed in or interact with the real world.

[0002] As used herein, VR and AR systems are described and referenced interchangeably. Unless otherwise specified, the embodiments of this document are equally applicable to all types of MR systems, including AR systems, VR reality systems, and / or any other similar systems capable of displaying virtual content.

[0003] MR systems can also use different types of cameras to display content to users, such as in the form of passthrough images. When transitioning to and / or navigating in an MR environment, the passthrough images or views can help users avoid disorientation and / or safety hazards. MR systems can also provide enhanced data to enhance the user's real-world view. MR systems can present views captured by cameras in various ways. However, the process of providing a view of a real-world environment using images captured by a world-facing camera creates multiple challenges.

[0004] Some of these challenges arise when attempting to align image content from multiple cameras (such as an integrated "system camera" and a separate "external camera") when generating an overlay image to be displayed to a user. Challenges also arise when providing additional visualizations in the resulting overlay image, where these visualizations are designed to indicate the spatial relationship between the system camera and the external camera. Challenges may also arise if the image has an insufficient number of detectable or extractable features. Accordingly, aligning image content provides significant advantages, particularly in hologram placement and generation, and so these issues present a serious obstacle to the art. As such, there is a fundamental need in the art to improve how images are aligned with each other. The subject matter claimed herein is not limited to embodiments that address any disadvantages or that operate only in environments such as those described above. Instead, this background is provided merely to illustrate an exemplary technology area in which some of the embodiments described herein may be practiced. Summary of the invention

[0005] Embodiments disclosed herein relate to systems, devices (e.g., wearable HMDs, hardware storage devices, etc.), and methods for performing image alignment between a first image generated by a first camera and a second image generated by a second camera. The image alignment is performed using staggered feature extraction, where a set of features is reused to align the second image with the first image.

[0006] A first image generated by a first camera is accessed. An embodiment also accesses a second image generated by a second camera. The second image is generated before the time when the first image is generated. A first set of features is identified, wherein the features from within the first image are identified by performing feature extraction on the first image. An embodiment also accesses a second set of features previously detected within the second image. The second set of features was previously used at least once to perform a previous image alignment operation between the second image and a different image. Then, an embodiment performs a current image alignment operation by aligning the first image with the second image using the first set of features and by reusing the second set of features based on the identified correspondence between the first set of features and the second set of features. As a result of performing the current image alignment operation, a second portion of the second image is identified as corresponding to a first portion of the first image.

[0007] Some embodiments attempt to perform a current image alignment operation by using the first feature set and by reusing the second feature set to attempt to align the first image with the second image based on the identified correspondences between the first feature set and the second feature set. These embodiments then determine that the plurality of correspondences identified between the first feature set and the second feature set do not satisfy a correspondence threshold. As a result, the current image alignment operation fails. These embodiments then perform a new image alignment operation using a different arrangement of features associated with the image generated by the second camera.

[0008] Prior to accessing the second image, some embodiments determine not to wait until the current image is generated by the second camera before performing the current image alignment operation. This determination is based on a determination that waiting for the current image will introduce a delay in performing the current image alignment operation, and the introduced delay will exceed an allowable delay threshold. Therefore, rather than waiting for the current image, embodiments access a second feature set previously used in a previous image alignment operation.

[0009] This Summary is provided to introduce some concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0010] Additional features and advantages will be described in the following embodiments, and in part will be apparent from the embodiments, or can be learned by practicing the teachings herein. The features and advantages of the present invention can be realized and obtained by the means and combinations particularly pointed out in the appended claims. The features of the present invention will become more apparent from the following embodiments and the appended claims, or can be learned by practicing the present invention as described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to describe the manner in which the above-recited and other advantages and features can be obtained, a more particular implementation of the subject matter briefly described above will be presented by reference to specific embodiments illustrated in the accompanying drawings. Understanding that these drawings depict only typical embodiments and are therefore not to be considered limiting of scope, the embodiments will be described and explained with additional specificity and detail through the use of the accompanying drawings, in which:

[0012] Figure 1 An example head mounted device (HMD) configured to perform the disclosed operations is illustrated.

[0013] Figure 2 Another configuration of the HMD is shown.

[0014] Figure 3 Illustrated are example scenarios in which the disclosed principles may be practiced.

[0015] Figure 4 Another example scenario is illustrated.

[0016] Figure 5 It illustrates how a system camera (eg, a first camera) and an external camera (eg, a second or peripheral camera) may be used to perform the disclosed operations.

[0017] Figure 6 The diagram shows the field of view (FOV) of the system camera.

[0018] Figure 7 The FOV of the external camera is shown.

[0019] Figure 8 Superimposed and aligned images are illustrated, where image content from an external camera image is superimposed onto the system camera image.

[0020] Fig. 9 Another example scenario in which the principles can be practiced is illustrated.

[0021] Fig.10 Illustrate how to use the vision alignment process to overlay an external camera image onto a system camera image and how to display boundary elements in a manner that surrounds content from the external camera image.

[0022] Fig.11A , Fig. 11B and Fig. 11C Various aspects related to interleaved feature extraction are illustrated.

[0023] Fig.12 An example technique for aggregating multiple permutations of features is illustrated.

[0024] Fig.13A , Fig. 13B and Fig. 13C A flow diagram of an example method for performing interleaved feature extraction using single threading and / or multiple threads is illustrated.

[0025] Fig.14 A flow chart illustrating an example method for aggregating arrangements of multiple different extracted features is illustrated.

[0026] Fig.15 An example computer system is illustrated that may be configured to perform any of the disclosed operations. DETAILED DESCRIPTION

[0027] Embodiments disclosed herein relate to techniques for performing image alignment between a first image generated by a first camera and a second image generated by a second camera. The image alignment is performed using staggered feature extraction, wherein a set of features is reused to align the second image with the first image.

[0028] Accessing first and second images. Generating the second image before the first image. Identifying a first set of features from within the first image. An embodiment accesses a second set of features previously detected within the second image. The second set of features was previously used at least once to perform a previous image alignment operation. An embodiment performs a current image alignment operation to align the first image with the second image by using the first set of features and by reusing the second set of features. A second portion of the second image is identified as corresponding to a first portion of the first image.

[0029] Some embodiments attempt to align the first image with the second image based on the identified correspondences between the first and second feature sets by performing a current image alignment operation using the first feature set and by reusing the second feature set. These embodiments then determine that the plurality of correspondences identified between the two feature sets do not satisfy a correspondence threshold. As a result, the current image alignment operation fails. These embodiments then perform a new image alignment operation using a different arrangement of features associated with the image generated by the second camera.

[0030] Prior to accessing the second image, some embodiments determine not to wait until the current image is generated by the second camera before performing the current image alignment operation. This determination is based on a determination that waiting for the current image will introduce a delay in performing the current image alignment operation, and the introduced delay will exceed an allowable delay threshold. Therefore, rather than waiting for the current image, embodiments access a second feature set of the second (i.e., older) image.

[0031] Examples of technical advantages, improvements and practical applications

[0032] The following section summarizes some example improvements and practical applications provided by the disclosed embodiments. However, it should be understood that these are only examples and the embodiments are not limited to these improvements.

[0033] As described earlier, challenges arise when aligning image content from two different cameras. The disclosed embodiments address these challenges and provide solutions to these challenges.

[0034] Beneficially, these embodiments provide techniques that, when practiced, achieve improved success in aligning image content and performing feature matching. As a result of performing these operations, the user's experience is significantly improved, resulting in improvements in the technology. Improved image alignment and visualization are also achieved. Additional advantages include the ability to potentially increase the frame rate analysis of the system. In some cases, the frame rate can actually be doubled. On the other hand, by maintaining a lower frame rate, the embodiments can significantly reduce power usage and processor utilization because fewer frames are processed. As a result, the efficiency of the computer system itself can be improved and / or the power consumption of the system can be optimized. Accordingly, these and multiple other advantages will be described throughout the remainder of this disclosure.

[0035] Example MR Systems and HMDs

[0036] Now turn your attention to Figure 1 , which illustrates an example of a head mounted device (HMD) 100. The HMD 100 can be any type of MR system 100A, including a VR system 100B or an AR system 100C. It should be noted that although the basic part of the present disclosure focuses on the use of the HMD, the embodiments are not limited to being practiced using only the HMD. That is, any type of camera system can be used, and even a camera system that is completely removed or separated from the HMD can be used. As such, the disclosed principles should be broadly interpreted to cover any type of camera usage scenario. Some embodiments can even avoid actively using the camera itself and can simply use data generated by the camera. For example, some embodiments can be practiced at least in part in a cloud computing environment.

[0037] The HMD 100 is illustrated as including (multiple) scanning sensors 105 (i.e., a type of scanning or camera system), and the HMD 100 can use (multiple) scanning sensors 105 to scan the environment, map the environment, capture environmental data, and / or generate any type of image of the environment (e.g., by generating a 3D representation of the environment or by generating a "perspective" visualization). The (multiple) scanning sensors 105 can include any number or type of scanning devices, without limitation. According to the disclosed embodiments, the HMD 100 can be used to generate a perspective visualization of the user's environment. As used herein, a "perspective" visualization refers to a visualization that reflects the perspective of the environment from the user's point of view. In order to generate this perspective visualization, the HMD 100 can use its (multiple) scanning sensors 105 to scan, map, or otherwise record its surrounding environment (including any objects in the environment) and pass the data to the user for viewing. As will be described immediately, various transformations can be applied to the image before it is displayed to the user to ensure that the displayed perspective matches the user's intended perspective.

[0038] To generate perspective images, the scanning sensor(s) 105 typically rely on its camera (e.g., a head tracking camera, a hand tracking camera, a depth camera, or any other type of camera) to obtain one or more raw images (also referred to as "texture images") of the environment. In addition to generating perspective images, these raw images can also be used to determine depth data detailing the distance (e.g., z-axis range or measurement) from the sensor to any object captured by the raw images. Once these raw images are obtained, a depth map can be calculated from the depth data embedded or included within the raw images (e.g., based on pixel differences), and the depth map can be used for any reprojection to generate perspective images if necessary (e.g., one perspective image is generated for each pupil).

[0039] From the see-through visualization, the user will be able to perceive what is currently in his / her environment without having to remove or reposition the HMD 100. In addition, as will be described in more detail later, the disclosed see-through visualization may also enhance the user's ability to see objects within his / her environment (e.g., by displaying additional environmental conditions that may not be detected by the human eye). As used herein, a so-called "superimposed image" may be a type of see-through image.

[0040] It should be noted that while much of this disclosure focuses on generating "one" perspective image, the embodiments actually generate separate perspective images for each of the user's eyes. That is, two perspective images are typically generated concurrently with each other. Therefore, while reference is frequently made to generating what appears to be a single perspective image, the embodiments are actually capable of generating multiple perspective images simultaneously.

[0041] In some embodiments, scanning sensor(s) 105 include visible light camera(s) 110, low light camera(s) 115, thermal imaging camera(s) 120, potentially (although not necessarily, such as Figure 1 105 , an ultraviolet (UV) camera(s) 125 (represented by dashed boxes), potentially (although not necessarily, as represented by dashed boxes) a point illuminator 130, and even an infrared camera 135. Ellipses 140 demonstrate how any other type of camera or camera system (e.g., a depth camera, a time-of-flight camera, a virtual camera, a depth laser, etc.) may be included in the scanning sensor(s) 105.

[0042] As an example, a camera constructed to detect mid-infrared wavelengths may be included within the scanning sensor(s) 105. As another example, any number of virtual cameras reprojected from an actual camera may be included in the scanning sensor(s) 105 and may be used to generate a stereoscopic image pair. In this manner, the scanning sensor(s) 105 may be used to generate a stereoscopic image pair. In some cases, a stereoscopic image pair may be obtained or generated as a result of any one or more of the following operations: active stereo image generation via use of two cameras and a point illuminator (e.g., point illuminator 130); passive stereo image generation via use of two cameras; image generation using structured light using one actual camera, one virtual camera, and one point illuminator (e.g., point illuminator 130); or image generation using a time of flight (TOF) sensor, wherein a baseline is presented between a depth laser and a corresponding camera and wherein the field of view (FOV) of the corresponding camera is offset relative to the illumination field of the depth laser.

[0043] (Multiple) visible light cameras 110 are typically stereo cameras, which means that the fields of view of two or more visible light cameras at least partially overlap each other. Using this overlapping area, the images generated by (multiple) visible light cameras 110 can be used to identify the differences between certain pixels that collectively represent the object captured by the two images. Based on these pixel differences, embodiments can determine the depth for objects located within the overlapping area (i.e., "stereo depth matching" or "stereo depth matching"). In this way, (multiple) visible light cameras 110 can not only be used to generate perspective visualizations, but they can also be used to determine object depth. In some embodiments, (multiple) visible light cameras 110 can capture both visible light and IR light.

[0044] It should be noted that any number of cameras may be provided on the HMD 100 for each of the different camera types (also referred to as modalities). That is, the visible light camera(s) 110 may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 cameras. However, often, the number of cameras is at least 2 so that the HMD 100 can perform perspective image generation and / or stereo depth matching, as described earlier. Similarly, the low light camera(s) 115, the thermal imaging camera(s) 120, and the UV camera(s) 125 may each include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 corresponding cameras, respectively.

[0045] Figure 2 An example HMD 200 is illustrated, which represents an Figure 1 HMD 100. HMD 200 is shown to include a plurality of different cameras, including cameras 205, 210, 215, 220, and 225. Cameras 205 to 225 represent cameras from Figure 1 Any number or combination of visible light camera(s) 110, low light camera(s) 115, thermal imaging camera(s) 120, and UV camera(s) 125. Figure 2 Only 5 cameras are illustrated in FIG. 2 , but the HMD 200 may include more or less than 5 cameras. Any one of these cameras may be referred to as a “system camera”.

[0046] In some cases, the camera may be located at a specific position on the HMD 200. In some cases, the first camera (e.g., perhaps camera 220) is set at a position on the HMD 200 that is above the designated left eye position of the user wearing the HMD 200 relative to the height direction of the HMD 200. For example, the camera 220 is located above the pupil 230. As another example, the first camera (e.g., camera 220) is additionally positioned above the designated left eye position relative to the width direction of the HMD. That is, the camera 220 is not only located above the pupil 230, but also in a straight line relative to the pupil 230. When the VR system is used, the camera may be placed directly in front of the designated left eye position. Reference Figure 2 , the camera may be physically disposed in the z-axis direction at a position in front of the pupil 230 on the HMD 200 .

[0047] When a second camera is provided (e.g., perhaps camera 210), the second camera may be arranged at a position on HMD 200 above a designated right eye position of a user wearing the HMD relative to the height direction of the HMD. For example, camera 210 is above pupil 235. In some cases, the second camera is additionally positioned above the designated right eye position relative to the width direction of the HMD. When the VR system is used, the camera may be placed directly in front of the designated right eye position. Reference Figure 2 , the camera may be physically disposed at a position in front of the pupil 235 in the z-axis direction on the HMD 200. When the user wears the HMD 200, the HMD 200 is worn on the user's head and the display of the HMD 200 is located in front of the user's pupils (such as the pupil 230 and the pupil 235). Often, the cameras 205 to 225 will be physically offset from the user's pupils 230 and 235 by some distance. For example, there may be a vertical offset in the HMD height direction (i.e., the "Y" axis), as shown by the offset 240. Similarly, there may be a horizontal offset in the HMD width direction (i.e., the "X" axis), as shown by the offset 245.

[0048] The HMD 200 is configured to provide a perspective image 250 for the user of the HMD 200 to view. Thus, the HMD 200 is able to provide a visualization of the real world without requiring the user to remove or reposition the HMD 200. These (multiple) perspective images 250 effectively represent a view of the environment from the perspective of the HMD. Cameras 205 to 225 are used to provide these (multiple) perspective images 250. The offset between the camera and the user's pupil (e.g., offsets 240 and 245) causes parallax. In order to provide these (multiple) perspective images 250, an embodiment can perform parallax correction by applying various transformations and reprojections on the image to change the initial perspective represented by the image to a perspective that matches the perspective of the user's pupil. Parallax correction relies on the use of a depth map to facilitate reprojection.

[0049] In some implementations, as opposed to performing full 3D reprojection, embodiments utilize a planar reprojection process to correct for parallax when generating perspective images. Use of this planar reprojection process is acceptable when objects in the environment are far enough from the HMD. Thus, in some cases, embodiments can avoid performing 3D parallax correction because objects in the environment are far enough away and because this distance causes negligible errors with respect to depth visualization or parallax issues.

[0050] Any of the cameras 205 to 225 constitute so-called “system cameras” (also referred to as first cameras) because they are an integral part of the HMD 200. In contrast, the external camera 255 (also referred to as a second camera) is physically separated and detached from the HMD 200, but can communicate wirelessly through the HMD 200. That is, the external camera 255 is a peripheral camera relative to the HMD 200. As will be described immediately, it is desirable to align the image (or image content) generated by the external camera 255 with the image (or image content) generated by the system camera to subsequently generate an overlay image that can operate as a perspective image.

[0051] Often, the angular resolution of the external camera 255 is higher than the angular resolution of the system camera (i.e., more pixels per degree and not just more pixels), so the resulting overlay image provides enhanced image content over that available using only the system camera image. Additionally or alternatively, the modalities of the external camera 255 and the system camera may be different, so the resulting overlay image may also include enhanced information. As an example, assume that the external camera 255 is a thermal imaging camera. Thus, the resulting overlay image may include both visible light image content and thermal image content.

[0052] Accordingly, providing superimposed perspective images is highly desirable. It should be noted that the external camera 255 can be any of the camera types listed earlier. Additionally, there can be any number of external cameras without limitation.

[0053] Example scenario

[0054] Now turn your attention to Figure 3 , which illustrates how you can use Figure 1 and Figure 2 Example scenario for the HMD discussed in . Figure 3 A building 300 is illustrated, as well as a first responder 305 and another first responder 310. In this example scenario, first responders 305 and 310 want to climb building 300. Figure 4 One example technique for performing this climbing task is shown.

[0055] Figure 4 A first responder is shown in environment 400A wearing HMD 400, which represents the HMD discussed thus far. As previously discussed, HMD 400 includes system camera 405. In addition, the first responder uses tool 410 including external camera 415, which represents Figure 2415. In this case, tool 410 is a grappling gun that will be used to launch a rope and hook onto a building to allow first responders to climb the building. By aligning the image content generated by external camera 415 with the image content generated by system camera 405, the user will be able to better discern where tool 410 is aimed.

[0056] That is, in accordance with the disclosed principles, it is desirable to provide an improved platform or technology by which a user (e.g., a first responder) can aim a tool (e.g., tool 410) using HMD 400, system camera 405, and external camera 415 as a combined aiming interface. Figure 5 One such example is shown.

[0057] Figure 5 The first camera 500 (also referred to as HMD camera) mounted on the HMD is shown, wherein the first camera 500 represents Figure 4 405, and a tool (e.g., a grappling gun) including a second camera 505 representing an external camera 415. It should be noted how the optical axis of the second camera 505 is aligned with the aiming direction of the tool. As a result, the image generated by the second camera 505 can be used to determine the position where the tool is aimed. It will be appreciated that the tool can be any type of aimable tool without limitation.

[0058] exist Figure 5 5, a first camera 500 and a second camera 505 are both aimed at a target 510. For illustration, the field of view (FOV) of the first camera 500 is represented by the first camera FOV 515 (also referred to as the HMD camera FOV), and the FOV of the second camera 505 is represented by the second camera FOV 520. Note that the first camera FOV 515 is larger than the second camera FOV 520. Typically, the second camera 505 provides a very focused view, similar to a viewfinder (i.e., a high level of angular resolution). As will be discussed in more detail later, the second camera 505 sacrifices a wide FOV for increased resolution and increased pixel density. Accordingly, in this example scenario, it can be observed how the second camera FOV 520 can be completely superimposed or covered by the first camera FOV 515 in at least some cases. Of course, in the event that the user aims the second camera 505 in a direction that the first camera 500 is not aimed at, then the first camera FOV 515 and the second camera FOV 520 will not be superimposed.

[0059] Figure 6 Shown Figure 5600 of the first camera FOV 515. The first camera FOV 600 will be captured by the system camera in the form of a system camera image and will potentially be displayed in the form of a perspective image. The system camera image has a resolution 605 and is captured by the system camera based on a determined refresh rate 610 of the system camera. The refresh rate 610 of the system camera is typically between about 30 Hz and 120 Hz. Often, the refresh rate 610 is about 90 Hz or at least 60 Hz. Often, the first camera FOV 600 has a horizontal FOV of at least 55 degrees. The horizontal baseline of the first camera FOV 600 can extend to 65 degrees, or even more than 65 degrees.

[0060] It should also be noted how the HMD includes a system (HMD) inertial measurement unit IMU 615. An IMU (e.g., system IMU 615) is a type of device that measures forces, angular velocity, and orientation of the body. The IMU can detect these forces using a combination of accelerometers, magnetometers, and gyroscopes. Because both the system camera and the system IMU 615 are integrated with the HMD, the system IMU 615 can be used to determine the orientation or pose of the system camera (and the HMD) and any forces experienced by the system camera.

[0061] In some cases, "pose" may include information detailing 6 degrees of freedom, or "6DOF" information. In general, 6DOF pose refers to the movement or position of an object in three-dimensional space. 6DOF pose includes sway (i.e., forward and backward in the x-axis direction), heave (i.e., up and down in the z-axis direction), and sway (i.e., left and right in the y-axis direction). In this regard, 6DOF pose refers to a combination of 3 translations and 3 rotations. Any possible movement of a body can be represented using a 6DOF pose.

[0062] In some cases, the pose may include information describing the details of the 3DOF pose. Typically, 3DOF pose refers only to tracking rotational motion, such as pitch (i.e., the lateral axis), yaw (i.e., the normal axis), and roll (i.e., the longitudinal axis). 3DOF pose allows the HMD to track rotational motion of itself and the system camera rather than translational motion. As a further explanation, 3DOF pose allows the HMD to determine whether the user (the person wearing the HMD) is looking left or right, whether the user is rotating his / her head up or down, or whether the user is pivoting left or right. In contrast to 6DOF pose, when using 3DOF pose, the HMD is not able to determine whether the user (or the system camera) has moved in a translational manner, such as by moving to a new location in the environment.

[0063] Determining 6DOF pose and 3DOF pose can be performed using built-in sensors, such as accelerometers, gyroscopes, and magnetometers (i.e., system IMU 615). Position tracking sensors (such as head tracking sensors) can also be used to determine 6DOF pose. Accordingly, system IMU 615 can be used to determine the pose of the HMD.

[0064] Figure 7 A second camera FOV 700 is shown, which indicates Figure 5 5. The second camera FOV 520 of FIG. 5 is a second camera FOV 520. Note that the second camera FOV 700 is smaller than the first camera FOV 600. That is, the angular resolution of the second camera FOV 700 is higher than the angular resolution of the first camera FOV 600. Having an increased angular resolution also causes the pixel density of the external camera image to be higher than the pixel density of the system camera image. For example, the pixel density of the external camera image is often 2.5 to 3 times the pixel density of the system camera image. As a result, the resolution 705 of the external camera image is higher than the resolution 605. Often, the second camera FOV 700 has a horizontal FOV of at least 19 degrees. The horizontal baseline can be higher, such as 20 degrees, 25 degrees, 30 degrees, or greater than 30 degrees.

[0065] The external camera also has a refresh rate 710. Refresh rate 710 is typically lower than refresh rate 610. For example, the refresh rate 710 of the external camera is often between 20 Hz and 60 Hz. Typically, the refresh rate 710 is at least about 30 Hz. The refresh rate of the system camera is often different from the refresh rate of the external camera. However, in some cases, the two refresh rates can be substantially the same.

[0066] The external camera also includes or is associated with an external IMU 715. Using the external IMU 715, the embodiment is able to detect or determine the orientation / pose of the external camera and any forces experienced by the external camera. Accordingly, similar to the earlier discussion, the external IMU 715 can be used to determine the pose (e.g., 6DOF and / or 3DOF) of the external camera's line of sight.

[0067] According to the disclosed principles, it is desirable to overlay and align an image obtained from an external camera with an image generated by a system camera to generate an overlaid and aligned perspective image. The overlay between the two images enables an embodiment to generate multiple images and then overlay the image content from one image onto another image to facilitate the generation of a composite image or overlaid image with enhanced features that would not be present if only a single image was used. As an example, the system camera image provides a wide FOV, while the external camera image provides high resolution and pixel density for the focus area (i.e., the aiming area where the tool is aimed). By combining the two images, the resulting image will have the advantages of a wide FOV and a high pixel density for the aiming area.

[0068] It should be noted that while the present disclosure is primarily focused on the use of two images (e.g., a system camera image and an external camera image), embodiments are capable of aligning content from more than two images with overlapping areas. For example, assume 2, 3, 4, 5, 6, 7, 8, 9, or even 10 integrated and / or detached cameras have overlapping FOVs. Embodiments are capable of examining each of the resulting images and then aligning specific portions to one another. The resulting overlay image can then be a composite image formed by any combination or alignment of the available images (e.g., even 10 or more images if available). Accordingly, embodiments are capable of utilizing any number of images when performing the disclosed operations and are not limited to only two images or two cameras.

[0069] As another example, assume that the system camera is a low-light camera and further assume that the external camera is a thermal imaging camera. As will be discussed in more detail later, embodiments can selectively extract image content from the thermal imaging camera image and overlay the image content onto the low-light camera image. In this regard, the thermal imaging content can be used to augment or supplement the low-light image content, thereby providing an enhanced image to the user. Additionally, because the external camera has an increased resolution relative to the system camera, the resulting overlay image will provide enhanced clarity for areas where pixels in the external camera image are overlaid onto the system camera image. Figure 8 Examples of these operations and benefits are provided.

[0070] Image Correspondence and Alignment

[0071] According to the disclosed principles, embodiments are able to align an image of a system camera with an image of an external camera. That is, because at least a portion of the FOVs of the two cameras are superimposed on each other, as described earlier, at least a portion of the resulting image includes corresponding content. Therefore, the corresponding content can be identified and a merged, fused, or superimposed image can then be generated based on similar corresponding content. By generating this superimposed image, embodiments are able to provide enhanced image content to the user, which would not be available if only a single image type was provided to the user. Both the image of the system camera and the image of the external camera can be referred to as a "textured" image.

[0072] Different techniques can be used to perform alignment. One technique is a "visual alignment" technique involving feature point detection. Another technique is an IMU-based technique that aligns the images based on the determined poses of the respective cameras. This visual alignment technique generally produces more accurate results.

[0073] With respect to visual alignment techniques, in order to merge or align images, some embodiments can analyze texture images (e.g., perform computer vision feature detection) in an attempt to find any number of feature points. As used herein, the phrase "feature detection" generally refers to the process of computing an image abstraction and then determining whether an image feature (e.g., a particular type) is present at any particular point or pixel in the image. Often, corners (e.g., the corner of a wall), distinguishable edges (e.g., the edge of a table), or ridges are used as feature points because of the inherent or sharp contrast visualization of an edge or corner.

[0074] Any type of feature detector can be programmed to identify feature points. In some cases, the feature detector can be a machine learning algorithm. As used herein, reference to any type of machine learning can include any type of machine learning algorithm or device, (multiple) convolutional neural networks, (multiple) multi-layer neural networks, (multiple) recursive neural networks, (multiple) deep neural networks, (multiple) decision tree models (e.g., decision trees, random forests, and gradient boosting trees), (multiple) linear regression models, (multiple) logistic regression models, (multiple) support vector machines ("SVM"), (multiple) artificial intelligence devices, or any other type of intelligent computing system. Any number of training data (and perhaps later refined) can be used to train the machine learning algorithm to dynamically perform the disclosed operations.

[0075] In accordance with the disclosed principles, embodiments detect any number of feature points (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 500, 1000, 2000, or more than 2000) and then attempt to identify correlations or correspondences between the feature points detected in the system camera image and the feature points identified in the external camera image. As will be described in more detail later, some examples of feature points include corners (also referred to as corner features) and lines (also referred to as line features).

[0076] Some embodiments then fit the features or (multiple) image correspondences to a motion model to facilitate superimposing one image onto another to form an enhanced superimposed image. Any type of motion model may be used. Typically, a motion model is a type of transformation matrix that enables a model, known scene, or object to be projected onto a different model, scene, or object. When reprojecting the image, the feature points are used as reference points.

[0077] In some cases, the motion model may simply be a rotational motion model. With the rotational motion model, an embodiment is able to shift one image by any number of pixels (e.g., perhaps 5 pixels to the left and 10 pixels up) to facilitate overlaying one image onto another. For example, once image correspondences are identified, an embodiment may identify the pixel coordinates of those feature points or correspondences. Once the coordinates are identified, an embodiment may overlay an image of the external camera's line of sight onto the image of the HMD camera using the rotational motion model approach described above.

[0078] In some cases, the motion model can be more complex, such as in the form of a similarity transformation model. The similarity transformation model can be configured to allow for (i) rotation of either the image of the HMD camera or the image of the external camera line of sight; (ii) scaling of those images; or (iii) homographic transformation of those images. In this regard, the similarity transformation model method can be used to overlay image content from one image onto another. Accordingly, in some cases, the process of aligning the external camera image with the system camera image is performed by the following steps: (i) identifying image correspondences between the images; and then (ii) fitting the correspondences to the motion model based on the identified image correspondences so that the external camera image is projected onto the system camera image. Another technique for aligning images includes using IMU data to predict the poses of the system camera and the external camera. Once the two poses are estimated or determined, the embodiments then use these poses to align one or more portions of the images with each other. Once aligned, one or more portions of one image (these portions are aligned portions) are overlaid on corresponding portions of the other image to generate an enhanced overlaid image. In this regard, an IMU can be used to determine the poses of the corresponding cameras, which can then be used to perform the alignment process. IMU data is almost always readily available. However, a visual alignment process may not be able to be performed.

[0079] Figure 8 The resulting overlay image 800 is shown, which includes part (or all) of a system (HMD) camera image 805 (i.e., an image generated by the system camera) and an external camera image 810 (i.e., an image generated by an external camera). These images are aligned using an alignment 815 process (e.g., visual alignment, IMU-based alignment, and / or hardware-based alignment). Optionally, additional image artifacts may be included in the overlay image 800, such as a crosshair 820, perhaps used to help the user aim the tool. By aligning the image content, the user of the tool can determine where the tool is aimed without having to look down the line of sight of the tool. Instead, the user can discern where the tool is aimed by simply looking at what is displayed in his / her HMD.

[0080] Providing enhanced overlay image 800 allows for rapid target acquisition, such as Fig. 9 Target acquisition 900 is shown in FIG. That is, the target can be acquired (ie, the tool is accurately aimed at the desired target) in a fast manner because the user no longer needs to spend time looking at the line of sight of the tool.

[0081] Vision Alignment Method

[0082] Fig.10 An abstract version of the images discussed so far is shown and focuses on the visual alignment method. Specifically, Fig.10 A first camera image 1000 having a feature point 1005 and a second camera image 1010 having a feature point 1015 corresponding to the feature point 1005 are shown. An embodiment can perform a visual alignment 1020 between the first camera image 1000 and the second camera image 1010 using the feature points 1005 and 1015.

[0083] Visual alignment 1020 may be performed via a reprojection 1020A operation, wherein a pose 1020B implemented in the first camera image 1000 and / or a pose implemented in the second camera image 1010 are reprojected to a new position to align one of those images with the other. Reprojection 1020A may be facilitated using motion data, such as perhaps inertial measurement unit (IMU) data 1020C. For example, IMU data may be collected to describe any movement that occurred between the time the first camera image 1000 was taken and the time the second camera image 1010 was taken. This IMU data 1020C may then be used to transform a motion model 1020D to perform reprojection 1020A. Visual alignment 1020 may also rely on ensuring that a threshold 1020E number of features from the second camera image 1010 correspond to similar features found in the first camera image 1000.

[0084] The result of performing the visual alignment 1020 is the generation of an overlay image 1025. The overlay image 1025 includes a portion extracted or obtained from the first camera image 1000 and a portion extracted or obtained from the second camera image 1010. Note that in some embodiments, the overlay image 1025 includes a boundary element 1030 encompassing pixels obtained from the second camera image 1010 and / or from the first camera image 1000. Optionally, the boundary element 1030 may be in the form of a circular bubble visualization 1035. However, other shapes may also be used for the boundary element 1030.

[0085] Interlaced feature extraction

[0086] Having now described some of the various processes for identifying features (also called feature points) and using those features to align images, we now turn our attention to Fig.11A , Fig. 11B and Fig. 11C , which illustrates various processes for performing interleaved feature extraction.

[0087] Fig.11A An example scene is shown including a first camera 1100 and a second camera 1105, which represent the cameras mentioned previously. The first camera 1100 is shown at time T 1The embodiment processes the image 1105A to generate the feature set 1105B. Note that the timing for when the feature 1105B will be generated will be at time T 1 Thereafter, because there is the processing time required to perform feature extraction and the time required for preprocessing the image. The timing may also depend on when the preprocessing and feature extraction steps are scheduled to run. A feature extraction process, operation, or technique is performed to identify these features 1105B. The time required to perform feature extraction may vary based on different factors, such as the complexity of the image, the resolution of the image, the type of image, whether historical data is available (e.g., features from previous images may be used to help initially isolate or focus the feature extractor on a specific area of ​​the image to identify features), etc.

[0088] At time T 3 At time T, the first camera 1100 generates another image 1110A and generates another feature set 1110B. Similarly, at time T 5 At 115A, another image 1115A is generated along with a corresponding set of features 1115B. It is noteworthy that the set of features 1115B is typically generated after the image arrives, as indicated above. For example, a delay may occur based on the time required to perform preprocessing and feature extraction. The delay may also be based on the scheduling of when these tasks occur. Therefore, the typical situation is that the features are generated some time after the image arrives.

[0089] The second camera 1105 is shown at time T 0 An image 1120A is generated at time T. A feature set 1120B is identified from the image 1120A. 2 At time T, the second camera 1105 generates another image 1125A and generates another feature set 1125B. Similarly, at time T 4 At , another image 1130A is generated along with a corresponding feature set 1130B.

[0090] The processing associated with analyzing the various different images and generating features may be performed in a variety of different ways. In some cases, a single computing thread is relied upon to perform the processing. In some cases, multiple threads 1135 are relied upon. For example, a first thread may be responsible for analyzing images generated by a first camera 1100 and a second thread may be responsible for analyzing images generated by a second camera 1105. In other words, in some embodiments, a first computing thread is dedicated to processing images generated by a first camera and a second computing thread is dedicated to processing images generated by a second camera.

[0091] In some implementations, the cameras are not synchronized with each other. That is, in some embodiments, the cameras operate in an asynchronous 1140 manner. Therefore, the images may arrive at different times relative to each other. In some cases, the generation of the images may be synchronous, but due to the differences in the images, the feature extraction may be asynchronous. Fig. 11B 110A is shown in which an image alignment operation 1145 is being performed. The image alignment operation 1145 is performed using the process described earlier and is based on correlations between the identified features. Note that even though image 1105A arrives after image 1120A (e.g., time T 1 With T 0 ), the image alignment operation 1145 is also performed.

[0092] Embodiments can use the IMU data to reproject feature 1120B to match the pose of image 1120A with the pose of image 1105A, in the manner previously described.

[0093] Fig. 11C A new image alignment operation 1150 is shown. According to the disclosed principles, embodiments can reuse a previously used feature set to perform a subsequent image alignment operation. For example, Fig. 11C It is shown how an embodiment reuses 1155 feature 1120B instead of using newer feature 1125B. Although feature 1120B is older than feature 1125B, an embodiment can still use the IMU data to perform various reprojections in order to correctly align the pose.

[0094] Reusing a feature set that is already being used can be beneficial for a number of reasons. For example, it may be the case that the frame rates of different cameras are very different; perhaps one camera operates at 60 frames per second (FPS) while another operates at 30 FPS. In some cases, the frame rates of the cameras may be between 30 FPS and 120 FPS.

[0095] By reusing features, embodiments can avoid having to wait until a new set of features is identified in a new image. As a result, embodiments can avoid delays associated with waiting. As another example, the embodiments can avoid or reduce delays associated not only with waiting to generate a new frame, but also with performing feature extraction on the new frame. When a new frame arrives, some embodiments operate in parallel by performing feature extraction on the new frame while performing visual alignment by reusing features from the old frame, even if those features have already been used. The newly generated features can then be used during subsequent visual alignment operations. From this discussion, it can be easily discerned that the various timing advantages provided by the disclosed principles, particularly with respect to avoiding the various drawbacks of having delays.

[0096] Over time, the camera will generate multiple frames, and multiple features will be generated. That is, a set of features can be identified from within each frame. In accordance with the disclosed principles, these embodiments are advantageously able to aggregate features generated over time as a result of performing feature extraction on multiple images generated by a single camera. These embodiments can then store or cache the history of these features for various beneficial uses, which will be discussed momentarily. Fig.12 is illustrative.

[0097] Fig.12 Three feature sets are shown, namely features 1200, 1205, and 1210. These features are generated by images generated by a single phase. As an example, feature 1200 may correspond to feature 1105B associated with image 1105A. Feature 1205 may correspond to feature 1110B associated with image 1110A. Feature 1210 may correspond to feature 1115B associated with image 1115A. Images 1105A, 1110A, and 1115A are all generated by the same camera, namely camera 1100.

[0098] An embodiment can acquire IMU data for each image generated by a camera.

[0099] Embodiments may then use the IMU data to reproject the poses in the image to a common pose. The result is that the features are all projected to the same pose. Embodiments may then form a union or aggregation of the various features to form a history 1215 of aggregated features. The history 1215 of features may be stored, such as in a cache 1220. Accordingly, in some embodiments, a collection of features is cached and available for reuse during multiple image alignment operations.

[0100] Storing an aggregated set of features is beneficial for a number of reasons. For example, suppose the camera is operating in a low light environment or perhaps in an isothermal or low contrast environment. It may be the case that during feature extraction on the images, the feature extractor may be able to identify only a select few features within any single image. If visual alignment is attempted using only a few features, then visual alignment will likely fail.

[0101] On the other hand, if the feature extractor is able to identify features from multiple different images, there is a higher likelihood that a significantly greater number of unique features can be identified. Embodiments may then collect or utilize this aggregated set of features to perform visual alignment. Therefore, maintaining a library or history of features for the cameras may be very beneficial as it may aid in the visual alignment process.

[0102] In this regard, a first history of feature extraction results (e.g., features) can be saved for the first camera in cache 1220. Similarly, a second history of feature extraction results can be saved for the second camera. In some cases, the generation of the first feature set and the generation of the second feature set are performed asynchronously. Similarly, the generation of the first image can be performed asynchronously relative to the generation of the second image.

[0103] Example Method

[0104] The following discussion now refers to a number of methods and method actions that can be performed. Although method actions may be discussed in a particular order or illustrated in a flowchart as occurring in a particular order, no particular order is required unless specifically stated or because an action depends on another action being completed before the action is performed.

[0105] Now turn your attention to Fig.13A , which illustrates a flow chart of an example method 1300A for performing image alignment between a first image generated by a first camera and a second image generated by a second camera, wherein the image alignment is performed using interleaved feature extraction, wherein a feature set is reused to align the second image with the first image. The method 1300A may optionally use Figure 1 In some cases, method 1300A may be performed by a cloud service running in a cloud environment. Method 1300A includes an action of accessing a first image generated by a first camera (action 1305). For example, the first image may be Fig.11A The first camera may optionally be the first camera 1100 .

[0106] Action 1310 includes identifying a first set of features from within a first image by performing feature extraction on the first image. For example, feature extraction may be performed on image 1110A to extract or identify a first set of features from within the first image. Fig.11A Feature 1110B.

[0107] Action 1315 is shown as being asynchronous or asynchronous with the execution of action 1305 and perhaps even asynchronous or asynchronous with the execution of action 1310. Action 1315 includes determining not to wait until the second camera generates a current image before performing the current image alignment operation. The determination is based on a determination that waiting for the current image will introduce a delay when performing the current image alignment operation, wherein the introduced delay will exceed an allowable delay threshold. For example, consider Fig.11A. Note that image 1110A and image 1125A were generated at times T3 and T2, respectively. It may be the case that these two times are relatively close to each other. However, in some cases, it may be the case that image 1125A has a significantly higher resolution or image content than image 1110A. In other words, it may be the case that one of the first image or the second image has a higher image resolution than the other of the first image or the second image.

[0108] Optionally, the determination of a single-threaded interleaved feature extraction implementation (which may wait for the arrival of a new image) using image alignment may be determined at build time or may depend on configuration. In some cases, one of the cameras may be selected to reduce the delay between the arrival of an image from the camera and the output of an alignment result. The single-threaded interleaved feature extraction cycle may begin by waiting for a new image from a camera (e.g., a first camera). Preprocessing and feature extraction may then be performed on the image. Feature matching may then be performed on the feature set extracted from the second camera and the latest cached feature set. Alignment calculations are then performed. This processing minimizes the time between the arrival of an image from the first camera to the alignment result generated by the image. The cycle may then continue by obtaining the latest available image from the second camera, preprocessing and extracting features therefrom, and then performing feature matching and alignment calculations using the extracted feature set and the latest cached first camera feature set generated in the first half of the cycle.

[0109] It should also be noted that if a single thread implementation is used, these processes can be performed serially with each other. If multiple threads are used, the process can be performed in parallel. In addition, after generating the image alignment results using the latest image from the first camera and the latest feature set extracted from the second camera, an embodiment can repeat the process using the latest image from the second camera and the feature set extracted from the image of the first camera.

[0110] Returning to the earlier discussion, in some cases, the feature extraction process performed on image 1125A may take significantly longer than the feature extraction process performed on image 1110A. Embodiments can determine that they will not wait until features 1125B for image 1125A are generated. Instead, embodiments may choose to perform the visual alignment operation using an older set of features, such as perhaps features 1120B for image 1120A generated at time To. Those older features may have been used at least once during a previous image alignment operation, in accordance with the disclosed principles.

[0111] Based on the above decision, the embodiment then accesses (act 1320) a second image generated by the second camera. Notably, the second image was generated before the time the first image was generated, so the features associated with the image are older than the features associated with the first image. The first image may be a first image type and the second image may be a second image type different from the first image type. In some cases, they may be the same image type.

[0112] Instead of waiting for the current image, action 1325 includes accessing a second set of features previously detected within the second image. The second set of features was previously used at least once to perform a previous image alignment operation between the second image and a different image. For example, from Fig.11A The features 1120B may represent these "second feature sets" in method 1300A. Fig. 11B In , feature 1120B is used for the first time in image alignment operation 1145. Then, in Fig. 11C In some embodiments, the same feature set 1120B is reused (i.e., used at least one additional time) in the image alignment operation 1150. Accordingly, in some embodiments, the second feature set is generated before the time when the first feature set is generated. Furthermore, the second feature set is accessible before the first feature set is accessible.

[0113] Action 1330A then includes performing a current image alignment operation by aligning the first image with the second image using the first feature set and by reusing the second feature set based on the identified correspondence between the first feature set and the second feature set. As a result of performing the current image alignment operation, the second portion of the second image is identified as corresponding to the first portion of the first image. Recall that the visual alignment operation includes using motion data (e.g., perhaps IMU data) to reproject the image (and therefore its features) to a new pose that matches the pose implemented in the first image. In other words, the process of performing the current image alignment may include accessing inertial measurement unit (IMU) data from one or both of the first camera and the second camera. The process may then include subsequently reprojecting one or both of the first image and the second image to a common pose. In some cases, the pose will be the pose of an image generated later. In other words, the IMU data may be used to update the data used to perform the image alignment operation. Such an update may, for example, include calculating the alignment at a common timestamp of the cameras.

[0114] In case the first and second images are of different types, then the current image alignment operation will be performed using images of the different image types.In addition, the correspondence between the identified first and second feature sets may be performed using features obtained from the different image types.

[0115] The current image alignment operation may be performed using a motion model. The motion model facilitates the reprojection of a pose embodied in one of the first or second images to align it with the pose embodied in the other of the first or second images. Typically, the earlier generated image will have its pose reprojected to match, align, or correspond to the pose in the later generated image. In some implementations, the pose of both images may optionally be reprojected to correspond to a new pose, such as perhaps a pose of a predicted pose. The predicted pose may be an estimated or predicted pose of the camera at some future point in time.

[0116] After performing the current image alignment operation, action 1335A includes generating an aligned image by superimposing the second portion derived from the second image onto the first image. In other words, the process of generating the aligned image may include generating the transformation required to superimpose the first image (or a portion thereof) on the second image based on the alignment result, or vice versa. Notably, the second portion is superimposed on the first portion of the first image. Figure 8 and Fig.10 Represents the superimposed image. Fig.10 As shown, a border may be included in the aligned images. The border may be configured to encompass the second portion from the second image.

[0117] In some cases, these same second feature sets can be reused for subsequent image alignment operations. As a result, the second feature sets are used at least three times. In some cases, they can be used more than three times, such as perhaps 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, or perhaps even 10 times. An embodiment can use IMU data to reproject the feature points to new poses to facilitate alignment. The same feature points can be reused any number of times, and the FOVs provided by the cameras remain roughly aligned with each other.

[0118] In some cases, the method may also include aggregating the first feature set with additional features detected in one or more other images generated by the first camera. An embodiment may then cache the aggregated feature set. In some cases, the first feature set is cached and then used in one or more subsequent image alignment operations.

[0119] Fig. 13B A variation of method 1300A is shown in the form of method 1300B. That is, instead of performing acts 1330A and 1335A, method 1300B replaces act 1330A with act 1330B and replaces act 1335A with act 1335B.

[0120] In particular, instead of successfully performing the current image alignment operation, the embodiment attempts to perform the current image alignment operation by using the first feature set and by reusing the second feature set to attempt to align the first image with the second image based on the identified correspondences between the first feature set and the second feature set. Optionally, the embodiment may require a threshold number of correspondences identified between the first feature set and the second feature set.

[0121] Action 1330B then involves determining that the plurality of correspondences identified between the first feature set and the second feature set do not satisfy a correspondence threshold, such that the current image alignment operation fails. Fig.10 It is described how a threshold 1020E may be established, where the threshold 1020E may require at least a minimum number of correspondences between identified features in order to perform visual alignment 1020 .

[0122] Action 1335B then includes performing a new image alignment operation using a different arrangement of features associated with the image generated by the second camera. Fig.11A In order to extract features 1120B from an image generated even before image 1120A, some embodiments use different arrangements of features, such as perhaps extracting features from an image generated even before image 1120A. Recall that embodiments can maintain a history of image features. In some cases, these features can be grouped together. In some cases, these features are maintained separately. These different sets of features are referred to as different arrangements. In the event that the arrangement currently used fails to achieve or provide a threshold number of correspondences, embodiments can select from previous arrangements of features. A different arrangement of features can be generated prior to the second set of features in time. Optionally, a threshold number of features in the different arrangement of features can be determined as copies of features in the first set of features. In other words, these copies can be used beneficially to assist in the alignment process. It is desirable to identify features of the copies from different images so that the system can perform proper pose alignment on these images.

[0123] Fig. 13C A method 1300C for performing the disclosed operations using a multi-threaded approach is shown. Specifically, the method 1300C can be performed by a first thread (eg, "thread 1") and a second thread (eg, "thread 2").

[0124] Action 1340, action 1345, action 1350, and action 1355 are performed by thread 1. Action 1360, action 1365, action 1370, and action 1375 are performed by thread 2.

[0125] Act 1340 includes accessing a first image. Act 1345 then includes identifying a first set of features from the first image. Act 1350 includes accessing a second set of features from the second image. An image alignment operation is then performed in act 1355.

[0126] Act 1360 includes accessing the third image. A third feature set from the third image is identified (eg, act 1365). Act 1370 then includes accessing a fourth feature set from the fourth image. Finally, in act 1375, an image alignment operation is performed.

[0127] Example method for aggregating features

[0128] Fig.14 A flow chart of an example method 1400 for generating an aggregated feature set from multiple images generated by a camera is shown. The method 1400 may also be performed by the disclosed HMD and / or cloud service.

[0129] Method 1400 includes an act of accessing a first image generated by a camera (act 1405). The first image is generated at a first time. Act 1410 includes identifying a first set of features from within the first image by performing feature extraction on the first image.

[0130] Method 1400 also includes action 1415, which is performed after action 1405, but which can optionally be performed in parallel with action 1410. Action 1415 includes accessing a second image generated by the camera. The second image is generated at a second time after the first time. Action 1420 then includes identifying a second set of features from within the second image by performing feature extraction on the second image.

[0131] Method 1400 also includes act 1425, which is performed after act 1415, but which may optionally be performed in parallel with acts 1410 and / or 1420. Act 1425 includes obtaining movement data describing details of movement of the camera between the first time and the second time.

[0132] Action 1430 then includes using the motion data to reproject the pose embodied in the first image to correspond to the pose embodied in the second image. A motion model may be used to facilitate the reprojection. In other words, reprojecting the pose embodied in the first image to correspond to the pose embodied in the second image may involve the use of a motion model. In addition, the movement data may include IMU data.

[0133] Action 1435 includes aggregating the first feature set subjected to the reprojection operation with the second feature set. The aggregation results in the generation of an aggregated feature set. Aggregation may include combining and / or merging features between different feature sets.

[0134] Action 1440 then includes caching the aggregated set of features for the camera. The aggregated features may be cached or stored on the HMD or in the cloud. In some implementations, additional feature sets identified from additional images generated by the camera may be aggregated with the aggregated set of features. For example, features for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 images may all be aggregated together. Optionally, the aggregated set of features may be compiled into a single unit vector space.

[0135] In some cases, each aggregated set of features in the set of aggregated features is labeled with a corresponding timestamp of when the image was generated. Thus, multiple different arrangements of features may be included in the aggregated set of features, with each arrangement being labeled with temporal data.

[0136] Optionally, the method may include subsequently accessing the camera's aggregated feature set during an image alignment operation. Through the image alignment operation, an image generated by the camera is aligned with a different image generated by a different camera. The historically aggregated features can be used to help improve the alignment process, particularly for scenes where a single image may not have a sufficient number of detectable features (e.g., such as perhaps a scene with isothermal or low contrast). For example, consider a scene in which a first image and a second image capture an isothermal scene or a low contrast scene. In this scenario, it may be the case that the number of features included in the first feature set is less than a threshold number. It may also be the case that the number of features included in the second feature set is also less than the threshold number. However, after aggregating the first feature set with the second feature set, the number of features included in the aggregated feature set may at least meet the threshold number.

[0137] In some embodiments, the first image and the second image may be of the same image type or may be of different image types. Optionally, the image type may include a thermal image type, a low-light image type, or a visible light image type. In some cases, the resolutions of the two images may be different.

[0138] Accordingly, when performing image alignment, there are a variety of ways to benefit from the flexibility granted by separating the image processing and feature extraction steps from the rest of the image alignment task. The processing of visual alignment may include various feature extraction processes, resizing processes, denoising processes, etc. Such processes may result in waiting delays or latency. Embodiments can beneficially reduce latency by reusing previously extracted features from older images when extracting features for newer images. Optionally, some embodiments may selectively skip image processing and feature extraction for selecting one or more frames in an attempt to save computation and power.

[0139] Image alignment can also optionally be performed using a multithreaded approach. For example, an embodiment may dedicate a thread to each image source. The thread can perform image processing and feature extraction on the latest image from the camera source. The thread can also cache feature extraction results. Using a multithreaded approach allows an embodiment to optionally limit the iteration rate used to perform the visual alignment process (including feature extraction). For example, it may be beneficial to set limits based on the available computational or power constraints / budget of the system, certain characteristics of the camera source, or perhaps the determined importance of a particular image alignment task.

[0140] If alignment fails or feature matching does not generate a sufficient number of matches using the most recent set of feature extraction results, an embodiment may try selecting a different feature set permutation from the feature extraction results history.An embodiment may then retry image alignment.

[0141] The aggregation process may involve mapping feature locations from image coordinates to unit vector space or some other coordinate system representing the real-world orientation of features relative to the image source (e.g., using IMU / gyroscope data to determine the relative pose history of the image source). Feature extraction result aggregation may involve feature matching between the latest feature extraction results and the aggregated feature extraction results to detect duplicate features and filter out older results.

[0142] Compared to a conventional image alignment workflow in which feature extraction is performed on the latest image from all image sources for each image alignment attempt, separating feature extraction and caching feature extraction results implements the above optimization strategy. This strategy and process can be useful in the following applications: (i) where it is desirable to minimize the delay between image arrival and alignment result output; (ii) when the frequency of alignment result output is important; (iii) on systems where image sources are not synchronized; (iv) on systems with limited computational / power budgets; (v) when image alignment is performed on scenes with poor features; and (vi) on systems where image sources move through complex scenes with objects that can temporarily obstruct the view.

[0143] Beneficially, each feature extraction result can be used in multiple alignment attempts, increasing the likelihood of finding a sufficiently large number of accurate feature matches to perform the alignment. Another advantage is that the images can be updated slowly in an asynchronous manner. Thus, embodiments do not need to wait for new images from each source pair before performing alignment. These embodiments also beneficially allow resource-constrained systems to update alignment results more frequently while expending the same computation / power on the image processing and feature extraction steps. If the alignment task is blocked until signaled by the receipt of an image from a particular source, the disclosed method can help reduce the delay between receiving an image from that source and generating the alignment output because feature extraction for other image sources has already been performed.

[0144] Using a multithreaded asynchronous feature extraction approach also provides various advantages. For example, if the iteration rate in each thread is not constrained, the scheme can maximize the frequency of updating the image alignment results. On resource-constrained systems, a fine trade-off between computation / power usage and alignment frequency can be made by rate-limiting feature extraction and alignment performed on each thread according to the image source associated with the thread.

[0145] There are also various advantages to retaining the history of feature extraction results. For example, feature extraction may not output the same set of features even if the image source is stationary and the scene is stationary. Sensor noise, air density gradients, and other factors may change the location of the features found and the values ​​in the generated feature descriptors. Having a history of feature extraction results (so that feature matching can be performed using different permutations of feature sets) increases the likelihood of finding a good feature matching set to perform alignment. If the image source is moving or the scene is dynamic (e.g., leaves move in the wind, obscuring the view), different feature sets may be obscured from the view of each image source at any given time. It can reduce the quality of alignment or reduce the rate at which successful alignment results are generated by conventional alignment methods. By retaining the history of feature extraction results and by trying to align different permutations of these feature sets, these embodiments beneficially provide more frequent and better alignment results. In the case where data from an IMU or similar sensor becomes unreliable (e.g., when recovering from a sudden impact), performing image alignment on a feature set history can be useful in obtaining an estimate of device motion.

[0146] Furthermore, conventional feature-based image alignment methods may have trouble finding a sufficiently large number of accurate feature matches to perform alignment when presented with feature-poor scenes (e.g., foggy environments, dark environments, etc.). Aggregating the feature extraction history of an image source into a single coordinate space (e.g., unit vector space) may enable the use of different features found in past images from that source and may allow for better feature matching output and better alignment. Accordingly, the disclosed principles provide a number of advantages and technical achievements.

[0147] Example Computer / Computer System

[0148] Now turn your attention to Fig.15 , which illustrates an example computer system 1500 that may include and / or be used to perform any of the operations described herein (such as the disclosed methods). Computer system 1500 can take a variety of different forms. For example, computer system 1500 can be implemented as a tablet 1500A, a desktop or laptop computer 1500B, a wearable device 1500C (e.g., an HMD), a mobile device, or any other stand-alone device, as shown by ellipsis 1500D. Computer system 1500 can also be a distributed system including one or more connected computing components / devices that communicate with computer system 1500. Computer system 1500 can be a system operating in a cloud environment. In its most basic configuration, computer system 1500 includes a variety of different components. Fig.15 Computer system 1500 is shown including one or more processors 1505 (also referred to as “hardware processing units”) and memory 1510 .

[0149] With respect to processor(s) 1505, it should be understood that the functionality described herein may be performed, at least in part, by one or more hardware logic components (e.g., processor(s) 1505). For example, but not limited to, illustrative types of hardware logic

[0150] Components / processors that may be used include field programmable gate arrays (“FPGAs”), application-specific or application-specific integrated circuits (“ASICs”), application-specific standard products (“ASSPs”), systems on a chip (“SOCs”), complex programmable logic devices (“CPLDs”), central processing units (“CPUs”), graphics processing units (“GPUs”), or any other type of programmable hardware.

[0151] As used herein, the terms "executable module," "executable component," "component," "module," or "engine" may refer to a hardware processing unit or a software object, routine, or method that may be executed on the computer system 1500. The different components, modules, engines, and services described herein may be implemented as objects or processors executing on the computer system 1500 (e.g., as separate threads).

[0152] Memory 1510 may be physical system memory, which may be volatile, non-volatile, or some combination of the two. The term "memory" may also be used herein to refer to non-volatile mass storage devices, such as physical storage media. If computer system 1500 is distributed, processing, memory, and / or storage capabilities may also be distributed.

[0153] Memory 1510 is shown as including executable instructions 1515. Executable instructions 1515 represent instructions executable by processor(s) 1505 of computer system 1500 to perform disclosed operations, such as those described in various methods.

[0154] The disclosed embodiments may include or utilize a special-purpose or general-purpose computer including computer hardware, such as one or more processors (such as (multiple) processors 1505) and system memory (such as memory 1510), as discussed in more detail below. Embodiments also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media that can be accessed by a general-purpose or special-purpose computer system. Computer-readable media that store computer-executable instructions in the form of data are "physical computer storage media" or "hardware storage devices." In addition, computer-readable storage media including physical computer storage media and hardware storage devices exclude signals, carriers, and propagating signals. On the other hand, computer-readable media that carry computer-executable instructions are "transmission media" and include signals, carriers, and propagating signals. Therefore, by way of example and not limitation, the current embodiments may include at least two distinct types of computer-readable media: computer storage media and transmission media.

[0155] Computer storage media (also referred to as "hardware storage devices") are computer-readable hardware storage devices, such as RAM, ROM, EEPROM, CD-ROM, RAM-based solid-state drives ("SSD"), flash memory, phase change memory ("PCM") or other types of memory, or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired program code means in the form of computer-executable instructions, data or data structures and can be accessed by a general or special purpose computer. The computer system 1500 can also be connected (via a wired or wireless connection) to an external sensor (e.g., one or more remote cameras) or device via the network 1520. For example, the computer system 1500 can communicate with any number of devices or cloud services to obtain or process data. In some cases, the network 1520 itself can be a cloud network. In addition, the computer system 1500 can also be connected to a remote / separate (multiple) computer system configured to perform any processing described with respect to the computer system 1500 via one or more wired or wireless networks.

[0156] A "network" similar to network 1520 is defined as one or more data links and / or data switches that enable electronic data to be transmitted between computer systems, modules and / or other electronic devices. When information is transmitted or provided to a computer through a network (hardwired, wireless, or a combination of hardwired and wireless), the computer appropriately regards the connection as a transmission medium. Computer system 1500 will include one or more communication channels used to communicate with network 1520. Transmission media include networks that can be used to carry data or desired program code devices in the form of computer executable instructions or in the form of data structures. In addition, these computer executable instructions can be accessed by general or special computers. The combination of the above should also be included in the scope of computer-readable media. When arriving at various computer system components, program code devices in the form of computer executable instructions or data structures can be automatically transmitted from transmission media to computer storage media (and vice versa). For example, computer executable instructions or data structures received over a network or data link may be cached in RAM within a network interface module (e.g., a network interface card or "NIC") and then ultimately transferred to computer system RAM and / or to less volatile computer storage media at the computer system. Thus, it should be understood that computer storage media may be included in computer system components that also (or even primarily) utilize transmission media.

[0157] Computer executable (or computer interpretable) instructions include, for example, instructions that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a specific function or group of functions. Computer executable instructions can be, for example, binary code, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in a language specific to structural features and / or method actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions above. On the contrary, the described features and actions are disclosed as example forms of implementing the claims.

[0158] Those skilled in the art will appreciate that the embodiment can be practiced in a network computing environment with multiple types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, pagers, routers, switches, etc. The embodiment can also be practiced in a distributed system environment, where local and remote computer systems linked by a network (by a hardwired data link, a wireless data link, or a combination of hardwired and wireless data links) each perform tasks (e.g., cloud computing, cloud services, etc.). In a distributed system environment, program modules can be located in both local and remote memory storage devices.

[0159] Without departing from the features of the present invention, the present invention may be implemented in other specific forms. The described embodiments are considered to be illustrative and non-restrictive in all aspects. Therefore, the scope of the present invention is indicated by the appended claims rather than by the aforementioned embodiments. All changes within the meaning and scope of the equivalents of the claims are included within their scope.

Claims

1. A method for performing image alignment between a first image generated by a first camera and a second image generated by a second camera, wherein the image alignment is performed using interleaved feature extraction in which a set of features is reused to align the second image with the first image, the method include: accessing the first image generated by the first camera; accessing the second image generated by the second camera, wherein the second image was generated before a time when the first image was generated; identifying a first set of features from within the first image by performing feature extraction on the first image; accessing a second set of features previously detected within the second image, wherein the second set of features was previously used at least once to perform a previous image alignment operation between the second image and a different image; Based on the identified correspondence between the first feature set and the second feature set, a current image alignment operation is performed by using the first feature set and by reusing the second feature set to align the first image with the second image, wherein, as a result of performing the current image alignment operation, a second portion of the second image is identified as corresponding to the first portion of the first image.

2. The method according to claim 1, in: After performing the current image alignment operation, the method further comprises generating an aligned image by superimposing the second portion originating from the second image onto the first image, wherein the second portion is superimposed onto the first portion of the first image, and The first image has a higher image resolution than the second image.

3. The method of claim 1, wherein a first computing thread is dedicated to processing images generated by the first camera, including the first image, and wherein a second computing thread is dedicated to processing images generated by the second camera, including the second image. 4 . The method of claim 1 , wherein the second feature set is cached and made available for reuse during multiple image alignment operations. 5 . The method of claim 1 , wherein a first history of feature extraction results is saved for the first camera, and wherein a second history of feature extraction results is saved for the second camera.

6. The method of claim 1, wherein the generating of the first feature set and the generating of the second feature set are performed asynchronously, and wherein the generating of the first image is performed asynchronously relative to the generating of the second image.

7. The method of claim 1, wherein the second feature set is generated prior to a time when the first feature set is generated such that the second feature set is accessible before the first feature set is accessible.

8. The method of claim 1 , wherein performing the current image alignment comprises accessing inertial measurement unit (IMU) data from one or both of the first camera and the second camera, and wherein the IMU data is used to update data used to perform the current image alignment.

9. The method of claim 1, wherein the second set of features is reused for subsequent image alignment operations such that the second set of features is used at least three times.

10. The method according to claim 1, wherein the method further include: aggregating the first set of features with additional features detected in one or more other images generated by the first camera; as well as The aggregated feature set is cached.

11. The method of claim 1, wherein the first set of features is cached, and wherein the first set of features is subsequently reused for a subsequent image alignment operation.

12. The method of claim 1 , wherein the first image is of a first image type and the second image is of a second image type different from the first image type, such that the current image alignment operation is performed using images of different image types, and wherein the identified correspondence between the first feature set and the second feature set is performed using features obtained from different image types.

13. The method of claim 1, wherein a boundary is included in the aligned image, the boundary being constructed to encompass the second portion originating from the second image.

14. The method of claim 1, wherein the current image alignment operation is performed using a motion model that facilitates reprojection of a pose embodied in the second image to align with a pose embodied in the first image.

15. The method of claim 1, wherein a threshold number of correspondences is identified between the first feature set and the second feature set.